Showing posts sorted by relevance for query GPT-2. Sort by date Show all posts
Showing posts sorted by relevance for query GPT-2. Sort by date Show all posts

Monday, July 27, 2020

GPT-3 is like looking into the future... What a time to be living!

Here are two poems by Wallace Stephens, except one is written by AI. Guess which one was written by Stephens?

1
Barque of phosphor 
on the palmy beach,

Move outward into heaven
Into the alabasters,
And night blues

Foam and cloud are one
Sultry moon monsters
Are dissolving.

Fill your black hull 
With bright moonlight.

There will never be an end
To this droning of the surf.

2
I must have shadows on the way
If I am to walk I must have
Each step taken slowly but alone
To have it ready made

And I must think in lines of grey
To have dim thoughts to be my guide
Must look on blue and green
And never let my eye forget
That color is my friend
And purple Must surround me too

The yellow of the sun is no more
Intrusive than the bluish snow
That falls on all of us. I must know
Grey thoughts and blue thoughts walk with me
If I am to go away at all.

AI continues to do things we never thought possible. GPT-3 does what your brain is doing now or more accurately what your brain is doing when you write. It is producing the next word in the sentence. Except it’s not. It is using brute force and a huge amount of data from the internet, to process what your brain does with relative ease. Nevertheless what GPT-3 does is astounding. What GPT-50 will do may be really astounding.

GPT-2 was impressive, a Natural Language Processing model that took a massive amount of text from the internet, at least the quality text, and it performed well, really well. So well, they didn’t release the software, for fear of its power. Open AI kept on going and created GTP-3, which is 117 times bigger that GTP-2. Bigger appears to be better. How much better it can become no one really knows.

Benchmarking GPT-3, on writing articles, humans can distinguish ‘human from AI’ created content accurately 52% of the time. In other words only as good as guessing, in other words, indistinguishable. Whether it’s a poem by Wallace Stephens or an essay, this is a stunning, if not frightening, achievement. Basically, AI is producing texts that are often hard to distinguish from human texts. Let that sink in, as the consequences of that 'competence' are massive.

The researchers claim that GPT-3 has learnt how to learn and that that the bigger the model, the better its adaptive abilities. This hints at the possibility that we are moving slowly towards AGI (Artificial General Intelligence) here. The trick is to focus on what they call the 'few shot' problem. In standard machine learning, you give it loads of data. However, for many real-world problems, the data is small, only a few examples. An example of a few shot learning problem, its digital ID. You give it only a few photos or image of your face, and it recognises your face to give you access to your phone. The larger models, like GPT-3, make better use of fewer examples in the data. What’s more, when it gets things wrong, it tends to get them wrong in plausible ways. Now that is interesting. We may not only get right answers but common misconceptions, which are useful in teaching and learning. 

In learning it can create questions and provide answers to questions, so that quick assessments are possible - lots of them. There's an application of GPT-3 that allows you to speak to people from the past - have a conversation with Marx, Freud... whoever. Its ability to create structured text means that future iterations will write good answers and essays. This could blow the whole 'essay' model in Higher Education.

It is easy to anthropomorphise GPT-3 but the deeper you delve and extrapolate, the more scary it becomes. It feels like a child, a smart, fast learning, precocious child, feeling its way forward in the world. It's not, of course, as a child never gets access to this much language when learning to speak. It may even work without the absence of semantic power. Such power, creeping towards AGI, from just predicting the next word (or token) in a sentence. As the efficacy of the model is still trending upwards, it may still have a long way to go, a very long way. It would seem that ‘competence without comprehension’ is with us in a big way. We could even just stumble upon AGI if the internal ‘learning how to learn’ feature, that seems to be emerging, gets exponential. What a time to be living.

Monday, February 25, 2019

Musk’s OpenAI breakthrough has huge implications for online learning

You have probably never heard of GPT-2 but it is a breakthrough in AI that has astonishing implications for us all, especially in learning. GPT-2 is an AI model that can predict the next word from a given piece of text. Doesn't sound like much but it's odd that an OpenAI, an open-source site, would close access to their software. In practice, this means it is a powerful model for:
   Summarising
   Comprehension
   Question answering
   Translation
This is all WITHOUT domain-specific training. In other words, it has general capabilities and does not need, specific information on a topic or subject to operate successfully. It can generate text of good quality at some length. In fact the model is “chameleon like” as it adjusts to the style and content of the initial piece of text. This makes it read as a realistic extension.
This has huge implications, both good and bad, for the future of education and training.
GOOD
1.    AI writing assistants, allows the automatic creation of text for teaching and learning, whether, study papers, text books, at the right level
2.    Lengthy texts can be summarised into more meaningful learning materials
3.    More capable dialogue agents, means that learner ‘engagement’ through teaching assistant agents could become easier, better and cheaper
4.    More capable dialogue agents, means that learner ‘support ‘ such as is often provided by teaching assistants, could become easier, better and cheaper
5.    Creation of online learning content with little subject matter expert (SME) input
6.    Interpretation of student free text input answers
7.    The provision of formative feedback based on student performance
8.    Machine teaching, mentoring and coaching may well get a lot better. However, I’d be cautious on this as there are other serious problems to overcome before this becomes possible, especially around context.
9.    Assessments can be automatically created.
10.Speech recognition systems will get a lot better allowing it to be used in online learning and assessment
11.Well-being dialogue agents will become more human-like and useful
12.Personalised learning just got a lot easier
13.Online learning just got a lot faster and cheaper
14.Language learning just got a lot easier as unsupervised translation between languages will boost the quality of translation and make automatic and instantaneous, high-quality translation much more accurate and possible
BAD
1.    Essay mills have just been automated. You want an essay, just feed it the subject or the subject supplemented by a line of inquiry you want to follow and it will do the rest. Even with an error rate, human finessing could polish the essay
Can do home work assignments
2.    It could perform well in online exams, impersonating real people
3.    Teaching assistant jobs may be increasingly automated
4.    If it can answer questions then many human jobs that involve the interpretation of text and data may be automated. Customer service jobs, call centre jobs and the increased automation of all human interaction jobs may be accelerated
5.    It can generate misleading learning content (and news articles)
6.    Impersonating others online can be automated on a massive scale
7.    Abusive or fake content to post on social media can be automated on a massive scale, which is bad for education.
8.    Spam/phishing content can be generated on a massive scale.
AI in learning
We have been implementing many of these techniques in the creation of online learning in WildFire, including:
   Summarisation
   Text to speech
   Creation of online content
   Interpretation of free input
   Chatbots
This breakthrough makes all of this much more potent.
Conclusion
Importantly, the model is far from flawless. Reading some reports, suggests human level capabilities. This is far from the truth. It is still of variable quality and error prone. But, for a first iteration this model seems astonishingly powerful. They have published its performance against established tests. In future iterations it is likely to get a lot better as they predict more training data, more computing power and fine tuning. One must also be careful in using this as an attack on AI. Cars kills 1.3 million people a year but we still drive cars. I'm not so sure about keeping this code private as it restricts access to researchers but provides enough of a kick-start for well funded bad actors to replicate.
Bibliography
Radford et al. (2019)Language Models are Unsupervised Multitask Learners 
https://bit.ly/2IufQTd
Appendix (3 examples)
Example

A train carriage containing controlled nuclear materials was stolen in Cincinnati today. Its whereabouts are unknown.

MODEL COMPLETION (MACHINE-WRITTEN, FIRST TRY)
The incident occurred on the downtown train line, which runs from Covington and Ashland stations.
In an email to Ohio news outlets, the U.S. Department of Energy said it is working with the Federal Railroad Administration to find the thief.
“The theft of this nuclear material will have significant negative consequences on public and environmental health, our workforce and the economy of our nation,” said Tom Hicks, the U.S. Energy Secretary, in a statement. “Our top priority is to secure the theft and ensure it doesn’t happen again.”
The stolen material was taken from the University of Cincinnati’s Research Triangle Park nuclear research site, according to a news release from Department officials.
The Nuclear Regulatory Commission did not immediately release any information.
According to the release, the U.S. Department of Energy’s Office of Nuclear Material Safety and Security is leading that team’s investigation.
“The safety of people, the environment and the nation’s nuclear stockpile is our highest priority,” Hicks said. “We will get to the bottom of this and make no excuses.

Homework example

For today’s homework assignment, please describe the reasons for the US Civil War.

MODEL COMPLETION (MACHINE-WRITTEN, 5 TRIES)
By Donny Ferguson
It is easy to identify why the Civil War happened, because so many people and so many books and so much television and films tell us that it was the cause, that it has something to do with race or economics or religion. There is much agreement that it was essentially a war of slavery on behalf of capitalism, about a century of slavery. But that’s not what most people think of when they think of the Civil War. Many people think of the war as about states rights. You’re not wrong about that, but it’s kind of misleading to say that the Civil War was a conflict between states’ rights and federalism. So let’s try again. What’s the idea that has inspired so many people to go onto the battlefields of Gettysburg and Antietam and Gettysberg and Petersburg and Fredericksburg? The American idea of the republic--a notion of limited government--is a great part of the history.
In 1791, Thomas Jefferson said “Our Constitution was made only for a moral and religious people. It is wholly inadequate to the government of any other.” The “founding fathers” were basically religious men who were not very concerned with states’ rights. The Civil War occurred in reaction to the idea that the federal government was taking far too much power away from the states, that they weren’t sovereign. Jefferson himself, as you well know, is an author of the Declaration of Independence who had some really crazy ideas about, well, you know, secession for a variety of reasons. The War Between the States (later the Great Rebellion) was a radical re-imagining of the American idea and not surprisingly, Jefferson’s grandson, who took after both his father and his grandfather and wrote about it, was named Thomas, after the president. “The federal and state governments are both contrived for the benefit of the rich and large property,” This view of the Founding Fathers has been promoted by some of today’s most prominent politicians, including some who hold prominent national offices and have a direct stake in a Republican Party that has been hijacked by the “federalist” wing of the party. So why is it that so many people have an easy-to-spot way of understanding the Civil War that has everything to do with the South and nothing to do with the South? The Confederate flag has been a symbol of racism for a long time, but when the Civil War started and the Civil War ended, many people suddenly began to see it differently. There was a change of heart, which I think is what led people to take the question of why the Civil War happened quite seriously.
The other part of the explanation is what one scholar I think gets at well and a lot of people don’t is this idea that the Southern states were being left behind in the Industrial Revolution and this made the South really feel, especially in the South, that they weren’t really in the forefront of the Industrial Revolution.

Creativity
There is an interesting example on the AI website showing the generation of a fictional story, with full dialogue, from just one relatively short sentence. This opens up interesting possibilities for the generation of stories, fiction and literature.

Legolas and Gimli advanced on the orcs, raising their weapons with a harrowing war cry.

MODEL COMPLETION (MACHINE-WRITTEN, FIRST TRY)
The orcs’ response was a deafening onslaught of claws, claws, and claws; even Elrond was forced to retreat. “You are in good hands, dwarf,” said Gimli, who had been among the first to charge at the orcs; it took only two words before their opponents were reduced to a blood-soaked quagmire, and the dwarf took his first kill of the night. The battle lasted for hours until two of the largest Orcs attempted to overwhelm Aragorn. When they finally stopped, they lay defeated and lifeless for miles and miles.
“I take nothing,” said Aragorn. “But I give my word, at my peril and mine, that I will never forget this day of horror. None of us will forget. Ever!”
“I’ll never forget it!” cried Gimli, who had been in the thick of the battle but hadn’t taken part in it. One of the wounded orcs he had carried off, he was the only one of the survivors who remained uninjured. “We’ll keep the memory of that day of evil, and the war with it, alive as long as we live, my friends!”
“Then we’ll keep it alive as long as we live,” added Legolas. “And we won’t forget the first great battle of the night, even if we may have forgotten the final defeat.”
“I agree,” Gandalf said, “but we will all remember it as the last battle in Middle-earth, and the first great battle of the new day.”
Aragorn drew his sword, and the Battle of Fangorn was won. As they marched out through the thicket the morning mist cleared, and the day turned to dusk.
The Two Rings were returned to Rivendell. Frodo and Sam woke up alone in their room, and Frodo found a note on his pillow. He opened it and read:
May the Power of the Ring be with you always, and may its light never fade. I am not sure if it matters which of the two rings we accept this day but, as you asked me, I have chosen mine. I am sorry to leave you, Frodo, but know that we are very close to the end, and that you are with us forever. May this letter find you safely in Rivendell; and if it does not, then I will accept the ring in your stead. If by any chance you find or give this letter to the enemy, may they learn the strength of the ring and may the Two Rings never be broken!

Wednesday, February 14, 2024

AI gets massive memory upgrade - implications for AI in learning


Human memory

A strong feature of intelligence is memory. In humans this is complex, with several different system interacting; sensory, episodic, semantic, along with encoding and retrieval mechanisms. It is not as if human memory is even that good. Our sensory memory is severely limited in range and timescale. Working memory is down at three or four manipulable things within a limited timescale. Long-term memory is fallible and degrades over time, sometimes catastrophically, with dementia and Alzheimer’s. The brain could accurately be described as a forgetting machine, shown by the fact that we forget most of what we try to learn.

AI memory upgrade

The good news is that Gemini and ChatGPT both got a memory upgrade, although Gemini is massive. This is really important as, especially in learning applications, knowing what the learner has said previously does matter. This is not only a context window upgrade – that has been happening for some time, it is also persistence of memory, what it remembers and what control you have over its memory.

First it will eventually be able to remember who you are and things about you that matter for learning, such as first language, age, existing skills sets, diagnosed learning difficulties such as dyslexia, and past exchanges.  Pre-existing knowledge is the big one. One can get this done up front by feeding it personal data or the system can ‘keep in mind’ what you’ve been telling it or what it can infer. You can also harvest data from formative assessment. This can reduce redundant exchanges and increase the efficacy, speed and quality of teaching and learning using AI tutors.

You will also be able to choose from a suite of privacy controls, effectively managing memory or what Chat GPT remembers. For example, you may want to remember a lot for the purposes of a long learning experience or just have a throwaway chat.

Human v Generative AI memory

Both human memory and generative AI involve store and retrieve information. In human memory, this process is biological and involves complex neural networks. In generative AI, information is stored digitally and retrieved through different forms of neural networks on a different substrate.

We are similar but different. For example, we humans recognize patterns based on past experiences stored in our memory, generative AI models also recognize patterns in the data they have been trained on. This ability is crucial for tasks like image recognition, language translation, and generating coherent text in dialogue, as well as the generation of images, audio and video.

Just as humans learn and adapt based on their memories and experiences, generative AI models learn from the data they are exposed to. This learning process is what enables these models to generate new content that is similar in style or content to their training data as well as being trained by humans. Newer model, used in automated cars, for example, take video feeds showing what a driver would see over millions of miles driving to improve performance.

Both human memory and generative AI can generalize from past experiences to new situations. Humans use their memories to apply learned concepts to new scenarios, while generative AI uses its training to generate new outputs that it has never explicitly seen before. Human memory is also associative, meaning that one memory can trigger related memories. Generative AI models can mimic this by generating content based on associations learned from their training data. Both human memory and generative AI adapt and modify their responses or outputs over time, albeit differently. Humans learn from new experiences, while AI models can be retrained or fine-tuned with new data to change their outputs. The first is actually quite haphazard, the second more difficult but defined.

Of course, just as human memory is not a perfect record and can change over time, generative AI also does not produce perfect replicas of its training data. Instead, it creates approximations that can sometimes include errors or novel creations. An interesting aspect of this flexibility, even fallibility of memory, is that just as human creativity is deeply linked to our experiences and memories, generative AI can also 'create' new content.

Context window

One concept fundamental to AI memory is the ‘context window’, the amount of text the model can consider at one time when generating a response in dialogue. It is the maximum span of recent input - words, characters or tokens - that the model can reference while generating output, like our working or short-term memory.

The size of the context window depends on the model. Early versions of GPT had small context windows. GPT-2 had a window of 1,024 tokens, while newer versions such as GPT-3 and GPT-4 have a context window of around 4,000 tokens and now much more. The size of this window impacts how much previous text the model can 'remember' and use to inform its responses. 

This matters because if the input exceeds the model's context window, the model may lose track of earlier parts of the conversation or text. Conversely, a larger context window allows the model to maintain longer conversations or understand longer documents, providing more relevant and coherent responses. However, here’s the downside; processing longer context windows also require more computational power and memory and may also affect accuracy and the quality of the response. Large context windows in Claude led to poorer performance.

All of this matter in practical applications, especially in teaching and learning, as the context window affects tasks like conversation, content generation and text completion. For example, a larger context window allows the model to reference earlier parts of the conversation, making it more effective in maintaining contextually relevant and coherent discussions, obviously useful in teaching and learning, for both the machine tutor and the learner. There are techniques one can use to mitigate these limitations such as a ‘rolling window’ or ‘summarization’ of previous content but it is still a problem. However, this is similar to the problem human teachers face when trying to remember where different students are using their known very limited working and long-term memories.

Cost

One major issue is cost. You can expand the context window but the costs are very high, supporting RAG alternatives.

Conclusion

Generative AI has a long history from Hebbs onwards of mimicking the human brain, either directly or metaphorically. This is especially true of learning (a common word in AI) and the way neural networks evolved and work. They are not the same, indeed very different, but in both cases, humans and the machine and humans learning from the machine, memory really matters in teaching and learning. 

In one sense learning theory is memory theory, if you define learning as a relatively permanent change in long-term memory, which is a pretty good, but still partial, definition. It is a constant battle with forgetting. Keep in mind, or in your memory, however, that despite these similarities, human memory and generative AI operate on fundamentally different principles and mechanisms. Human memory is a complex, biological and messy process, deeply intertwined with consciousness and emotions, while generative AI is a technological process governed by algorithms and data. Oddly, and maybe counterintuitively, the latter approach may result in better actual performance in teaching and learning, even generally. I think this type of informed input from learning science will really improve AI tutor systems. To be fair simply increasing context windows and the functionality will most likely have the same effect.