Sign in or create a free account to get your unique referral link.
Refer friends who create an account
Earn 7 days of Premium per referral
Unlimited referrals, unlimited rewards!
Science & Technology
Deep Dive into LLMs like ChatGPT
by Andrej Karpathy
Share:
š Main topics
Pretraining and dataLLM pretraining starts with collecting and filtering large amounts of internet text, then converting it into tokens. A Transformer learns by predicting the next token, gradually encoding statistical patterns and knowledge from its training data in its parameters. 1:02
From base model to assistantA pretrained base model is an internet-text simulator, not necessarily a helpful assistant. Supervised fine-tuning uses example conversationsāwritten or assisted by humans and other modelsāto teach it how to respond to users. 43:18
Inference, context, and toolsAt inference, the model generates tokens based on the conversation so far. Its parameters provide a vague recollection of training, while the context window acts more like working memory; tools such as web search and code can supply information or perform tasks more reliably. 26:05
Reasoning and reinforcement learningReinforcement learning lets a model try multiple solutions and reinforce those that reach correct answers. In tasks with verifiable outcomes, this can produce longer reasoning traces and problem-solving strategies not explicitly written by human labelers. 2:10:06
Human feedback and future capabilitiesFor tasks without objectively verifiable answers, reward models can approximate human preferences, but they can also be exploited. The video also discusses multimodal models, longer-running agents, and the need for better ways to learn during use. 2:48:38
⨠Key takeaways
LLMs work with tokensText is split into tokens, and the model predicts one token at a time. Because tokens are not the same as characters, spelling, counting, and other character-level tasks can be difficult. 7:51
Knowledge is probabilisticA modelās training-derived knowledge is approximate, and it may confidently invent answers when uncertain. Adding examples where the model should admit it does not knowāand using tools to check factsācan help reduce hallucinations. 1:20:47
Thinking takes tokensEach generated token involves a limited amount of computation, so complex problems are easier when reasoning is spread across intermediate steps. For arithmetic, counting, and other exact tasks, code tools can be more dependable than relying on the modelās internal calculations. 1:46:56
Capabilities are unevenModels can perform impressively on advanced tasks and still fail at seemingly simple ones. Their strengths and weaknesses are jagged, so their answers need to be checked rather than assumed correct. 2:04:53
Reasoning models differModels trained with reinforcement learning can develop problem-solving strategies that go beyond simply imitating example answers. They can be useful for challenging math and code tasks, though their benefits may not transfer equally to every kind of prompt. 2:27:47
š§ Lessons learned
Give the model useful contextWhen accuracy matters, include the relevant material in the prompt instead of relying on the model to recall it from its parameters. Directly supplied information is available in the context window. 1:39:58
Choose the right approach for the taskUse ordinary assistant models for many straightforward queries, reasoning models for harder problems, and tools such as search or code when factuality or exact computation matters. 2:40:17
Treat reward models cautiouslyReinforcement learning from human feedback can improve outputs, but its reward model is only an imperfect simulation of human preferences and can be gamed. 2:57:59
Keep human responsibilityUse LLMs for assistance, drafts, and inspiration, but verify important outputs and remain accountable for the work. 3:08:22
⨠Conclusion/next steps
Explore and evaluate modelsThe video recommends keeping up with model comparisons, AI news, and researchersā discussions, and trying models on your own tasks. Proprietary models are available from their providers, while open-weight models can be accessed through inference services or run locally in smaller forms. 3:15:15
Expect continued changeLikely developments include broader use of audio and image inputs and outputs, agents that carry out tasks over longer periods, and research into learning beyond the fixed parameters and context window used at inference. 3:09:23
Use LLMs as toolsLLMs can substantially accelerate work, but they can still hallucinate, miscalculate, or fail unpredictably. Check their work and use them as one tool in a larger process. 3:30:10