0:05 Hello everyone and welcome to lecture two of CME295. So today is an exciting day because we'll be talking about large language models and in particular how they relate to the transformer architecture that we saw last lecture. But before we start, we will recap what we saw last time and we will do that at every lecture to see how this connect to what we'll see today. So if you remember last time we were
0:38 wondering how models could process text. So models they don't understand text as it is they understand numbers. So if you remember one thing that we looked at was first of all how to divide our text into indivisible units and this process is called tokenization. And we saw a few kinds of tokenization algorithms. And we said that the most common tokenization algorithm these days is the subword tokenization because it
1:10 has a nice tradeoff between the vocabulary that it was building and the sequence length that input text would have. But this was not it because once you divide a text into tokens the next step is to compute embeddings. So if you remember we saw Word2vec as a way to compute embeddings but they were not aware of the context. So then we talked about another class of models that were
🔒 The full, searchable transcript is available with Pro.
Common questions about this video