Images vs. Text in AI Models
Efficiency and Cost Reduction
Technical Explanation
Practical Experimentation
Cautions and Recommendations
0:04 You see that face? You know what that face is? That that is the smug face of somebody who was proven correct. And the best part, it's right after caught in 8K. You're even told by your friends you're wrong. So, you're probably thinking, "Okay, what what are we talking about? What are you even trying to say? What were you proven right about?" Well, my friends, it turns out after many hurt feelings and being told I was wrong, images might actually be better for models than text. I knew it.
0:35 Don't make it do image recognition on that, dude. >> No, no, you just do this. What are you talking about? >> Hey, >> now this isn't just me saying this, of course. This is actually DeepSeek releasing a new paper that suggests that if you provide images, you could pull out 10 text token from a single image token with near 100% accuracy. In other words, a model's internal representation of an image is 10 times efficient as its internal representation of text.
🔒 The full, searchable transcript is available with Pro.
Common questions about this video