Open Models and Local Deployment
Tools for Running LLMs Locally
Choosing the Right Tool
Advancements in Local LLMs
Learning Opportunities
0:00 Open models like Gwen, Kimmy, and the GLM family are now a strong enough that you don't always need a hosted API. You can run them on your own laptop, so no one sees your conversation or data. Here are five tools to run LLMs locally. llama.cpp is a C++ inference engine that runs on CPUs, GPUs, and Apple silicon. It started as a side project to run llama on a MacBook and grew into the foundation most other local tools are built on.
0:32 llama.cpp also introduced a standard file format for local models, GGUF. A GGUF file packs the weights, tokenizer, and metadata into one file, and supports quantization down to 4-bit and lower, which is what makes large models fit on consumer hardware. You download a GGUF from Hugging Face, run llama.cpp and provide the model and your prompt, and you get tokens back. Use llama.cpp when you want the lightest possible
1:02 runtime, or when you are deploying to constrained hardware, like an edge device or a laptop without a dedicated GPU. Ollama is a wrapper around llama.cpp that turns it into a developer tool. It handles model downloads, quantization choices, and starting a local server so you can chat with any LLM. You run Ollama run Gemma 4, it pulls the weights, it starts a local server, and gives you a chat prompt.
🔒 The full, searchable transcript is available with Pro.