If you've ever wondered why some AI models need a data center while others run on a gaming PC, why "70B" is a big deal, or what "quantized" means โ this page connects the dots, in plain language.
1. Parameters: the "billions" in every model name
A large language model is, at its core, an enormous grid of numbers called parameters (or weights). These numbers are what the model learned during training โ its knowledge, its reasoning patterns, its sense of language, all encoded as billions of tiny numeric values.
When you see a model called "7B", "70B", or "405B", that's the parameter count: 7 billion, 70 billion, 405 billion. As a rough rule:
- More parameters โ more capability. Bigger models hold more knowledge, follow more complex instructions, reason through longer chains of logic, and produce stronger, more refined output.
- Fewer parameters โ simpler behavior. Small models can still write, summarize, and chat โ but they make more mistakes, lose track of complex tasks sooner, and struggle with multi-step reasoning, subtle code bugs, and nuance.
2. Why parameters cost memory: the VRAM math
Every parameter has to be loaded into memory (ideally fast GPU memory, called VRAM) to run the model. Stored at full training precision (16-bit, i.e. 2 bytes per parameter), the math is simple:
| Model size | Memory at 16-bit | What can hold it |
|---|---|---|
| 7B | ~14 GB | One consumer GPU (RTX 4090/5090) |
| 70B | ~140 GB | 2ร server GPUs (e.g. H200) |
| 405B | ~810 GB | A rack of 6โ8 server GPUs working together |
| ~1T+ (frontier) | multiple TB | Clusters of many GPU servers |
And that's just to load the model โ you also need extra memory for the conversation context, and enough compute to push tokens through all those billions of parameters quickly.
3. The cloud giants: Claude, Gemini, ChatGPT
The big cloud models you use every day run on data-center GPUs โ hardware in a completely different league from anything consumer:
- An NVIDIA H200 has 141 GB of ultra-fast HBM memory (~4.8 TB/s bandwidth). A B200 pushes that to 192 GB at ~8 TB/s. Compare: a top-end gaming card (RTX 5090) has 32 GB at ~1.8 TB/s.
- These GPUs are deployed in clusters of eight or more per server, and many servers per model, linked by specialized interconnects so they act like one giant GPU.
- That's what lets frontier models run at hundreds of billions to trillions of parameters โ at full quality โ and still respond in seconds while serving millions of people simultaneously.
This is the core trade you make with cloud AI: you don't own the hardware, you pay per token (that's what the index compares) โ but you get maximum capability and speed with zero setup.
4. Quantization: compressing the model
What if you want to run a model on hardware that can't fit it? You quantize it โ store each parameter with fewer bits. It's a form of compression:
| Precision | Bytes/param | 70B model needs | Quality |
|---|---|---|---|
| 16-bit (full) | 2.0 | ~140 GB | Original quality |
| 8-bit | 1.0 | ~70 GB | Nearly indistinguishable |
| 4-bit | ~0.55 | ~40 GB | Small but real quality loss |
| 2โ3-bit | ~0.3โ0.4 | ~25โ30 GB | Noticeable degradation |
The catch: every bit you shave off rounds the model's knowledge more coarsely. At 8-bit you barely notice. At 4-bit, the model is still very usable but starts slipping on exactly the things that need precision โ complex reasoning, tricky code, subtle instructions. Below 4-bit, quality drops fast.
5. Running models locally: the honest picture
Local models are a fantastic and fast-improving world โ private, free per token, always available, fully under your control. But set expectations correctly:
- A single consumer GPU (16โ32 GB) comfortably runs models in the 7Bโ32B range, quantized. That's excellent for chat, summarization, drafting, simple coding help, and automation.
- To approach cloud-frontier quality and speed you'd need multiple server-grade GPUs (H200-class) โ tens of thousands of dollars each, plus power and cooling. Short of that, local setups won't match the big cloud models on complex, multi-step "neural heavy-lifting" tasks.
- Speed matters too: even if a big model fits (e.g. partially in regular RAM), it may generate only a few tokens per second โ too slow for real work.
6. Putting it together: agents, tools, and scripts
The most effective real-world systems don't ask one model to do everything. They combine AI with ordinary deterministic code โ regular scripts that are fast, free, and never hallucinate:
- Use a script for anything with an exact answer: math, date handling, database lookups, file operations, API calls, validation. Code is always right about these; models sometimes aren't.
- Use a model for what code can't do: understanding messy human language, making judgment calls, writing, summarizing, deciding which tool to use next.
- An agent is exactly this combination: a model in a loop with a set of tools it can call. Some tools are AI (a small classifier model, a vision model), some are plain code (a calculator, a web scraper, a database query) โ the model orchestrates, the tools execute.
This is why the "small local model" story matters: an agent doesn't need frontier intelligence in every component. A well-designed pipeline might use a cloud model for the hard reasoning, a tiny fine-tuned local model for a high-volume repetitive step, and plain scripts for everything deterministic. Each part does what it's best at โ and that's how you get the most out of all of it.
7. Decision models: a probability instead of a paragraph
A lot of the steps in an agent aren't writing problems at all. Did the tests pass? Which tool is next? Is this command safe to run? Asking a big chat model means waiting for it to talk its way to an answer, then parsing the talk.
A decision model skips the talking. You give it a state (a log, a diff, a ticket, a web page) and typed questions: pick one of these options, yes or no, or where on this scale. It reads everything once and returns a probability for every option. Nothing is generated, so there's no text to parse and no answer outside the options you gave it.
- They're small and fast: most are 0.5Bโ12B fine-tunes answering in tens to hundreds of milliseconds, cheap enough to run on every step of a loop.
- The probabilities are usable numbers. A well-calibrated model that says 90% is right about 90% of the time, so you can set a threshold: act above it, escalate to a bigger model or a human below it.
- They're a component, not a replacement. The decision model is the reflex; a chat model still does the reasoning and writing when a step needs it.
The Decision models page ranks about 80 of them by smarts, calibration, speed and cost per 1,000 decisions.
8. Cheat sheet
| Term | Plain meaning |
|---|---|
| Parameters (B) | The model's learned numbers, in billions. More โ more capable, more memory, more compute. |
| VRAM | GPU memory. The model must fit in it to run fast. Rough rule: GB needed โ params ร bytes-per-param. |
| Quantization | Storing parameters with fewer bits to shrink the model. 8-bit โ free; 4-bit = mild loss; below that, real loss. |
| Tokens | The chunks of text models read and write (~ยพ of a word each). Cloud pricing is per million tokens โ see the index. |
| Context window | How much text the model can consider at once, measured in tokens. |
| Fine-tuning | Additional training that specializes a model for your specific task or style. |
| Inference | Actually running the model to generate output (vs. training it). |
| Decision model | A model that returns a probability for each option you give it instead of writing text. Fast, cheap, easy to threshold. |
| Agent | A model in a loop with tools โ some AI, some plain code โ that it calls to get real work done. |
Ready to compare the cloud side? The LLM Index ranks coding models by real per-request cost and blended quality, and the subscriptions page compares flat-rate coding plans by tokens per dollar.