Michael Limberger
Need me? Email mike@limberger.ca
AI
Quantization Guide
Quantization Guide
Before we download Magidonia, we need a clear choice about quantization. That is where hardware meets the model, and the choice follows you through everything else.
What Is Quantization?
Quantization is compression for model weights. A neural network stores millions of numbers it uses to generate the next token. Those weights often start as full-precision floating point values (FP16), which is two bytes per weight. Quantization stores them in fewer bits (4-bit, 6-bit, 8-bit, and so on). Smaller numbers mean a smaller file, less RAM, and usually faster loading.
The tradeoff is precision. You lose a little quality when you shrink the weights. With Magidonia, community testing showed you can keep most of that quality if you pick the right level for your machine.
The Spectrum
For Magidonia-24B, these are the main options people actually use:
Q4_K_M ~14.3 GB ~90% quality vs FP16 - Unnecessary compromise at 96GB
Q5_K_M ~16.8 GB ~94% quality - Still leaving quality on the table
Q6_K ~19.3 GB ~98% quality - Excellent second choice
Q8_0 ~25.0 GB ~99.5% quality - YOUR PICK. Near-lossless.
Those percentages come from community testing on coherence, consistency, and creative quality across a lot of generations. Around 90%, people notice weird jumps, forgotten details, and contradictions. Around 98%, you usually need a blind test to spot the difference. At about 99.5%, you are splitting hairs.
Why Q8_0 For Your 96GB Mac
On a 96GB Mac, the Q8_0 build is about 25GB, which leaves roughly 71GB of headroom. The model needs that space in memory, and generation also builds a KV cache for attention. On a 24B model with a long context window, that cache can get large.
With about 25GB for the model and plenty left over, you can keep 32,000+ token context windows without swapping to disk. Swap kills latency. Dropping to Q6_K saves a few GB and costs a little quality you do not need to give up. Dropping to Q4_K_M frees even more RAM, but coherence suffers enough that it is the wrong trade when you have this much memory. At this hardware level, Q8_0 is the sweet spot.
About Bartowski'S Gguf Quantizations
You will see bartowski's name a lot next to GGUF files. He quantizes popular models with iMatrix, a technique TheDrummer and Ollama docs both point people toward.
iMatrix is an importance matrix. Not every weight matters equally when you shrink the model. iMatrix keeps the important weights at higher precision and compresses the less important ones harder. You lose less quality per GB than with a blunt, equal shrink. Think of it like smart image compression: blur the parts that do not matter, keep the sharp edges where they count.
If you are downloading Magidonia for this tutorial, use the bartowski GGUF. That is what we pull next.
When To Go Lower
Lower quants make sense in other situations. If you run several models at once, Q6_K can free room for a second load. On a 16GB laptop, Q4_K_M is often the floor for usable quality, and you manage context carefully. The consensus from people with ample Mac memory stays simple: use Q8_0 when quality matters and RAM is not the bottleneck.
Quick Reference Table
| Quantization | File size | Quality loss | Use case |
|---|---|---|---|
| Q4_K_M | 14.3 GB | ~10% | 16GB RAM (minimum) |
| Q5_K_M | 16.8 GB | ~6% | 24GB RAM (portable) |
| Q6_K | 19.3 GB | ~2% | 32GB RAM (good backup) |
| Q8_0 | 25.0 GB | ~0.5% | 64GB+ RAM (this tutorial) |
The files differ mainly in weight precision. Q4_K_M uses about 4 bits for main weights. Q8_0 uses about 8 bits. More bits means more space and more quality kept.
Moving Forward
For this tutorial we use Q8_0. When you download the model in the next section, look for that tag in the name. That is the build you want.