AI

Local AI memory guide: which open models fit 24GB to 128GB

· Geeknewz Author

Two DDR5 memory modules lying on a dark surface, one showing its memory chips

Microsoft and Nvidia spent Wednesday’s Windows and Surface event in San Francisco selling one idea: run your AI on your own laptop instead of renting it from the cloud. The new Surface Laptop Ultra starts at 24GB of unified memory and goes up to 128GB, and Nvidia’s RTX Spark page promises you can “run models up to 120B parameters.” Apple makes a similar pitch for the MacBook Pro, which tops out at 128GB too.

The spec sheets don’t answer the question you actually face at checkout, though, which is which models will fit in the memory you’re paying for. So we pulled the real download sizes of eight popular open models from Hugging Face and worked out which memory tier each one needs. The short version is that 24GB is tighter than it sounds, 32GB is the practical floor for today’s best mid-size models, and 128GB is the only tier that gets you into 120-billion-parameter territory.

Nvidia DGX Spark mini computer with a gold metal-foam front on a wooden desk
Nvidia’s DGX Spark, the 128GB desktop sibling of the RTX Spark laptop chip. Photo by Daniel Lu via Wikimedia Commons (https://commons.wikimedia.org/wiki/File:Nvidia_DGX_Spark_oblique_view_dllu.jpg). CC BY-SA 4.0.

How much memory a model actually takes

A model’s weights are just numbers, and the precision you store them at sets the size. At 16-bit precision each parameter takes 2 bytes, at 8-bit it takes 1 byte, and the popular 4-bit builds land a little above half a byte once you count the extra data they carry. Qwen3.8 27B shows how that works in practice. Its official 16-bit files add up to 55.6GB, the 8-bit build from Unsloth is 29.0GB, and the common 4-bit build (Q4_K_M) is 16.5GB, which works out to about 0.59 bytes per parameter.

The weights aren’t the whole bill, because the model also keeps a cache of the conversation so far. Qwen3.8 27B’s published config uses full attention on 16 of its 64 layers, with 4 key-value heads of 256 dimensions each, so every token of context costs about 64KB at 16-bit. That’s roughly 2.1GB for a 32,000-token conversation, 8.6GB at 128,000 tokens, and 17.2GB if you use the full 262,144-token window, which is more than the 4-bit model itself. Then Windows or macOS, your browser and whatever else you have open need their share.

To keep the comparison honest, we used one simple rule for every model: add 15% to the file size for the cache and runtime (about a 32,000-token conversation on Qwen), and assume you can hand the model about 75% of your total memory. That leaves 18GB usable on a 24GB laptop, 24GB on 32GB, 36GB on 48GB, 48GB on 64GB and 96GB on 128GB. It’s a rule of thumb rather than a guarantee. Microsoft’s own demo showed about 110GB of a 128GB Surface addressable by the GPU, according to Notebookcheck, so a big machine has a little more room than our rule assumes, while long conversations eat into it fast.

Which open models fit where

Here’s what we found. File sizes come from each model’s official Hugging Face repository for the 16-bit and native versions and from Unsloth’s widely used GGUF builds for the 4-bit and 8-bit ones, all checked on October 7, 2026.

Model Parameters 4-bit file 8-bit file 16-bit or native Smallest tier, 4-bit Smallest tier, 8-bit
Gemma 4 12B 12.0B 7.1GB 13.1GB 23.9GB 24GB 24GB
gpt-oss-20b 20.9B (3.6B active) 13.8GB, ships in its own 4-bit format 24GB n/a
Gemma 4 26B A4B 25.8B (about 4B active) 16.9GB 27.3GB 51.6GB 32GB 48GB
Qwen3.8 27B 27.8B 16.5GB 29.0GB 55.6GB 32GB 48GB
Gemma 4 31B 31.3B 18.3GB 33.2GB 62.6GB 32GB 64GB
gpt-oss-120b 116.8B (5.1B active) 65.3GB, ships in its own 4-bit format 128GB n/a
Nemotron 3 Super 120B A12B 123.6B (12B active) 82.5GB 128.5GB 247.2GB 128GB Doesn’t fit
DeepSeek V4.1 Flash 763.2B 510.3GB native, 365.7GB even at 2-bit Doesn’t fit Doesn’t fit

The 24GB tier is the one to be careful with, since it’s the configuration in the $2,599.99 base Surface Laptop Ultra and the base 16-inch MacBook Pro. It runs the 12B-class models comfortably, even at 8-bit, and it’s a good home for gpt-oss-20b. The 27B to 31B models that many people consider the sweet spot right now need about 19 to 21GB at 4-bit under our rule, which is just past the 18GB we budget on a 24GB machine. You can squeeze one in with a short conversation and little else open, but you’ll be fighting for memory the whole time.

At 32GB, those mid-size models fit at 4-bit with room for a reasonable conversation, which is why we’d call it the real starting point for local AI. Going to 48GB or 64GB mostly buys you the 8-bit versions of the same models, which are a bit more accurate, rather than a new class of model. The jump that changes what you can run is 128GB. That’s where gpt-oss-120b, at 65.3GB, and the 4-bit Nemotron 3 Super, at 82.5GB, become practical, so Nvidia’s “up to 120B parameters” claim holds up for 4-bit models. The newest giant open models are another matter. DeepSeek V4.1 Flash is 510GB in its native format and still about 366GB at 2-bit, and models like the 1-trillion-parameter Mistral Large 4 we compared with Reflection Beam this morning are server hardware territory, whatever laptop you buy.

Memory decides what runs, bandwidth decides how fast

Fitting a model is only half the story. When a model writes each word, it has to read its active weights from memory, so memory bandwidth sets a hard ceiling on speed. You can estimate that ceiling by dividing bandwidth by the bytes read per token. Apple publishes these numbers on its MacBook Pro spec page: 153GB/s for the M5, 307GB/s for the M5 Pro, and 460GB/s or 614GB/s for the two M5 Max versions. For the 16.5GB 4-bit Qwen3.8 27B, that works out to ceilings of about 19 tokens per second on an M5 Pro (307 ÷ 16.5) and 37 on the top M5 Max (614 ÷ 16.5).

Nvidia doesn’t list memory bandwidth for RTX Spark, which is the single most useful number missing from Wednesday’s launch. Its desktop sibling, the DGX Spark, is rated at 273GB/s with the same 1 petaflop of FP4 compute and 128GB of memory. If the laptop chip matches that, which Nvidia hasn’t confirmed, the same Qwen model would top out around 16.5 tokens per second. Real-world speeds land below these ceilings, so treat them as upper limits for comparing machines rather than predictions.

This is also why mixture-of-experts models suit big-memory laptops so well. gpt-oss-120b has 116.8 billion parameters but only activates 5.1 billion for each token, so it reads about 2.85GB per word (65.25GB × 5.1 ÷ 116.8) instead of the whole file. At 273GB/s that’s a ceiling near 96 tokens per second, far faster than a dense model a quarter of its size. You need the memory to hold all of it, but you only pay the bandwidth cost for the slice in use.

Our take on which tier to buy

This is Geeknewz’s opinion. If local AI is a real reason you’re shopping for a Surface Laptop Ultra, an RTX Spark laptop from another brand or a MacBook Pro, skip the 24GB models and treat 32GB as the minimum. If you mainly want a capable assistant that works offline, 32GB with a 4-bit 27B to 31B model is the best value right now. Pay for 128GB only if you specifically want 120B-class models like gpt-oss-120b or Nemotron 3 Super, and expect speed to depend on bandwidth as much as capacity.

For RTX Spark in particular, we’d wait for two things before spending top-tier money. One is Nvidia publishing the laptop chip’s memory bandwidth, and the other is independent token-per-second tests once machines ship on October 16, because the 2.1 times faster time-to-first-token Nvidia claims over an M5 Pro MacBook Pro is the company’s own preliminary figure. Memory is also the one part you can’t upgrade later on any of these laptops, so it’s worth getting right the first time.

Data note: model file sizes come from Hugging Face model repositories (official repos for 16-bit and native files, Unsloth GGUF builds Q4_K_M and Q8_0 for 4-bit and 8-bit, and a community 2-bit build for DeepSeek), checked on October 7, 2026. Context-cache math uses the model’s published config.json. Active-parameter counts for gpt-oss come from OpenAI’s model card.