AI

Mistral Large 4 vs Reflection Beam: open-weight models compared

· Geeknewz Author

A technician points a handheld tester at networking equipment mounted in a server rack

Two labs outside China announced big open-weight models within a day of each other this week, and both are asking you to wait a few weeks before you can download anything. Reflection unveiled Beam on October 5, a 501-billion-parameter coding and agent model it promises to release under Apache 2.0 later this month. Mistral followed on October 6 with a public preview of Mistral Large 4, a trillion-parameter multimodal model it nicknamed "le Chonk," with weights due by the end of October.

Because neither set of weights is out yet, most coverage has stuck to each company's benchmark charts. We went through Reflection's launch post, Mistral's announcement, Mistral's model card and its EU technical documentation to line the two up on the things that decide whether you can actually use them: size, hardware, license, price and the few benchmarks both companies report.

A person installs a graphics card into a desktop computer case next to a CPU cooler
Photo by JESHOOTS.COM via Unsplash (https://unsplash.com/photos/sMKUYIasyDM). Unsplash License.

The two models side by side

Mistral Large 4 (preview)Reflection Beam (preview)
Total parameters1.05 trillion501 billion
Active per token49B (52B counting embeddings and output layers)23B
InputsText and images (1.6B vision encoder)Text only
Context window1M tokens1M tokens (extended in midtraining)
Training hardware3,800 Nvidia Grace Blackwell GPUs in Mistral's own European datacenters6,144 GB300 GPUs for pretraining, 10,500 GB300 GPUs for four weeks of RL
License for weightsNot named yetApache 2.0 (promised)
Weights expected"End of the month" (reporters were told Oct. 27)"Later this month"
Can you use it today?Yes, preview API on Mistral StudioWaitlist for select users
API list price$1.36 in / $4.18 out per million tokensNot announced
DeepSWE v1.1 (vendor-reported)61.7%44.4%
SWE-Atlas Codebase QnA (vendor-reported)59.4%34.6%

The two benchmark rows are the only ones where both companies published a score on the same named test, and Mistral leads both by a wide margin. Treat that gap with some care, though. Each lab ran its own harness, neither model's weights are public for anyone to rerun, and Mistral says its reinforcement learning run is still going, so the preview you can call today isn't the final checkpoint. For a sense of where those numbers sit, Reflection's own chart lists GLM 5.3 at 61.0 on DeepSWE v1.1 and Kimi K3 at 68.0, which puts Large 4's 61.7 right in the pack of China's best open models.

What the size difference costs you

The interesting part is how similar the two designs are underneath. Large 4 activates 49 billion of its 1,050 billion parameters on each token, about 4.7%, and Beam activates 23 of 501 billion, about 4.6%. They are equally sparse mixture-of-experts models, and Large 4 is simply about twice as big in every direction.

That doubling shows up in two bills. The first is compute per token. Reflection's post estimates generation cost with a simple rule, roughly 2 times the active parameters in floating-point operations per generated token. Using that same rule, Large 4 needs about 98 billion operations per token (2 × 49B) and Beam about 46 billion (2 × 23B), so Large 4 does roughly 2.1 times the work for every word it writes. Beam's whole pitch is that it also writes fewer tokens to reach an answer, which would widen that gap further, but that claim is Reflection's and needs independent testing.

The second bill is memory, which decides what hardware you need to self-host. As a rough rule, weights take about one byte per parameter at 8-bit precision and half a byte at 4-bit, before you add room for the context cache. That puts Large 4 at roughly 1.05 TB at 8-bit or 525 GB at 4-bit, and Beam at roughly 501 GB or 250 GB. A common server node with eight 80 GB H100s has 640 GB of GPU memory in total, so Beam at 8-bit fits on one node with about 139 GB to spare for long contexts, while Large 4 only fits there if you quantize it to 4-bit and accept a tighter context budget. Those are back-of-envelope numbers rather than vendor guidance, but they explain why a 500B model is a much easier thing to run in-house than a 1T one.

Training budgets tell a similar story about where each lab spent its money. Beam's pretraining used 6,144 GB300 GPUs for under four weeks, which caps it at about 4.1 million GPU-hours (6,144 × 672 hours), and its reinforcement learning phase used 10,500 GPUs for four weeks, about 7.1 million GPU-hours. In other words, Reflection spent more GPU time teaching Beam to reason and use tools than it spent on pretraining. Mistral hasn't published a duration, but VentureBeat reports roughly two months on its cluster, which would work out to around 5.5 million GPU-hours at 3,800 GPUs (3,800 × about 1,440 hours). The chips aren't identical, so don't read that as a precise comparison.

The license question matters more than the benchmarks

If you plan to build a product on one of these, the license is the line in the table to watch. Reflection says plainly that Beam's weights will ship under Apache 2.0, which lets you use, modify and redistribute the model commercially. Mistral calls Large 4 "open-weight" but hasn't named a license, and its EU technical documentation currently lists self-deployment only under a confidential bespoke agreement, with a note that the document will be updated if open weights are released. VentureBeat reports the weights are expected under a custom Mistral license. That would be a change from Mistral Large 3, which shipped under Apache 2.0 last December, so it's worth reading the terms before assuming Large 4 works the same way.

Price is easier to pin down for Mistral because you can already pay for it. At the list rate on the model card, a team sending 50 million input tokens and 10 million output tokens a month would pay about $110 (50 × $1.36 plus 10 × $4.18), and cached input drops to $0.14 per million. Reflection hasn't said how it will price hosted access, and with Apache 2.0 weights, third-party hosts will likely set their own prices anyway.

Our take: who should try which

This is Geeknewz's opinion, not either company's. If you need image understanding, document work or an API you can test this week, Large 4 is the only one of the two you can actually use right now, and its early coding scores look strong. If you want to run a model on your own hardware with no strings attached, Beam is the more practical bet on paper, since it needs about half the memory and compute and comes with the clearer license. Either way, we'd hold off on committing until the weights and technical reports land later this month and independent evaluators have rerun these benchmarks. Until then, both charts are the companies grading their own homework.