AI

The $3.5M receipt: AI labs are livestreaming their OOMs now

· Geeknewz Author

Laptop with code on screen in a dim workspace

Original Geeknewz editorial — analysis from public releases and reporting, not a single-outlet rewrite.

For years, frontier AI felt like a magic show with the curtains welded shut. Models appeared. Benchmarks appeared. The bill, the broken runs, and the “we deleted that dataset because the rollouts looked cursed” notes stayed backstage.

Rows of illuminated server racks in a data center
Photo via Unsplash (https://unsplash.com/photos/1558494949-ef010cbdcc31). Unsplash License.

This week the backstage door cracked. Xiaomi open-sourced its MiMo-V2.6 omnimodal family and—more importantly—ran a public reinforcement-learning dashboard for roughly six days while the meters were still spinning. Combined spend on the streamed Pro and Flash RL runs landed near $3.47 million. The log did not only celebrate pass-rate bumps. It announced GPU OOMs from expert load imbalance, cluster-to-grader network failures, a three-hour undetected dataset glitch, and restarts mid-run.

That is not a press kit. That is a receipt with coffee stains.

Close-up of colorful source code on a monitor
Photo via Unsplash (https://unsplash.com/photos/1516116216624-53e697fedbea). Unsplash License.

Open weights were table stakes; open mess is the flex

Shipping MIT-tagged weights to Hugging Face is now a crowded party. Community GGUF and MLX conversions appear before marketing finishes the champagne. What still separates a release is whether outsiders can see the process: how much compute burned, where the trainer face-planted, which datasets got yanked because rollout logs smelled wrong.

MiMo-V2.6’s streamed board—about 30 steps per run, on the order of tens of billions of tokens, ~750,000 trajectories each—turned those questions into a spectator sport. Relative training-task gains of ~25% (Flash) and ~12% (Pro), plus DeepSWE-style progress markers, matter. The OOM notices matter more for culture. Closed labs occasionally publish postmortems. Live meters with failure banners while spend climbs are a different genre.

What a public log still cannot prove

Transparency theater has limits, and the honest read admits them. Dashboards are self-reported. Outsiders cannot audit the counters. Vendor harness scores are progress markers, not gospel. Restarts at step 15 or 17 mean the published “30 steps” can be the surviving tail of a longer, more expensive job—with discarded compute nowhere on the invoice. Open weights are not an open pipeline: licenses cover files, not the full data recipe or trainer code.

Benchmark tables themselves tell a nuanced story. MiMo’s own comparisons can look ferocious on one cyber eval column and merely mid-pack on exploit-construction benches versus closed giants. Finding known weaknesses in controlled settings and building working exploits are different skills. A public log that shows both the brag and the gap is still rarer than a leaderboard screenshot.

Why labs do this anyway

Because reputation is a scarce resource in an industry drowning in claims. Publishing the ugly restarts signals engineering seriousness to researchers who have lived those failures. It also needles rivals who treat training as classified. Hacker News reactions to the MiMo drop fixated less on absolute capability and more on process—people calling the realtime board a teaching tool, others reading it as a shot across American frontier labs. Both readings can be true at once.

There is a competitive angle, too. When U.S. labs debate pacing the frontier and antitrust headlines swirl around coordinated slowdowns, a Chinese OEM livestreaming RL spend and bad days is a soft power flex: we will show the work. Whether that raises safety or just marketing is a separate argument. The information is public either way.

What would make “train in public” durable

Three bar raises, none of them romantic:

  • Independent evals on harnesses the vendor does not control.
  • Reproductions that hit the same failure classes, not just the same model card adjectives.
  • A second live run that still publishes bad news when the bad news is worse than an OOM restart—because that is when disclosure actually costs something.

Until then, treat streamed training boards as a new genre of open-source evidence: better than silence, not the same as an audit.

Geeknewz read

The industry’s next credibility race is not only who tops a private index. It is who can stand next to a live cost counter when the cluster hiccups. MiMo-V2.6 will not settle the frontier by itself. It did something rarer this week: it made the mess visible on purpose. If more labs copy the habit—and survive the screenshots—buyers, researchers, and regulators all get a clearer map of what “frontier” actually costs.