Tech

iPhone 18 Pro’s A20 Pro runs a 27B model at 2× the 17 Pro—then hits the memory wall

· Geeknewz Author

Smartphone held in hand against a soft background

On-device AI just got a louder benchmark—and a clearer ceiling. According to Wccftech’s Sep 20 report, an iPhone 18 Pro running Apple’s new A20 Pro chip demoed the Bonsai 27B language model locally at roughly double the token-generation speed of last year’s iPhone 17 Pro and its A19 Pro.

The clip making the rounds comes from developer Adrien Grondin, who posted that the 18 Pro beat his expectations for pocket inference. The hardware story underneath is why Apple finally spent silicon budget on a dual Neural Engine.

Close-up of a computer circuit board
Photo via Unsplash (https://unsplash.com/photos/1518770660439-4636190af475). Unsplash License.

What’s actually new in the SoC

Wccftech notes the A20 Pro is Apple Silicon’s first dual-16-core Neural Engine—an NPU layout aimed squarely at AI workloads rather than a modest generational bump. Pair that with 12GB of faster 96-bit LPDDR5X on both iPhone 18 Pro and Pro Max, plus reported unified memory bandwidth around 115.2 GB/s, and you get a phone that can keep a heavily compressed 27-billion-parameter model fed without constant cloud round-trips.

For NPU-bound work, Wccftech cites peak Neural Engine throughput that can outpace the SoC’s 7-core GPU—a reminder that “AI phone” marketing is increasingly about the accelerator mix, not just CPU scores.

Person using a smartphone outdoors
Photo via Unsplash (https://unsplash.com/photos/1512941937669-90a1b58e7e9c). Unsplash License.

The catch: 1-bit fits, 2-bit doesn’t

Bonsai 27B in the demo is a 1-bit quantized model. Aggressive compression is why it can run on devices with as little as 4GB of RAM and fly on 8GB or 12GB handsets. The awkward follow-up on X asked why anyone would keep running Bonsai after Bonsai 2 shipped. The disappointing answer, per Wccftech: Bonsai 2’s denser 2-bit quantization is simply too big to fit comfortably in iPhone 18 Pro local memory, and performance falls apart when you try.

That is the real product lesson. Token speed headlines are sexy. Memory capacity decides which model generation you can actually ship offline.

What this does—and does not—prove

A doubled Bonsai 27B run is evidence of headroom for future on-device Apple Intelligence features. It is not proof that Siri, every AI app, or the whole phone is universally twice as fast. Quantized demos also are not desktop-class assistants living permanently in your pocket. They are existence proofs with footnotes.

Still, the trajectory matters. Last year’s A19 Pro already showed that 27B-class compressed models could crawl onto phones. This year’s dual Neural Engine plus fatter, faster RAM turns the crawl into a sprint—until the next denser model refuses to fit.

Why the benchmark landed now

Independent on-device demos have been stacking for months. PrismML and other compression shops already showed related 27B-class models crawling on last year’s iPhone 17 Pro hardware. Grondin’s 18 Pro clip is the sequel with better silicon: same model class, roughly twice the tokens per second, and a Neural Engine layout that finally looks purpose-built for local LLM work rather than photo-enhancement leftovers.

That timing matters for buyers. Apple Intelligence features still lean on a mix of on-device and Private Cloud Compute paths depending on the task and region. A phone that can sprint a 1-bit 27B model offline does not automatically unlock every Siri AI skill without a network—but it does raise the floor for what Cupertino can keep private on the handset when the model fits.

It also reframes the upgrade pitch. Camera bumps and battery tweaks are easy to screenshot. Dual Neural Engines and 12GB of faster LPDDR5X are harder to market in a keynote—until a developer posts a side-by-side token race and the internet does the advertising for free.

Geeknewz take

Treat the Grondin clip as a memory-wall progress report, not a Siri revolution. Apple is buying local-AI ceiling with NPU and bandwidth; denser weights will keep bouncing off RAM until Cupertino ships bigger memory configs. Cloud agents will keep eating the hard prompts. Phones will keep eating the private, offline, low-latency ones that fit. The interesting fight is which workloads cross that line first.

Source: Wccftech — Apple’s A20 Pro Demonstrated To Be An On-Device AI Beast (Omar Sohail, Sep 20, 2026).