PrismML Wants to Shrink AI Down to Fit in Your Pocket
If PrismML isn't on your radar yet, it probably should be — not because it has raised piles of cash (it hasn't, just a modest $22.25 million seed round), but because of the technical pedigree behind it and the potentially industry-shifting technology it's building.
The company's core bet is contrarian: capable, high-performing, reasoning-capable large language models don't actually have to be large. PrismML is compressing reasoning models so aggressively that they can run on everyday PCs and smartphones. The startup is even rumored to be in talks with Apple, though CEO Babak Hassibi declined to confirm that to TechCrunch.
Bonsai 2 27B: A 27B Model in Under 6 GB
On Thursday, PrismML released Bonsai 2 27B, the latest in its Bonsai family of compressed models. It takes Qwen3.8 27B — a widely used open-source model from Alibaba — and squeezes it down to just 5.9 GB. That's small enough to fit on a standard PC and, potentially, on a high-end smartphone. In memory terms, it's a 9x to 10x reduction versus the original model.
The Team Behind the Compression
PrismML was founded by a group of Caltech researchers and is led by Hassibi, a Caltech professor and an expert in compression technologies. The startup also counts Ion Stoica as an advisor. Stoica is a co-founder of Databricks and several other companies, and he directs UC Berkeley's famed Sky Computing Lab — the research group behind a long line of technologies and startups, from Letta to SGLang.
PrismML's backers include Khosla Ventures, Cerberus Capital, and Caltech itself.
Not the Only Player, But Claiming a Different Edge
PrismML is far from alone in the LLM compression race. Multiverse Computing, founded by a well-known professor from Spain's Donostia International Physics Center, is another notable contender — and it has raised significantly more capital.
Hassibi argues that PrismML's compression technology is distinctive because its models lose virtually no performance relative to the originals. Bonsai 2 matches 98% of Qwen's aggregate benchmark scores, up from the first Bonsai released in March, which hit 95%. That original model has already been downloaded more than 11 million times, and PrismML's even smaller models have racked up another 2.6 million downloads, according to the company.
The trajectory is clear: each generation gets closer to the original. Whether PrismML can ever reach 100% benchmark parity remains an open question. Hassibi acknowledges that compression will likely always have some impact.
Why 98% Is Probably Good Enough
Still, perfect benchmark parity is largely an academic concern. LLMs aren't so precise in their uncompressed form — and benchmarks aren't so perfectly representative of real-world tasks — that a 2% degradation would meaningfully change how a model performs in practice. On top of that, the surrounding software — the harness a model runs inside of — matters enormously for real-world accuracy.
The Bigger Picture for 2026
PrismML's progress lands at a moment when the AI industry's center of gravity is shifting from raw model scale toward efficiency, on-device inference, and deployment economics. As inference costs dominate AI budgets and privacy-conscious users push back against cloud dependence, the ability to run a genuinely capable reasoning model locally — on a laptop or a phone — is no longer a novelty but a strategic advantage. If PrismML's compression claims hold up under real-world workloads, the company's tiny models could quietly reshape how and where AI gets used.
via TechCrunch AI
