Subscribe Sign in

Culture

PrismML hopes its small LLM will change how we all use AI

TechCrunch
3 min read Rewritten in plain language

Artificial intelligenceStartupsLlms

Show what we removed Rules applied: A1 A3×5 A4 A9 D2×4 D3×5 D4×2 E3×3 F2×2 all 30 rules
  • If AI lab PrismML isn’t on your radar yet, it should be — not because it’s raised gobs of money (it hasn’t yet, just a $22.25 million seed round), but because of the technical minds involved and the potentially industry-changing tech it’s developing.
  • On Thursday, PrismML released Bonsai 2 27B, its latest in a family of models, which compresses Qwen3.8 27B, a widely used open source model from Alibaba, down to 5.9 GB.
  • PrismML was founded by a group of Caltech researchers and is led by Hassibi, a Caltech professor and an expert in compression technologies.
  • PrismML is also backed by investors Khosla Ventures, Cerberus Capital, and Caltech.
  • PrismML’s approach, called “ternary” weights, simplifies that down to three: +1, −1, or 0.

5 sentences from our version of the report, chosen to cover it. Nothing here is written; every line is in the article below. How

On Thursday, PrismML released Bonsai 2 27B, its latest in a family of models, which compresses Qwen3.8 27B, a widely used open source model from Alibaba.

The report’s most important sentence, shortened and in plain words. How

Headline check

Headline as published: PrismML hopes its tiny LLM will change how we all use AI

There is nothing in this headline a machine can check against the report: no figure, no name and no quotation.

Nothing was measured here, so nothing is claimed. How this is checked

Read the full reportHide the full report3 min

If AI lab PrismML isn’t on your radar yet, it should be — not because it’s raised gobs of money (it hasn’t yet, just a $22.25 million seed round), but because of the technical minds involved and the potentially industry-changing tech it’s developing.

PrismML is betting that capable, high-performing, reasoning large language models don’t, in fact, have to be large.

It is making reasoning models so small they can fit on PCs and smartphones. (It’s even rumored to be in talks with Apple, though CEO Babak Hassibi declined to comment on that to TechCrunch.)

On Thursday, PrismML released Bonsai 2 27B, its latest in a family of models, which compresses Qwen3.8 27B, a widely used open source model from Alibaba, down to 5.9 GB. That’s small enough to fit on a PC and, possibly, a high-end smartphone. It’s a 9x to 10x reduction in memory versus the original.

PrismML was founded by a group of Caltech researchers and is led by Hassibi, a Caltech professor and an expert in compression technologies. The startup also counts Ion Stoica as an adviser. Stoica is a co-founder of Databricks (and other companies) and the director of Berkeley’s famed Sky Computing Lab, which has birthed many technologies and startups, from Letta to SGLang.

PrismML is also backed by investors Khosla Ventures, Cerberus Capital, and Caltech.

This startup is not the only company working on LLM compression tech. Multiverse Computing, founded by a well-known professor from Spain’s Donostia International Physics Center, is another. (And Multiverse Computing has raised gobs of cash.)

But Hassibi says that PrismML’s compression tech is unique because its LLMs have lost no performance compared with the originals. Bonsai 2 matches 98% of Qwen’s aggregate benchmark scores. That’s up from the first Bonsai, released a couple of months ago in March, that matched 95%. That original model has already been downloaded over 11 million times, and PrismML’s even smaller models have been downloaded another 2.6 million times, the company says.

So this shows that PrismML’s compression results have improved from one release to the next. Whether it could ever get to 100% benchmark performance parity is a question that remains to be seen. Compression will likely always have some impact, Hassibi says.

Still, benchmark parity is academic anyway. LLMs are not so accurate in their uncompressed form, and benchmarks not reflective of actual tasks, that a 2% degradation would likely meaningfully affect how a model performs in actual use. (Plus, the surrounding software — the harness a model runs inside of — matters a lot when it comes to accuracy, too.)

PrismML says it achieves this by shrinking the “weights” that make up a model — weights are the information a model learns and stores during training. Normally, each weight requires 16 bits. PrismML’s approach, called “ternary” weights, simplifies that down to three: +1, −1, or 0. With far smaller values to store for each weight, the model takes up less space. (For a deeper dive on the compression technique, here’s the project’s Hugging Face page.)

The startup’s next goal is to apply this compression technique to even bigger models. “The next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I expect it will be easier to retain the intelligence there,” Hassibi told TechCrunch.

As model size grows, he added, “There is more room to be able to compress them without losing the intelligence. So I would just say, as a general trend, for larger models, it’s easier to get to 100%.”

Stoica tells us that he’s excited for this tech because it’s making it possible for advanced models to run on users’ devices. “You are going to have intelligence at your fingertips, and it’s going to be free because it’s going to run on the device you already bought. It’s also going to be private, because you’re not going to send it to the cloud.”

The AI graveyard: a running list of projects and startups that didn’t make it Lauren Forristal

You are reading our version, not theirs. This is TechCrunch's report with its verdicts and loaded words taken out. Plain description stays, and so do adjectives that carry a fact, such as "former" or "federal". The reporting, the facts and the quotations are theirs — quotations are never edited — and the indicators beside it measure this version. Hover or tap Adjectives to see every one left in the text.

How this outlet filed it, and how we rewrote it

No other newsroom we read has filed on this event, so there is nothing to compare it with yet.

Outlet Niral ScoreAdjectivesSourcingSentimentHappiness
TechCrunchas they published this story 15.5 17 40 -0.1 47.4
Mundane Readneutralized from TechCrunch 10.3 16 40 0.0 54.8

Sign in to react.

Comments

Nothing here yet.

Sign in to comment.

Questions

Readers can ask a question about this story here. Questions and answers are for subscribers. Sign in to read them.

Comments are read before they appear where anything in them needs a person to look. Nothing posted here is ever deleted; a comment taken down keeps its text and the reason, so the decision can be looked at again. How this works