Cactus Compute’s Needle Crams a Foundation Model Into 29 Megabytes for Your Watch

Sanket Chaukiyal

October 5, 2026

TL;DR

  • Cactus Compute released Needle, an 8 to 29 MB foundation model binary built for phones, wearables, robots, cars, and microcontrollers.
  • It skips general chat entirely and focuses on tool calling, structured JSON extraction, text embeddings, and speech processing.
  • Cactus claims Needle beats models ten times its size on mobile tool calling and matches models two to three times its size on structured extraction.
  • A companion speech model called Whistle ships at 16.9 MB, claimed to be 8.6x smaller than Whisper Base and Moonshine.

A Foundation Model That Fits in a Text Message

Cactus Compute uploaded a GitHub repository called smallAI_18mb that documents its third-generation compact model, Needle. The whole thing ships as a single binary ranging from 8 MB to 29 MB depending on configuration, small enough to sit next to a handful of photos on your phone.

The README describes Needle’s architecture as a Laddered Simple Attention Network. That’s a mouthful, but the pieces matter: Monarch Hadamard MLP layers, grouped query attention (GQA), an engram n-gram memory module, and multi-lane hyper-connections. The model can also scale its own depth, running anywhere from 2 to 20 subnetwork layers depending on what the device can spare.

Needle quantizes down to 2.125 bits per weight, stored in Cactus’s own.cact container format. That’s an aggressive compression target, well below the 4-bit or 8-bit quantization most mobile LLM projects settle for.

The project doesn’t try to be a chatbot. Per the repo: “The whole model is a single 8-29 MB binary built on our SimpleAttention Network, and we trade general chat capacity to beat models 10x its size on mobile tool calls and match 2-3x bigger models on extraction.” Translation: don’t ask it to write you a poem. Ask it to parse a calendar invite or call an API, and it apparently holds its own against models many times its weight class.

Alongside Needle sits Whistle, a 16.9 MB speech-to-text model that shares the same C++ engine, runtime, and quantization container. The repo lists parameter configurations around 50M and 121M for the family, though it doesn’t spell out which number maps to which variant.

Our Take

I’ve watched a lot of “small model” announcements promise the moon and deliver a rounding error. This one at least picks a fight it can specify: tool calling and extraction, not general intelligence. That’s a narrower claim, and narrower claims are easier to actually hit.

Think of it like a Swiss Army knife that ditched the corkscrew and the toothpick to make the main blade sharper. Needle isn’t trying to be your conversational companion. It’s trying to be the thing that reads a text message, pulls out a flight number, and hands it to another system without you noticing it happened at all.

The competitive numbers are the real headline here. Cactus claims Needle beats models ten times its size on mobile tool calling and keeps pace with models two to three times larger on extraction. Whistle, meanwhile, goes after Whisper Base and Moonshine, two names that actually matter in the speech-to-text space, claiming an 8.6x smaller footprint while staying in the same ballpark on transcription. If that holds up under independent testing, it’s a real shot at the assumption that small models have to sacrifice everything to get small.

But here’s the catch nobody in a README ever volunteers: these are self-reported benchmarks from the team that built the model. There’s no third-party leaderboard result attached, no Hugging Face eval card, nothing from Whisper’s or Moonshine’s maintainers confirming the comparison. Is an 8.6x size reduction impressive? Sure. Is it meaningful without someone outside Cactus Compute running the same test suite? Not yet.

And the 2.125-bit quantization deserves a raised eyebrow too. That’s an unusually aggressive compression ratio, lower than the 4-bit schemes most mobile AI projects use, which makes me want to see accuracy numbers at that bit depth before getting too excited about the size alone.

Why Edge Devices Needed This

Running a large language model on a phone is one thing. Running one on a microcontroller with a few hundred kilobytes of RAM is an entirely different problem, and it’s the problem Cactus Compute built Needle to solve. Most foundation models, even the “small” ones making headlines this year, still measure in the hundreds of megabytes or low gigabytes once quantized. That puts them firmly out of reach for wearables, smart home sensors, and the embedded chips found in cars and robots.

The push toward compact, specialized models has been building for a while. Instead of cramming general intelligence onto tiny silicon, developers increasingly carve off a narrow slice of capability, tool calling, extraction, transcription, and optimize hard for that slice alone. But does narrower really mean better, or just smaller ambitions dressed up as a feature? Needle fits squarely into that trend either way. It isn’t a smaller version of a chatbot. It’s a different kind of object altogether, built from the ground up to live inside a battery budget measured in milliwatts rather than a cloud budget measured in GPU hours.

Three Things That Will Decide If Needle Sticks

Watch for independent benchmarking first. Self-reported numbers against models ten times the size are a bold claim, and the open source community has a habit of reproducing these tests within weeks of a release landing on GitHub. If Needle’s tool-calling results hold up outside Cactus’s own test harness, that’s when this stops being a repository curiosity and starts being a reference point other teams cite.

Watch adoption of the.cact container format too. A proprietary quantization format only matters if other developers actually build around it. If it stays locked to Cactus’s own runtime, Needle risks becoming a neat demo rather than real infrastructure.

And keep an eye on how Whistle performs against Whisper Base and Moonshine once someone outside the project runs a head-to-head transcription test. Speech recognition has plenty of entrenched players, and “8.6x smaller” only wins the argument if the accuracy gap stays small too.

Editor's Note

I keep coming back to that 2.125-bit quantization number. Most teams stop at 4-bit because accuracy falls off a cliff below it, so either Cactus found something clever or they're leaving out the part where performance suffers. What I actually want to see is someone outside the project running Needle against a real device battery drain test, not a benchmark table. Small model claims are cheap until a wearable runs it for a week straight.

– Sanket Chaukiyal, founder, SmartChunks

FAQ

What is Needle?

Needle is an 8 to 29 MB foundation model built by Cactus Compute, documented in the smallAI_18mb GitHub repository, designed to run tool calling, structured extraction, embeddings, and speech tasks directly on mobile, wearable, and microcontroller hardware.

How does Needle differ from a typical chatbot model?

Needle trades general conversational ability for narrow, specialized performance. According to the repo, it's built to beat models ten times its size on mobile tool calling and match models two to three times larger on structured extraction, rather than handling open-ended chat.

What is Whistle?

Whistle is Cactus Compute's companion speech-to-text model, a 16.9 MB file that shares Needle's C++ engine, runtime, and.cact quantization container. The repo claims it runs 8.6x smaller than Whisper Base and Moonshine while staying competitive on transcription.

Are Needle's performance claims independently verified?

Not yet, based on what's in the repository. The benchmarks come from Cactus Compute itself, with no outside leaderboard or third-party eval cited in the README, so the comparisons against larger models should be treated as a claim to test rather than a settled result.

Sanket Chaukiyal — Editor at Smart Chunks

Sanket Chaukiyal

Technology editor • 12+ years in editorial

Sanket is the founder and editor of Smart Chunks. He spent over six years at Autocar India (Haymarket SAC Publishing) as Sub Editor and Senior Copy Editor, and later served as Account Director (Content) at Rite Knowledge Labs. He holds a Master's in Media and Communication from the Symbiosis Institute of Media and Communication.

All articles → • LinkedIn