TL;DR
- NaiveAI shipped final weights and the model card for Naive-N0.5-Flash, a 309B-parameter mixture-of-experts model with only 15.5B active parameters, under the MIT license.
- The model runs a native 1M-token context window through a hybrid attention setup: 39 sliding-window layers plus 9 DeepSeek-style sparse attention layers, with no full-attention layers at all.
- API access lands at $0.10 per million input tokens, $0.40 per million output tokens, and just $0.01 per million cached tokens.
- Built on the MiMo-V2.5 base and trained across 3.25 trillion tokens, the model claims inference speeds up to 2,000 tokens per second in its Ultrafast mode via NaiveRT.
What NaiveAI Just Shipped
NaiveAI didn’t tease this one for weeks. It pushed a single commit, titled “PUBLISH FINAL MODEL CARD AND ASSETS,” to its Hugging Face repository on September 27, and that commit contains the whole package: weights, inference code, and a model card laying out everything from parameter counts to per-token pricing.
Naive-N0.5-Flash is a mixture-of-experts model with 309 billion total parameters, though only 15.5 billion of those activate on any given pass. That’s the trick MoE architectures play: build something massive, then route each token through a small slice of it. The model card puts it plainly: “Naive-N0.5-Flash is an open-weight 309B MoE model with 15.5B active parameters, built for coding and AI R&D.”
What sets it apart isn’t just the size ratio. It’s the attention mechanism. Across 48 transformer layers, 39 run sliding-window attention and 9 run DeepSeek-style sparse attention (DSA). None run full attention. NaiveAI is betting that combination can still deliver a genuine 1M-token context window without the memory and compute costs full attention usually demands at that scale.
The weights are released under the MIT license, meaning developers can fork, fine-tune, and deploy commercially without asking permission first.
And there’s an API for anyone who’d rather not host 309 billion parameters themselves.
The Sparse Attention Bet
Here’s the wager NaiveAI is making: you don’t need every layer paying full attention to every token to get a usable million-token context window. Full attention scales quadratically, so cost explodes as context grows. Sliding-window attention caps what each layer can see nearby, while DSA lets a smaller set of layers hunt for relevant tokens across the whole sequence, sparsely. Thirty-nine layers doing local sweeps, nine doing the expensive long-range search, zero doing a full scan of everything.
Picture a warehouse with a million shelves. A full-attention model insists on checking every shelf every single time someone asks for one item. NaiveAI’s approach hands most of the floor to thirty-nine workers who only check the shelves directly in front of them, and trusts nine specialist scouts to run the long aisles whenever something needs finding further away. It’s a bet that you can cover a warehouse that size without ever hiring a worker who reads the whole thing cover to cover.
Does it hold up as well as full attention at that scale? The model card doesn’t say, and nothing in what NaiveAI has published so far offers independent proof that it does under real workloads.
NaiveAI is positioning the model against GLM-5.3-Flash, DeepSeek-V4.1-Flash, Step-5-preview, and Gemini-3.6-Flash, running comparisons across coding, systems optimization, and research-style tasks. That’s a deliberate set of rivals. DeepSeek popularized the sparse attention idea Naive-N0.5-Flash borrows from. GLM and Step represent the crowded field of open-weight labs racing to match frontier capability at a fraction of the cost. Gemini-3.6-Flash is Google’s own efficiency-tier model, the kind of closed alternative NaiveAI needs developers to skip in favor of something they can download and run themselves.
I’ve watched enough of these open-weight drops promise the moon on paper only to underperform once real developers get their hands on them, so the benchmark comparisons here matter less to me right now than what happens over the next month of actual usage.
The pricing tells its own story. Cheap tokens, no ceremony. At $0.10 per million input tokens and $0.40 per million output tokens, with cached reads at a single cent per million, NaiveAI is undercutting what most closed frontier labs charge for comparable context windows. Cheap tokens paired with an MIT license make a specific argument: build on this instead of renting from someone else.
The MiMo Lineage
Naive-N0.5-Flash doesn’t start from scratch. It’s built on top of MiMo-V2.5, an existing open-weight base model, then pushed through 3.25 trillion tokens of additional training across three distinct stages: 50 billion tokens of indexer warmup, 3 trillion tokens of sparse attention training, and a final 200 billion token learning-rate decay phase.
That staged approach makes sense once you consider what NaiveAI was actually trying to teach the model. Sparse attention isn’t something you bolt onto a finished model and expect to work cleanly. The indexer warmup phase likely taught the DSA layers how to pick which tokens matter before the bulk of training even began, and the long middle stage is where most of the real capability got built.
This fits a broader pattern in open-weight releases this year: labs aren’t just training bigger models, they’re re-architecting attention itself to make long context affordable. DeepSeek’s sparse attention work is the most visible precedent here, and NaiveAI’s approach reads like a direct continuation of that lineage rather than a wholly separate idea.
What Developers Should Track Next
The first thing worth watching is whether independent testers can reproduce the inference speeds NaiveAI is claiming: 50 tokens per second per user in Standard mode, and up to 2,000 tokens per second in Ultrafast mode through its NaiveRT runtime. Those numbers read well in a model card. Do they hold when a thousand developers hit the API at once instead of one clean benchmark run?
Second, keep an eye on how the sparse attention architecture behaves at the far end of that 1M-token window. Plenty of models advertise huge context lengths that quietly degrade once you actually fill them. Whether Naive-N0.5-Flash’s 39-SWA, 9-DSA split holds coherence at 900,000 tokens in, not just 10,000, is the real test of this whole design.
Third, watch how GLM, DeepSeek, and Google respond on pricing. NaiveAI just set a low bar at ten cents per million input tokens. Someone in this group is going to try to undercut it within weeks.
Editor's Note
What catches my eye here isn't the parameter count, it's the pricing paired with an MIT license. A cent per million cached tokens is aggressive enough that I expect smaller labs to feel pressure within weeks, not months. I'm skeptical of the inference speed claims until someone outside NaiveAI runs them at scale, benchmark numbers on a model card have burned me before. What I'll be watching is whether the sparse attention holds up past half a million tokens, because that's where most long-context claims quietly fall apart.
– Sanket Chaukiyal, founder, SmartChunks
FAQ
What is Naive-N0.5-Flash?
It's an open-weight mixture-of-experts language model from NaiveAI with 309 billion total parameters and 15.5 billion active parameters per token, released under the MIT license with a native 1M-token context window built for coding and AI R&D.
How does its attention mechanism work?
It uses a hybrid setup across 48 transformer layers: 39 layers run sliding-window attention and 9 run DeepSeek-style sparse attention (DSA), with no full-attention layers at all, which is meant to let it handle long context without the usual quadratic compute cost.
How much does the API cost?
NaiveAI prices it at $0.10 per million input tokens, $0.40 per million output tokens, and $0.01 per million cached tokens.
What is Naive-N0.5-Flash trained on top of?
It's built on the open-weight MiMo-V2.5 base model, then trained through 3.25 trillion additional tokens across three stages: a 50 billion token indexer warmup, 3 trillion tokens of sparse attention training, and a 200 billion token learning-rate decay phase.
Source: Hugging Face
