Upstage’s Solar Mini 4 Hits 208 Tokens/Sec Free In Cline, Tops AAII Score
Cline launched Upstage‘s Solar Mini 4 free in its AI coding assistant on Oct. 9, giving users the highest AAII score of any free model available in the tool.
Key Takeaways
- Cline launched Upstage’s Solar Mini 4 free in its AI coding assistant on Oct. 9
- Solar Mini 4 has a 524,000-token context window and runs at 208 tokens per second
- The model has 35 billion total parameters, with 3 billion active at any one time
- Cline did not disclose pricing for non-free tiers or expansion to other coding assistants
Cline said the model runs a 524,000-token context window at 208 tokens per second, built on a mixture-of-experts architecture with 35 billion total parameters but only 3 billion active at any one time.
A mixture-of-experts model splits its parameters into specialized sub-networks and activates only a handful for each query, letting it match the output quality of a much larger dense model while running faster and cheaper. The AAII benchmark referenced by Cline measures agentic coding performance, testing how well a model plans and executes multi-step programming tasks rather than just answering single prompts.
Why Active Parameter Count Is The Real Story Behind Solar Mini 4
The 3 billion active parameter figure matters more than the 35 billion total because inference cost and latency scale with active parameters, not total model size.
Also Read: MetaMask’s Agent Wallet Lets AI Bots Trade On Aave: Kulechov Confirms Rollout
That is why it can hit 208 tokens per second, fast enough for real-time coding suggestions while drawing on a much larger pool of specialized knowledge.
Upstage’s Push Into Agentic Coding Tools
Upstage has built a smaller but increasingly visible presence among Asian AI labs releasing open and semi-open models aimed at developer tools rather than consumer chat.
The model follows a string of mixture-of-experts releases from Chinese and Korean labs this year that prioritize efficiency over raw parameter count, driven by the cost of running dense frontier models at scale inside coding assistants used thousands of times a day by a single developer team.
What To Watch As Adoption Spreads
Cline did not disclose pricing for non-free tiers or whether the model will expand to other coding assistants beyond its own platform. The real test comes from sustained agentic coding benchmarks over the coming weeks rather than a single score at launch.
Read Next: Agentic RL Breakthrough Turns Checkable Tasks Into Rewards
