Quasar 438B benchmark results showing Europe's highest-scoring AI model on the Artificial Analysis index

Quasar 438B Beats Every European AI Model With a 43-Point Score

Quasar 438B, the first large reasoning model from Spain’s Multiverse Computing, scored 43 on the Artificial Analysis Intelligence Index and now ranks as the highest-scoring AI model built anywhere in Europe.

Key Takeaways

  • Quasar 438B scored 43 on the Artificial Analysis Intelligence Index, against 38 for Nvidia Nemotron 3 Ultra and 30 for Mistral Medium 3.5
  • It returned 500 tokens including reasoning in 15.3 seconds, faster than every model scoring above it
  • The model runs in English and Spanish and is priced at $0.60 per million input tokens and $1.80 per million output tokens
  • Claude Opus 5 scored 63 on the same index, so the result is a European record rather than a frontier one

Why Quasar 438B Came From A Compression Firm

Quasar 438B is the first large model Multiverse Computing has shipped under its own name. The San Sebastian company spent its first six years selling a technique rather than a model.

Its CompactifAI platform uses tensor networks, borrowed from quantum physics, to remove 50 to 80 percent of an existing model’s parameters while holding accuracy close to the original.

The company announced the model as a reasoning system for enterprise agents and coding work, served through the CompactifAI API.

Founded in 2019 and now around 480 people, it announced a Series C in July targeting up to $570 million at a reported $1.7 billion valuation.

What The Quasar 438B Score Actually Measures

The Artificial Analysis Intelligence Index is a composite of nine separate evaluations rather than a single test, covering knowledge, graduate-level reasoning, code, tool use and long-context work. A model cannot win it by specialising.

Also Read: Multiverse’s Quasar 438B is First European Model to Top AI Index

Also Read: Chinese AI Models Crushing US Rivals With 90% Price Cuts

Quasar 438B was strongest on long context, scoring 75.0 on AA-LCR, level with Grok 4.6 and within a point of Claude Opus 5 at 75.7. On Terminal-Bench v2.1, the agentic coding test, it scored 69.3, leading Mistral Medium 3.5 by 18.7 points.

Speed Is The Number That Matters For Agents

The figure the company leans on is 15.3 seconds for 500 tokens with reasoning included. Artificial Analysis measured output at roughly 183 tokens per second and a first token in 1.06 seconds, against a 2.06 second median for reasoning models.

For an agent chaining dozens of calls before a human sees anything, latency compounds in a way a benchmark score does not.

Where Quasar Still Trails The Frontier

Claude Opus 5 scored 63 against Quasar 438B’s 43 and led the coding test by 19.8 points. The model is text only, with no image or audio input, and it is proprietary rather than open weights.

Artificial Analysis also recorded unusually high output volume during testing, roughly 350 million tokens, which points to verbose reasoning. At $1.80 per million output tokens, that verbosity is the cost a buyer should model before the headline rate.

Europe’s best-scoring model no longer comes from the continent’s most famous AI startup. Whether that holds past the next release from Paris or Munich is the open question.

Read Next: OpenAI’s Astra Model Crosses Critical Cyber Risk Threshold

Similar Posts