AI Agents Unlock 657-Fold Cost Cuts, New Benchmark Finds
AI Agents were the focus of a new benchmark published Tuesday that tests whether language models can turn their own expensive capabilities into cheap, reusable shortcuts instead of running full queries every time.
Key Takeaways
- BOTTLED tests whether AI agents can convert expensive capabilities into cheap, reusable tools under fixed time, compute and API budgets
- Across ten models and three tasks, raw problem-solving skill did not predict how well models could bottle skills cheaply
- In 48 of 60 test runs, bottled versions scored below the lower bound of the original model’s performance range
- Opus 5 retained roughly 82% of its original accuracy while cutting costs by about 657 on query-product relevance classification
AI Agents And Bottling
The benchmark, called BOTTLED, gives AI agents an entire unlabelled workload and forces them to complete it under fixed time, compute and API budgets, researchers wrote in a paper posted to arXiv.
The approach is named bottling, an agent’s ability to convert a general skill, like classifying text, into a specific, cheap tool such as a small trained model or a reusable script, rather than calling an expensive large model for every single instance.
Large language models are AI systems trained on vast text datasets to generate humanlike responses, and running them repeatedly at scale, sometimes millions of times for near-identical tasks, is one of the industry’s biggest hidden costs.
Also Read: Two Words Nearly Double A Language Model’s Math Score
AI Agents’ Bottling Results
Across ten models and three tasks, the researchers found that a model’s raw problem-solving skill did not predict how well it could bottle that skill cheaply. In 48 of 60 test runs, the bottled version scored below the lower bound of the original model’s own performance range.
But when bottling worked, the savings were large.
On a query-product relevance classification task, Anthropic‘s Opus 5 retained roughly 82% of its original accuracy score while cutting costs by a factor of about 657. The model also held up against Jev, a smaller system built specifically for cheap repetitive inference, recovering about 94% of Jev’s accuracy at a quarter of its projected cost.
The paper’s authors note that 31 of 60 bottling runs still underperformed simpler small-model distillation baselines, meaning the technique is inconsistent rather than a guaranteed win.
Whether future agent frameworks build bottling into default workflows, or whether it stays a task-specific optimization, depends on results from a wider range of models than this initial study covered.
Read Next: Claude Models Unlock Access For Approved Cyber Defenders
