Ant Group Open-Sourced Its Finance AI — And a Benchmark to Grade It

Ant Group has open-sourced Ling-3.0-flash-Fin, a finance-tuned model built to carry out an analyst’s workflow rather than answer questions about it, in an announcement made Tuesday at the 2026 Inclusion Conference on the Bund in Shanghai.

The model has 124 billion total parameters with around 5.1 billion active per token. Its weights are available under an MIT licence through Hugging Face and ModelScope, and the model also runs on OpenRouter and Vercel

Ant says it was developed with financial institutions and industry experts, and that users can deploy it privately and connect it to search, Python, databases and spreadsheets. TechNode reported in late August that Ant planned to open-source the model within a week of its launch.

From Answering Questions To Doing The Analysis

Ant is targeting four tasks: sourcing information from official filings, reasoning across multiple documents, valuation modeling and report generation.

The demonstrations show the chain rather than the individual steps. In one, the model reconstructed four disclosures on Google’s monthly token volume using 23 tool calls, prioritizing primary sources and maintaining an evidence trail. 

In another it made 55 tool calls across 10-K filings, earnings releases and 8-K disclosures for Spire, rebuilding continuing-operations earnings and validating $17.8 million in historical profit. A third updated a workbook of seven sheets and more than 5,000 formulas while leaving the file editable.

That is closer to the work of a junior analyst than to a question-and-answer chatbot.

Also Read: Mistral Valuation Hits $24 Billion in Europe’s First MEGA Round

A Benchmark That Grades The Process

Ant is also open-sourcing FinFIRST, a benchmark for financial search agents built with the investment banking team at China International Capital Corporation. The first version contains 123 expert-authored tasks, 701 atomic criteria and 12,300 rubric points, and scores the full research process rather than matching a final answer.

That design matters. Judging an agent on its answer alone rewards a system that reaches the right number by the wrong route, which is the failure mode that makes banks wary of automation.

Also Read: OpenAI’s Breakthrough Navier-Stokes Claim Draws Mathematician Rejection

Where The Limits Sit

Ant reports competitive results across FinFIRST, APEX-Agents, SpreadsheetBench and τ³-Banking, but every score is vendor-reported and no independent evaluation has been published. 

Ling’s own team says expert review remains necessary for key assumptions, valuation outputs and investment conclusions, and that the outputs do not constitute investment advice. 

The model can trace report figures, calculate valuations, and run sensitivity analysis. It provides the rigorous data foundation needed, leaving the critical validation of economic assumptions to human experts. 

Read Next: Anthropic Debuts Claude Design on Sept. 7 as Its Quietest Launch

Similar Posts