Ant Group has shipped Ling-3.0-flash-Fin, an open-weights model built on its Ling-3.0-flash base and tuned for financial work: source checking, valuation spreadsheets, drafting reports. Ant says financial institutions and industry experts helped build it, without naming which ones.

Artificial Analysis put the new model through its Intelligence Index and recorded a score of 23, the same mark MiniMax-M2.7 posts. Fin reaches that number while activating just 5.1 billion parameters per token, roughly half of MiniMax-M2.7’s 10 billion, and both models carry 124 billion total parameters. That is a parameter story, not a compute story. Fin needed about 67,000 output tokens on average to complete an Intelligence Index task, more than triple MiniMax-M2.7’s 21,000 and about a third above the 50,000 that Ant’s own Ling-3.0-flash-VL used for the same run. A smaller active-parameter count only lowers inference cost if the model does not talk its way back to parity with extra tokens, and here it does.

On finance-specific work, Fin also scored 24 on Artificial Analysis’s Finance and Accounting Index, tying VL exactly. The tie hides a real split: Fin’s business knowledge accuracy came in at 17 percent against VL’s 11 percent, but Fin’s non-hallucination rate on that same test fell to 67 percent, well below VL’s 81 percent. A model that gets more financial facts right, then confabulates more often when it gets one wrong, is a tougher sell to a compliance desk than a matching headline score implies.

Fin trails VL on the tasks closest to actual analyst output. On GDPval-AA v2, which grades agents on professional knowledge work, Fin scored 1171 Elo against VL’s 1225, though it still beat MiniMax-M2.7’s 1087. On AA-Briefcase, where agents turn document sets into memos and spreadsheets, Fin posted 967 to VL’s 986 and cleared fewer rubric checks along the way, 23.5 percent versus 24.9 percent. Fin’s one edge was presentation quality, coming in slightly ahead of VL despite lacking VL’s image inputs, a result Artificial Analysis called surprising rather than a capability Ant had claimed.

Agentic reliability shows the widest gap. AutomationBench-AA, which tests business-app workflows under guardrails, gave Fin 7 percent against VL’s 16 percent. Both models flatlined at 0 percent on Terminal-Bench v4.0, a harder terminal-use benchmark, so neither is close to running command-line agent tasks unsupervised.

The model is text-only, holds a 256,000-token context window, ships under an MIT license, and is available on OpenRouter with a rate-limited free tier. A finance team weighing Fin against MiniMax-M2.7 or its own sibling VL should price in both numbers the headline score leaves out: a hallucination rate worse than VL’s and a token bill triple MiniMax-M2.7’s, since a cheaper-looking model that writes more and checks itself less is not a savings until someone verifies its output.

Based on reporting and benchmark data published by Artificial Analysis on September 16, 2026.