Surge AI, the AI training data and evaluation company, has published a benchmark called DAYJOB: Finance that hands AI agents 80 real-world finance assignments built from actual supporting documents, then checks whether the write-up an agent hands back would hold up to a professional’s scrutiny.
One sample task included in the benchmark’s public grading rubric shows how demanding that bar is. The scenario asks an agent to advise a furniture retailer on whether to fund a new store buildout, using financial statements that contain planted errors: sign mistakes in the “returns and allowances” line that quietly inflate reported profit. To score well, an agent has to catch those errors, recalculate earnings correctly, run a discounted cash flow analysis for each candidate location, and check the results against a real debt covenant, a minimum cash reserve, and a required shareholder payout, then deliver a written recommendation a partner could act on.
Surge AI has not published which models it tested or how they scored. The rubric alone signals where most “finance agent” demos are likely to fail: not on formatting, but on the buried arithmetic errors that only a careful analyst would catch.
Reported by Surge AI on its DAYJOB: Finance benchmark page; no publication date was given.