FinRCA-Bench tests LLM reasoning on financial reconciliation tasks
New benchmark evaluates how well LLMs retrieve evidence across invoices, POs, and ledgers for financial root-cause analysis—separating true reasoning from document access.
Research paper evaluates frontier LLMs (Claude, ChatGPT) on IPO financial analysis tasks using automated rubric generation, extending Finance Agent v2 benchmarks to SEC S-1 filings and SpaceX IPO c...
Continue reading