FinRCA-Bench tests LLM reasoning on financial reconciliation tasks
New benchmark evaluates how well LLMs retrieve evidence across invoices, POs, and ledgers for financial root-cause analysis—separating true reasoning from document access.
arXiv paper introduces Herculean, the first benchmark evaluating agentic AI on financial intelligence tasks beyond static Q&A, testing whether agents can reliably perform actual financial professio...
Continue reading