TaxCalcBench tests whether AI can autonomously file taxes
New benchmark evaluates LLMs' ability to handle tax filing tasks end-to-end, revealing capability gaps in real-world tax preparation workflows.
Kepler develops novel double-entry evaluation framework combining code-based metrics and LLM assessment to grade accounting task quality, extracting most value from disagreements between evaluation...
Continue reading