Study reveals LLM agent credit signals fail to identify causally important steps

Research auditing step-level credit assignment in LLM agents finds that LLM-judge scores, logprob ratios, and confidence signals fail to identify which steps actually matter better than chance when...

18 more from arXiv: AI + Accounting

Researchers propose LEDGER framework to audit LLM agent workflows

LEDGER introduces claim-to-evidence trace graphs to make LLM agent outputs auditable, addressing the challenge of verifying correctness in autonomous workflows that execute code, edit files, and ge...

Mandato brings cryptographic audit trails to AI agent actions

Researchers propose Mandato, a protocol-level governance framework enforcing digitally signed mandates on AI agent actions with cryptographically chained audit trails—addressing authorization and a...

New XAI method aids auditing of opaque AI systems in accounting

Researchers propose 'Rule of Thumb' explainability framework to interpret AI decisions using partial information, with applications to LLM classification and AI system audits—relevant for accountan...

Auditing AI recommendations against complete census reveals blind spots

Research framework for auditing AI recommendation systems by comparing outputs against complete data census, addressing critical accountability gaps in AI-driven decision-making for high-stakes dom...

Stay ahead of AI in accounting

Get the latest news on agentic AI for accounting, audit, and tax delivered to your inbox. Curated by AI, reviewed by professionals.

Subscribe to Newsletter