Claude Opus 4.7 tops GPT-5.4 on accounting AI benchmark
Claude Opus 4.7 scored 79.2% on DualEntry's accounting task benchmark, surpassing GPT-5.4 within hours of release—signaling continued AI model competition in accounting-specific use cases.
Claude Opus 4.7 scored 79.2% on DualEntry's accounting task benchmark, surpassing GPT-5.4 within hours of release—signaling continued AI model competition in accounting-specific use cases.
Researchers propose an explainable ensemble learning approach using Shapley values to detect financial fraud while meeting OCC and Federal Reserve transparency requirements, addressing $32B annual ...
arXiv paper introduces Cognitive Core, a decision substrate designed to prevent silent errors in AI systems handling regulatory compliance and institutional decisions like clinical triage and appeals.
Academic research on 1,801 participants shows humans consistently over-attribute responsibility to human decision-makers in AI-assisted lending workflows, raising compliance and ethics risks for fi...
arXiv paper proposes formal framework for auditability of LLM agents that act on external systems, addressing accountability and responsibility assignment in deployed agent systems.
Researcher proposes using generative AI to automatically construct formal, evidence-linked argument graphs for regulatory compliance and certification audits, addressing traceability and accountabi...
De Jure pipeline uses iterative LLM self-refinement to automatically convert dense regulatory documents into structured, machine-readable rules—eliminating manual expert annotation in compliance wo...
ArXiv study warns that agentic LLMs with local machine access can leak credentials and redirect transactions via prompt injection—a threat beyond standard jailbreak tests.
Academic study reveals personal AI agents with elevated privileges—like OpenClaw—vulnerable to prompt injection attacks that could leak credentials or redirect financial transactions, exposing gaps...
Analysis questions whether LLMs can reliably answer tax questions, highlighting accuracy gaps in a core accounting use case.
Unpaved introduces an audit framework to evaluate bias in AI developer tools used in Global South contexts, addressing how automation platforms may perpetuate inequality in emerging markets.
PoliTax Split introduces a benchmark dataset using presidential tax returns to evaluate AI models' ability to extract and categorize tax form data, advancing automation in tax document processing.
Job postings for accountants mentioning AI/ML skills jumped 67% YoY, marking the largest rise among any profession, signaling growing industry demand for AI-capable accounting talent.
Researchers propose a measure-theoretic Markov framework to evaluate reliability, ambiguity, and governance costs when autonomous AI agents replace deterministic workflows in organizations.
Researchers propose LineMVGNN, a spectral graph neural network designed to detect suspicious transactions and accounts in AML systems with improved accuracy over rule-based methods while maintainin...
Academic paper identifies critical gap in AI agent security: lack of pre-execution authorization controls for tool calls like fund transfers and database queries, proposing policy-based enforcement...
Academic paper proposes evaluation metrics beyond model accuracy to assess whether human-AI teams can collaborate safely and effectively, addressing miscalibrated reliance in AI deployments.
New open benchmark extends agent evaluation to knowledge-intensive document retrieval and full-duplex voice, with GPT-5.2 reaching ~25% on complex policy navigation tasks.
Job postings for accounting roles mentioning AI skills jumped from 18% to 30% year-over-year, signaling rapid industry shift toward AI-literate talent.
DualEntry's benchmark tested 19 AI models on 101 accounting tasks; GPT-5.4 led with 77.3% accuracy but failed on roughly 1 in 4 real-world accounting workflows.