Frontier AI models fail buried audit evidence despite high benchmark scores
Audit of frontier LLMs reveals shallow document-reading ability: models score well on standard benchmarks but accuracy drops sharply when evidence is buried, with increased hallucinations, tool cal...