Researchers propose audit framework for verifying AI agents' self-reported ca...
Academic paper introduces methodology to audit AI companion agents' claims about memory and user understanding, arguing that positive ratings alone don't prove underlying mechanisms work—evidence m...