Researchers design knowledge-gated tasks to test LLM agents on professional c...
arXiv paper introduces protocol for constructing verifiable LLM agent tasks that isolate domain-specific knowledge from task execution, enabling better evaluation of whether agents truly understand...