HANDBOOK.md introduces a benchmark specifically designed to evaluate long-context agentic instruction following in language-model agents. This benchmark assesses how well agents adhere to a system prompt, policy file, or skills document during their operations.
arxiv.org
2 min
7/29/2026
HANDBOOK.md introduces a benchmark specifically designed to evaluate long-context agentic instruction following in language-model agents. This benchmark assesses how well agents adhere to a system prompt, policy file, or skills document during their operations.
arxiv.org
2 min
7/29/2026
HANDBOOK.md introduces a benchmark specifically designed to evaluate long-context agentic instruction following in language-model agents. This benchmark assesses how well agents adhere to a system prompt, policy file, or skills document during their operations.
arxiv.org
2 min
7/29/2026
No more articles to load