Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#claude#ai-ethics#code-generation#ai-safety#openai#anthropic#discussion

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

© 2026 Themata.AI • All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
llmsai-agentsbenchmarkinginstruction-following

Handbook.md shows that long policy documents do not reliably govern agents

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

arxiv.org

July 29, 2026

2 min read

🔥🔥🔥🔥🔥

60/100

Summary

HANDBOOK.md introduces a benchmark specifically designed to evaluate long-context agentic instruction following in language-model agents. This benchmark assesses how well agents adhere to a system prompt, policy file, or skills document during their operations.

Key Takeaways

  • The HANDBOOK.md benchmark consists of 65 agentic tasks designed to evaluate how language-model agents follow long, binding policy documents in professional environments.
  • The benchmark spans five domains: finance, medical billing, insurance, logistics, and HR, with tasks modeled on the behavior of enterprise employees following company handbooks.
  • The best-performing model configuration achieved a pass rate of 36.2% under strict grading criteria, while most configurations remained below 25%.
  • Common failure patterns include agents allowing plausible in-environment requests to override standing policies and losing rule details over long task durations.
Read original article

Community Sentiment

Mixed

Positives

  • Local inference could magically fix LLM defects, giving users control and potentially transforming how we manage complex models.
  • Models like Kimi K3 are pushing the boundaries of frontier performance, suggesting exciting advancements in LLM capabilities.
  • Using a limited context window (like 50%) seems to lead to better outcomes, showing that managing input can enhance model effectiveness.

Concerns

  • Claiming 1M tokens in context is misleading; users face severe limitations due to quantization and poor sampling methods.
  • Local models still exhibit the same defects as cloud models, debunking the idea that local hosting is a panacea for LLM issues.
  • The inability of models to retain information over extended interactions is a significant flaw, undermining their effectiveness in following complex policies.

Related Articles

Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

Evaluating AGENTS.md: are they helpful for coding agents?

Feb 16, 2026

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Study: Self-generated Agent Skills are useless

Feb 16, 2026

A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

Feb 10, 2026

Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

Jul 13, 2026

Latent Programming Horizons in Coding Agents

Coding agents think ahead of time

Jul 14, 2026