Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#discussion#claude#ai-ethics#code-generation#ai-safety#openai#trending

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

© 2026 Themata.AI • All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
llmsai-agentsbenchmarkinginstruction-following

Handbook.md shows that long policy documents do not reliably govern agents

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

arxiv.org

July 29, 2026

2 min read

🔥🔥🔥🔥🔥

63/100

Summary

HANDBOOK.md introduces a benchmark specifically designed to evaluate long-context agentic instruction following in language-model agents. This benchmark assesses how well agents adhere to a system prompt, policy file, or skills document during their operations.

Key Takeaways

  • The HANDBOOK.md benchmark consists of 65 agentic tasks designed to evaluate how language-model agents follow long, binding policy documents in professional environments.
  • The benchmark spans five domains: finance, medical billing, insurance, logistics, and HR, with tasks modeled on the behavior of enterprise employees following company handbooks.
  • The best-performing model configuration achieved a pass rate of 36.2% under strict grading criteria, while most configurations remained below 25%.
  • Common failure patterns include agents allowing plausible in-environment requests to override standing policies and losing rule details over long task durations.
Read original article

Community Sentiment

Mixed

Positives

  • Local inference could magically fix LLM defects, giving users control and potentially transforming how we manage complex models.
  • Models like Kimi K3 are pushing the boundaries of frontier performance, suggesting exciting advancements in LLM capabilities.
  • Using a limited context window (like 50%) seems to lead to better outcomes, showing that managing input can enhance model effectiveness.

Concerns

  • Claiming 1M tokens in context is misleading; users face severe limitations due to quantization and poor sampling methods.
  • Local models still exhibit the same defects as cloud models, debunking the idea that local hosting is a panacea for LLM issues.
  • The inability of models to retain information over extended interactions is a significant flaw, undermining their effectiveness in following complex policies.

Related Articles

Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

Evaluating AGENTS.md: are they helpful for coding agents?

Feb 16, 2026

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

Agentic Context Management: Memory and Cost as Architecture Problems

Aug 26, 2026

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Study: Self-generated Agent Skills are useless

Feb 16, 2026

A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

Feb 10, 2026

Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

Jul 13, 2026