Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#discussion#trending#claude#ai-ethics#code-generation#ai-safety#openai

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

© 2026 Themata.AI • All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
ai-agentshuman-in-the-loopcode-generationai-safety

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

Humans missed 1 in 3 threats approving AI agent commands across 40,000 plays

scalex.dev

August 6, 2026

6 min read

🔥🔥🔥🔥🔥

63/100

Summary

Humans missed 1 in 3 threats while approving AI agent commands in a browser game involving 40,000 plays. The game simulated a human-in-the-loop scenario where players approved or denied commands under time pressure, revealing vulnerabilities in command approval processes.

Key Takeaways

  • Players missed 1 in 3 threats when approving AI agent commands, achieving an average accuracy of 66.3%.
  • 32.9% of game sessions resulted in a negative score due to penalties from approved threats and blocked safe commands.
  • The command "npm run analyze" was the most frequently missed threat, approved 64.7% of the time by players.
  • Commands categorized as exfiltration or code execution had a miss rate of 33.4%, indicating a significant risk of undetected threats.
Read original article

Community Sentiment

Mixed

Positives

  • The stats from the AI agent game reveal fascinating insights about human decision-making under pressure — 40k runs and still 1 in 3 threats missed is eye-opening.
  • The AI agent's sandboxing approach is an innovative step, eliminating permission prompts and allowing for more autonomous operations in a controlled environment.

Concerns

  • Relying on users to constantly approve commands is a laughable security model — history shows it just doesn't work over time.
  • The inconsistency in what prompts are considered dangerous makes the entire testing process fundamentally flawed and the results meaningless.
  • The permission prompts are a terrible model that only leads to fatigue and mistakes; we need better methods for AI decision-making.

Related Articles

Patterns and problems in multiagent systems

Patterns and problems in emerging multi-agent systems

Aug 16, 2026

The looming AI clownpocalypse · honnibal.dev

The Looming AI Clownpocalypse

Mar 2, 2026

UK Cyber Test: AI Agent Attempted to Social Engineer Open So...

Mythos Attempted to Social Engineer Open Source Maintainer to Merge Malware

Aug 7, 2026

A GitHub Issue Title Compromised 4,000 Developer Machines

A GitHub Issue Title Compromised 4k Developer Machines

Mar 5, 2026

An AI Agent Published a Hit Piece on Me – More Things Have Happened

An AI Agent Published a Hit Piece on Me – More Things Have Happened

Feb 14, 2026