Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#claude#ai-ethics#code-generation#ai-safety#openai#anthropic#discussion

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

© 2026 Themata.AI • All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
ai-agentshuman-in-the-loopcode-generationai-safety

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

Humans missed 1 in 3 threats approving AI agent commands across 40,000 plays

scalex.dev

August 6, 2026

6 min read

🔥🔥🔥🔥🔥

58/100

Summary

Humans missed 1 in 3 threats while approving AI agent commands in a browser game involving 40,000 plays. The game simulated a human-in-the-loop scenario where players approved or denied commands under time pressure, revealing vulnerabilities in command approval processes.

Key Takeaways

  • Players missed 1 in 3 threats when approving AI agent commands, achieving an average accuracy of 66.3%.
  • 32.9% of game sessions resulted in a negative score due to penalties from approved threats and blocked safe commands.
  • The command "npm run analyze" was the most frequently missed threat, approved 64.7% of the time by players.
  • Commands categorized as exfiltration or code execution had a miss rate of 33.4%, indicating a significant risk of undetected threats.
Read original article

Community Sentiment

Mixed

Positives

  • The stats from the AI agent game reveal fascinating insights about human decision-making under pressure — 40k runs and still 1 in 3 threats missed is eye-opening.
  • The AI agent's sandboxing approach is an innovative step, eliminating permission prompts and allowing for more autonomous operations in a controlled environment.

Concerns

  • Relying on users to constantly approve commands is a laughable security model — history shows it just doesn't work over time.
  • The inconsistency in what prompts are considered dangerous makes the entire testing process fundamentally flawed and the results meaningless.
  • The permission prompts are a terrible model that only leads to fatigue and mistakes; we need better methods for AI decision-making.

Related Articles

The looming AI clownpocalypse · honnibal.dev

The Looming AI Clownpocalypse

Mar 2, 2026

A GitHub Issue Title Compromised 4,000 Developer Machines

A GitHub Issue Title Compromised 4k Developer Machines

Mar 5, 2026

An AI Agent Published a Hit Piece on Me – More Things Have Happened

An AI Agent Published a Hit Piece on Me – More Things Have Happened

Feb 14, 2026

The VibeSec Reckoning

The VibeSec Reckoning

May 27, 2026

The setup

What happened after 2k people tried to hack my AI assistant

Jun 26, 2026