Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#claude#ai-ethics#code-generation#ai-safety#openai#anthropic#discussion

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

© 2026 Themata.AI • All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
llmsai-safetyopenaideveloper-tools

Don't ask an LLM for a confidence score

Why I Hate LLM Confidence Scores

justinflick.com

July 28, 2026

16 min read

🔥🔥🔥🔥🔥

48/100

Summary

LLM confidence scores often misrepresent the certainty of generated outputs. Relying on these scores can lead to misunderstandings about the reliability of AI responses.

Key Takeaways

  • Asking a large language model (LLM) to generate a confidence score for its responses is deemed completely useless and lacks scientific validity.
  • Current research indicates that LLMs cannot reliably quantify their own confidence due to their highly unreliable and context-dependent internal states.
  • Reflective reasoning patterns may help LLMs catch some errors, but they do not provide sufficient understanding for models to accurately quantify their confidence.
  • The concept of confidence scores in LLMs is criticized as a psychological safety trick that does not enhance the trustworthiness of the output.
Read original article

Community Sentiment

Mixed

Positives

  • With the right prompting strategy, LLMs can verbalize probabilities that are better calibrated than their own estimates, which opens up new avenues for better decision-making.
  • For many applications, like finance, high or low confidence scores have proven to be accurate about 90% of the time, making them practically useful despite some edge cases.

Concerns

  • The claim that asking LLMs for confidence scores is completely useless is overly strong, especially when research shows they can be useful under certain conditions.
  • Models often lack a good sense of their own capabilities, leading to unreliable self-assessments when asked about the meaningfulness of their outputs.

Related Articles

Arguing With Agents

Arguing with Agents

Apr 16, 2026

Agentic test processes, LLM benchmarks, and other notes on agentic coding from Galapagos Island

Agentic coding notes from Galapagos Island

Jul 4, 2026

The Future of Everything is Lies, I Guess

The Future of Everything Is Lies, I Guess

Apr 8, 2026

Experts Have World Models. LLMs Have Word Models.

Experts Have World Models. LLMs Have Word Models

Feb 8, 2026

Do LLMs pass the mirror test?

Do LLMs pass the mirror test?

Jun 28, 2026