
arxiv.org
May 2, 2026
2 min read
52/100
Summary
Conversational large language models are fine-tuned for instruction-following and safety, allowing them to comply with benign requests while refusing harmful ones. Research indicates that the refusal behavior in these models is mediated by a single directional mechanism.
Key Takeaways
Community Sentiment
Positives
Concerns

Language Model Contains Personality Subnetworks
Mar 2, 2026

Why Large Language Models Fail at Tabular Prediction
Aug 4, 2026
Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
Feb 5, 2026

VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO
Jun 23, 2026

Chain-of-Thought Reasoning in the Wild Is Not Always Faithful (2025)
Aug 19, 2026