
coderabbit.ai
September 5, 2026
7 min read
44/100
Summary
Some of the hardest work in code review happens outside the changed lines. A change can look correct in isolation and still break code elsewhere in the system. That is what makes our early results for OpenAI's GPT-6 Astra most interesting. In our evaluation, Astra caught approximately 4% more labeled bugs through actionable findings than GPT-5.6 Sol, and 22% more than Opus 5. The biggest jump come...