1 story found
We spent two weeks running Claude 4.5 Sonnet and GPT-5 through SWE-bench Verified, real refactors, and long-context debugging. The winner surprised us.