Compare
A · run_51736avsB · run_f4a155code-reviewSwap A and B- Duration
- 1m 1s +42%
- A: 43.0s
- Cost
- $0.093 +69%
- A: $0.055
- Tokens
- 17.0k -9%
- A: 18.6k
- Steps
- 11 +120%
- A: 5
- Model changed
- Claude Sonnet
- A: GPT-4.1
Steps, aligned
| Change | Step | A | B |
|---|---|---|---|
| Faster | Input policy check | 1.1s | 0.6s |
| Same | Plan the task | 13.5s | 10.9s |
| Faster | search_codebase | 9.5s | 1.8s |
| Faster | Retrieve context · top 8 | 3.6s | 2.1s |
| Changed | Decide next step | 12.7s | 9.8s |
| Added | run_tests | — | 4.1s |
| Added | search_codebase | — | 4.9s |
| Added | Decide next step | — | 7.7s |
| Added | get_diff | — | 4.3s |
| Added | run_tests | — | 3.1s |
| Added | Write final answer | — | 8.3s |
Output diff
− 0 words + 0 words 100% unchangedReviewed 14 files. Three issues: the cache key ignores the locale, a test is skipped without a reason, and a server action is missing input validation. Tests pass. Suggested changes posted as comments, one marked blocking.