Compare
A · run_f4a155vsB · run_51736acode-reviewSwap A and B- Duration
- 43.0s -30%
- A: 1m 1s
- Cost
- $0.055 -41%
- A: $0.093
- Tokens
- 18.6k +10%
- A: 17.0k
- Steps
- 5 -55%
- A: 11
- Model changed
- GPT-4.1
- A: Claude Sonnet
Steps, aligned
| Change | Step | A | B |
|---|---|---|---|
| Slower | Input policy check | 0.6s | 1.1s |
| Same | Plan the task | 10.9s | 13.5s |
| Slower | search_codebase | 1.8s | 9.5s |
| Slower | Retrieve context · top 8 | 2.1s | 3.6s |
| Changed | Write final answer | 9.8s | 12.7s |
| Removed | run_tests | 4.1s | — |
| Removed | search_codebase | 4.9s | — |
| Removed | Decide next step | 7.7s | — |
| Removed | get_diff | 4.3s | — |
| Removed | run_tests | 3.1s | — |
| Removed | Write final answer | 8.3s | — |
Output diff
− 0 words + 0 words 100% unchangedReviewed 14 files. Three issues: the cache key ignores the locale, a test is skipped without a reason, and a server action is missing input validation. Tests pass. Suggested changes posted as comments, one marked blocking.