Compare
A · run_254f49vsB · run_974046code-reviewSwap A and B- Duration
- 35.5s -4%
- A: 36.8s
- Cost
- $0.040 +215%
- A: $0.013
- Tokens
- 16.5k +81%
- A: 9.2k
- Steps
- 11 +120%
- A: 5
- Model changed
- GPT-4.1
- A: Claude Haiku
Steps, aligned
| Change | Step | A | B |
|---|---|---|---|
| Faster | Input policy check | 1.0s | 0.4s |
| Faster | Plan the task | 12.2s | 4.8s |
| Same | search_codebase | 2.7s | 2.6s |
| Faster | Retrieve context · top 8 | 5.2s | 1.4s |
| Changed | Decide next step | 13.4s | 5.2s |
| Added | run_tests | — | 2.7s |
| Added | search_codebase | — | 1.9s |
| Added | Decide next step | — | 4.8s |
| Added | get_diff | — | 2.9s |
| Added | run_tests | — | 2.2s |
| Added | Write final answer | — | 4.5s |
Output diff
− 0 words + 0 words 100% unchangedReviewed 14 files. Three issues: the cache key ignores the locale, a test is skipped without a reason, and a server action is missing input validation. Tests pass. Suggested changes posted as comments, one marked blocking.