Compare
A · run_cb8850vsB · run_54892acode-reviewSwap A and B- Duration
- 4.9s -89%
- A: 45.8s
- Cost
- $0.030 -15%
- A: $0.035
- Tokens
- 25.0k +211%
- A: 8.1k
- Steps
- 11 +22%
- A: 9
- Model changed
- Claude Haiku
- A: GPT-4.1
Steps, aligned
| Change | Step | A | B |
|---|---|---|---|
| Faster | Input policy check | 0.6s | 0.1s |
| Faster | Plan the task | 8.9s | 0.7s |
| Faster | search_codebase | 2.6s | 0.3s |
| Faster | Retrieve context · top 8 | 2.1s | 0.2s |
| Faster | Decide next step | 6.3s | 0.9s |
| Faster | run_tests | 2.6s | 0.2s |
| Faster | search_codebase | 2.7s | 0.2s |
| Faster | Decide next step | 6.2s | 0.7s |
| Changed | get_diff | 11.0s | 0.1s |
| Added | run_tests | — | 0.4s |
| Added | Write final answer | — | 0.7s |
Output diff
− 0 words + 0 words 100% unchangedReviewed 14 files. Three issues: the cache key ignores the locale, a test is skipped without a reason, and a server action is missing input validation. Tests pass. Suggested changes posted as comments, one marked blocking.