Compare
A · run_974046vsB · run_aabd22code-reviewSwap A and B- Duration
- 11.0s -69%
- A: 35.5s
- Cost
- $0.028 -32%
- A: $0.040
- Tokens
- 10.9k -34%
- A: 16.5k
- Steps
- 7 -36%
- A: 11
- Model changed
- Gemini 2.5 Pro
- A: GPT-4.1
Steps, aligned
| Change | Step | A | B |
|---|---|---|---|
| Faster | Input policy check | 0.4s | 0.1s |
| Faster | Plan the task | 4.8s | 2.4s |
| Faster | search_codebase | 2.6s | 1.2s |
| Faster | Retrieve context · top 8 | 1.4s | 0.6s |
| Faster | Decide next step | 5.2s | 2.5s |
| Faster | run_tests | 2.7s | 1.0s |
| Changed | Write final answer | 1.9s | 2.5s |
| Removed | Decide next step | 4.8s | — |
| Removed | get_diff | 2.9s | — |
| Removed | run_tests | 2.2s | — |
| Removed | Write final answer | 4.5s | — |
Output diff
− 0 words + 0 words 100% unchangedReviewed 14 files. Three issues: the cache key ignores the locale, a test is skipped without a reason, and a server action is missing input validation. Tests pass. Suggested changes posted as comments, one marked blocking.