- steps
- 17
- duration
- 1m 20s
- errors
- 0
- files
- 5
Build a unit converter
Build a unit converter as a static web app in this directory: index.html, style.css, convert.js. Categories: length, mass, temperature, data size, time. Instant two-way conversion between any pair in a category, swap button, precision selector, recent-conversions list in localStorage, keyboard-first UX, dark theme, and tests.html + tests.js exercising the conversion math including temperature edge cases. Run the tests headless with node and fix what fails.
3 sessions from the same published task, on one clock. Scrub to any moment and compare what existed, or open a pane to read its full trajectory.
- steps
- 26
- duration
- 2m 38s
- errors
- 1
- files
- 5
- steps
- 24
- duration
- 3m 36s
- errors
- 0
- files
- 5
| gpt-5.5 · low | gpt-5.5 · medium | gpt-5.5 · high | |
|---|---|---|---|
| steps | 17 | 26 | 24 |
| tool calls | 7 | 9 | 8 |
| duration | 1m 20s | 2m 38s | 3m 36s |
| tokens out | 1k | 3k | 5k |
| reasoning | ~7 | ~63 | ~129 |
| compactions | 0 | 0 | 0 |
| subagents | 0 | 0 | 0 |
| errors | 0 | 1 | 0 |
| files | 5 | 5 | 5 |
Token and context figures come from the harness's own per-turn reporting; sessions captured before telemetry landed show a dash. Reasoning marked ~ is estimated from the transcript's thinking text where the harness doesn't report it separately.
Verdict
Pairwise picks on the outputs above, from the current Arena tally.
Arena picks do not collect a reason or judging criterion, so no qualitative verdict is inferred here. The bars report only the current comparison tally.