One prompt. Different decisions.

See what more
thinking buys you.

The same brief. The same starter. No follow-ups. Compare finished apps, not a leaderboard, and decide which model and effort earned its usage.

Explore the builds
Benchmark rules

One frozen prompt. Same starter, zero steering, complete builds, and usage attached to every result.

How it works
The arena

Choose. Then use it.

Select one captured run and use the finished app.

Change build

Inspect this run

Reference

Details when you need them.

Plan context Messages, tokens, and credits

Understand the usage estimates.

01Message

One request or task you send to Codex.

02Tokens

Pieces of code and text Codex reads or generates.

03Credits

The billing unit applied to that token usage.

Plus3k credits

≈ 15–90 messages

Pro 5x15k credits

≈ 75–450 messages

Pro 20x60k credits

≈ 300–1,800 messages

Credit estimates use the token-based rate card checked July 14, 2026. Allowances and message ranges are a measured planning snapshot, not fixed quotas. Consumption varies with the model, token mix, task complexity, codebase size, and conversation length.

Current Codex rate card ↗
How to read this Your eyes are the evaluator

Separate visible quality from engineering evidence.

Visual quality is not reduced to a synthetic score. Engineering reviews inspect source and behavior; your eyes still judge design, product decisions, and whether the result feels finished.

  • Look for hierarchy, polish, useful states, and coherent decisions.
  • Then check whether extra time and tokens created visible value.
  • Remember one run is evidence, not a universal model ranking.
Frozen brief

Weekender v1

Focused comparison

Two builds. One decision.

Build A
Build B
Engineering review

Review evidence