See what more
thinking buys you.
The same brief. The same starter. No follow-ups. Compare finished apps, not a leaderboard, and decide which model and effort earned its usage.
One frozen prompt. Same starter, zero steering, complete builds, and usage attached to every result.
How it worksChoose. Then use it.
Select one captured run and use the finished app.
Change build+
Find the efficient frontier.
Higher is better. Further left uses less. The connecting line marks configurations that improve the best score available as usage increases.
The run matrix
Scores review engineering, not visual design. Select a run, then open its evidence.
A first read, not a verdict.
Measured highlights from one captured run per configuration.
Use Luna or Terra at Light or Medium for fast, lower-cost iteration. Increase effort when the preview gives you a reason.
For this prompt, it produced the most source lines at nearly the same time and estimated credits as Sol High.
Engineering scores are one frozen review snapshot. Compare the working previews and add repeated runs before treating a pattern as stable.
Details when you need them.
Plan context
Messages, tokens, and credits
Understand the usage estimates.
+
One request or task you send to Codex.
Pieces of code and text Codex reads or generates.
The billing unit applied to that token usage.
≈ 15–90 messages
≈ 75–450 messages
≈ 300–1,800 messages
Credit estimates use the token-based rate card checked July 14, 2026. Allowances and message ranges are a measured planning snapshot, not fixed quotas. Consumption varies with the model, token mix, task complexity, codebase size, and conversation length.
Current Codex rate card ↗
How to read this
Your eyes are the evaluator
Separate visible quality from engineering evidence.
+
Visual quality is not reduced to a synthetic score. Engineering reviews inspect source and behavior; your eyes still judge design, product decisions, and whether the result feels finished.
- Look for hierarchy, polish, useful states, and coherent decisions.
- Then check whether extra time and tokens created visible value.
- Remember one run is evidence, not a universal model ranking.