Stop Using Skills: Chart Assets¶
These figures were built from the VLIW without-indices cohorts linked in the
draft of Stop Using Skills. They use Export
Studio's aggregate Groups view, following the compact condition-level chart
pattern in Dan Luu's token-use analysis.
Figures¶
Claude Opus 5¶

The figure excludes nine no-skill runs whose best result in the four-hour window remained at or above 1,900 cycles; all nine finished between 2,091 and 2,094 cycles. The resulting comparison selects 26 runs: 24 have both a scored best candidate and working-token observation inside the window, and two are reported as missing rather than silently dropped. The exact exclusion rule and run names are recorded in the data file.
GPT-5.6 Sol¶

All 24 selected Codex runs contribute. High and max each contain six skill and six no-skill observations.
Method¶
- Source snapshot: the production
vliw / without-indicesreport on 2026-08-10. - Window: first four active hours of each run.
- Observation: each run's best scored candidate with a working-token value.
- Opus exclusion: runs whose four-hour best remained at or above 1,900 cycles are omitted as stalled instruction-failure outcomes; this removes nine no-skill runs.
- Conditions: experimental skill usage, split by effort level.
- Summary: arithmetic mean for both score and tokens; bars show full min-max ranges; labels show the contributing sample size.
- Direction: fewer cycles is better, so the Y axis is inverted.
- Token scope: working tokens exclude cache reads. ScoreBench cost estimates do include provider-reported cache reads when pricing data is available.
- Skill scope: required and coordination skills such as
scorebench,progress-logging, and Arbor agent helpers do not count as an experimental optimization or knowledge skill.
The PNG files are 4800x2700. The SVG files are resolution-independent and are the preferred source for publication layouts.