Tutorial Tooling
Which model, and when
Every model we pay for, what to hand it and what to keep away from it, how the tools compare, what the benchmarks say about effort, and the habits that keep quota and mistakes down.
12 min · Added 1d ago · Part of Tutorials
The 30-second version
Every model family now comes in the same three shapes: a flagship for the hardest work, a standard tier for daily building, and a fast tier for cheap, high-volume tasks. The names rotate every few months. The tiers do not.
- Default setup for this repo: ZCode on GLM-5.3, effort high. Escalate the hard 10% to a flagship, downshift the boring 50% to Flash.
- Plan with a flagship, build with a standard tier, polish with a fast one. Never do all three with one model out of laziness.
- On subscriptions you burn quota, not tokens: flagships empty your window fastest, and the fast tiers stretch the same plan several times further. The smallest model that clears the task is the right model.
- When a task matters, raise the effort setting before you raise the tier. The same model at a higher effort beats a bigger model at a lower one more often than not.
- The biggest quality saver is not a better model. It is verifying every milestone: run lint, typecheck, and the app, then commit.
Model by model: use it for this
The map is the overview: faster models on the left, ones that can carry harder tasks higher up, and bigger dots eat your plan quota faster. Pick the smallest dot that sits high enough for the job. Below it, every model we pay for with its actual assignment.
Anthropic (Claude Code, the Claude apps). Anthropic's own docs say: if unsure, start with Opus 5.5.
- Sonnet 5.5 → the daily driver. Use for new features, refactors, tests, migrations, most PRs, in that order. Start every task here and escalate only when it has failed the same thing twice.
- Opus 5.5 → the “this one matters” pick. Use for planning architecture, reviewing big diffs, bugs that survived two attempts, and long agentic runs. It holds an architecture across many files better than anything below it.
- Fable 5.1 → the ceiling. Use for the rare task where Opus at max effort already failed, and for very long agent runs where the extra depth pays for itself. Not for anything routine: it is the slowest and hungriest slot we have.
- Haiku 4.5 → chores at speed. Use for commit messages, changelogs, doc passes, quick lookups, and classification. Smallest context window we pay for (200K), so keep its inputs trimmed.
OpenAI (Codex in the ChatGPT app and CLI). Sol is the builder there; the rest are support acts.
- GPT-6 Sol → OpenAI's own “use for complex coding and agentic workflows” model, and the Codex default since September 22, 2026. Use it the way you use Sonnet: features and multi-step agent work.
- GPT-6 Luna → the chores lane in Codex. Use for scripts, bulk find-and-replace class edits, summaries, and anything high-volume; it is nearly free on credits and very fast.
- GPT-5.6 Terra → the balanced middle when Sol quota is tight. Use for everyday tasks you would rather not spend flagship rates on.
- GPT-6 Astra → OpenAI's ceiling. Use for the hardest end-to-end work and research-heavy tasks. The slowest GPT; treat it like Fable, not like Sol.
GLM (ZCode, on the coding plan). Reasoning is always on, with an effort dial from low to max, and both models read images.
- GLM-5.3 → the default in this repo. Use for full features end to end, planning at high or max effort, and most of what Sonnet and a fair amount of what Opus get used for, at standard quota.
- GLM-5.3-Flash → the volume lane. Use for polish loops, screenshot-driven UI fixes, drafts, and chores. It carries 3x the quota and Z.ai positions it near Opus 4.8 class, so it is far more capable than “fast tier” suggests.
One more GLM exists, FlashX (200 tokens per second, faster still), but it is API-only for now and not on the plan. Ignore it until that changes.
Benchmarks, with effort levels
Benchmarks tell you whether two models are in the same class; they cannot tell near-identical scores apart, because a few points is inside the noise of prompts, harness settings, and effort. Read them like height charts, and read the effort tags: each score was reported at a specific reasoning effort, and the pairs of rows (same model, two efforts) are the most useful bars on the chart.
What the chart actually says. Opus 5.5 at default effort already beats every model Anthropic compared it with on FrontierCode. On CursorBench, effort alone moves it 52.5 to 57.8, five points for free. On Z.ai's own bench, GLM-5.3 at high effort passes Opus 4.8 while spending roughly half the output tokens of its max run. And the flagship tier is not one thing: Fable 5 clearly ahead of GLM-5.3, but GLM-5.3 clearly ahead of last generation's Opus.
Treat every vendor's number as marketing until your own repo confirms it: run the same task on two tiers and compare the diffs. That five-minute experiment beats any leaderboard.
The tools, and when to use which
The tools are converging on the same shape: an agent that reads your repo, edits files, runs commands, and asks before it does anything risky. What differs is the model behind it and the workflow around it.
- ZCode is the desktop agent this repo is set up for: AGENTS.md, skills, plan mode, subagents, browser control, scheduled runs, and an idle-time queue that runs work when compute is cheap. It runs GLM-5.3 and Flash on the Z.ai coding plan, quota is points-based, and weekend calls cost half points.
- Claude Code is the terminal agent for Anthropic models. Its
/modelpicker has a pattern worth stealing even outside it:opusplanplans with Opus and then hands the build to Sonnet, which is the plan-with-a-flagship pattern built in. It also exposes the effort dial and auto-compacts long sessions. - Codex is OpenAI's agent, in the ChatGPT app and as a CLI. Since September 22, 2026 it carries GPT-6 Sol for complex coding and GPT-6 Luna for cheap high-volume work, and its cloud tasks run several jobs in parallel while you do something else.
/usageshows where your tokens go. - The plain terminal is still the right tool for anything deterministic: grep, git, builds, deploys. All three agents also run headless (
claude -p,codex exec), which is how you put a one-shot prompt in a script or a git hook.
How to hand over a task
The failure mode to design against is the one-shot mega task: one enormous prompt, forty minutes of autonomous work, one enormous diff nobody read. That is how wrong models get used for wrong tasks and how mistakes compound silently. Instead, hand work over in milestones.
The loop: write the brief, let the agent plan, approve the plan, build one milestone, verify it yourself (run the lint, the typecheck, the app), commit, then point at the next milestone. Each cycle is minutes, not an hour, and a wrong turn costs one step instead of the whole task.
- Write briefs like acceptance criteria. “Add a save button that disables while the request is in flight and shows the error inline on failure” beats “improve the form.”
- Plan before build on anything architectural. Plan mode is read-only and cheap; a wrong plan caught there costs nothing.
- Commit between milestones. The commit is your rollback. An agent that can run
git resetto a good state is safe to let fail; one buried under an hour of uncommitted work is not. - One concern per session. A session that fixed the slider, then the header, then the build error carries all three contexts into every later decision. New task, new session.
Tokens, caching, and quota
First, what “cache vs normal price” actually means, because agents make it matter. An agent turn never sends just your last message: it re-sends the whole story every time. System rules, AGENTS.md, every file it opened, every exchange so far, all of it travels again on every turn. That is why a twenty-turn session costs much more than twenty separate questions.
To make that survivable, vendors keep a temporary copy of what you have already sent, called the cache. When the next turn starts with exactly the text the cache has seen, that part is reread at a fraction of the normal price: about a tenth at Anthropic and OpenAI, about a fifth at Z.ai. Only the genuinely new part, the latest message, the newest file reads, and the answer, is billed at full price. The one catch: the cached prefix has to stay identical. Edit your rules file mid-session, or paste something different at the top of the conversation, and the cache resets: every reread is full price again.
- Keep repo rules in AGENTS.md, not in prompts. Rules that live at a stable path get cached once and reused. The same rules pasted into every prompt are paid for every time.
- Fresh session per task. Dead-end context is the most expensive context: you pay to reread it every turn, and it steers the model wrong.
- Let subagents do the searching. A search subagent reads fifty files and returns one answer; the main session never pays for the fifty.
- Right-size the tier mid-task. Planning and review on the flagship, the build on standard, cleanup on fast. Switching is one command in every tool.
- Use the cheap lanes for batch work. ZCode weekends cost half points and its idle queue runs when compute is otherwise free. Anthropic and OpenAI both take 50% off for async batch API calls.
The big mistakes
Most model mistakes reduce to using the wrong tier for the job or trusting output nobody verified. The list below is the short version of everything above, inverted.
- Flagship for everything. It is slower, it burns quota, and it overthinks one-line changes. Escalate on evidence, not on anxiety.
- Fast tier for everything. The opposite error. Small models fail big refactors quietly: plausible diffs, invented APIs, tests that test nothing.
- Low effort on a hard task. The hidden third error. If the result is bad, check the effort dial before you blame the model.
- The mega task. One prompt, one hour, one unread diff. Milestones and commits exist so this never happens.
- The rotting session. Thirty turns of dead ends teach the model your dead ends. Restart with a two-line summary of what worked.
- Trusting the diff you did not run. Code that compiles is not code that works. Lint, typecheck, run the app, then accept.
- Secrets in prompts. Keys pasted into a chat are keys in a log. Keep them in env files the agent reads on its own, and rotate anything that ever lands in a prompt.