AI Learning Series · Lesson 1 of 4
Most AI results are decided before the first prompt is typed. The model you choose sets the ceiling on quality, the floor on cost, and how much context the task can carry — so choose it on purpose, not by habit.
Status: outline. The on-screen walkthrough hasn't been recorded yet — this post is the plan it will follow, and it gets the real walkthrough once the lesson airs.
| Lesson | Focus | Question it answers |
|---|---|---|
| 1 — Model Usage | Model tier, context, tokens | Which model does this task need? |
| 2 — Prompting, Project, and Context | Prompt structure, standing context | What goes in front of the model? |
| 3 — Documentation, Rules, and Skills | Repository structure | How does a repo stay consistent across sessions? |
| 4 — Cutting Cost | Measurement, levers, guardrails | How do you spend less without losing quality? |
What You'll Be Able to Do
- Look at a task and name the model tier it needs — with a reason
- Estimate how much context to send before sending it
- Estimate roughly what the run will cost before it runs
Not "which model is best" — there isn't one. Which model fits this task.
Out of scope here: prompt wording and project setup (lesson 2), rules files and skills (lesson 3), and cutting spend (lesson 4). This lesson teaches the vocabulary of cost — tokens, context, input vs. output — so lesson 4 has something concrete to cut.
The Model Decision
Model Tier — Capability vs. Speed vs. Price
- Capability, speed, and price move together: stronger reasoning is usually slower and more expensive per token
- The same task runs on a small model and a large one, side by side, so the tradeoff is visible instead of asserted
Task Fit — Match the Job, Not the Headline
- Lookup, classification, drafting, multi-step reasoning, and long-running agent work don't need the same model
- Three real tasks from this blog's own repository each get a model and a stated reason — never "because it's the newest"
| Task type | Typical fit | Why |
|---|---|---|
| Lookup / classification | Small, fast tier | Short input, one right answer, easy to verify |
| Drafting / rewriting | Mid tier | Needs fluency more than deep reasoning |
| Multi-step reasoning / debugging | Large tier | Errors compound — capability pays for itself |
| Long-running agent work | Large tier, with limits | Many steps, many tokens — needs a budget and a stop condition |
Context Window — What Fits, What Falls Off
- More context isn't automatically better — past a point, you pay for tokens the model barely uses
- Past the window's limit, older content drops off silently
- On screen: a long file pasted whole, then only the relevant section, with output and cost compared
Tokens — Input vs. Output
- Input and output tokens are counted separately, and often priced separately
- A real prompt from this project and its response both get counted, so "what will this cost" stops being a guess
Effort and Reasoning Settings
- Some models expose a setting for how much they reason before answering
- One genuinely hard task runs at a low setting and a high one — compared on results, not just latency
The Decision Checklist
▶ show code▼ hide code
# Model Decision — before you send
1. What kind of task is this? (lookup / draft / reasoning / agent)
2. What does "correct" look like, and how will you check it?
3. What is the smallest tier that can plausibly pass that check?
4. What context does it actually need? (sections, not whole files)
5. Rough token estimate: input + expected output
6. Is there an effort/reasoning setting — and does this task need it?
7. Check failed? Change one thing (tier, context, or setting) and rerun
Verify Before You Rely On It
Model names, context-window sizes, and prices change too often to publish as fact, so this lesson names none of them. Before the lesson records, each one is re-checked against the provider's own documentation:
- Current model names and the tier each one sits in
- Context window size per model
- Input and output price per model — shown only if checked that same week
- Whether an effort or reasoning setting exists on the demonstrated models, and what it's called
Need the specifics now? Check Anthropic's docs, OpenAI's docs, and Hugging Face's model cards — the same sources the lesson checks before going live.
Common Questions
"Which model is best?" None, in the abstract. The right model is the smallest one that passes your check for this task.
"Should I send the whole file?" Usually not. Send the section the task depends on — whole files cost more and can bury the part that matters.
"Where do I find current prices and limits?" The provider's own documentation, checked the week you need it — not a blog post, including this one.
Next: Lesson 2
Choosing the model is half the setup. Lesson 2 covers what goes in front of it — prompt structure, acceptance criteria, and standing project context. For the advanced version of choosing where AI work runs, the hybrid AI setup guide routes work between Hugging Face, Claude, and a locally run Ollama model.