Back to blog

AI Learning Series, Lesson 1: Model Usage

September 23, 20266 min
Dev Tools

AI Learning Series · Lesson 1 of 4

Most AI results are decided before the first prompt is typed. The model you choose sets the ceiling on quality, the floor on cost, and how much context the task can carry — so choose it on purpose, not by habit.

Status: outline. The on-screen walkthrough hasn't been recorded yet — this post is the plan it will follow, and it gets the real walkthrough once the lesson airs.

LessonFocusQuestion it answers
1 — Model UsageModel tier, context, tokensWhich model does this task need?
2 — Prompting, Project, and ContextPrompt structure, standing contextWhat goes in front of the model?
3 — Documentation, Rules, and SkillsRepository structureHow does a repo stay consistent across sessions?
4 — Cutting CostMeasurement, levers, guardrailsHow do you spend less without losing quality?

What You'll Be Able to Do

  • Look at a task and name the model tier it needs — with a reason
  • Estimate how much context to send before sending it
  • Estimate roughly what the run will cost before it runs

Not "which model is best" — there isn't one. Which model fits this task.

Out of scope here: prompt wording and project setup (lesson 2), rules files and skills (lesson 3), and cutting spend (lesson 4). This lesson teaches the vocabulary of cost — tokens, context, input vs. output — so lesson 4 has something concrete to cut.


The Model Decision

Model Tier — Capability vs. Speed vs. Price

  • Capability, speed, and price move together: stronger reasoning is usually slower and more expensive per token
  • The same task runs on a small model and a large one, side by side, so the tradeoff is visible instead of asserted

Task Fit — Match the Job, Not the Headline

  • Lookup, classification, drafting, multi-step reasoning, and long-running agent work don't need the same model
  • Three real tasks from this blog's own repository each get a model and a stated reason — never "because it's the newest"
Task typeTypical fitWhy
Lookup / classificationSmall, fast tierShort input, one right answer, easy to verify
Drafting / rewritingMid tierNeeds fluency more than deep reasoning
Multi-step reasoning / debuggingLarge tierErrors compound — capability pays for itself
Long-running agent workLarge tier, with limitsMany steps, many tokens — needs a budget and a stop condition

Context Window — What Fits, What Falls Off

  • More context isn't automatically better — past a point, you pay for tokens the model barely uses
  • Past the window's limit, older content drops off silently
  • On screen: a long file pasted whole, then only the relevant section, with output and cost compared

Tokens — Input vs. Output

  • Input and output tokens are counted separately, and often priced separately
  • A real prompt from this project and its response both get counted, so "what will this cost" stops being a guess

Effort and Reasoning Settings

  • Some models expose a setting for how much they reason before answering
  • One genuinely hard task runs at a low setting and a high one — compared on results, not just latency

The Decision Checklist

▶ show code
# Model Decision — before you send
1. What kind of task is this? (lookup / draft / reasoning / agent)
2. What does "correct" look like, and how will you check it?
3. What is the smallest tier that can plausibly pass that check?
4. What context does it actually need? (sections, not whole files)
5. Rough token estimate: input + expected output
6. Is there an effort/reasoning setting — and does this task need it?
7. Check failed? Change one thing (tier, context, or setting) and rerun

Verify Before You Rely On It

Model names, context-window sizes, and prices change too often to publish as fact, so this lesson names none of them. Before the lesson records, each one is re-checked against the provider's own documentation:

  • Current model names and the tier each one sits in
  • Context window size per model
  • Input and output price per model — shown only if checked that same week
  • Whether an effort or reasoning setting exists on the demonstrated models, and what it's called

Need the specifics now? Check Anthropic's docs, OpenAI's docs, and Hugging Face's model cards — the same sources the lesson checks before going live.


Common Questions

"Which model is best?" None, in the abstract. The right model is the smallest one that passes your check for this task.

"Should I send the whole file?" Usually not. Send the section the task depends on — whole files cost more and can bury the part that matters.

"Where do I find current prices and limits?" The provider's own documentation, checked the week you need it — not a blog post, including this one.


Next: Lesson 2

Choosing the model is half the setup. Lesson 2 covers what goes in front of it — prompt structure, acceptance criteria, and standing project context. For the advanced version of choosing where AI work runs, the hybrid AI setup guide routes work between Hugging Face, Claude, and a locally run Ollama model.

Get in Touch

Interested in a topic? Drop a note and select a category. I'm also available for a free consultation meeting — reach out and we'll set something up.