Back to blog

AI Learning Series, Lesson 4: Cutting Cost With an AI Workflow

September 23, 20267 min
Dev Tools

AI Learning Series · Lesson 4 of 4

You can't cut what you haven't measured. Every AI cost claim worth making starts with a real baseline, changes one thing at a time, and counts quality every time it counts spend — a cheaper run that fails the check is a loss, not a saving.

Status: outline. The on-screen walkthrough hasn't been recorded yet, and this lesson's whole premise is a measured result — so there are no numbers here yet. None will be invented to make the post feel finished; the real baseline and final table replace this note once the lesson airs.

LessonFocusQuestion it answers
1 — Model UsageModel tier, context, tokensWhich model does this task need?
2 — Prompting, Project, and ContextPrompt structure, standing contextWhat goes in front of the model?
3 — Documentation, Rules, and SkillsRepository structureHow does a repo stay consistent across sessions?
4 — Cutting CostMeasurement, levers, guardrailsHow do you spend less without losing quality?

What You'll Be Able to Do

  • Measure what one real AI task costs
  • Find where that spend actually goes
  • Cut it with a change you can name and defend — without making the result worse

The claim this lesson is allowed to make is narrow on purpose: less for the same result, measured on screen. Not a percentage borrowed from somewhere else. Whatever the measured number turns out to be is the number that gets published, even if it's small.

Out of scope here: provider negotiation, enterprise discounts, self-hosting to avoid API spend entirely, and any price or savings percentage stated as settled fact. Every number on screen comes from real usage that week, shown with its source.


The Four Levers

Cost comes last because the levers are what lessons 1–3 teach. Skip them and there's nothing to pull.

LeverTaught in
Right-size the model for the jobLesson 1
Send less, better-chosen contextLessons 1 and 2
Get it right the first time — acceptance criteria up frontLesson 2
Load only what a task needs — layered docs, rules, skillsLesson 3

The Method

Baseline — Measure Before Touching Anything

  • One task that genuinely repeats, run as-is today
  • Tokens in and out, retry count, and pass/fail recorded before any change
▶ show code
# Baseline Log — one task, before any change
Task:           (one you actually repeat)
Model/setting:
Input tokens:
Output tokens:
Retries:
Result:         pass / fail against the stated check
Source/date:    where each number came from, and when

Breakdown — Where the Spend Actually Goes

  • The baseline gets split by step: which part of the task, which context, which retries
  • The expensive part isn't always the part that feels expensive

One Lever at a Time — Re-Measure Every Change

  • Change one thing, re-measure, and keep the change only if quality holds
  • A cheaper run that fails the check stays in the table as a loss — never quietly dropped

Workflow Reshaping — Cheap Steps on Cheap Models

  • Split the task so routine steps run on a small model and only the genuinely hard step reaches for an expensive one
  • Two worked examples of the same lever: the hybrid AI setup guide routes a coding workflow between Hugging Face and Claude, and OpenTechnologyApp's chatbot vendor routing sends a chatbot's easy questions to a cheap model and hard ones to an expensive one automatically

Guardrails — Stop Conditions and Spend Budgets

  • Anything that runs in a loop gets a stop condition and a spend budget, so one mistake can't run up a bill unattended
  • On screen: a loop deliberately overruns and gets stopped by its own limit

The Final Table

Filled in with measured values once the lesson records — shown here as the shape it will take.

ChangeCost vs. baselineQuality checkKept?
Baseline———
Lever 1———
Lever 2———
Lever 3———

The Rules for the Results

  • Quality is measured every time cost is. A cheaper run that fails the check is a loss, full stop.
  • One lever per re-run. Otherwise the table can't say which change did what.
  • No-saving levers stay in. Cutting them to make the results look cleaner would misrepresent the method.

Verify Before You Rely On It

  • Current price per model — and whether input, output, and cached input bill differently — from the provider's pricing page that same week, with the date on screen
  • Whether prompt caching or batch processing exist for the demonstrated model and tool, what they're called, and their constraints
  • How the tool reports per-request token usage, so the baseline is measured, not estimated
  • That the chosen task genuinely repeats and is safe to run live
  • Any spend-limit or usage-cap feature in the tool, for the guardrails segment

Common Questions

"How much will this save me?" Unknown until you measure your own baseline. Any percentage quoted without one is someone else's workload.

"Should I just switch everything to the cheapest model?" Only the steps that still pass their check on it. A cheap run that fails costs you the retry, the review, and the fix.

"I came here for cost first — can I skip lessons 1–3?" You can read this first, but the levers only work once the earlier setup is in place.


The Full Series

Pick the model. Write the prompt and set up the project. Structure the repository. Then spend less on all of it without losing quality. None of this replaces a provider's own documentation for what's current — it's an order of operations, worked through on a real task instead of slides. Start from lesson 1, or browse the whole series on the Learn AI page.

Get in Touch

Interested in a topic? Drop a note and select a category. I'm also available for a free consultation meeting — reach out and we'll set something up.