Back to blog

Your Chatbot Doesn't Have to Run on One AI

September 22, 20265 min
Open Technology App

New to this? You're in the right place. This post assumes you've chatted with something like ChatGPT before, but never set up an AI-powered chatbot for a business tool. If a word needs explaining, it's explained right where it shows up.

The problem: one AI, every question

Imagine a help desk where every single question — "what's the WiFi password" and "why did our biggest client's account just get suspended" — goes to the same senior specialist. That specialist can handle both, but paying their hourly rate to answer the WiFi question is wasteful. A smarter setup routes easy questions to a junior staffer and saves the specialist for the hard calls.

That's exactly the problem with pointing a chatbot at one AI model for everything. The best AI models (the ones that are best at hard reasoning) also tend to be the most expensive per question. If every simple lookup — "what's the status of this task?" — goes through the expensive model, you're paying premium prices for questions that don't need it.

What "routing" means

Routing is automatically sending a question to a different AI model depending on how hard the question looks. OpenTechnologyApp, the project-management tool this blog also covers, builds this in: an administrator picks up to three "vendors" (the companies providing an AI model — Anthropic, OpenAI, and Hugging Face are the three currently supported) and assigns each one a role:

  • Primary — the everyday choice for most questions.
  • Simple — a cheaper model reserved for easy, low-effort questions.
  • Fallback — used automatically if the primary model's service has a hiccup.

A behind-the-scenes system decides which role a question falls into, and sends it there — you don't have to pick manually each time.

How the system decides "easy" vs. "hard"

It looks for patterns in what you typed. Words that suggest you want something changed ("update," "delete"), something spanning multiple projects, code, a reasoning question ("why," "compare"), or several steps at once — all of those push a question toward "harder." A plain lookup ("what is," "show me") on a single project pushes it toward "easier." The system adds up these signals into a score, and two adjustable numbers (a "low" and a "high" line) decide which of the three roles above actually handles it.

You don't need to know the exact math to use this — an administrator can leave the default settings alone and it works reasonably well out of the box. If you want the actual formula and real worked examples, the deeper guide Choosing Complexity Thresholds covers it.

Picking vendors: what's realistic today

Four vendors show up in OpenTechnologyApp's routing settings:

  • Hugging Face — generally the cheapest hosted option, and a reasonable pick for the "simple" role.
  • OpenAI and Anthropic — the higher-quality, higher-cost options, better suited to the "primary" and "fallback" roles.
  • Ollama — a way to run an AI model entirely on your own computer, for free, with nothing sent to an outside company. It's a real, working option in the routing settings — with one important catch: it only works when the app itself can reach your computer. That's true for local development and for a self-hosted deployment on your own server, but not for the hosted, Vercel-deployed version of the app — a cloud server has no way to reach localhost on your laptop. The admin screen shows an amber warning to this effect whenever Ollama is picked for any role, so you won't find out the hard way. (This blog's own hybrid-AI post covers Ollama for a related but different use case — routing a coding assistant's work, not the in-app chatbot.) As of this writing, this support lives on an OpenTechnologyApp pull request that's open for review but not yet merged — real, working code, just not deployed to every organization's app instance yet.

A realistic example

Say a small team's OpenTechnologyApp chatbot fields about 100 questions a day. Most are quick status checks — "what's overdue," "who's assigned to this." A handful each day are genuinely complex — "reorganize these 40 items across three projects and explain your reasoning." Routing everything to a premium model charges premium rates for the 90-odd simple questions that didn't need it. Splitting the traffic — cheap model for the quick checks, premium model for the genuinely hard handful — cuts the AI bill substantially while keeping the hard questions answered by the model best equipped for them.

Try it

If you're an administrator on an OpenTechnologyApp organization, look for "Org AI — Routing Config" in your admin settings. The setup guide walks the actual screen field by field, with a minimal working configuration to start from if you'd rather not tune anything on day one.

Get in Touch

Interested in a topic? Drop a note and select a category. I'm also available for a free consultation meeting — reach out and we'll set something up.