Somebody has to decide which model handles a given call, and the fastest way to ship that decision is a hardcoded string in a config file. It's the right call at the time: there's no reason to build routing infrastructure before you know you need it. But that string doesn't update itself when a new model ships, when a price changes, or when a provider updates something behind an unchanged name. Someone has to notice each of those changes, re-test, and edit the config, one service at a time. That's recurring work, and it scales with how many agents you run.
Neither a hardcoded model nor a static router checks itself
A classifier plus a table is a reasonable first upgrade: sort the task into a cheap/general/hard tier, map the tier to a model, send it. It's more granular than one model for everything, but it has the same underlying issue. The assignment is fixed when it's written and stays fixed until someone edits it. Nothing in a static table checks whether the cheap tier is actually clearing its tasks, or logs what a call cost after a retry, so a model that keeps failing a category of task keeps getting routed that same category anyway.
What changes with Minima
Minima replaces the table with a lookup that has outcomes attached to it:
Task in → Minima: recommend → Model call (your own client) → Result
↑ |
└───────────── feedback: outcome + cost ───────────┘POST /v1/recommend takes a task description, checks past outcomes on similar tasks, and returns the cheapest model predicted to clear a quality bar you set with a single value (cost_quality_tradeoff, 0–10). You call that model the same way you already do, with the same provider and the same client. The lookup sits beside the call rather than in front of it, so no proxy gets added to the request and no new point of failure sits on the part that actually generates output. The only addition is a recommendation lookup (~100–300ms) alongside it.
POST /v1/feedback reports what happened, and that record is what the next recommendation for a similar task gets checked against. Without it, the picks stay at a general estimate instead of getting specific to your traffic.
The tiebreaker only fires when evidence is thin
When recalled history is sparse or the top two candidates are close, Minima asks a cheap reasoning model to help rank them and combines that with the deterministic score. This only happens under those two conditions. If the reasoning model fails, Minima uses the deterministic result instead of returning an error.
A dependency failure degrades the pick, not the call
The part that stores outcome history and checks similarity between tasks is a separate system underneath Minima. If it's unreachable or recall times out, the recommend call still returns; it falls back to a general estimate with a warning attached instead of erroring out or hanging. And because Minima never proxies the model call itself, none of this touches your request path either way.
How this compares to building it yourself
A table you build yourself is accurate on the day you write it and stays fixed until someone edits it again. Minima's accuracy comes from what you report, so it improves the more it's used instead of staying flat between manual updates. The cost structure differs the same way: the table costs the initial build plus every future edit, while Minima costs an integration pass up front: a lookup before the call, a report after it, with the call in the middle unchanged. After that you're not the one maintaining it. That holds whether you're wrapping one service or rolling it out gradually across a fleet.
What's actually underneath it
The part that makes any of this work (storing what happened, checking which past tasks look like this one, deciding how much to trust a given past result) isn't something Minima builds itself. That's Mubit, the memory system Minima is built on. Minima is one specific, narrow way of using it, for one decision: which model should run this call. None of this removes the need to think about model selection once, but it removes the need to personally re-derive the right answer every time the landscape shifts, because the remembering underneath is Mubit's job, not yours.