Almost everyone raving about Fable 5 is a developer on the Max plan, building whole apps with it. The rest of us, on free or Pro, burn through our usage by Tuesday.
You do not have to choose Fable 5 or GPT-5.6 Sol. Pair them: let Fable 5 plan and review, let GPT-5.6 Sol do the actual building, and builders report cutting the AI token bill by about half for similar output. It is the same principle brisk. teaches for every AI setup: match the model to the job, do not pay genius rates for onion work.
I do not code, and never will. My job for 13 years was translating between developers and the people who sign off on their work. This is that translation, in plain English, before you touch a single setting.
A caveat, up front: I have not run this setup for six months myself. I found it, followed the steps, and wrote the clean version below. The "half the bill" figure is what builders running this pattern report, not a number I measured personally. Treat it as a strong lead, not a guarantee.
Why Does Using One AI for Everything Cost So Much?
Because you end up paying premium, genius-level rates for simple tasks a cheaper model could do just as well.
Picture a Michelin-star kitchen where the head chef also chops onions, peels potatoes, and washes pans. The food is fine. The bill is insane, because you are paying genius rates for onion work.
That is what happens when one expensive AI model handles everything, from hard planning down to boring copy-paste tasks. It works. It just quietly burns money on jobs that never needed the smartest model in the building.
Stop making one AI do everything: the expensive habit versus the split.
The rule: match the model to the job. Genius work gets the genius model. Onion work gets the cheap, fast one.
What Are the Two Roles in an AI Cost-Saving Setup?
The planner and the worker. Split any AI task into these two roles and the cost drops without the quality dropping.
- Fable 5 is the boss. The expensive, deep-thinking model. It plans the work and reviews the result. It does not do the grunt work itself.
- GPT-5.6 Sol is the worker. The fast, cheaper model. It takes the plan and builds the thing, then fixes it when the boss sends it back.
Developers call this the architect-worker pattern: the smart model draws the blueprint, the fast one lays the bricks. You do not need the jargon, just the split.
You are not choosing between the expensive AI and the cheap one. You are hiring both, for the jobs each is actually good at.
Match the model to the job: the boss plans and reviews, the worker builds and fixes.
What Is a Token, and Why Does Splitting the Work Save Money?
A token is a small chunk of text, roughly a word. Every time an AI reads your request and writes an answer, it spends tokens, and smart models charge more coins per word than fast ones.
If one expensive model does everything, you pay premium coins on every word, including the boring ones. Split the work, and the expensive model only spends coins on the few things that need real thought, while the cheap model handles the bulk of ordinary work at a lower rate.
The saving hides in frequency: you call the expensive brain a handful of times per task, and the cheap hands dozens of times. Builders running this pattern report cutting the total bill by around half for similar output. That is their number, not one I have measured myself.
Where the money actually goes: a few expensive calls versus many cheap ones.
Where the money goes: the expensive model should touch your work a few times, not a few hundred. If it is doing the boring bulk, you are overpaying.
How Does the /route Command Actually Work?
You type /route, describe the job, and the two-model setup runs itself: Fable 5 plans, GPT-5.6 Sol builds, Fable 5 reviews, and the loop repeats until the boss approves.
Here is one run, in plain English:
- You type /route and describe the job, for example "add a check to the signup form so people can't leave the email blank."
- The boss plans. Fable 5 restates the job in one line and writes a short, step-by-step plan, saved before anything gets built.
- The worker builds. GPT-5.6 Sol takes the plan and does the actual work.
- The boss reviews. Fable 5 reads back what the worker built, checking for what is wrong, missing, or sloppy. It does not rubber-stamp the first try.
- The loop repeats. Anything wrong goes back to the worker to fix. Build, check, fix, again, until the boss is happy.
The /route loop: Anthropic's best runs OpenAI's best, on repeat until it's approved.
/route only runs when you call it. Ordinary questions and quick edits never trigger the expensive boss, so it does not burn your premium limits on a normal Tuesday.
A cheap worker makes mistakes. That is fine. The setup expects it, which is why the expensive boss checks every plate before it leaves the kitchen.
Why Set This Up Now, Before Fable 5 Goes Metered?
Because Fable 5 is only free inside the Claude subscription for a limited window, and it already eats your allowance at roughly double the rate of the older top model.
When the preview window closes, Fable 5 goes API-only and metered: about $10 per million tokens in, $50 per million out. The model you were happily throwing every task at becomes the model you count every call of.
That gap is the whole reason the boss-and-worker split stops being a nice idea and becomes the setup. Let Fable 5 plan and review a handful of times. Let the cheap worker do the hundred small things. You get the smart model's judgment without paying smart-model rates on every keystroke.
Do this now: wire it up while Fable 5 is still included in your plan, so the day it flips to metered, your bill barely moves.
Where Does This Setup Break, and Who Should Skip It?
It is a developer setup that runs in a terminal, so if opening a terminal already makes you tense, this is not your Tuesday-morning tool yet.
A few more limits worth naming:
- The model names will change. GPT-5.6 Sol is the worker model as things stand today. Model names get retired and renamed, so do not be surprised when the label shifts.
- The cheap worker is not magic. It gets things wrong, which is exactly why the boss reviews every result. Skip the review to save a few more coins and you are back to hoping, and hoping is expensive.
- It is manual, on purpose. /route only fires when you call it. That is a feature, not a bug, but it means this is not a system that runs your whole job while you sleep.
If you do not work in coding tools and have no wish to, take the principle, not the install: match the model to the job. That saves money in any AI tool you already use, including the tool-agnostic setups covered in Locked in.
Where Do You Get the Exact Setup?
The exact commands live in a short, free guide, on purpose. Exact setup steps belong somewhere you can copy-paste without a typo, not buried in an article.
Inside the guide:
- The four commands that install the piece connecting the two models
- The one config line that makes the cheap worker the default
- The complete /route skill file, ready to paste in
- A demo prompt to run the loop end to end and watch it work
It takes under five minutes if you already have the tools installed, and the guide starts from zero if you do not. Subscribe to the brisk. newsletter and it lands in your inbox.
The free brisk. setup guide: wiring Fable 5 and GPT-5.6 Sol into one /route command.
Sources and further reading
- Locked in. - the brisk. guide to AI setups that do not lock you into one tool
- Anthropic and OpenAI pricing pages for current Fable 5 and GPT-5.6 Sol rates (reconfirm before you rely on it, pricing moves fast)
You would never pay a Michelin chef to peel potatoes. You would have her design the menu and taste the plates, and let the cheap, fast hands do the rest. Your AI bill works the same way. Subscribe to brisk., grab the setup guide, and go run your own kitchen.