Why I ran this test in the first place
I sell homeware on Takealot—kitchen gadgets, storage boxes, the occasional lamp. About sixty orders a week when load-shedding isn't scaring people away from online shopping. Returns are part of the job: someone orders a set of glass containers, one arrives cracked, they want a refund or a replacement, and I need to reply fast or Takealot's seller rating takes a hit.
For two months I'd been using Kryotta's Compare arena for every customer service reply. I'd paste the query, run it through Claude Sonnet 4.5 and Gemini Flash, pick the better answer, edit it, send it. It worked, but it was slow. Three clicks per reply, ten seconds of reading both outputs, another five seconds deciding. When you're handling fifteen return emails before 9 a.m., that adds up.
Auto routing promises to skip the decision: you send one prompt, Kryotta picks the best model for that specific task, you get one answer. I wanted to know if it actually worked, and which model it would choose for the three types of query I see most often—policy questions, EFT refund instructions, and angry customers who think I've personally wronged them.
So I tracked 200 queries over two weeks. Every time Auto routing returned an answer, I noted which model it had selected, checked whether the answer was usable without edits, and compared the cost to what I would've spent running Compare. Here's what happened.
Policy questions: Gemini Flash won almost every time
Sixty-eight of the 200 queries were policy questions. "Can I return this if I've opened the box?" "How long do I have to request a refund?" "Do you cover return shipping?" Straightforward, factual, zero emotion required.
Auto routing sent sixty-three of those to Gemini Flash. The answers were accurate, polite, and short—exactly what you want when someone just needs to know the rule. Flash quoted Takealot's return policy correctly, included the seven-day window, and kept the tone helpful without over-explaining. Cost per query: about 15 cents in tokens.
The five it sent to Claude were slightly more ambiguous questions—one customer asked if they could return an item they'd bought as a gift but the recipient didn't want, and another wanted to know if a "manufacturing defect" (a wobbly table leg) counted as a return or a warranty claim. Claude handled the nuance better, gave a two-sentence explanation of the difference, and suggested the customer send a photo. Those replies took 30 cents each, but they were worth it because the customer didn't come back with a follow-up.
The lesson: if the query is a yes-or-no policy lookup, Auto routing will pick the cheaper, faster model and you won't notice a quality drop. If there's any grey area, it escalates to Claude. That's exactly what I'd do manually, but Auto routing does it in half a second.
EFT refund instructions: Claude for clarity, Gemini for speed
Forty-one queries needed EFT refund instructions. Takealot processes most refunds automatically, but if a customer paid via EFT or there's a pricing dispute, I handle it manually. The customer needs my banking details, an explanation of the timeline (two to three business days), and reassurance that the refund is coming.
Auto routing split these almost evenly: twenty-two to Claude, nineteen to Gemini Flash. I couldn't see a clear pattern at first, so I compared the prompts. The ones that went to Claude included extra context—"Customer is upset, third email, wants confirmation"—or the customer's original message was long and emotional. The ones that went to Gemini were shorter, more transactional: "Customer paid via EFT, needs refund of R487, send details."
Claude's answers were warmer and more structured. It opened with an apology for the inconvenience, gave the banking details in a clean list (account name, bank, account number, reference), explained the timeline, and ended with "Let me know once you've made the payment and I'll confirm on my side." Gemini's answers were faster and cheaper—10 cents versus 28 cents—but they felt more like a form letter. The banking details were correct, but the tone was flat.
I started editing Gemini's EFT replies to add one warm sentence at the top. That brought the quality up to where I needed it, and the total cost (10 cents for the generation plus ten seconds of my time) was still lower than running Compare. If the customer was already annoyed, though, Claude's version needed no edits and that mattered more than saving 18 cents.
Angry customers: Claude every time, and rightly so
Thirty-four queries were what I'll politely call "escalated." Caps lock, accusations, threats to leave a one-star review, the occasional swear word. One customer received a cracked vase and wrote: "This is the second time you've sent me broken junk. I'm reporting you to Takealot and telling everyone on Facebook."
Auto routing sent every single one of these to Claude Sonnet 4.5. Not Gemini, not DeepSeek, not even once. And the replies were good—calm, empathetic, solution-focused, never defensive. Claude apologised without over-apologising, acknowledged the customer's frustration, explained what I'd do to fix it (full refund plus a R50 Takealot voucher as a gesture), and gave a timeline. The tone was warm but professional, the kind of reply that makes an angry customer feel heard.
I tested this by running a few of the same prompts through Compare afterward, just to see. Gemini's replies were polite but stiff—"We apologise for the inconvenience"—and they didn't match the emotional temperature of the original message. DeepSeek was even worse: technically accurate, but it read like a chatbot. Claude understood that an angry customer doesn't want a policy lecture; they want to know you care and you'll fix it.
Cost: 40 to 60 cents per reply, depending on how much context I included. Worth every cent, because three of those customers came back later to say thanks and one left a five-star review.
The tasks where Compare still makes more sense
Auto routing isn't perfect for everything. I still use Compare for three types of work:
Product descriptions. I want to see Claude's version and Gemini's version side by side, because one might focus on features and the other on benefits, and I'll often merge the two. Auto routing picks one model and you're done—faster, but you lose the option to cherry-pick the best bits.
First-time queries I haven't seen before. A customer asked if I could ship an item to Gqeberha via Pargo locker instead of door-to-door. I'd never written that reply before, so I ran it through Compare to see how each model explained the Pargo process. Gemini was clearer. If I'd used Auto routing, I might've gotten Claude and never known Gemini's version was better for that specific task.
Anything where tone is subjective. If I'm writing a promotional WhatsApp message or a response to a supplier, I want to compare two versions and pick the one that feels right. Auto routing makes a choice for you, and sometimes I don't agree with it.
The rule I've settled on: Auto routing for customer service queries I've handled fifty times before, Compare for anything new or creative.
The cost difference over two weeks
Two hundred queries. Auto routing cost me R87 in tokens. If I'd run every query through Compare (two models per query, pick one, discard the other), I'd have spent R214. I saved R127, which doesn't sound like much until you multiply it over a year—that's R3,300, or about four months of my Kryotta subscription.
The time saving mattered more. Compare took me an average of eighteen seconds per query (paste, run, read both, pick, edit). Auto routing took six seconds (paste, run, edit). That's twelve seconds saved per query, 2,400 seconds over two weeks, or forty minutes. I spent that forty minutes listing new products instead of reading AI outputs, and one of those listings has already made back the R87 I spent on tokens.
When to trust the router and when to override
Auto routing isn't magic. It's a routing algorithm that looks at your prompt, estimates which model will do the job well for the lowest cost, and sends it there. Most of the time it's right. Sometimes it's not, and you need to know when to override.
Trust it when: the task is repetitive, the quality bar is "good enough," and you've seen the output type before. Policy questions, refund instructions, order confirmations, "Where's my package?" replies.
Override it when: the customer is very upset, the query is ambiguous, you're writing something public (a response to a Takealot review), or the cost difference doesn't matter because you need the best possible answer. In those cases, either use Compare or manually pick Claude.
You can't override Auto routing mid-task—it's one-shot, you get the answer from whichever model it chose—but you can paste the same prompt into Compare afterward if you're not happy. I did that eleven times out of 200, always for the angry-customer queries where I wanted to see if Gemini could've matched Claude's tone. It couldn't.
The checklist: Auto routing or Compare?
Here's the decision tree I now use for every Takealot query:
Use Auto routing if:
– You've answered this type of query at least ten times before.
– The customer's tone is neutral or mildly annoyed, not furious.
– You're fine with one answer and a quick edit, rather than comparing two versions.
– You're processing more than five queries in one sitting and speed matters.
Use Compare if:
– It's a new type of query and you want to see how different models handle it.
– The customer is very upset and the reply needs to be perfect.
– You're writing something public or permanent (a return policy page, a Takealot Q&A answer).
– Tone is subjective and you want options.
Pick Claude manually if:
– The customer used emotional language or you're apologising for a mistake.
– The query involves a judgment call, not just a policy lookup.
– You'd rather pay 30 cents for a reply that needs zero edits than 10 cents for one that needs two minutes of rewriting.
Pick Gemini manually if:
– It's a simple, factual question and you want the answer in under three seconds.
– You're doing high-volume work (fifty return confirmations) and cost matters more than warmth.
I don't pick DeepSeek manually for customer service. It's excellent for other tasks—I use it for inventory spreadsheets and supplier emails—but for Takealot returns, Claude and Gemini cover everything I need.
Questions people ask
Does Auto routing always pick the cheapest model?
No. It picks the model it thinks will do the job well for the lowest cost, but "well" is part of the equation. I've seen it route simple queries to Claude when the prompt included emotional language, even though Gemini would've been cheaper. It's optimising for quality and cost together, not cost alone.
Can I see which model Auto routing chose after it's answered?
Yes. Kryotta shows you the model name at the top of the output. I kept a spreadsheet for this test, but you don't need to—just glance at the label if you're curious.
What if Auto routing picks a model I don't want to use?
You can't override mid-task, but you can paste the same prompt into Compare or manually select a different model and run it again. I did that eleven times in 200 queries, always when I suspected Claude would handle an angry customer better than whichever model Auto routing had chosen.
Is Auto routing worth it for a small Takealot store?
If you're handling fewer than twenty queries a week, Compare is probably fine—you're not spending enough time on AI replies for the speed difference to matter. Above twenty queries, Auto routing saves you real time and real money, and the quality drop is smaller than I expected.
I'm keeping Auto routing switched on for returns and switching to Compare for product descriptions. That's the balance that works for a sixty-orders-a-week homeware store in Johannesburg, and I think it'll work for most Takealot sellers who aren't running a warehouse operation. If you want to try the same test with your own queries, Kryotta's Compare and Auto routing tools are both available on the same workspace—you don't need two subscriptions, just two buttons.



