I used AI to reply to 200 MTN MoMo payment queries in one week: which model understood cedis

I tested Claude, Gemini and DeepSeek on 200 real payment questions from my Kumasi phone shop during a sale week. Here's which model actually understood cedis and didn't invent refund policies.

KKryotta TeamProduct & research · · 10 min read
Phone accessories shop display with cases, protectors and cables on shelves, smartphone showing messages on counter
Phone accessories shop display with cases, protectors and cables on shelves, smartphone showing messages on counter

I tested this during a sale week that nearly broke me

I run a phone accessories shop in Kumasi. Cases, screen protectors, charging cables, power banks—the sort of stock that moves fast when students head back to campus or someone drops their phone on the tro-tro floor. I sell through WhatsApp, Instagram and a small Jumia storefront. Nearly everyone pays with MTN MoMo. During a normal week I might get 30 customer messages. Half are "Do you have iPhone 14 cases?", the other half are payment questions: "I sent GH₵45 but the reference number isn't working", "Can you refund GH₵80 to this number?", "I paid yesterday, where's my order?"

Last month I ran a back-to-school sale. Twenty percent off all cases, free screen protector with every phone purchase over GH₵200. I posted it on Instagram Friday evening, and by Saturday morning I had 200 messages waiting. A third of them were payment queries. Someone sent GH₵120 to the wrong number. Someone else paid GH₵65 instead of GH₵85 and wanted to know if they should send the balance or start over. Another person's transaction failed but MoMo debited them anyway, and they were (understandably) furious.

I couldn't keep up. I'd answer five messages, ten more would arrive. By Sunday afternoon I was copying and pasting the same explanation about failed transactions, making mistakes with amounts, and getting snappy with people who didn't deserve it. I needed help, but hiring someone for one week made no sense.

I'd been testing AI models on Kryotta for product descriptions, and I knew it had Auto routing—send a task to a fast, cheap model first, then escalate to something stronger if the answer isn't good enough. I decided to try it for MoMo queries. I used Claude Sonnet 4.5, Gemini Flash and DeepSeek V3, let Auto routing pick which one handled each message, and tracked which model actually understood cedis, kept replies natural in Ghanaian English, and didn't invent refund policies I don't have.

Over seven days I processed just over 200 payment questions this way. Some models were brilliant. Others confidently told customers things that were completely wrong.

The models I tested and what I was looking for

I stuck to three: Claude Sonnet 4.5 because everyone said it was the most reliable, Gemini Flash because it's fast and cheap (which matters when you're answering 200 messages), and DeepSeek V3 because I'd heard it was good with numbers and I needed something that wouldn't confuse GH₵45 with GH₵54.

I wasn't looking for perfection. I wanted a model that could:

  1. Read a MoMo transaction reference and match it to the right order. If someone says "I paid GH₵65, reference MTN-24-03-XYZ", the model needs to find that payment in my records (I keep a Google Sheet) and confirm whether it matches the order amount.

  2. Handle cedis accurately. Don't round GH₵47.50 to GH₵48. Don't write "47.5 cedis" when every Ghanaian writes "GH₵47.50". Don't add amounts wrong.

  3. Sound like a real person on WhatsApp. Ghanaian customers don't say "I hope this message finds you well." They say "Please boss, I sent the money but no response." The reply needs to match that tone—friendly, direct, no corporate fluff.

  4. Know when to escalate. If someone's asking about a refund for a transaction from three weeks ago that I have no record of, the model should flag it for me instead of inventing an answer.

I set up Auto routing to try Gemini Flash first (it's the cheapest), then escalate to DeepSeek if the confidence score was below 80%, then escalate to Claude if DeepSeek also struggled. I wrote a simple prompt with my refund policy, a link to my Google Sheet of transactions, and examples of good replies.

Claude understood context but overexplained everything

Claude was the best at reading between the lines. If someone wrote "I paid yesterday evening, still no confirmation", Claude would check the transaction time, see that "yesterday evening" was 7 p.m., cross-reference my usual confirmation time (I send confirmations by 9 p.m. the same day), and reply with something like: "I can see your GH₵85 payment came through at 7:14 p.m. yesterday—reference MTN-24-11-4721. I'll send your tracking number by 6 p.m. today. Sorry for the delay, the sale had us backed up."

That's a good reply. The problem was Claude did this for every message, even the simple ones. Someone would ask "Did you get my GH₵30?" and Claude would write four sentences explaining the payment process, confirming the amount, apologizing for any confusion, and offering to help with anything else. On WhatsApp, that reads like a bot. People just want "Yes, got it, sending your order now."

Claude also handled cedis perfectly—never rounded, always formatted amounts as GH₵47.50, never wrote "47.5 Ghana cedis" or other awkward phrasings. It caught one case where a customer said they'd paid GH₵120 but the transaction in my sheet showed GH₵102. Claude flagged it, I checked, and it turned out the customer had paid in two parts (GH₵50, then GH₵52) and forgotten about the split. That kind of attention saved me from a messy argument.

But Claude was slow. Each reply took 8–12 seconds, and when 40 messages arrived in an hour, that added up. I used it for complicated queries—disputed amounts, refund requests, anything involving multiple transactions—but not for the straightforward "Did my payment go through?" questions.

Gemini was fast and sounded human, but couldn't count

Gemini Flash was fast. Most replies came back in under three seconds. The tone was perfect for WhatsApp—casual, direct, no fluff. When someone wrote "Boss, I send the money oo, why you no reply?", Gemini replied with "I see your GH₵50 payment, sorry for the delay! Packing your order now, I'll send the tracking number this evening." That's exactly how I'd reply myself.

The problem was numbers. Gemini added GH₵40 and GH₵25 and got GH₵75 (it's GH₵65). It once told a customer their balance was GH₵15 when they'd overpaid by GH₵5. It mixed up two transactions with similar amounts—someone paid GH₵80 for a phone case, someone else paid GH₵85 for a power bank, and Gemini confirmed the wrong order for both of them.

I caught most of these because I was double-checking every reply for the first two days, but one slipped through. A customer asked if I'd received their GH₵120 payment. Gemini said yes and confirmed the order. The actual payment was GH₵102. The customer screenshotted Gemini's reply and sent it back to me, confused, and I had to apologize and explain the mistake. Not a disaster, but embarrassing.

Gemini was brilliant for simple confirmations where I just needed to say "Yes, got your payment" or "Your order ships tomorrow", but I stopped using it for anything involving calculations or multiple transactions.

DeepSeek was accurate with amounts but sounded stiff

DeepSeek got the numbers right. Every time. It handled cedis formatting perfectly, never confused GH₵47.50 with GH₵45.70, and when I asked it to calculate a refund after a partial order (customer paid GH₵150, I only had GH₵110 worth of stock, so I owed them GH₵40 back), it got it right on the first try.

The replies, though, felt like they'd been written by someone who learned English from a textbook. "Your payment of GH₵65 has been received and confirmed. Your order will be dispatched within 24 hours. Thank you for your patronage." No one says "patronage" on WhatsApp in Ghana. It's not wrong, it's just… stiff.

I edited a lot of DeepSeek's replies before sending them. I'd take its answer, confirm the numbers were right, then rewrite the tone to sound more like me. That added time, but it was still faster than doing the whole thing myself, and I trusted DeepSeek's math in a way I didn't trust Gemini's.

DeepSeek also did well with edge cases. One customer had paid GH₵200 across four separate transactions (GH₵50, GH₵80, GH₵50, GH₵20) over two days, and they wanted to know if I'd received the full amount. DeepSeek listed all four transactions, confirmed the total, and matched it to the order. Claude would've done the same, but DeepSeek was faster.

Auto routing saved me when I didn't know which model to use

Here's the thing: I didn't choose which model handled each message. Kryotta's Auto routing did. I set it to try Gemini first because it's cheap and fast, escalate to DeepSeek if the query involved numbers or multiple transactions, and escalate to Claude if the customer seemed frustrated or the situation was ambiguous.

About 60% of messages were handled by Gemini and never escalated. These were simple: "Did you get my payment?", "When will my order arrive?", "Can I pay in two parts?" Gemini answered them in under three seconds, sounded natural, and I barely had to edit.

Another 25% escalated to DeepSeep. These were the ones with calculations, refunds, or multiple MoMo references in one message. DeepSeek got the numbers right, I fixed the tone, done.

The remaining 15% went to Claude. These were the messy ones—disputed amounts, customers who'd paid weeks ago and only just followed up, people who were angry and needed a careful reply. Claude handled them well, even if it took longer.

The workflow I settled on: let Auto routing pick the model, read the reply, check any amounts against my Google Sheet, edit the tone if it sounds robotic, send. For 200 messages that took me about 12 hours total across the week. If I'd done it all myself, I'd still be replying now.

The step-by-step workflow you can copy

  1. Keep a simple payment log. I use a Google Sheet: date, customer name, MoMo reference, amount, order details, status (paid / confirmed / shipped). Update it every evening. The AI needs this to check transactions.

  2. Write a short prompt with your refund policy and tone. Mine was three paragraphs: how I handle refunds (within 48 hours to the same MoMo number), examples of good replies in Ghanaian English, and a note that if the customer sounds angry or the transaction is unclear, flag it for me.

  3. Set Auto routing to try Gemini first, escalate on numbers. Gemini handles 60% of queries fast and cheap. If the message mentions multiple amounts or asks for a calculation, escalate to DeepSeek. If it's ambiguous or emotional, escalate to Claude.

  4. Check every amount. Don't trust the AI with cedis until you've tested it for a few days. Cross-reference every figure against your sheet. Gemini will get some wrong.

  5. Edit the tone before you send. If the reply sounds like a corporate email, rewrite it. Your customers message you on WhatsApp because you're a real person, not a call center.

  6. Track which model answered what. After a week, look at which queries escalated and why. If Claude is handling 40% of your messages, your prompt might be too vague—tighten it so Gemini and DeepSeek can handle more.

I'm still using this system. The sale ended, but I still get 50–70 messages a week, and about half are payment questions. Auto routing handles most of them now, I check the numbers, fix the tone if needed, and I'm done. It's not perfect, but it's a lot better than drowning in WhatsApp notifications at 11 p.m. on a Sunday.

Questions people ask

Which model should I start with if I've never used AI for customer support?
Try Gemini Flash for simple confirmations—it's fast, cheap, and sounds natural. Just don't trust it with calculations until you've tested it on your own queries for a few days.

Can I connect this to WhatsApp Business directly or do I have to copy-paste?
Right now I copy-paste. Kryotta has a Gmail connector, but WhatsApp isn't integrated yet. I paste the customer's message into Kryotta, get the reply, check it, then send it through WhatsApp Business. Takes about 30 seconds per message.

What if the AI gives a wrong answer and I don't catch it?
It happens. I missed one wrong amount in 200 messages. The customer caught it, I apologized and fixed it. It's embarrassing, but it's not worse than the mistakes I made when I was replying to 40 messages in a row at midnight and could barely see straight.

Does this work for Vodafone Cash or AirtelTigo Money, or just MTN MoMo?
The workflow works for any mobile money service—you're just checking transaction references and amounts against your log. I only use MTN because that's what most of my customers have, but the same setup would work for Vodafone Cash.

If you're selling on WhatsApp or Instagram in Ghana and MoMo queries are eating your evenings, this is worth trying. Set up a payment log, write a simple prompt, let Auto routing handle the easy questions, and check the numbers before you send. You'll still need to read every reply for the first week, but after that it gets faster. You can test it free at Kryotta—no credit card, just sign up and try it on your next ten messages.

K
Written by
Kryotta Team
Product & research

Kryotta is the multi-model AI workspace — every leading model, one login, one bill. Try it free →

Related reading