I used AI to reply to 100 WhatsApp Business orders during wedding season: which model kept up

I tested Claude, Gemini, and DeepSeek on 100 real WhatsApp Business orders during peak wedding season. Here's which model handled customer replies best—and where each one stumbled.

KKryotta TeamProduct & research · · 9 min read
Cluttered small business desk with smartphone, order notes, and gift hamper boxes in natural light
Cluttered small business desk with smartphone, order notes, and gift hamper boxes in natural light

I tested three models on the same 100 WhatsApp Business orders

I run a small gifting business from Surat. We do customised hampers, trousseau packaging, wedding favours—things people order in bulk when someone's getting married or Diwali's two weeks out. October through December is when we make most of our money, and it's also when my phone becomes a 16-hour-a-day WhatsApp Business terminal.

Last season I got 100 orders in three weeks. That's triple our usual volume. Every order came with questions: "Can you swap the almonds for cashews?" "Will it reach Rajkot by the 18th?" "I sent ₹8,500 by UPI but the screenshot shows ₹8,050, the rest is processing." I was copying and pasting replies at midnight, getting tone wrong, forgetting to confirm payment amounts, and once—mortifyingly—sending a delivery update to the wrong customer.

This year I decided to test whether AI could handle the replies. I set up a simple workflow: every new WhatsApp Business message went into a Google Sheet via Zapier, then I fed batches of ten messages into Claude Sonnet 4.5, Gemini Flash, and DeepSeek V3. I wrote the first reply myself as a template, then asked each model to generate replies for the rest based on the customer's message, order details, and payment status. I tracked three things: how many replies I could send without editing, how many customers came back confused, and whether the model hallucinated details when I didn't give it enough context.

Here's what worked, what didn't, and which model I'm using this wedding season.

Claude kept tone consistent but needed explicit instructions for Hinglish

Claude Sonnet 4.5 wrote the most reliable replies when I gave it a clear template and told it exactly what to include. I'd paste the customer's message—usually a mix of English and Hindi, sometimes just "Bhaiya, order kab aayega?"—and Claude would match the tone without overdoing the Hinglish or sounding like a chatbot.

The prompt I used: "Reply to this WhatsApp Business message. Confirm the order details, give the delivery date, and ask them to send the UPI screenshot if they haven't already. Keep it warm and conversational, use light Hinglish only if the customer used it first, and don't add emoji unless I tell you to."

Where Claude struggled: it wanted to be helpful to the point of inventing information. If I didn't explicitly say "delivery date is 22nd November," it would sometimes write "your order will reach you soon" or—worse—pick a date that sounded plausible. I caught one reply that said "your hamper will arrive by Diwali" when Diwali was four days away and we hadn't even started packing. After that I made sure every prompt included the actual delivery date, payment amount, and tracking number if we had one.

I used Claude for 35 of the 100 orders. I edited six replies, mostly to remove a sentence that was technically correct but would've started a long conversation. ("If the payment doesn't reflect in 24 hours, please contact your bank"—true, but not what a bride's cousin wants to hear three days before the wedding.)

Gemini handled UPI screenshots and messy queries faster

Gemini Flash was better at dealing with the chaos of real customer messages. People don't send tidy queries. They send a UPI screenshot, then a voice note asking if we do gift wrapping, then another message two hours later saying they meant red wrapping, not gold, and can we add a handwritten card.

I'd paste the whole thread into Gemini—sometimes five or six messages from the same customer, out of order—and it would pull out the key details: payment amount from the screenshot, the request to swap colours, the delivery pincode. Then it would draft a reply that addressed everything without sounding like a list.

The prompt: "Here's a WhatsApp conversation with a customer. They've sent a UPI payment screenshot and asked about delivery and customisation. Write a reply that confirms the payment amount, answers their questions, and tells them we'll send an update once the order's packed. Match their tone."

Where Gemini tripped up: it sometimes assumed details I hadn't given it. A customer asked, "Can you deliver to Anand by the 15th?" and Gemini wrote, "Yes, we deliver to Anand, your order will reach by the 15th." We do deliver to Anand, but I hadn't confirmed the date yet because I didn't know if the supplier would send the almonds on time. I had to edit that one to "We deliver to Anand—let me confirm the 15th and get back to you by tomorrow."

I used Gemini for 40 orders. I edited nine replies, mostly to soften a commitment or add a line about GST invoices, which Gemini kept forgetting to mention even though I'd put it in the prompt.

DeepSeek was fast and cheap but needed heavy editing for tone

DeepSeek V3 is the model I wanted to love. It's fast, it costs almost nothing per query, and when I gave it a straightforward order confirmation to write—"Customer paid ₹12,000 for 50 wedding favours, delivery to Vadodara by 20th November"—it nailed it in one go.

But DeepSeek couldn't handle the conversational messiness that Gemini and Claude managed. If a customer wrote, "Bhai, mera order ka kya scene hai? Payment toh ho gaya tha," DeepSeek would reply in stiff, formal English: "Your order is currently being processed. Payment has been received. You will receive a tracking update shortly." Technically accurate, but it sounds like a bot, and in wedding season people want to feel like they're talking to a human who cares whether the hampers arrive on time.

I used DeepSeek for 25 orders—mostly the simple ones where the customer had already sent payment, asked one clear question, and I just needed to confirm details. I edited 18 of those 25 replies. That's a 72% edit rate, which defeats the point of automation.

Where DeepSeek did help: I used it to draft bulk messages when I needed to send the same update to fifteen customers at once. ("Your order's packed and will be dispatched tomorrow. Here's the tracking number.") For that, the formal tone didn't matter, and DeepSeek churned out fifteen variations in under a minute so the messages didn't look copy-pasted.

The checklist I'm using this season

Here's the workflow I've landed on after testing all three models on 100 real orders:

1. Use Claude for first replies and anything involving money.
If it's the customer's first message, or if they're asking about payment, refunds, or whether we received their UPI transfer, I use Claude Sonnet 4.5. I give it the order amount, payment status, and delivery date in the prompt, and I tell it not to guess. Edit rate: about 15%.

2. Use Gemini when the customer's sent multiple messages or a screenshot.
Gemini's better at pulling signal from noise. If someone's sent a UPI screenshot, a question about customisation, and a follow-up about delivery, I paste the whole thread into Gemini Flash and let it draft a reply that ties everything together. Edit rate: around 20%, mostly to add details Gemini assumed.

3. Use DeepSeek for bulk updates, not conversations.
If I need to send fifteen customers the same tracking update or a Diwali delivery deadline, I'll use DeepSeek V3 to generate variations so it doesn't look like a broadcast. But I won't use it for one-on-one replies unless the query's extremely simple and I'm fine editing the tone.

4. Never let the model invent dates, amounts, or tracking numbers.
This one's non-negotiable. If I don't have the tracking number yet, I tell the model "tracking number not available" in the prompt, and I make it write "we'll send the tracking number once it's dispatched." I've caught all three models inventing plausible-sounding details when I left gaps in the context.

5. Keep a swipe file of good replies and feed them back into the prompt.
After two weeks I had about twenty replies that customers responded well to—short, warm, no confusion, no follow-up questions. I now paste one or two of those into the prompt as examples when I'm asking the model to draft something similar. It's made the biggest difference to tone consistency, especially with Claude.

6. Check every reply before you send it, even if it looks perfect.
I know this defeats some of the speed gains, but I've learned the hard way. It takes me ten seconds to skim a reply and catch a wrong date or an overpromise. It takes an hour to fix the problem if I send it.

What I'd do differently next time

I didn't test Llama or Mistral because I wanted to start with the models I'd heard other small-business owners mention. I also didn't set up a proper auto-reply system—I was still copying replies from the AI and pasting them into WhatsApp Business manually, which is only a half-step better than writing them myself.

Next season I'll either use Kryotta's team workspace to share the workflow with my sister (who helps with orders during peak season) or look into a proper WhatsApp Business API integration so the replies go out automatically for simple queries like "Did you receive my payment?" For anything more complex, I'll keep the human-in-the-loop approach. Wedding orders are too high-stakes to let a model run unsupervised, even when it's getting 80% of replies right.

The other thing I'd track: customer satisfaction. I measured reply accuracy and edit rates, but I didn't ask customers whether they felt the replies were helpful or whether they could tell they were AI-generated. A few people did say "thanks for the quick response," which I'm taking as a good sign, but I'd want to be more deliberate about it next time.

Questions people ask

Can I use AI to reply to WhatsApp Business messages without sounding like a bot?
Yes, but you need to give the model examples of your real replies and tell it to match the customer's tone. Claude and Gemini both handled conversational Hinglish well when I showed them what "warm but professional" looked like for my business. DeepSeek struggled with tone unless the query was very straightforward.

Which model is cheapest for high-volume WhatsApp replies during wedding season?
DeepSeek V3 is the cheapest per query, but I ended up editing 70% of its replies, which ate into the time savings. Gemini Flash was the best balance of cost and usability for me—fast, accurate enough that I only edited one in five replies, and good at handling messy multi-message threads.

Should I let AI auto-reply to payment confirmations and delivery questions?
Not without checking the reply first. All three models occasionally invented details when I didn't give them enough context—wrong delivery dates, assumed tracking numbers, overpromises about customisation. I'd use AI to draft the reply, then skim it before sending. It still saves time, but it won't cost you a customer.

Does AI work for Hinglish customer messages on WhatsApp Business?
Claude and Gemini both handled light Hinglish well—they matched the customer's tone without overdoing it or dropping into pure Hindi when it wasn't natural. DeepSeek defaulted to formal English even when the customer wrote in Hinglish, which made replies feel stiff. If your customers mix languages, test the model on a few real messages before you commit to it.


I'm still using this workflow three months later. It's not perfect, but it's turned 16-hour WhatsApp days into something manageable, and I've only had one customer come back confused (my fault—I forgot to update the delivery date in the prompt). If you're drowning in WhatsApp Business orders this wedding season and you've thought about trying AI, start with ten messages and see which model matches your tone. You can set up the same workflow in Kryotta and test Claude, Gemini and DeepSeek side by side without switching tools.

K
Written by
Kryotta Team
Product & research

Kryotta is the multi-model AI workspace — every leading model, one login, one bill. Try it free →

Related reading