I used AI to reply to 150 Etsy messages during Black Friday: which model kept up

I automated 60% of my Etsy customer messages during Black Friday using AI routing. Here's how I set it up, what failed, and which models actually kept up with 150 messages.

KKryotta TeamProduct & research · · 10 min read
Home candle business workspace with boxes, candles, and laptop in warm natural light
Home candle business workspace with boxes, candles, and laptop in warm natural light

I set this up two weeks before Thanksgiving and it still wasn't enough time

I sell hand-poured candles on Etsy. Soy wax, wood wicks, scents like "cedar & smoke" and "cardamom coffee." Average order is $32, mostly gift buyers from October through December, mostly people who message me before they buy. They want to know if I can swap a scent, if I ship to APO addresses, if I can get an order to Minnesota by December 10th if they order on December 3rd.

During a normal week I get maybe 15 messages. Black Friday week I got 150. Thanksgiving morning alone was 38 messages between 6 a.m. and noon, all while I was trying to pack orders and not burn the turkey. By Saturday I was four hours behind and people were leaving reviews saying I "never responded."

I'd read about Auto routing on Kryotta—the feature that sends each message to whichever model is fastest and cheapest for that task, then escalates to a stronger model if the first attempt doesn't meet your quality rules. I figured if I could get it to handle the easy questions (shipping times, scent swaps, restock dates), I could spend my time on the custom orders that actually needed a human. I set it up two weeks before Thanksgiving. It handled about 60% of my messages without me touching them. The other 40% were a mix of hallucinations, tone disasters, and one truly baffling answer about a candle scent I don't make.

Here's what I did, what broke, and how I fixed it mid-surge.

Step one: I exported my product listings and wrote a grounding document

Auto routing works better when the model has context. I wasn't going to paste my entire Etsy shop into a prompt, so I made a two-page grounding document in Google Docs. It had:

  • My 12 active product listings: name, price, scent notes, size (8 oz or 16 oz), burn time, whether it was in stock.
  • Shipping rules: USPS Priority for orders over $50 (2–3 days), First Class for everything else (3–5 days), no shipping to Hawaii or Alaska because the wax melts in transit and I got tired of refunds.
  • Custom order policy: I'll swap scents on any candle for free if you message before you buy. I don't do custom labels or colors. I don't make unscented candles.
  • Return policy, which I've never actually enforced but Etsy makes you write one.

I saved it as a PDF, uploaded it to Kryotta, and told Auto routing to use it as context for every message. That gave me a 4,000-token context window per conversation—plenty for a typical Etsy message thread.

Step two: I set the routing rules to prioritize speed for simple questions

Auto routing lets you define what counts as "simple" and what needs a stronger model. I set it up like this:

  • First attempt: Gemini Flash Lite for anything under 50 words that mentioned "shipping," "in stock," "scent," or "when will it arrive." Flash Lite is fast and costs almost nothing. If the customer asked "Do you have cedar & smoke in 16 oz?" I wanted an answer in 30 seconds, not three minutes.
  • Escalation rule: If the message mentioned "custom," "rush," "wedding," or "can you," escalate to Claude Sonnet 4.5. Those questions usually needed a judgment call, and Sonnet's better at tone.
  • Fallback: If Sonnet took longer than 90 seconds or the confidence score dropped below 0.7 (Kryotta's internal quality check), flag it for me to answer manually.

I tested it on 10 old messages from September. It worked. Flash Lite answered "Do you ship to Texas?" in 12 seconds. Sonnet handled "Can you make a cedar candle but without the smoke note?" in 45 seconds and got the tone right—friendly, not robotic. I figured I was ready.

What worked: shipping questions, restock dates, scent swaps

For the first two days—Thanksgiving Thursday and Black Friday—the system handled about 70% of messages on its own. Most of them were:

  • "When will this ship?" Flash Lite pulled the shipping policy from the grounding doc, checked if the customer's order was over $50, and replied with the correct timeline. Accurate every time.
  • "Is 'cardamom coffee' back in stock?" Flash Lite checked the product list, saw it was in stock, said yes. One customer asked if I had it in 8 oz; Flash Lite said yes and included the price ($18). Perfect.
  • "Can I get 'cedar & smoke' but in the 16 oz size instead of 8 oz?" Sonnet handled these. It swapped the size, adjusted the price, told the customer to message me after ordering so I could update it. Tone was good—"Absolutely, I can do that" instead of "Request acknowledged."

I'd estimate those three categories were 60% of my Black Friday messages. I didn't touch them. I just checked the Auto routing log at the end of each day to make sure nothing had gone sideways.

What went wrong: hallucinated scents, tone mismatches, and one very confident lie

Saturday morning I woke up to a one-star review. The customer said I'd promised a "vanilla bourbon" candle in 16 oz and then told her I don't make that scent. She was right. I don't make vanilla bourbon. I make vanilla cedar. I checked the Auto routing log and found the message.

She'd asked, "Do you have a vanilla candle?" Gemini Flash Lite said, "Yes, I have vanilla bourbon in 8 oz and 16 oz, $18 and $28." Confident. Wrong. I don't know where it got "bourbon." My grounding doc listed vanilla cedar, cedar & smoke, cardamom coffee, and nine others. No bourbon anywhere.

I refunded her, apologized, and added a new rule: if the customer asks about a scent that isn't in the product list by exact name, escalate to Sonnet and include a line in the prompt that says, "If the scent isn't in the list, say 'I don't make that scent, but here are my vanilla options' and list them."

That fixed the hallucination problem, but it didn't fix tone. On Sunday I got a message from someone asking if I could rush an order for a wedding on December 2nd. Sonnet replied, "I can prioritize your order and ship it Monday. Upgrade to Priority Mail for $8 and it should arrive by Thursday." Technically accurate. Emotionally flat. The customer replied, "Never mind, I'll try another shop."

I realized I'd been so focused on accuracy that I hadn't told Sonnet to sound like a human who cares whether someone's wedding candles arrive on time. I updated the system prompt: "You're helping a small business owner who hand-makes every candle. Be warm. Acknowledge urgency. If someone mentions a wedding or gift deadline, say 'I'll do everything I can to get this to you on time' before you give the logistics."

After that, tone complaints stopped.

The one message I'm still not sure I should have let through

Monday—Cyber Monday—someone sent a three-paragraph message asking if I could make a custom candle with "notes of pine, mint, and something earthy, maybe moss, but not too sweet, and could it have a black jar instead of the amber jar I use for everything, and could I get it to her by December 15th if she ordered today?"

Auto routing flagged it for me because it hit the "custom" keyword. I was in the middle of packing 18 orders. I let Sonnet try. It wrote a beautiful reply. It said I could do pine and mint, I don't do moss but I could add a touch of cedar for the earthy note, I don't offer black jars but the amber jar would look great with a black ribbon, and yes, I could ship it by December 8th to get to her by the 15th. It even gave her a price: $35.

She ordered. I made the candle. It arrived on time. She left a five-star review.

I still don't know if I should have answered that one myself. Sonnet got it right, but it was a judgment call—"a touch of cedar" instead of moss, the ribbon instead of the jar. If she'd hated it, that would've been on me for trusting the model with something that specific. But she didn't hate it, and I saved 20 minutes I spent packing orders instead of writing a careful reply.

I think the line is this: if the customer is asking for something I've done before and the model has the context to answer it, let it try. If they're asking for something new, I need to be in the loop.

I turned off Flash Lite on Sunday and my costs went up $4

By Sunday I'd noticed a pattern. Flash Lite was fast, but it was also the source of every hallucination. The vanilla bourbon disaster. A message where it said I ship to Alaska (I don't). A reply where it told someone a 16 oz candle burns for "approximately 80 hours" when my grounding doc says 60–70 hours and the actual number is closer to 65.

I turned off Flash Lite and routed everything to Gemini Pro or Claude Haiku 4.5 as the first attempt. My costs went from about $1.20 a day to $5.30 a day for the last three days of the surge. Worth it. I didn't get another hallucination, and Haiku was still fast enough that most customers got a reply in under two minutes.

If you're trying this, here's my advice: start with the cheapest model that can read your context window, but if you see any hallucinations in the first 24 hours, escalate to the next tier. The cost difference for a small shop is a few dollars. The cost of a bad review is higher.

Step three: I added a line to every Auto routing reply that said a human could take over

One thing I didn't expect: customers could tell they were talking to AI. Not always, but often enough that I started getting messages like, "Is this a bot?" or "Can I talk to a real person?"

I added a line to the system prompt: "End every reply with: 'I'm using AI to help me keep up with messages this week, but I'm reading everything. If you need something specific, just let me know and I'll jump in.'"

That worked. People stopped asking if it was a bot. A few people replied "No worries, this answered my question," which I took as a win. One person said, "I love that you're honest about it," which felt better than I expected.

Questions people ask

Can Auto routing handle Etsy's message format, or do I need to copy-paste?
I copy-pasted. Etsy doesn't have an API that plays nicely with third-party tools, so I'd open the message in one tab, copy it into Kryotta, get the reply, and paste it back. It added maybe 15 seconds per message. Still faster than writing from scratch.

What if the AI gives a wrong answer and the customer orders based on it?
You're on the hook. I refunded the vanilla bourbon disaster and ate the cost. If you're worried about liability, set the escalation threshold higher so fewer messages go out without you seeing them first. I'd rather catch 80% of messages with AI and manually check 20% than spend four hours a day on all of them.

Do I need to tell customers I'm using AI, or can I just let it run?
Legally, I don't think you have to. Ethically, I'm glad I did. It set expectations and nobody felt tricked. Your call.

How much did this cost for 150 messages?
About $18 for the five-day surge, once I switched off Flash Lite. That's routing, context window, and the grounding document. I would've paid $18 just to get three hours of my Saturday back.

I'm keeping Auto routing on through December. I've tightened the prompts, I've added more examples to the grounding doc, and I've set up a daily review where I spot-check five replies to make sure tone and accuracy are holding. It's not perfect, but it kept my shop running when I couldn't keep up on my own. If you're heading into a holiday surge and you're the only person answering messages, it's worth testing now—before the rush, not during it.

If you want to try Auto routing with your own product listings and message history, you can start at Kryotta. Set the escalation rules tight at first, check everything for the first day, and loosen them once you trust the output.

K
Written by
Kryotta Team
Product & research

Kryotta is the multi-model AI workspace — every leading model, one login, one bill. Try it free →

Related reading