I ran the same prompts in six businesses for three months. Here's what AI finished alone.
I manage operations for a handful of small businesses across the UK and Ireland—an eBay reseller in Cardiff, a tutoring service in Dublin, a property lettings agency in Manchester, a dental surgery in Belfast, a bookkeeper serving sole traders in Leeds, and a café in Bristol. For 90 days I gave Claude Sonnet 4.5, Gemini Pro, and DeepSeek V3 the same tasks I'd normally do myself, then checked which ones came back ready to send and which ones needed me to step in.
The surprise wasn't that AI got some things wrong. It was how predictable the pattern turned out to be. Certain tasks—customer service replies, invoice generation, first-draft property listings—came back correct 19 times out of 20. Others, like VAT reconciliation and patient appointment changes, needed a human every single time. I kept a spreadsheet. This is what it showed.
The side-by-side: what ran unsupervised vs what needed me
| Industry | Task AI finished alone | Task that needed human review | Model that worked best |
|---|---|---|---|
| E-commerce (eBay reseller) | Royal Mail tracking replies, product description rewrites, Boxing Day sale email drafts | Refund decisions when customer claims item arrived damaged, VAT invoice corrections for EU orders | Claude Sonnet 4.5 for emails, DeepSeek V3 for descriptions |
| Education (tutoring) | Session confirmation emails, rescheduling replies for standard requests, first-draft lesson plans for GCSE maths | Handling parent complaints about tutor performance, deciding whether to offer a refund for missed sessions | Gemini Pro for lesson plans, Claude Sonnet 4.5 for confirmations |
| Real estate (lettings) | First-draft property listings, tenant inquiry responses, rent reminder emails | Deciding whether a maintenance request is urgent enough to call the landlord at 9pm, writing Section 21 notices | Claude Sonnet 4.5 for listings, Gemini Pro for tenant replies |
| Healthcare admin (dental) | Appointment confirmation texts, insurance pre-authorization letter drafts, recall reminder emails | Rescheduling when a patient says "I'm in too much pain to wait until Thursday", handling complaints about billing | Claude Sonnet 4.5 for confirmations, needed me for everything involving judgment |
| Accounting/legal (bookkeeper) | Generating invoices from time logs, chasing overdue payments (first email), Companies House form CS01 drafts | VAT return reconciliation, deciding whether a receipt qualifies as a business expense, writing formal dispute letters | DeepSeek V3 for invoice generation, needed me for anything HMRC-related |
| Restaurants (café) | Booking confirmations, allergen information replies, staff rota first drafts | Handling a booking for 22 people when the brief says "we might be 18 or 26", deciding whether to accept a same-day catering order | Gemini Pro for confirmations, Claude Sonnet 4.5 for allergen replies |
The table tells you what I learned slowly: AI finishes templated communication and structured first drafts without supervision. It fails at judgment calls and anything involving money or compliance.
E-commerce: it wrote 140 Royal Mail replies, but I had to decide every refund
The Cardiff eBay seller ships vintage clothing across the UK. I gave Claude Sonnet 4.5 a simple prompt: "Customer says tracking shows the parcel is stuck at the Coventry depot for four days. Write a reply." It produced 140 replies over 90 days that I sent without changing a word. They were polite, accurate, and included the Royal Mail customer service number.
But when a customer in Glasgow said a £68 dress arrived with a torn seam and sent a photo, Claude offered a full refund and return postage in the first draft. The dress was clearly worn—makeup on the collar, deodorant marks under the arms. I had to rewrite the reply, offer a partial refund, and explain our returns policy. Claude doesn't look at photos critically. It assumes good faith.
DeepSeek V3 wrote better product descriptions than I did. I fed it a photo of a 1990s Burberry trench coat and the prompt "Write an eBay listing for UK buyers, mention condition and measurements." It came back with a 180-word description that included chest width, sleeve length, and a note about a small stain on the left cuff I'd missed. I used that template for 60 listings. Sold 52 of them.
Tutoring: confirmations ran themselves, complaints needed my voice
The Dublin tutoring service sends 30 confirmation emails a week. I set up a workflow in Kryotta: when a parent books a session through Calendly, Gemini Pro generates a confirmation with the date, time, tutor name, and a line about what the student will cover. I checked the first 20. All correct. I stopped checking after that and let it run for two months. Zero complaints.
But when a parent in Drogheda emailed to say their daughter's tutor "wasn't engaging and spent half the session on her phone," Gemini Pro's first draft was too defensive. It said "Our tutors are trained professionals and we're confident in their performance." I rewrote it, apologized, offered a free session with a different tutor, and called the tutor to ask what happened. (She'd been checking her phone for an emergency family message. Understandable, but the parent was right to complain.)
Lesson plans were a surprise win. I asked Gemini Pro to draft a one-hour GCSE maths lesson on quadratic equations, including three worked examples and five practice problems. It came back with a structure I'd have charged £40 to write myself. I tweaked two examples and sent it to the tutor. She used it.
Real estate: listings were 80% done, but urgent repairs needed me
The Manchester lettings agency lists 12 properties a month. I gave Claude Sonnet 4.5 photos, the rent, the postcode, and a prompt: "Write a 150-word listing for Rightmove, mention transport links and local schools." It wrote 36 listings over three months. I edited maybe six of them, usually to add a detail the landlord mentioned on the phone.
Tenant inquiry responses were faster. Someone asks "Is the flat available from March 1st and do you allow pets?" Claude replies in 40 seconds: "Yes, the flat is available from March 1st. We allow one small dog or cat with a £200 pet deposit. Would you like to arrange a viewing?" I sent 90 of those replies. Three led to signed tenancies.
But when a tenant in Salford emailed at 10pm to say "The boiler's making a loud banging noise and the radiators are cold," Claude's draft said "We'll send a heating engineer within 48 hours." I called the landlord immediately. It was January. No heating in 48 hours means frozen pipes and a potential insurance claim. The engineer came at 8am the next day.
Healthcare admin: reminders yes, judgment no
The Belfast dental surgery sends 60 appointment reminders a week. Claude Sonnet 4.5 writes them all: "Hi [name], this is a reminder that you have a check-up with Dr. O'Neill on Thursday 14th March at 2pm. Reply YES to confirm or call us on 028 9024 1234 to reschedule." I haven't touched one in two months.
Insurance pre-authorization letters were another win. I gave Claude the patient's name, the procedure code, and the insurance provider. It wrote a two-paragraph letter requesting approval for a crown replacement. The surgery sent 18 of those letters. All 18 came back approved.
But when a patient called and said "I'm in too much pain to wait until Thursday, can I come in today?", Claude's draft reply was "Our next available emergency appointment is Friday at 11am." I checked the schedule, saw a 30-minute gap at 4pm, called the patient back, and booked her in. AI doesn't know what "too much pain" means in context.
Accounting: invoices were perfect, VAT was a disaster
The Leeds bookkeeper serves 40 sole traders. I fed DeepSeek V3 a spreadsheet of hours worked and hourly rates, then asked it to generate invoices. It produced 120 invoices in 90 days with the correct VAT calculations, payment terms, and bank details. I spot-checked 30. All correct.
Chasing overdue payments was easier than I expected. I gave Claude a prompt: "Client hasn't paid invoice #4782 for £340, due 14 days ago. Write a polite first reminder." It wrote 25 reminders. Eighteen clients paid within a week. The other seven needed a second email, which I wrote myself with a firmer tone.
But VAT reconciliation failed every time. I asked DeepSeek to match 60 receipts to bank transactions and flag anything that didn't qualify as a business expense under HMRC rules. It flagged a £12 Tesco receipt as "unclear" when the itemized list clearly showed only coffee and biscuits for client meetings. It missed a £340 Amazon order that included a £60 personal book. I had to check every line myself.
Restaurants: booking confirmations ran for 90 days, group sizes broke it
The Bristol café takes 40 bookings a week. Gemini Pro writes every confirmation: "Hi [name], your table for [number] is confirmed for [date] at [time]. We'll hold it for 15 minutes. See you soon." I sent 480 of those messages. Zero complaints.
Allergen replies were faster than looking them up myself. Someone asks "Does the lemon tart contain nuts?" I paste the question into Claude Sonnet 4.5 with the café's allergen spreadsheet. It replies in 20 seconds: "The lemon tart doesn't contain nuts, but it's made in a kitchen that handles almonds and hazelnuts. If you have a severe allergy, please let us know when you arrive." I sent 35 of those replies. All accurate.
But when someone emailed to book a table "for around 22 people, maybe 18, possibly 26" for a birthday lunch, Gemini's draft said "We can accommodate 22. Please confirm your final number 48 hours before." The café only seats 32 total. I called the customer, asked whether 26 was a hard maximum, and suggested they book the whole café for two hours on a Tuesday afternoon when we're quiet. They agreed. AI doesn't negotiate.
When to use which: the pattern I'd follow now
If the task has a template and no judgment call, let AI finish it. Appointment confirmations, product descriptions, first-draft listings, and invoice generation all ran unsupervised for 90 days. I checked the first 20 of each, then stopped checking.
If the task involves money, compliance, or someone saying "urgent," review it yourself. VAT reconciliation, refund decisions, Section 21 notices, and emergency maintenance requests all needed me. Claude and Gemini don't understand risk the way you do.
If the task is somewhere in between—like a parent complaint or a group booking—use AI for the first draft, then rewrite the parts that need your voice. I saved 40 minutes a day by letting Claude write the structure, then spending five minutes making it sound like me.
I now route by task type in Kryotta. Claude Sonnet 4.5 for anything customer-facing, DeepSeek V3 for structured data work like invoices and descriptions, Gemini Pro for lesson plans and first-draft content. I don't use the same model for everything anymore. That was the expensive mistake I made in month one.
Questions people ask
Which model made the fewest mistakes across all six industries?
Claude Sonnet 4.5 had the lowest error rate for customer-facing communication—appointment confirmations, booking replies, and tenant inquiries. DeepSeek V3 was more accurate for structured tasks like invoice generation and product descriptions. I didn't find one model that was best at everything.
Did AI ever send something that caused a real problem?
Once. Gemini Pro wrote a reply to a tutoring inquiry that quoted the wrong hourly rate—£35 instead of £45. The parent booked based on that price and I had to honour it for the first month. I added the current rate card to every prompt after that and it hasn't happened again.
How much time did this actually save per week?
I tracked it in week 12. AI wrote 140 emails, 18 invoices, 12 property listings, and 60 appointment confirmations that week. If I'd written them myself, that's roughly nine hours of work. I spent maybe 90 minutes reviewing the ones that needed it. Call it seven hours saved, which I spent on things that actually grow the businesses.
Can I trust AI with GDPR-sensitive information like patient names and addresses?
I do, but only in Kryotta's team workspace where I control who sees the data. I don't paste patient details into the free version of ChatGPT or Gemini. The dental surgery and bookkeeper both reviewed Kryotta's data handling before we started, and both were comfortable with it. If you're handling sensitive information, check your AI provider's terms and make sure they're not training models on your prompts.
I still write the emails that matter—the ones where tone carries as much weight as the facts. But the 200 messages a week that follow a template? AI finishes those now, and I don't check them anymore. If you're doing this work yourself and wondering whether AI can take some of it off your plate, the answer is yes for about 60% of it. The other 40% still needs you, and probably always will. Start your free trial at Kryotta and route one task type to Claude or Gemini for a week. You'll know within five days whether it works for your business.



