I used AI in 6 industries for 30 days: which tasks it actually finished vs which needed me

I tested AI across six industries for a month to see which tasks it could finish alone. The results: some are genuinely autonomous. Others fail in ways that cost money or trust if you don't catch them.

KKryotta TeamProduct & research · · 9 min read
Workspace with multiple laptops showing business dashboards, task lists, and calendars with morning light from window
Workspace with multiple laptops showing business dashboards, task lists, and calendars with morning light from window

The test: six sectors, one month, one question

I spent January testing AI across six industries I either work in or advise: e-commerce (a lighting store on Allegro), education (a German tutor with 40 students), real estate (a letting agent in Dublin), healthcare admin (a dental practice in Kraków), a CA firm in Brussels, and a restaurant in Lyon. The question was simple: which tasks did AI finish without me, and which ones still needed a human to check, edit, or override?

I'm not interested in what's theoretically possible. I wanted to know what actually worked when a customer emailed at 19:30, when an invoice was due, when someone left a one-star review. I used Claude Sonnet 4.5, Gemini Flash, and DeepSeek V3 in Kryotta's workspace, logged every task, and marked whether I could send the output as-is or whether I had to rewrite it.

The table below is what I found. Some tasks are genuinely autonomous now. Others still fail in ways that would cost you money or trust if you didn't catch them.

What AI finished vs what it didn't: the table

IndustryTaskAI completed it?Model that workedWhat broke when it didn't
E-commerce (Allegro)Product descriptions for LED ceiling lightsYesClaude Sonnet 4.5
E-commerceRefund email after Klarna disputeYesGemini Flash
E-commerceDeciding whether to accept a return (customer claims "wrong colour")NoAll threeAI approved a return for an item the customer had clearly used for two months; photos showed wear
EducationLesson confirmation emailsYesDeepSeek V3
EducationRescheduling three students after a sick dayNoClaude Sonnet 4.5Suggested times that clashed with other bookings; didn't check the calendar
Real estateViewing confirmation with deposit instructions (SEPA)YesClaude Sonnet 4.5
Real estateAnswering "Is the flat suitable for a family with two kids?"NoGemini FlashSaid yes because there are two bedrooms, but ignored the lease clause that bans children under 12
Healthcare adminAppointment reminders (SMS-length)YesGemini Flash
Healthcare adminExplaining why a procedure isn't covered by NFZNoClaude Sonnet 4.5Gave a plausible but incorrect reason; the actual policy had changed in December
CA / legalGenerating invoices with VAT breakdownYesClaude Sonnet 4.5
CA / legalDrafting a late-payment reminderYesDeepSeek V3
CA / legalAdvising a client on whether to register for VAT in GermanyNoAll threeSuggested the threshold wrong by €5,000; didn't account for distance-selling rules
RestaurantBooking confirmation (table for four, 20:00 Friday)YesGemini Flash
RestaurantReplying to a complaint about cold foodNoClaude Sonnet 4.5Apologised well but offered a €30 voucher without checking the manager's policy (it's €15 max)

The pattern is clear: AI finishes tasks that have a template, a clear input, and no judgement call. It struggles the moment you need it to check something external (a calendar, a contract, a policy document) or make a decision that could go two ways depending on context.

E-commerce: descriptions yes, returns no

Product descriptions are the easiest win. I gave Claude Sonnet 4.5 a spreadsheet of ceiling lights (model number, wattage, colour temperature, diameter) and the prompt: "Write a 60-word Allegro product description in Polish for each item. Mention energy rating, room size, and installation type. Neutral tone." It wrote 22 descriptions in four minutes. I changed two words across all of them.

Refund emails also work. A customer disputed a Klarna payment, claiming the item didn't match the listing. I fed Gemini Flash the order number, the dispute reason, and our refund policy. It wrote a two-paragraph email confirming the refund and explaining the timeline. I sent it as-is.

Returns are different. A customer said the lamp was "the wrong colour" and wanted a refund. I asked Claude to decide. It read the photos, read the listing, and approved the return. I looked at the photos myself: the lamp had scuff marks on the base and dust inside the shade. It had been used for weeks. AI saw "colour mismatch" and defaulted to yes. A human saw wear and said no.

When to use AI: Descriptions, booking confirmations, refund emails where the policy is binary. When you still need a human: Any return or complaint where the customer's claim needs to be weighed against evidence.

Education: confirmations yes, rescheduling no

Lesson confirmations are a template task. Student books a slot, AI sends an email with the time, the Zoom link, and a reminder to bring homework. DeepSeek V3 did this 40 times in January without a single error, and at €0.02 per email it's cheaper than the Zapier automation I used to run.

Rescheduling is harder. I got sick on a Tuesday and needed to move three lessons. I gave Claude Sonnet 4.5 the student names, their usual slots, and my availability for the rest of the week. It suggested times that clashed with other bookings. It didn't check my calendar because I hadn't connected it (and even if I had, it would need explicit instructions to cross-reference). I ended up doing it manually in ten minutes.

When to use AI: Confirmations, reminders, anything that repeats the same structure. When you still need a human: Anything that requires checking a live schedule or making a trade-off between two students' preferences.

Real estate: instructions yes, lease questions no

Viewing confirmations are perfect for AI. Someone emails asking to see a flat. I gave Claude the address, the viewing time, and the deposit amount (first month + €800 security, SEPA to this IBAN). It wrote a polite, clear email in under ten seconds. I sent 14 of those in January.

Lease questions are trickier. A couple asked if the flat was suitable for two kids under ten. Gemini Flash said yes because it has two bedrooms and a balcony. But the lease has a clause banning children under 12 (the building has strict noise rules and elderly residents). AI read the flat's features, not the contract. I caught it because I know that lease. If I hadn't, I'd have wasted their time and mine.

When to use AI: Viewing confirmations, deposit instructions, anything factual. When you still need a human: Questions that depend on contract clauses, building rules, or anything not in the listing itself.

Healthcare admin: reminders yes, policy explanations no

Appointment reminders are the lowest-risk task in this list. Gemini Flash took a spreadsheet of patient names, phone numbers, and appointment times, and wrote 80 SMS-length reminders in Polish. Zero mistakes.

Policy explanations are a different story. A patient asked why a specific dental procedure wasn't covered by NFZ (Poland's public health fund). I asked Claude Sonnet 4.5 to explain. It gave a confident, two-paragraph answer. The reason it gave was plausible but wrong. The actual policy had changed in December, and Claude was working from older training data. The patient would have been misinformed.

When to use AI: Reminders, confirmations, anything that just repeats information you've already verified. When you still need a human: Anything involving current policy, eligibility, or medical advice.

Invoice generation is a solved problem. I gave Claude Sonnet 4.5 a CSV of client names, services, amounts, and VAT rates (21% standard, 6% reduced for some services in Belgium). It produced 18 invoices in the correct format, with line items, totals, and payment terms. I checked every number. All correct.

VAT registration advice is not solved. A client asked whether they needed to register for VAT in Germany because they'd started selling on Amazon.de. I asked Claude, Gemini, and DeepSeek. All three gave answers. All three were wrong about the threshold (they said €22,000; it's €17,500 for distance selling from another EU country). None mentioned the new OSS scheme. I had to look it up myself.

When to use AI: Invoices, late-payment reminders, anything with a clear template and no interpretation. When you still need a human: Any question where the law has changed recently or where the answer depends on multiple variables.

Restaurant: bookings yes, complaints no

Booking confirmations are instant. Customer emails asking for a table for four on Friday at 20:00. Gemini Flash replies in 15 seconds with the confirmation, the address, and a line about calling if they're running late. I sent 60 of those in January.

Complaint replies are harder. A customer said their steak arrived cold. I asked Claude Sonnet 4.5 to write a reply. It apologised, explained what probably went wrong, and offered a €30 voucher. The tone was good. The problem: our manager's policy is €15 max for a single-dish complaint. AI didn't know that because I hadn't told it. If I'd sent that email, I'd have set a precedent I didn't want.

When to use AI: Confirmations, cancellations, thank-you emails. When you still need a human: Complaints, refund offers, anything where the response has a cost attached.

The GDPR question: which tasks are safe

Every task in the "yes" column is GDPR-safe if you follow basic rules: don't paste full customer records into a prompt, don't send personal data to a model that logs inputs for training (Kryotta doesn't, but check your tools), and don't use AI to make decisions about people (eligibility, creditworthiness, medical advice). Generating an invoice or a booking confirmation isn't a decision. Approving a refund or advising on VAT registration is.

If you're processing EU customer data, the safest tasks are the ones where AI is just formatting information you've already decided to share: confirmations, reminders, descriptions. The risky tasks are the ones where AI is interpreting policy or making a call. Those still need you.

Which model for which task

Claude Sonnet 4.5 is the best all-rounder for anything that needs tone and structure: emails, descriptions, invoices. It's expensive (around €0.80 per 100 messages), but it rarely needs editing.

Gemini Flash is faster and cheaper (about €0.15 per 100 messages) and works well for confirmations, reminders, and anything under 50 words. It's less reliable when the task has nuance.

DeepSeek V3 is the budget option (€0.02 per 100 messages). It handled lesson confirmations and late-payment reminders without trouble. I wouldn't use it for anything customer-facing that needs warmth.

Questions people ask

Can I use AI for GDPR-compliant customer support in the EU?
Yes, for confirmations, reminders, and templated replies. No, for decisions about refunds, complaints, or anything that interprets policy. Don't paste full customer records into prompts, and check that your AI tool doesn't log inputs for training.

Which tasks did AI finish without any edits?
Product descriptions, booking confirmations, appointment reminders, invoices, refund emails where the policy was clear. Anything with a template and no judgement call.

Which tasks still needed me to step in?
Returns (customer claims didn't match evidence), rescheduling (calendar conflicts), lease questions (contract clauses AI didn't see), VAT advice (outdated thresholds), complaint replies (offers that broke internal policy).

Do I need to check every AI output before sending it?
For the first 20–30 of any new task type, yes. After that, you'll know which tasks AI handles cleanly and which ones still trip it up. I now send booking confirmations without checking. I never send a complaint reply without reading it first.

I've been using Kryotta's workspace to test these tasks because I can switch models mid-flow and compare outputs without opening three browser tabs. If you're evaluating AI for your own repetitive tasks, the fastest way to know what works is to pick five real examples from last week and see which model finishes them without you. Start a free trial here and test it against your actual workload, not a demo.

K
Written by
Kryotta Team
Product & research

Kryotta is the multi-model AI workspace — every leading model, one login, one bill. Try it free →

Related reading