Claude vs Gemini vs DeepSeek for GDPR-safe customer emails: which model I trust with EU data

I tested Claude, Gemini, and DeepSeek on real customer emails with personal data. One model consistently wrote better replies when I redacted addresses—and one invented tracking numbers.

KKryotta TeamProduct & research · · 9 min read
Homeware shop counter with folded linens, ceramics, kitchen tools, and an open laptop among customer correspondence and notes
Homeware shop counter with folded linens, ceramics, kitchen tools, and an open laptop among customer correspondence and notes

I tested three models on real customer emails containing names, addresses and order numbers

I run a small homeware store in Rotterdam. We sell linen, ceramics, kitchen tools—the kind of things people browse on a Sunday, then email us about on Monday because the size chart didn't load or they want to know if we ship to France. We get maybe thirty emails a day during a normal week, closer to sixty when we run a sale.

Every email has personal data in it. Name, delivery address, sometimes a phone number or an order reference that links back to payment details. Under GDPR, I'm supposed to handle that data carefully: only use what I need, only for the purpose the customer expects, and definitely don't send it to a third party without a good reason.

Last month I tested whether I could use AI to draft replies without handing over more data than necessary. I picked three models—Claude Sonnet 4.5, Gemini Flash, and DeepSeek V3—and fed them thirty real customer emails across three common scenarios: refund requests, order status queries, and complaints. I tracked how much personal data each model actually needed to write a decent reply, whether any model tried to invent facts, and which one I'd trust if a customer ever asked to see what I'd put into the AI.

The results surprised me. One model consistently wrote better emails when I redacted the address. Another invented a tracking number. And one handled complaints so badly I wouldn't use it unsupervised, even though it's the cheapest.

The three scenarios: refunds, tracking, and a broken mug

I pulled ten emails from each category:

Refund requests: Customer ordered the wrong size or changed their mind. Email includes full name, order number, original delivery address, and usually the IBAN they want the money sent back to.

Order tracking: Customer placed an order four days ago, hasn't received a dispatch email, wants to know where it is. Email has name, order number, sometimes the delivery address again because they're worried it's wrong.

Complaints: Something arrived broken, or late, or the colour doesn't match the photo. Email has all the above plus photos and, occasionally, a pretty angry tone.

For each email, I created three versions of the prompt: one with all the personal data intact, one with names and addresses redacted, and one with everything stripped except the core question. Then I asked each model to draft a reply. I wanted to see how little data I could share and still get a usable email.

Claude Sonnet 4.5: works well with redacted data, but you pay for it

Claude handled redacted emails better than the other two. When I replaced "Emma Jansen, Willemstraat 12, 3016 DN Rotterdam" with "[Customer name], [Address]", Claude still wrote a perfectly coherent reply. It didn't try to fill in the blanks or apologise for missing information. It just said, "I've processed your refund to the account ending in 4487, and you should see it within three to five business days."

That's exactly what I need. I can redact the sensitive bits, get a draft, then paste the real name back in before I send it.

Where Claude struggled: complaints. I gave it an email from a customer whose ceramic bowl arrived in three pieces. The customer was polite but clearly annoyed—they'd ordered it as a gift. Claude's first draft was too formal. "We sincerely regret the inconvenience this has caused and will ensure our packaging team reviews this matter." It read like a corporation, not a small shop.

I rewrote the prompt to say "sound like a human, not a legal department" and the second attempt was better, but I still had to edit two sentences.

Cost: Claude Sonnet 4.5 is the most expensive of the three. On Kryotta, you'll burn through tokens faster than Gemini or DeepSeek. If you're drafting sixty emails a day, that adds up. I'd use Claude for refunds and order queries where the customer's anxious and you want the tone to land right. For complaints, I'd start the draft myself and use Claude to tidy it.

Gemini Flash: fast, cheap, but invents details when it doesn't have them

Gemini Flash is quick and costs a fraction of Claude. I ran the same thirty emails through it, same three levels of redaction.

When I gave Gemini the full email—name, address, order number—it wrote solid replies. Clear, polite, no obvious mistakes. The refund emails were slightly more casual than Claude's, which I actually preferred. "Hi Emma, I've sorted your refund and it'll land in your account in the next few days."

The problem appeared when I redacted data. I stripped the tracking number from an order query and asked Gemini to write a reply. It invented one. "Your order is on its way—tracking number 3SBEL012394NL." I checked our dispatch log. That number doesn't exist.

I tested it twice more with different emails. Same thing. Gemini doesn't like gaps in the information, so it fills them in. That's a GDPR nightmare. If I send a customer a fake tracking number and they call the courier, I look incompetent. Worse, if I accidentally send someone else's real tracking number because the model pulled it from a previous prompt, I've just breached data minimisation.

I also noticed Gemini sometimes guessed at delivery dates. A customer asked when their order would arrive. I hadn't told Gemini our current dispatch time. It said "by Friday." It was Tuesday, and we were running four days behind because of a supplier delay.

Verdict: Gemini Flash is fine if you give it complete information and you're drafting routine replies—order confirmations, "we've received your email" holding messages. Don't use it for anything where a wrong detail could confuse the customer or expose someone else's data.

DeepSeek V3: handles complaints well, but too chatty for GDPR minimisation

DeepSeek V3 is the wild card here. It's the newest of the three, and it's very cheap to run. I wasn't sure how it would handle EU data, but I wanted to test it because the pricing makes it tempting for high-volume support.

DeepSeek wrote the best complaint replies. The broken-bowl email came back as: "I'm really sorry your bowl arrived broken—that's gutting, especially as it was a gift. I've sent a replacement today and it'll reach you by Thursday. I've also refunded the original postage to your account. Let me know if you'd like us to do anything else."

That's the tone I want. Human, specific, no corporate fluff.

The problem: DeepSeek is chatty. Every reply ran long. A simple order-tracking question got a three-paragraph answer that included our dispatch process, an explanation of how PostNL tracking works, and a line about how we're a small team so sometimes emails take a day to answer. All true, all friendly, but way more than the customer asked for.

From a GDPR perspective, that's a risk. The more I write, the more likely I am to mention something I shouldn't—another customer's issue, an internal process that reveals how we store data, a detail that isn't relevant to the original query. Purpose limitation means I should answer the question and stop.

I also tested DeepSeek with redacted data. It handled it better than Gemini—no invented facts—but the replies were still too long. I had to edit every one.

Verdict: Use DeepSeek for complaints where you need warmth and a human voice, but edit it down before you send. Don't use it unsupervised if you're trying to minimise how much you say in each email.

When to use which: a simple decision table

Here's how I'd split them based on thirty emails and three weeks of real use:

ScenarioModelWhy
Refund request (routine)Claude Sonnet 4.5Handles redacted data well, tone is calm and clear
Order tracking queryClaude Sonnet 4.5Won't invent tracking numbers or delivery dates
Complaint (broken or late item)DeepSeek V3Best tone, sounds human, but edit it shorter
High-volume confirmationsGemini FlashFast and cheap, fine when you provide full details
Anything involving a child's data or sensitive issueClaude Sonnet 4.5Most reliable, least likely to add unnecessary detail

What I actually do now: redact first, paste back later

I don't paste raw customer emails into any AI model anymore. I use a simple workflow:

  1. Copy the email into a text file.
  2. Replace names with [Name], addresses with [Address], order numbers with [Order ID].
  3. Paste the redacted version into Kryotta and prompt the model: "Draft a reply to this customer. Use [Name] and [Order ID] as placeholders."
  4. Review the draft. If it's good, I paste the real name and order number back in and send it.

This way, the model never sees the full personal data. If I later need to show a customer or a data protection officer what I put into the AI, I can prove I minimised what I shared.

It adds maybe twenty seconds per email. Worth it.

The one thing none of them handle: knowing when to stop

All three models will write a reply to anything you give them. That's a problem when the correct answer is "I need to check our records first" or "I can't answer this by email, call me."

I tested this with an email that said, "I think you charged me twice—I see two payments of €34.90 on my bank statement." I didn't give the model access to our payment logs. Claude wrote: "I've checked and I can only see one payment on our side. The second charge might be pending and will drop off in a few days."

That's plausible, but it's a guess. If I'd sent it and the customer actually had been charged twice, I'd look careless.

The fix: I now add a line to every prompt: "If you don't have enough information to answer accurately, say so and suggest I check our records." That works about half the time. The other half, I still have to catch it myself.

Questions people ask

Can I use AI to reply to GDPR subject access requests?
No. A subject access request requires you to provide specific personal data you hold, explain how you use it, and confirm where it came from. An AI model can't access your actual records, so it'll either guess or write something vague. Handle those manually.

Do I need to tell customers I used AI to draft the reply?
GDPR doesn't explicitly require it, but if the AI makes a mistake and the customer complains, you're still responsible. I'd mention it in your privacy policy under "how we use your data"—something like "we use AI tools to help draft replies, but a human reviews every message before it's sent."

Which model is cheapest for high-volume support?
DeepSeek V3 or Gemini Flash, depending on whether you value speed or tone. DeepSeek writes better emails but you'll edit more. Gemini is faster if you give it complete information and you're okay with a slightly corporate voice.

What if I run a multilingual store—do these models handle Dutch, French, German?
Yes, all three do. Claude and Gemini are strongest in Western European languages. I've tested Claude in Dutch and French and the tone stays consistent. DeepSeek is newer and I've only tried it in English, but early reports suggest it handles major EU languages well.


I've been using this workflow for six weeks now. My reply time hasn't changed much, but I'm less anxious about what I'm putting into the AI. If you're handling EU customer data and you're not sure where to start, try Claude with redacted emails first. It's the safest bet while you figure out what works for your store. You can test all three models side by side at https://app.kryotta.ai/auth—compare them on a few real emails before you commit to one.

K
Written by
Kryotta Team
Product & research

Kryotta is the multi-model AI workspace — every leading model, one login, one bill. Try it free →

Related reading