The week I opened Stripe and saw 112 open disputes
I run a small agency. We manage Shopify stores for seven D2C brands—mostly apparel, supplements and home goods—and we handle their customer service through a shared inbox. Normal weeks, we see two or three Stripe disputes. Someone claims they never got the package, or their kid used their card, or they ordered the wrong size and decided a chargeback was easier than an email. We respond, upload the tracking number, usually win.
Then Black Friday happened. Sales tripled. So did disputes. By the Monday after Cyber Monday, I had 112 open cases in the Stripe dashboard. Most were "fraudulent" or "product not received" claims. A few were "product unacceptable"—someone bought a medium, wanted a large, filed a dispute instead of clicking the return label we'd already sent. Stripe gives you seven days to respond to most disputes, and the clock was running on all of them.
I had two options: hire someone for a week, or see if AI could write responses that actually won cases. I'd used Claude Sonnet 4.5 for email drafts before, so I knew it could write clearly. But a dispute response isn't a newsletter. It has to be factual, specific, cite Stripe's evidence requirements, reference the exact transaction and customer communication, and not sound like a robot apologizing for being a robot. Get it wrong and you lose $40, $80, $150—plus the product, plus the Stripe dispute fee.
I decided to test three models on the first thirty disputes: Claude Sonnet 4.5, Gemini Pro and DeepSeek V3. I'd write the responses, submit them, track which ones Stripe ruled in our favor, and use the winner for the remaining eighty-two.
What I gave each model—and what I didn't
For every dispute, I copied five things into the chat:
- The dispute reason Stripe showed ("fraudulent", "product not received", etc.)
- The customer's original order email and our confirmation reply
- Tracking info—carrier, tracking number, delivery date and signature if we had it
- Any support emails between us and the customer before the dispute
- The product description from the Shopify listing
I did not give the models a template. I wanted to see what they'd write from scratch, because I didn't have time to maintain a template library for every dispute type. My prompt was the same for all three:
"Write a dispute response for Stripe. The customer claims [reason]. Use the evidence below. Be specific, cite dates and tracking numbers, stay factual, don't apologize unless we actually made a mistake. 150–250 words."
Then I pasted the five pieces of evidence and hit send.
Claude: accurate, fast, occasionally too polite
Claude Sonnet 4.5 wrote responses that read like a calm human wrote them in a hurry. It pulled the tracking number, cited the delivery date, quoted the customer's own email when they'd confirmed their address. It didn't editorialize. When the evidence was thin—one case where the tracking showed "delivered" but no signature—it acknowledged the gap and suggested we offer a partial refund to settle.
I used Claude for ten disputes. We won seven. The three we lost were cases where the evidence genuinely wasn't strong enough: a package marked delivered to a mailroom with no signature, a customer who claimed the product was damaged and we had no photo proof it left our warehouse in good condition, and one "product unacceptable" case where the customer said the fabric was polyester and our listing said cotton. (It was a blend. Our copywriter had been sloppy. Claude caught it and told me we'd probably lose. We did.)
The problem with Claude: it was polite even when the customer was lying. One dispute claimed "fraudulent charge"—but we had three support emails from that customer asking when the order would ship, confirming the address, then saying the product arrived and asking how to wash it. Claude wrote, "We understand the customer's concern, however our records show..." I deleted "We understand" and just led with the emails. Stripe sided with us, but I didn't need the empathy cushion when someone was committing fraud.
Claude cost me about $0.80 in total across those ten responses. Fast, cheap, mostly right.
Gemini: confident, specific, wrong twice
Gemini Pro was the most confident. It wrote responses that sounded like they came from someone who'd done this a hundred times. It cited Stripe's own dispute policies by name—"per Stripe's requirements for product not received claims, delivery confirmation to the customer's verified address constitutes valid evidence"—and it structured every response the same way: evidence summary up front, timeline in the middle, conclusion at the end.
I used Gemini for ten disputes. We won six, lost four.
Two of the losses were my fault—weak evidence, same as Claude. But two were Gemini's fault. In one case, it cited a tracking number that didn't exist. I'd pasted the wrong tracking line from my spreadsheet (my mistake), and instead of saying "I don't see a valid tracking number," Gemini just invented one that looked plausible: a USPS format with the right number of digits. I didn't catch it until I went to upload the proof and realized the number didn't pull up anything on the USPS site. I had to rewrite the response by hand, submit it late, and we lost.
In another case, Gemini said the customer had confirmed delivery in an email. I checked. The customer had asked when it would be delivered, not confirmed it arrived. Gemini misread the thread. I caught that one before submitting, but it made me nervous. If I'm checking every fact the model claims, I'm not saving time.
Gemini was fast and cheap—maybe $0.60 for ten responses—but I couldn't trust it unsupervised.
DeepSeek: slow, blunt, and it won the most
DeepSeek V3 took longer to respond—twenty seconds instead of five—but it wrote the most effective disputes. No fluff, no empathy padding, just evidence and dates. When a customer claimed they never received a package that tracking showed delivered and signed for, DeepSeek wrote: "Tracking confirms delivery on November 28 at 2:14 PM to the address provided at checkout: [address]. Signature on file. The customer did not contact us before filing this dispute."
Blunt, but it worked. I used DeepSeek for ten disputes and won eight. The two losses were the same weak-evidence cases Claude and Gemini lost—no model can win a dispute when you don't have proof.
DeepSeek also caught things the other models missed. In one case, the customer claimed the product was "not as described." I'd pasted the product listing, which said "100% cotton," and the customer's email, which said "this feels like polyester." DeepSeek wrote: "The product listing states 100% cotton. The customer has not returned the product for verification. We offered a return label on November 30; the customer did not respond." Then it added a note in the chat: "If the product is actually a blend, you should settle this dispute and update the listing."
It was right. I checked the supplier spec sheet. It was 80/20 cotton-poly. I settled, updated the listing, and avoided more disputes on that product. Gemini and Claude had both written responses defending the "100% cotton" claim without questioning it.
DeepSeek cost about $0.50 for ten responses. Slower than Claude, but I won more.
What I did for the remaining 82 disputes
I used DeepSeek for the next seventy disputes and Claude for twelve cases where I wanted a softer tone—mostly "product unacceptable" disputes where the customer had a semi-reasonable complaint and I thought a polite response might get Stripe to split the difference.
DeepSeek won 61 of those seventy. Claude won eight of twelve. Combined, I won 76 of the original 112 disputes—a 68% win rate. Our normal win rate is around 60%, so this wasn't a miracle, but I wrote 100 responses in four days instead of two weeks, and I'm confident the responses were better than what I'd have written by hand at 11 PM on a Wednesday.
I also learned that AI doesn't fix bad evidence. If you don't have tracking, or you shipped to the wrong address, or your product listing lies, no model will save you. The wins came from cases where we had the proof and I just needed to format it clearly and submit it fast.
The mistakes I made—and the one rule I'd follow next time
I trusted Gemini's fake tracking number because it looked real. I didn't check the USPS site until I went to upload the screenshot. Now I verify every tracking number and every factual claim before I submit, even if the model sounds confident. Confidence is not accuracy.
I also wasted time switching between models. By dispute thirty, I knew DeepSeek was winning more, but I kept testing Claude and Gemini because I wanted a clean sample size. If I did this again, I'd test five of each model, pick the winner, and use it for everything.
One rule I'd follow: always paste the customer's exact words from their dispute reason. Stripe shows you a text box where the customer explains why they're disputing. I didn't paste that into my prompts at first—I just used Stripe's category label ("fraudulent," "product not received"). Big mistake. When I started including the customer's own explanation, the models wrote responses that directly rebutted their claims instead of giving generic answers. That alone probably added ten wins.
The setup I'm using now
I keep a Kryotta workspace open with one chat for DeepSeek V3. When a dispute comes in, I paste the five pieces of evidence and the customer's dispute reason, add the prompt, get the response in twenty seconds, read it once to check for hallucinations, copy it into Stripe. If the evidence is weak—no signature, no photo proof, customer has a reasonable complaint—I'll ask DeepSeek, "Do we have a strong case here?" It's been right every time.
I'm not using Gemini anymore. I'm keeping Claude around for the rare case where I want a softer tone or I need to write a settlement offer instead of a dispute response, but 90% of the time it's DeepSeek.
Total cost for 100 disputes: maybe $6. Time saved: about thirty hours. Win rate: slightly better than I'd get writing by hand, and way better than I'd get if I rushed through 112 responses in a panic.
Questions people ask
Can I use the same model for pre-dispute customer service emails?
Yes, but you'll want to adjust the tone. DeepSeek's bluntness works for disputes because Stripe wants facts, not apologies. For customer service, Claude's warmer voice usually lands better. I use Claude for "we're sorry, here's a return label" emails and DeepSeek for "here's the tracking proof you asked for" emails.
What if the model invents a fact I don't catch?
You lose the dispute, and possibly the customer's trust if Stripe shares your response with them. I check every tracking number, every date, every claim about what the customer said. It takes two minutes. Skipping that check cost me one dispute and taught me not to skip it again.
Do I need to write a different prompt for each dispute type?
Not really. I use the same prompt for "fraudulent" and "product not received" disputes—the evidence is what changes. For "product unacceptable" cases, I'll add "acknowledge the customer's concern if it's reasonable, but explain our return policy" to the prompt. That's it.
Does Stripe know I'm using AI?
They don't ask, and the responses don't sound like AI if you're using a decent model and checking the output. Stripe cares whether your evidence is valid, not whether you typed it yourself or had a model format it.
I'm not going back to writing these by hand. If you're sitting on a pile of Stripe disputes and you're not sure where to start, open Kryotta, pick DeepSeek or Claude, paste your evidence, and see what you get. You'll know in five minutes whether it's faster than doing it yourself—and you'll probably win more disputes than you expect.



