I used 6 prompt templates for 200 customer emails, invoices and Slack replies: which formula per task

I tested six prompt templates across 200 business tasks—customer emails, invoices, code reviews—and discovered that task type matters more than template elegance. Here's which formula works for what.

KKryotta TeamProduct & research · · 11 min read
Workspace showing laptop, notebook with templates and checkmarks, coffee mug, scattered invoice pages and chat message notes under warm desk lamp
Workspace showing laptop, notebook with templates and checkmarks, coffee mug, scattered invoice pages and chat message notes under warm desk lamp

Does a single prompt template really work across customer emails, invoices and code reviews?

No. I tested six templates across 200 business tasks—GDPR-compliant customer replies, multi-currency invoicing, Slack code reviews, WhatsApp order confirmations—and the same structure that made invoices clear turned customer emails into legal documents nobody wanted to read.

The promise sounds good: write one reusable prompt, plug in variables, get consistent output. I believed it until I watched a template that nailed invoice formatting ("You are a finance assistant; output must include VAT breakdown, SEPA payment details, and due date in bold") turn a simple "Where's my refund?" reply into three paragraphs of policy text that made the customer angrier.

I kept a log. Six templates, each tested across thirty to forty tasks in its category. I tracked which structure worked, which needed a human edit before sending, and which broke down so badly I rewrote the whole thing. The pattern that emerged: task type matters more than template elegance. A role + context + constraint formula works brilliantly for structured output (invoices, reports, technical comments) and fails for anything that needs warmth or judgement (apologies, negotiations, anything GDPR-sensitive where tone is half the compliance).

Here's what I learned, template by template, with the exact prompts I used and the moments they stopped working.

Template one: the invoice formatter (role + output structure + compliance constraint)

Prompt: "You are a billing assistant for a Shopify store selling outdoor gear across the EU. Generate an invoice for [order details]. Output must include: line items with VAT per item, total in euros, SEPA payment details, due date 14 days from invoice date, GDPR footer stating how we store payment data. Format as plain text, readable in email."

What it handled well: Thirty-five invoices for orders between €40 and €850. Every one came back with clean line items, correct VAT (I spot-checked the percentages for Germany, France, Poland), and the SEPA IBAN formatted properly. The GDPR footer was generic but accurate: "We store your payment details in accordance with GDPR Article 6(1)(b) for contract performance. Data is held for seven years per tax law, then deleted."

Where it broke: Variable VAT rates across countries. The template assumed one rate. When an order shipped to a customer in Portugal (23% VAT) versus one in Luxembourg (17%), I had to feed the correct rate manually or the model defaulted to a made-up average. I added "[VAT rate: X%]" to the prompt and that fixed it, but it's not truly automatic anymore.

Model used: Claude Sonnet 4.5. Gemini Flash sometimes dropped the GDPR footer or merged line items into a paragraph. Claude kept the structure every time.

Template two: the GDPR customer support reply (role + empathy instruction + data-handling constraint)

Prompt: "You are a customer support agent for a Dutch online bookshop. A customer asks: [question]. Reply in a warm, clear tone. If the question involves personal data (orders, payment, address), remind them we handle data per GDPR and offer to verify identity before sharing details. Keep it under 80 words."

What it handled well: Forty-three routine questions—"Where's my order?", "Can I change my delivery address?", "Do you ship to Belgium?" The tone was friendly, the GDPR reminder felt natural ("I can look that up for you—just to keep your account safe, can you confirm the email on your order?"), and I only edited two replies for phrasing.

Where it broke: Anything requiring judgement. A customer wrote, "You sent me the wrong book and now it's out of stock, I want a refund AND compensation for my time." The template gave me a polite, policy-perfect reply that would have made the customer escalate to Twitter. I rewrote it by hand, apologised properly, offered a €10 voucher, and moved on. The template can't read anger or decide when to bend a rule.

Model used: Claude Sonnet 4.5 again. I tried Gemini Flash for ten of these and it leaned too formal—"We regret the inconvenience" instead of "I'm really sorry about that." DeepSeek V3 was closer to Claude but occasionally invented a policy that didn't exist.

Template three: the code review comment (role + context + tone constraint)

Prompt: "You are a senior developer reviewing a pull request. The code is [description]. Write a comment that: explains the issue clearly, suggests a fix, stays constructive (no 'this is wrong' phrasing), and keeps it under 60 words. Assume the developer is competent but missed something."

What it handled well: Thirty-eight code reviews for a team building a React app for a Berlin-based logistics startup. The comments were clear, specific, and nobody felt attacked. Example: a junior dev wrote a function that didn't handle null values. The model wrote: "This will throw if shipmentData is null. You could add a guard clause at the top—if (!shipmentData) return null;—or default it in the function signature. Either way works; just needs a safety check before the map." Perfect.

Where it broke: Nuance. When a mid-level dev made an architectural choice I disagreed with (using local state instead of a context provider), the template gave a technically correct comment that implied their approach was a mistake. It wasn't—it was a trade-off. I rewrote it to say, "This works, though if we add more components that need this data, a context might save us some prop drilling. Worth discussing in standup?" The template doesn't do "both approaches are valid, here's the trade-off."

Model used: Claude Sonnet 4.5 for twenty-five, Gemini Flash for thirteen. Gemini was faster and almost as good, but Claude's tone was slightly warmer. DeepSeek V3 was too terse—it read like a linter error, not a human comment.

Template four: the WhatsApp order confirmation (role + key details + brevity constraint)

Prompt: "You are the owner of a small bakery in Lisbon taking orders via WhatsApp. A customer ordered: [items, total, delivery time]. Write a confirmation message that includes the total in euros, delivery time, and sounds like a quick text from a real person. Under 40 words."

What it handled well: Fifty-one confirmations. Every one felt human. "Got it—2 pastel de nata boxes and 1 bolo de arroz, €12.50 total. I'll have them to you by 3pm today. Thanks!" No customer replied asking if I got the order, which was the whole point.

Where it broke: It didn't. This template had a 100% success rate because the task is narrow and the tone is easy. The only edit I made regularly was swapping "I'll have them to you" for "I'll drop them off" when I knew the address, but that's a nitpick.

Model used: Gemini Flash. It's faster than Claude for short-form replies, and the quality difference is invisible at this length. I tested ten with Claude anyway—no meaningful difference. DeepSeek V3 was fine too, though it occasionally added an emoji I didn't ask for.

Template five: the multilingual product description (role + SEO constraint + translation note)

Prompt: "You are a copywriter for a Polish furniture shop selling on Allegro. Write a product description for [item]. Include key features, dimensions, material, and why it fits small apartments. Output in Polish. Keep it under 100 words. Natural tone, not a spec sheet."

What it handled well: Twenty-nine descriptions for chairs, shelves, and tables. The Polish was fluent (I had a native speaker check five of them), the tone was conversational, and the "small apartment" angle came through clearly. Example for a folding desk: "Biurko składane, idealne do małych mieszkań—rozłożysz je, gdy pracujesz, złożysz, gdy potrzebujesz miejsca. 80 cm szerokości, blat z litego dębu, nogi metalowe. Zmieści laptopa, notes i kawę. Wytrzyma 30 kg." (Translation: "Folding desk, perfect for small apartments—unfold it when you work, fold it when you need space. 80 cm wide, solid oak top, metal legs. Fits a laptop, notebook, and coffee. Holds 30 kg.")

Where it broke: Cultural nuance. A description for a minimalist lamp came back fine in Polish but used phrasing that felt more German than Polish—technically correct, slightly off. My native-speaking friend said, "It's like Google Translate wrote it, then a human fixed the grammar but not the vibe." I adjusted two words and it felt local again.

Model used: Gemini Flash. Claude Sonnet 4.5 was equally fluent but slower. I didn't test DeepSeek V3 here because I'd already hit my pattern: Gemini for speed on straightforward tasks, Claude when tone matters.

Template six: the Slack project update (role + audience + length constraint)

Prompt: "You are a project manager updating the team on [project]. The audience is developers, a designer, and the founder. Summarise progress, blockers, and next steps. Tone: clear, not corporate. Under 120 words."

What it handled well: Fifteen updates over three weeks for a SaaS project (a booking tool for coworking spaces in Amsterdam). The summaries were tight, the blockers were honest, and nobody replied asking for clarification. Example: "We shipped the calendar sync this morning—Google and Outlook both working. Still stuck on the Mollie payment integration; their API docs are missing the webhook signature bit, so I've emailed support. Design for the admin dashboard is in Figma, ready for review. Next: finish Mollie, start on user roles, and I'll need 30 minutes with everyone Thursday to lock down the pricing tiers."

Where it broke: It didn't, but it came close once. A week where nothing shipped, the model wrote, "No progress this week due to blockers." That's true but demoralising. I rewrote it: "Mollie integration took longer than expected—still waiting on their support team. Used the time to refactor the booking logic, which'll make the next three features faster." Same facts, better morale.

Model used: Claude Sonnet 4.5. I tried Gemini Flash for three updates and they were fine but slightly more formal—"We have completed" instead of "We shipped." For team communication, Claude's tone was worth the extra second of latency.

What I'd do differently if I ran this test again

I'd write two templates per task type: one for the routine 80% (invoice, order confirmation, simple code review) and one for the 20% that needs judgement (apology, architectural feedback, bad-news update). The single-template dream doesn't hold. I'd also stop trying to make one model do everything. Gemini Flash for speed, Claude Sonnet for tone, and I'd skip DeepSeek V3 for anything customer-facing until it stops inventing details.

I'd add a "check if this needs a human" step to every template. A simple rule: if the task involves an apology, a rule exception, or a decision that affects money or relationships, flag it for review. Ten seconds of reading beats one angry reply.

And I'd test every template in its actual channel. A prompt that works in Kryotta's editor might break when you paste the output into WhatsApp and the formatting vanishes, or into Slack where the tone reads differently under your avatar. I caught this with the project update template—looked great in my notes, felt slightly off in #general until I shortened two sentences.

Questions people ask

Can I use the same prompt template for emails in English and German?
Not without editing. I tested the customer support template in both languages (I speak enough German to spot problems). The structure works, but idioms don't translate—"I'll look into that" became "Ich werde das untersuchen," which is correct but sounds like a police investigation. You need a native speaker to adjust the phrasing, or you risk sounding like a bot. The role + constraint format is portable; the actual words aren't.

Do I need to mention GDPR in every prompt, or does the model know the rules?
Mention it. I ran five customer support prompts without the GDPR reminder and got replies that either skipped the data-handling note or made up a policy. ("We delete your data after 30 days" when our actual retention is seven years for tax purposes.) The models know GDPR exists but they don't know your specific compliance setup. Spell it out in the constraint: "Remind customers we handle data per GDPR Article 6(1)(b), stored for seven years, then deleted."

Which template structure works best for invoices with multiple VAT rates?
Role + output structure + a variable for VAT rate. My original template assumed one rate and broke when orders crossed borders. I added "[VAT rate: X%]" as a required input and the problem disappeared. If you sell across the EU and VAT varies by country, feed the rate manually or connect your prompt to a lookup table that maps country codes to rates. The model won't guess correctly.

How do I stop AI replies from sounding too formal in WhatsApp?
Add "sounds like a quick text from a real person" to the prompt, and set a tight word limit (under 40 words). I tested this with the bakery confirmation template—no limit, I got three-sentence paragraphs. With the 40-word cap, I got one or two sentences that felt like a human typed them on a phone. Brevity forces informality.

I've been running these six templates for a month now, editing maybe 15% of the output before I send it. The time saved is real, but so is the need to read every reply before it goes out. If you're testing prompt templates for your own business tasks, start with the structured stuff (invoices, reports, code comments) where the model can't do much damage, then move to customer-facing work once you've learned where your template breaks down. I built all six of these prompts in Kryotta, switching between Claude and Gemini depending on the task, and the compare feature helped me spot which model kept the tone I wanted versus which one drifted formal.

K
Written by
Kryotta Team
Product & research

Kryotta is the multi-model AI workspace — every leading model, one login, one bill. Try it free →

Related reading