Last month I sat down with three people on my vendor team—Priya handles customer queries on WhatsApp Business, Akash manages inventory across our Meesho and Flipkart stores, and Neha processes refunds—and tried to explain how the AI tools I'd been testing actually work. I'd spent two weeks using Claude and Gemini to draft replies to return requests, write product descriptions, and summarize daily sales reports. The models saved me maybe four hours a week, but every time I asked Priya to try one herself, she'd look at the screen, type three words, delete them, and go back to writing the message manually.
So I blocked out an afternoon. I brought my laptop, a notebook, and a plan to explain tokens, context windows, and hallucinations using examples from the work they already do. Half my analogies failed. The other half clicked immediately. And one grounding mistake—where I let the model invent a refund policy instead of pasting our actual terms—cost us ₹8,000 in wrong approvals before Neha caught it.
Here's what I tried, what worked, and the one thing I'd explain differently if I started over.
The restaurant analogy for tokens died in 90 seconds
I started with tokens. I told Priya that when she types a message into Claude, the model doesn't read words—it reads tokens, which are chunks of text about four characters long. A 100-word customer complaint might be 130 tokens. The model charges per token, so a longer message costs more.
I said it's like ordering food: each dish costs money, and if you order more dishes, the bill goes up.
Priya nodded politely. Akash checked his phone. Neha asked, "But why does it matter how long the message is if the answer is the same?"
Fair point. The restaurant analogy makes sense if you already understand that AI has a cost structure, but if you're used to typing into WhatsApp for free, the idea that length affects price feels arbitrary. I dropped it and tried again.
I opened Kryotta, pasted a 50-word return request from a customer in Surat, and showed them the token count at the bottom of the screen: 68 tokens. Then I pasted a 200-word complaint about a delayed dupatta set: 267 tokens. I said, "Longer input costs more because the model has to process more chunks. It's like data charges—send a 2 MB photo on WhatsApp versus a 10 MB video. Same idea."
That landed. Priya said, "So if I paste the whole chat history every time, it costs more?"
Yes. And that's why I don't paste 40 messages when the model only needs the last five.
UPI transaction limits made context windows obvious
Context windows were easier. I said every model has a limit to how much text it can hold in memory at once—like how UPI lets you send ₹1 lakh per transaction, but if you need to send ₹3 lakh, you split it into three payments.
Claude Sonnet 4.5 has a 200,000-token context window. Gemini Flash has 1,000,000. If you're summarizing a week of customer complaints, you need a bigger window. If you're writing one refund email, a smaller window is fine and cheaper.
Akash got it immediately. He said, "So if I want to upload our entire Meesho product catalog and ask the model to find duplicates, I need Gemini because the catalog is too big for Claude?"
Exactly. And if you're just asking the model to rewrite one product title, Claude is faster and costs less.
Neha asked whether the model "forgets" things if the conversation gets too long. I said yes—once you hit the token limit, the model drops the earliest messages and only keeps the recent ones. It's not like a person forgetting; it's more like a WhatsApp chat where you scroll up and the old messages aren't loaded yet.
She said, "So if I'm asking follow-up questions, I should keep the thread short?"
Right. Or paste the important part again if the model needs it.
GST invoice errors explained hallucinations better than anything
Hallucinations were harder. I said sometimes the model invents information that sounds correct but isn't—like when you ask it to summarize a return policy and it adds a clause that doesn't exist.
Priya said, "Why would it do that?"
I tried the "phone autocorrect" analogy: the model predicts the next word based on patterns it's seen before, and sometimes it predicts something plausible instead of something true. Like when your phone changes "Meesho" to "mesh" because it's never seen the word.
Blank stares.
I switched examples. I said, "Remember last Diwali when our accountant's software auto-filled a GST invoice and put the wrong HSN code on 40 dupatta sets? The code looked real, the format was right, but it was for cotton sarees, not synthetic dupattas. Same thing. The model fills in a detail that fits the pattern but isn't accurate."
Neha laughed. She said, "So it's just guessing?"
Sort of. It's predicting the most likely next word, and sometimes the most likely word is wrong.
I showed them a real example. I'd asked Claude to summarize our Meesho return policy, and it said customers could return items within 10 days of delivery. Our actual policy is 7 days. The model saw "10 days" in thousands of other return policies during training and assumed ours matched.
Akash said, "So how do we stop it?"
You paste the real policy into the prompt. You ground the model with facts.
The ₹8,000 mistake happened because I didn't ground the refund template
This is where I messed up. I'd built a prompt template for Neha to handle refund requests. The template said: "Read this customer message, check our return policy, and write a reply approving or denying the refund."
I didn't paste the return policy. I assumed the model knew it.
Neha used the template for two weeks. She processed maybe 60 requests. Then our Meesho dashboard showed we'd approved 12 refunds for orders older than 7 days—our cutoff. The total was ₹8,247.
I went back and checked the AI-generated replies. Every one said, "We're happy to process your refund within our 10-day return window." The model had hallucinated the policy, Neha trusted the output, and we ate the cost.
I rebuilt the template. Now it starts with: "Our return policy allows refunds within 7 days of delivery for items in original packaging. Read this customer message: [paste message]. Does it qualify? Write a reply."
Neha hasn't approved a wrong refund since.
What I'd explain differently: show the cost first, then the concept
If I ran this session again, I'd start with cost, not theory. I'd open Kryotta, run three prompts—one short, one long, one with a document attached—and show the token count and rupee cost for each. Then I'd say, "This is why length matters. This is why we don't paste everything. This is why I pick Gemini Flash for big jobs and Claude Sonnet for small ones."
Priya told me later that the UPI analogy stuck because she uses UPI 20 times a day. The GST invoice example worked because Neha had lived through that exact mistake. The restaurant thing failed because it wasn't her frame of reference.
The other thing I'd do: I'd show a hallucination live. I'd ask the model a question it can't answer—"What's our Flipkart commission rate for kurta sets?"—and let them watch it invent a number. Then I'd paste the real rate and show how the answer changes.
People trust AI less when they see it confidently wrong. That's useful.
We're three months in and the team uses two models daily
Priya now uses Gemini Flash to draft WhatsApp replies for return requests. She pastes the customer message, our return policy, and the order details, and the model writes a reply in under 10 seconds. She edits it, sends it, and moves to the next query. She's clearing 30 messages a day instead of 18.
Akash uses Claude Sonnet 4.5 to rewrite product titles when we add new inventory. He pastes the supplier's description, tells the model to make it Meesho-friendly (short, benefit-focused, with the fabric and size in the first line), and gets five variations in 15 seconds. He picks one, tweaks it, uploads it.
Neha still writes refund emails herself, but she uses the grounded template I built. She pastes the policy, the customer message, and the order date, and the model tells her yes or no. She's denied four refunds this month that she would have approved manually because the model caught the date issue.
I use Kryotta because I can switch between models without opening four tabs, and the token count shows up before I hit send. The team likes that they can see the cost in rupees, not dollars.
Questions people ask
Do I need to explain tokens to my team, or can I just say "don't paste long messages"?
You can skip the theory if they follow the rule, but I found that once Priya understood why length costs more, she started trimming her prompts herself. She'd delete the "Hi, I hope you're well" part and paste only the complaint. That saved tokens without me nagging her.
Which model should I use if I'm teaching someone AI for the first time?
Claude Sonnet 4.5. It's fast, it writes clearly, and it's hard to break. Gemini Flash is cheaper for long documents, but Claude is more predictable when you're learning.
How do I stop my team from trusting hallucinated answers?
Show them a wrong answer. Ask the model something you know it can't answer accurately, let them see it invent a plausible response, then paste the real information and show how the output changes. They'll stop assuming the first draft is correct.
What's the simplest way to ground a model so it doesn't hallucinate our policies?
Paste the policy into the prompt. Don't assume the model knows your return window, your pricing, or your shipping terms. If the fact matters, paste it.
The ₹8,000 mistake taught me that. Now every template my team uses starts with the ground truth, and we haven't had a wrong approval since. If you're running a vendor team and you're tired of writing the same WhatsApp replies 40 times a day, try Kryotta—it'll let you test Claude and Gemini side by side, and you'll see the token cost in rupees before you spend anything.



