Myth one: you pay per question, so keep prompts short
Most Indian business owners I talk to believe AI works like SMS credits: one question, one charge. So they write "product description" instead of "Write a 150-word product description for a handloom cotton kurta in size M, indigo dye, block-printed, selling on Flipkart for ₹1,850, highlight the natural fabric and traditional craft." They think the longer prompt costs more.
It doesn't. You pay for tokens—roughly every four characters of input and output combined. A short, vague prompt often costs more because the AI generates a useless answer and you have to ask again. That second attempt doubles your spend.
I tested this with WhatsApp Business replies. I run a small gifting business in Pune. A customer asks, "Do you deliver to Nashik?" I tried two approaches on Claude Sonnet 4.5. First: "reply to customer." Output was generic, 80 tokens, didn't mention my delivery zone. I had to retry. Total: 160 tokens. Second attempt: "Reply to this WhatsApp message: 'Do you deliver to Nashik?' Say yes, we deliver across Maharashtra within 3–5 days via Delhivery, ₹80 flat shipping, free over ₹1,500. Friendly tone." Output was perfect, 95 tokens, one attempt. I saved tokens by being specific.
The idea that brevity saves money is backwards. A detailed prompt gives the model enough context to get it right the first time. That's cheaper.
Myth two: context window is how much the AI remembers between chats
I've heard this from at least a dozen Shopify store owners: "I have to remind the AI what we talked about yesterday because it forgets." They think context window means memory across sessions, like a conversation history that gets wiped overnight.
Context window is how much text the model can hold in a single request—input plus output. Claude Sonnet 4.5 has a 200,000-token window; Gemini Flash has 1 million. That's roughly 150,000 to 750,000 words. It's not about remembering yesterday's chat. It's about how much you can paste into one prompt right now.
This matters when you're working with long documents. A friend runs a CA firm in Surat. He needs to extract line items from 40-page GST reconciliation PDFs and match them against invoices. He was pasting five pages at a time, asking the AI to process them, then pasting the next five. He thought the model couldn't handle more. It can. I showed him how to paste the entire 40-page export into Gemini Pro (which has a bigger window than Claude for document work). One prompt, one output, accurate line-item extraction. He cut his processing time by two-thirds.
The limit isn't memory. It's how much you can fit into the current conversation. Once you understand that, you stop breaking tasks into tiny chunks that lose context halfway through.
Myth three: hallucinations mean the model is broken or lying
"Hallucination" sounds sinister. Business owners hear it and assume the AI is malfunctioning or deliberately making things up. I've had people ask if they should avoid AI entirely because "it hallucinates."
A hallucination is when the model generates text that sounds confident but isn't grounded in the input you gave it or in factual reality. It's not a bug. It's how the model works: it predicts the next most likely word based on patterns, not truth. If you ask it to write a product return policy and it invents a "7-day return window" when you never mentioned one, that's a hallucination. The model filled a gap with plausible-sounding text.
This happens most often when you ask for facts the model wasn't trained on or that require real-time data. I sell organic skincare on Amazon.in. I once asked Claude to "write a comparison of my moisturiser against Mamaearth's bestseller." It invented features Mamaearth doesn't have—"contains bakuchiol and squalane"—because I didn't give it the actual ingredient list. The output sounded authoritative. It was wrong.
The fix isn't to avoid AI. It's to ground your prompts. Paste the real ingredient list. Include your competitor's product page text. Give the model source material to work from. When I redid the prompt with both ingredient panels pasted in, the comparison was accurate.
Hallucinations aren't random. They're predictable. If you don't give the model enough factual input, it will guess. If you do, it won't.
Myth four: open-weight models are free, so they're worse
There's a belief that "free" AI models—what people call open-source, though the correct term is open-weight—are lower quality or risky because they're not from Google or Anthropic. I've heard store owners say they only trust "paid" models because "you get what you pay for."
Open-weight models like Llama 3.3 70B and DeepSeek V3 are often as good as closed models for most business tasks, and sometimes better. You're not paying per token when you run them yourself, but on a platform like Kryotta they're priced lower than Claude or Gemini because the compute cost is different. That doesn't mean they're worse. It means the economics are different.
I tested product descriptions for a jewellery seller in Jaipur. She sells silver jewellery on Meesho and needed 60 listings written. We ran the same prompt—"Write a 120-word description for a silver oxidised jhumka with mirror work, ₹450, traditional Rajasthani style"—through Claude Sonnet 4.5, Gemini Flash, and Llama 3.3 70B. Llama's output was the most vivid. It used "tribal motif" and "festive elegance" without being flowery. Claude was accurate but flat. Gemini added details she didn't ask for, like "perfect for weddings," which wasn't the vibe.
She used Llama for all 60 listings. Cost her a fraction of what Claude would have. The model being open-weight had nothing to do with quality. It's about fit. Some tasks need Claude's nuance. Some don't.
Myth five: if you upload a file, the AI reads every word
This one surprised me. A WooCommerce store owner in Bangalore told me he uploads his supplier's 80-page Excel price list every week and asks the AI to "update my product prices." He assumed the model reads the whole sheet, finds his SKUs, and adjusts the prices. It doesn't work that way.
When you upload a file, the model can access it, but it doesn't automatically parse every cell or remember every detail unless you tell it what to look for. If your prompt is vague—"update prices"—the model might skim the first few rows, assume a pattern, and miss half your catalogue.
I helped him rewrite the prompt: "This Excel file has 1,200 rows. Column A is SKU, Column D is new wholesale price in ₹. My WooCommerce store uses SKU codes that match Column A. Extract every SKU and its new price, then output a CSV I can import into WooCommerce with columns: SKU, new_price. Double-check that all 1,200 rows are included."
The output was complete. Before, he was getting 300–400 rows and manually filling in the rest. The file was always there; the instruction wasn't clear enough.
Uploading a document isn't the same as giving the model a task. You still have to say what you want done with it, in detail.
What actually affects your AI bill (and results)
Tokens are the only thing you pay for. Input tokens (what you send) plus output tokens (what the model generates). Longer, specific prompts often cost less per successful result because you avoid retries. Context window size doesn't cost extra; it's a limit, not a charge. The model you choose matters—Claude Sonnet costs more per token than Gemini Flash or Llama, but it's better at some tasks. Open-weight models aren't inferior; they're differently priced.
Hallucinations happen when you don't give the model enough factual grounding. Uploading a file doesn't mean the AI will magically know what to do with it.
The expensive mistakes aren't about tokens or model choice. They're about vague prompts, unrealistic expectations, and not understanding that AI predicts text—it doesn't think or remember or read your mind. Once you know how tokens, context, and grounding actually work, you stop wasting money on retries and start getting usable output on the first attempt.
I've watched sellers spend ₹15,000 a month on AI and get mediocre product descriptions because they thought "write a description" was enough. I've also watched a solo founder in Indore generate 200 Instagram captions for ₹180 because she wrote detailed prompts and used the right model. The difference isn't the tool. It's knowing what you're actually paying for.
Questions people ask
Do I get charged if the AI gives a wrong answer?
Yes. You pay for the tokens generated, whether the answer is useful or not. That's why clear, grounded prompts matter—they reduce the chance of paying for output you can't use.
Can I reuse the same prompt across different models without changing it?
You can, but results vary. Claude is better at nuanced tone, Gemini handles long documents well, Llama is fast and cheap for straightforward tasks. If one model's output isn't right, try the same prompt on another before rewriting it.
If I paste a 50-page PDF, does that cost more than a two-sentence question?
Yes, because input tokens are part of the total. But if that 50-page paste gives you a complete, accurate answer in one go, it's cheaper than ten vague attempts that each need a retry.
Are open-weight models safe for business use, or is there a legal risk?
Open-weight models like Llama and DeepSeek are licensed for commercial use. The "weight" (the model itself) is open, but you're still responsible for what you generate. No extra legal risk compared to closed models—just read the licence.
If you're tired of guessing why your AI outputs are expensive or useless, try writing one prompt with full context, specific instructions and the actual source material pasted in. Test it on two models and compare. You'll see the difference in one attempt. Start on Kryotta—Claude, Gemini, Llama, DeepSeek and more, all in one workspace, billed in rupees.



