I explained tokens, context windows and grounding to my Nairobi team — what finally made sense

A Nairobi team struggled to understand AI costs and capabilities until we compared tokens to M-Pesa records and context windows to WhatsApp chat history. Here's what finally clicked.

KKryotta TeamProduct & research · · 10 min read
Person at desk with laptop and phone in morning light, workspace with scattered notes and papers
Person at desk with laptop and phone in morning light, workspace with scattered notes and papers

Tokens are M-Pesa transaction records: you pay for what you send and receive

When I first told my Nairobi team we were going to use AI to draft customer replies and write product descriptions, the first question was "how much does it cost?" Fair. I said it depends on tokens. Blank stares.

Here's what worked: think of tokens like M-Pesa transaction records. When you send KSh 500 to a supplier, Safaricom logs the send, the receive, the confirmation SMS. You don't pay per word in the SMS—you pay for the whole transaction. Tokens work the same way. When you paste a product description into Claude and ask it to rewrite it in Swahili, the model counts every word you send (the English version, your instruction, any examples you include) and every word it sends back (the Swahili version, the explanation if you asked for one). Both directions cost tokens. A short prompt might use 200 tokens. A long one with three examples and a 400-word draft can hit 1,500 tokens before the model even replies.

Most models charge per million tokens. Claude Sonnet 4.5 costs about $3 per million input tokens and $15 per million output tokens. If you're generating 50 product descriptions a day at 800 tokens each (input plus output), you're using 40,000 tokens—about 12 US cents, or roughly KSh 15. Doesn't sound like much until you realise your team is pasting entire WhatsApp chat logs into the prompt because they think more context is always better.

One thing to do today: Open your AI tool's usage dashboard. Check how many tokens your last ten prompts used. If any prompt is over 2,000 tokens, you're probably including stuff the model doesn't need—old drafts, redundant instructions, full email threads when a two-line summary would work. Cut the input by half and see if the output quality drops. It usually doesn't.

Context windows are WhatsApp chat history: the model can only scroll back so far

Second question from the team: "If I'm halfway through a conversation with Claude and I paste in a new customer email, will it remember what we talked about before?" Sometimes yes, sometimes no. Depends on the context window.

I explained it like this: WhatsApp lets you scroll back through a chat, but if the chat is two years old and has 10,000 messages, your phone takes forever to load it and sometimes just gives up. AI models have the same problem. The context window is how much conversation history the model can "see" at once. Claude Sonnet 4.5 has a 200,000-token window. Gemini Pro has 2 million tokens. Llama 3.3 70B has 128,000 tokens.

Sounds huge, but it fills up faster than you think. Let's say you're using Claude to draft replies to Jumia seller messages. You paste the first message (150 tokens), Claude replies (300 tokens), you paste the second message (200 tokens), Claude replies again (350 tokens). After 20 messages, you've used maybe 15,000 tokens. Still fine. But if you're also pasting your entire product catalogue into every conversation because you want Claude to reference stock levels, and that catalogue is 50,000 tokens, you've now used 65,000 tokens in one session. You're a third of the way through the window, and Claude is starting to "forget" the early messages when it answers question 25.

One thing to do today: If you're working on a long task—writing a full website's worth of product pages, drafting a week of Instagram captions—start a new conversation every ten outputs. Don't try to do all 50 captions in one thread. The model's answers get vaguer and more repetitive once the window is 60 per cent full, even if you haven't hit the technical limit.

Hallucinations are boda directions without Google Maps: the model invents when it doesn't know

Third question, from my Instagram manager: "I asked Gemini to write a caption about our new honey face mask and it said the mask contains 'organic Kakamega honey certified by the Kenya Bureau of Standards.' We don't have KBS certification. Where did it get that?"

It made it up. That's a hallucination. The model doesn't know what's in your product, doesn't know if you have certification, but it knows that "organic Kakamega honey certified by KBS" sounds like something a Kenyan skincare brand would say. So it writes it, confidently, with no asterisk or disclaimer.

I told her: it's like asking a boda driver for directions to a shop you heard about but can't remember the name of. If the driver doesn't know the shop, some will admit it. Others will take you somewhere that sounds right—"there's a shop like that near the junction"—because they'd rather guess than say they don't know. AI models are the second type of driver. They're trained to complete the sentence, not to say "I don't have that information."

Hallucinations happen most often when you ask for specifics the model can't verify: prices, dates, certifications, customer testimonials, stock levels, legal claims. I've seen Gemini invent a "Nairobi SME Grant 2024" that doesn't exist. I've seen Claude quote a "University of Nairobi study on shea butter absorption rates" that no one can find. Both models wrote the fake information in the same confident tone they use for real facts.

One thing to do today: Never publish AI-generated content that includes numbers, claims or credentials without checking them yourself. If the model says your product is "dermatologically tested," make sure it is. If it says you offer "next-day delivery to all 47 counties," make sure you do. Treat every factual claim like a boda driver's guess until you've confirmed it.

Grounding is checking stock before promising delivery: give the model real data or it'll improvise

The fix for hallucinations is grounding. Grounding means giving the model the actual information it needs before you ask it to write. Instead of saying "write a product description for our honey face mask," you say "write a product description for our honey face mask. Ingredients: shea butter, raw honey from Baringo, vitamin E oil, beeswax. No certifications yet. Price: KSh 1,200 for 50ml. Ships within Nairobi same-day, outside Nairobi 2–3 days."

Now the model has the facts. It won't invent a KBS certification because you told it there isn't one. It won't promise next-day delivery to Kisumu because you said 2–3 days. Grounding is just pasting the truth into the prompt.

I do this for every task now. If I want Claude to draft a reply to a Jumia customer asking when their order ships, I paste the tracking number, the courier name and the current status from my dashboard. If I want Gemini to write an Instagram caption for a restocked product, I paste the product name, price, size and how many units I have left. Takes an extra 30 seconds. Cuts hallucinations by about 90 per cent.

One thing to do today: Pick one task you use AI for regularly—product descriptions, customer replies, social captions. Write a two-sentence grounding template with the facts the model needs every time. Paste that template into every prompt. For example: "Product: [name]. Price: KSh [amount]. Ingredients: [list]. Stock: [number] units. Ships: [timeline]." Fill in the brackets, then add your instruction below.

Open-weight models are boda boda, closed models are Uber: different trade-offs, same destination

Last question, from the guy who handles our chama's savings records: "You keep saying Claude, Gemini, Llama. What's the difference? Are some better?"

Not better—different. I explained it like this: closed models (Claude, Gemini, GPT-OSS) are like Uber. You don't know which driver you'll get, you don't see the route algorithm, but the app works the same way every time and Uber handles the insurance and safety checks. Open-weight models (Llama, Mistral, DeepSeek) are like boda boda. You can see the bike, you can negotiate the route, you can even buy your own bike and ride it yourself if you want. Less hand-holding, more control.

For most Nairobi SMEs, closed models are simpler. You sign up, paste your prompt, get an answer. You're not managing servers or worrying about which version of the model to download. Open-weight models are worth it if you're processing sensitive data (customer M-Pesa records, chama member details) and you don't want it leaving your laptop, or if you're running thousands of prompts a month and want to cut costs by hosting the model yourself. But that means learning to install and run the model, which is a weekend project, not a Tuesday morning task.

I use Claude Sonnet 4.5 for anything customer-facing (replies, captions, product descriptions) because it's polite and rarely goes off-script. I use Gemini Flash for bulk work (tagging 200 product photos, sorting emails into folders) because it's faster and cheaper. I've tested Llama 3.3 70B for internal drafts, and it's fine, but I haven't bothered hosting it myself because Kryotta lets me switch between models in one workspace without juggling subscriptions.

One thing to do today: If you're only using one model, try the same prompt in two others and compare the outputs. Paste a product description task into Claude, Gemini and Llama. See which one matches your brand voice. You'll waste KSh 5 in tokens and save yourself from paying for the wrong model all year.

Test with KSh 200 before you commit to a monthly plan

Here's what I wish I'd done in January: spent KSh 200 on tokens and tested every model on the three tasks I do most often (product descriptions, customer replies, Instagram captions) before signing up for anything. Instead, I paid for a ChatGPT Plus subscription (about KSh 2,600 a month) because everyone said it was the best, used it twice, then realised Claude was better for my tone and Gemini was faster for photo tagging.

Most AI platforms let you pay as you go. Kryotta charges per token with no monthly minimum, so you can test Claude Sonnet 4.5, Gemini Pro, Llama and DeepSeek for the price of a lunch in Westlands, see which one works, then use only that one until you need something different. You're not locked in, and you're not paying for models you don't touch.

One thing to do today: Write down the three tasks you'd use AI for this week. Draft one example of each task by hand (one product description, one customer reply, one Instagram caption). Then paste each task into two different models and compare the outputs to your hand-written version. Pick the model that needs the least editing. That's your starting point.

Questions people ask

Q: If I close the chat and come back tomorrow, does the model remember what we talked about?
No. Every conversation starts fresh unless the tool explicitly saves chat history (some do, some don't). If you need the model to reference something from yesterday, paste it into today's prompt.

Q: Can I use AI to reply to customer M-Pesa payment confirmations automatically?
Technically yes, but I wouldn't. AI can draft the reply, but you should still check it before sending—especially if the message includes a refund amount or a promise about shipping. One wrong number and you've created a dispute.

Q: Which model is cheapest for writing 100 product descriptions?
Gemini Flash or DeepSeek V3. Both cost a fraction of Claude per token and handle bulk tasks well. You'll edit more, but if you're doing 100 in one sitting, the time you save on cost is worth it.

Q: Do I need to understand tokens to use AI, or can I just ignore it?
You can ignore it until your bill is KSh 3,000 and you're not sure why. Takes five minutes to check your usage dashboard and see where the tokens went. Do it once, and you'll know whether you're pasting too much or using the wrong model for the job.

I'm not saying AI will run your business for you. It won't check stock, pack orders or answer the phone. But if you're spending two hours a day writing product descriptions or replying to the same Jumia questions, and you're not sure whether Claude or Gemini or "tokens" or "context windows" are worth your time, start with one task and KSh 200. See what happens. You'll know by lunchtime whether it's useful or just another subscription you'll forget to cancel.

If you want to test Claude, Gemini, Llama and DeepSeek in one place without juggling accounts, Kryotta lets you switch models mid-project and compare outputs side by side. No monthly minimum, pay only for what you use, and the token counter is right there in the dashboard so you'll know exactly what you're spending before you hit send.

K
Written by
Kryotta Team
Product & research

Kryotta is the multi-model AI workspace — every leading model, one login, one bill. Try it free →

Related reading