I used 5 explainer prompts to teach my team tokens and context windows: which one clicked

I built five AI prompts to teach my team about tokens and context windows. Three worked perfectly. Here's which explanation finally made the concepts click—and what I'd do differently.

KKryotta TeamProduct & research · · 9 min read
Team of four people gathered around a laptop, reviewing content together in a bright office with natural lighting
Team of four people gathered around a laptop, reviewing content together in a bright office with natural lighting

I opened with "tokens are like text message character limits" and lost half the room

I manage a four-person team—two freelance writers, one Shopify VA, one email marketer—and we'd been using AI for maybe three weeks when I realized nobody understood why I kept switching models. They'd open Kryotta, see Claude Sonnet 4.5, Gemini Flash, and Llama 3.3 70B in the dropdown, pick whichever name sounded coolest, and paste in a 3,000-word product brief without checking if the model could even handle it. One writer burned through $4.80 in tokens generating a blog post that Gemini Flash would've written for $0.12. The Shopify VA kept asking Claude to "remember" details from a conversation two days earlier, then got frustrated when it forgot everything.

So I scheduled a 90-minute onboarding call. I built five prompt templates—short explanations I could paste into the chat and have the model explain itself—because I figured if the AI could teach the concepts, the team would trust it more than they'd trust me lecturing. Three of the five prompts worked. One confused everyone. And one accidentally taught my email marketer why hallucinations matter by letting the model invent a Stripe dispute policy that would've cost us $1,200 if she'd used it.

Here's each prompt, what I was trying to explain, and which one finally made the context window idea stick.

Prompt one: I asked Claude to explain tokens using their own work

I started with tokens because that's what shows up on the bill. I opened a chat in Kryotta, pasted this, and screen-shared the result:

"Explain what a token is to someone who writes product descriptions for a Shopify store. Use an example of a 150-word description they just wrote. Tell them how many tokens that is, why it matters for cost, and why some models charge more per token than others."

Claude came back with a clear answer: a 150-word description is roughly 200 tokens, one token is about four characters, and different models charge different rates because they're trained on different amounts of data. It compared Claude Sonnet 4.5 (more expensive, better at nuance) to Llama 3.3 70B (cheaper, faster, good enough for straightforward tasks). The two writers nodded. The Shopify VA asked if she could just use the cheapest model for everything.

That's when I realized the prompt worked to explain tokens to team members, but it didn't explain why you'd pay more. So I added a follow-up: I asked Claude to rewrite the same 150-word description twice—once as Claude, once as Llama—and show the difference. The Claude version caught a tone shift between the headline and body copy. The Llama version was fine but flat. The VA got it immediately.

Cost for that explanation: $0.02 in tokens. Time saved in future billing questions: hours.

Prompt two: I had the model calculate its own context window using a Black Friday email batch

Next I needed to explain context windows, because my email marketer kept splitting tasks across three separate chats when one long conversation would've been faster. I pasted this into a fresh Claude chat:

"I'm going to give you 40 customer emails from a Black Friday sale—complaints, refund requests, questions about shipping. Each email is 80–120 words. Calculate how many tokens that is, explain what a context window is, and tell me whether you can read all 40 emails in one conversation or if I need to split them into smaller batches."

I didn't actually paste 40 emails—I pasted eight, because I wanted to see if the model would do the math. Claude estimated 4,000–6,000 tokens for 40 emails, explained that its context window is 200,000 tokens (so yes, it could handle all 40 in one chat), and added that keeping everything in one thread means it remembers details from earlier emails when it writes replies to later ones.

My email marketer stared at the screen. Then she said, "Wait, so if I paste all the Cyber Monday complaints into one chat, it'll notice if three people are angry about the same shipping delay?" Yes. Exactly.

That prompt worked because it used her actual work. The AI context window explained concept didn't land until she saw a number—200,000 tokens—and compared it to the size of a real task she does every week. I should've led with this one.

Prompt three: I asked the model to hallucinate on purpose, then caught it

Hallucinations were harder. I couldn't just define the term, because everyone assumed "hallucination" meant the model would write obvious nonsense. They didn't realize it would invent plausible-sounding policies, prices, or deadlines that were completely wrong.

So I set a trap. I opened a chat with Gemini Flash and pasted this:

"A customer bought a $45 candle from our Etsy shop and wants a refund after 60 days. Write a reply explaining our refund policy."

I didn't give Gemini our actual refund policy. I wanted to see what it would invent. It came back with a polite email saying we accept returns within 90 days if the product is unused. Our real policy is 30 days, opened or not, and we charge a $5 restocking fee.

I screen-shared the result and asked the team: what's wrong with this email? The Shopify VA spotted it first—she'd processed enough returns to know we don't do 90 days. I explained that the model didn't lie on purpose; it predicted what a reasonable refund policy might look like based on patterns in its training data. That's a hallucination. It sounds confident. It's wrong.

Then I pasted our actual refund terms into the chat and asked Gemini to rewrite the reply. This time it got every detail right. That's grounding vs hallucination: you give the model the real information, and it uses that instead of guessing. One of my writers said, "So it's like when I write a blog post without checking the product specs first." Yes. Exactly like that.

This prompt is the one I'd use again. What are hallucinations AI becomes obvious when you watch the model make a confident, expensive mistake.

Prompt four: I had two models introduce themselves and explain the difference between open and closed

I wanted to explain open weight vs closed models without turning it into a licensing lecture, so I tried something goofy. I opened two chats side by side—Claude Sonnet 4.5 in one window, Llama 3.3 70B in the other—and pasted the same prompt into both:

"Introduce yourself. Explain whether you're an open-weight or closed model, what that means for someone using you in a small business, and one thing you're particularly good at."

Claude said it's a closed model built by Anthropic, which means the weights aren't public but the architecture is optimized for long, complex tasks like drafting legal emails or summarizing multi-page contracts. Llama said it's open-weight, built by Meta, and anyone can download and run it locally if they want—though most people use it through a platform like Kryotta. It's faster and cheaper than Claude for straightforward tasks like rewriting product descriptions or generating Instagram captions.

I expected this to be the prompt that clicked. It wasn't. My team didn't care about weights or licensing. What they cared about was speed and cost, and the prompt didn't make that concrete enough. I should've asked both models to rewrite the same email and timed them.

Still, one useful thing came out of it: my Shopify VA now picks Llama for bulk tasks (tagging 60 products, drafting five variants of the same discount code email) and saves Claude for the one or two emails per week that need a delicate tone. That's the behavior I wanted, even if the explanation didn't land cleanly.

Prompt five: I asked the model to estimate cost for a real workflow, and that's what stuck

The last prompt was the one I should've started with. I took a task my email marketer does every Monday—read 25–30 customer service emails from the weekend, write replies, flag anything that needs a refund or escalation—and asked Claude to estimate the token cost if she used AI for the whole workflow:

"I have 28 customer emails, average 100 words each. I'll paste them into one chat, ask you to draft replies to the simple ones, summarize the complaints, and flag refund requests. Estimate how many tokens that will use, calculate the cost if I use you (Claude Sonnet 4.5) versus Gemini Flash, and explain which model you'd recommend for this task."

Claude estimated 4,500 input tokens, 3,000 output tokens, and a total cost of about $0.18 using Claude or $0.03 using Gemini Flash. It recommended Gemini for this workflow because the replies don't require deep reasoning—they're mostly "sorry for the delay, here's your tracking number" or "I've processed your refund"—and Flash handles those in seconds.

My email marketer screenshotted that response. She now pastes her Monday batch into Gemini Flash, spends three cents, and saves herself 90 minutes. On the two or three emails per month that need a careful tone—someone's upset, or there's a Stripe dispute, or we messed up an order—she switches to Claude and pays the extra fifteen cents.

That's the prompt that made the cost structure make sense. It wasn't abstract. It was her actual Monday morning, priced out in dollars.

What I'd change: start with their tasks, not the definitions

If I ran this onboarding again, I'd flip the order. I'd start with prompt five (cost for a real workflow), move to prompt three (watch the model hallucinate, then ground it), and save the token and context window explanations for last. My team didn't need to understand the theory before they saw the tool work. They needed to see it work, then understand why it worked.

The one thing I wouldn't change: letting the model explain itself. Every time I tried to lecture—"a token is a chunk of text about four characters long"—I lost the room. Every time I pasted a prompt and let Claude or Gemini answer the question, someone nodded and said, "Oh, okay, that makes sense."

I also wouldn't skip the hallucination trap. That fake Etsy refund policy did more to teach grounding vs hallucination than any definition I could've written. One of my writers now pastes our brand voice guide into every chat before she drafts anything. The Shopify VA pastes our return policy. The email marketer pastes the current promo codes. They learned that the model is helpful, fast, and confident, but it doesn't know our business unless we tell it.

Questions people ask

How much did this onboarding cost in tokens?
About $0.40 total across five prompts and a dozen follow-ups. The fake refund email (prompt three) cost the most because I had Gemini rewrite it twice.

Which models should I use to teach AI to freelancers on my team?
Claude Sonnet 4.5 for the explanations—it's patient and clear—and then Gemini Flash or Llama 3.3 70B for the hands-on examples, because your team will use those models more often and the cost difference matters.

Do I need to explain open-weight vs closed models if my team just wants to write emails faster?
Probably not. I included it because I thought it mattered, but in practice my team cares about speed and cost, not licensing. If someone asks, explain it then.

What if my team is scared the AI will replace them?
Show them the workflow cost prompt (number five). When they see that the model costs three cents to draft replies but still needs a human to check tone, catch errors, and decide which emails to escalate, it's less scary. The tool makes their Monday faster; it doesn't do their job.

I've been running this same onboarding process with every new contractor I bring on, and the five-prompt structure takes about an hour. You can try all five in Kryotta—Claude, Gemini, and Llama are all there, and the auto-routing will pick the cheapest model that can handle the task. If your team is still writing everything manually, one hour and forty cents in tokens might be the best training budget you spend this quarter.

K
Written by
Kryotta Team
Product & research

Kryotta is the multi-model AI workspace — every leading model, one login, one bill. Try it free →

Related reading