Today is Class 0. Tomorrow, to whatever extent it happens, is Class 0.5.
The first real class is Monday. This is warm-up — no grade, no artifact, no wrong answers.
Set expectations up front. Some classes will get the full deck across two sessions, some just today, some most of it and some of tomorrow. Nobody is behind. Real course routine — the grading, the paper, the calendar — starts Monday.
Questions I've been asked so far
Are you going to teach us how to hack?
Why do I have to take computer science?
Is this the final seating chart?
How and when will lab groups get chosen?
Do we have to code?
Are we in this class all year?
What is happening in the spring semester?
Some get answered today. Some Monday. All of them get answered.
The receipt slide — students see they were heard. Read them aloud, acknowledge, don't answer yet. Today's deck partially answers "hack?" and "do we have to code?"; Monday handles "seating chart / lab groups / all year / spring." Add any question that came up in an earlier section that isn't here.
Day Zero · Two-class arc
One of these is in this room.
Two AI systems. The same questions. You predict which one wins — and say why.
Non-graded. No artifact leaves the room. Two classes: today is Round 1 (Familiar); next class is Rounds 2 & 3 (Recent, then Fresh). Have LM Studio and the Gemini tab already open behind this slide.
What you already know
You've used AI. Name the ones you know.
ChatGPT · Claude · Gemini · Copilot · Meta AI · Perplexity — brand names first, then what's under them.
60–90 seconds. Let students throw names at you. Board what they say. You're finding out who uses what and what the shared vocabulary is. Do NOT correct anyone yet — the sorting happens on the next slide.
Two words for two things
The company makes the model. You reach it through the company's app.
Model you talk to
Company
ChatGPT (GPT-4o, o1, …)
Claude (Sonnet, Opus, …)
Gemini (Pro, Flash, …)
The variants — Sonnet vs. Opus, 4o vs. o1, Pro vs. Flash — get swept under the rug today. What matters: company, model. Same shape everywhere.
Left column is filled — the models students already named. Right column is blank on purpose: have the room supply the company for each. Answers: OpenAI · Anthropic · Google. If they blank on one, that's fine — the point is that the distinction (model vs. company) is stable across every AI they use. Same shape you'll see on slide 7 for the two contestants (Gemma 3 4B in LM Studio :: Gemini at gemini.google.com).
What is a model's training data?
“The Cat in the ______.”
What is a model's knowledge base?
So what does a model “know”?
Three questions, one at a time — don't dump the slide. Q1 (training data): blank stares are fine. Show the Cat in the ______ line: the room will chorus Hat! — that IS training data. Every student in this room has "The Cat in the Hat" somewhere in their own training data. Models read billions of patterns like this. Q2 (knowledge base): what accumulates from all that reading. Q3 (does it know): the trick — patterns aren't knowledge. Sets up "cutoff" and "does it know YOU" without your saying either.
Local vs cloud — a case you already know
Local · Microsoft Word
The program is on your computer.
Works with the wifi off.
Nothing you type leaves the machine.
Cloud · Google Docs
The program is in Google's building.
Needs the internet — no connection, nothing.
Everything you type is sent across the country.
Same question, both times: where is the program actually running — your computer, or someone else's?
Land this before applying it to AI. Every student has used both. The question "where is the program running" is the whole point of Day Zero — you're just naming it here in a case they already know.
This laptop is a small version of that
Same building blocks. Different scale.
This 4060 laptop
One GPU. An RTX 4060.
8 GB of graphics memory.
On that desk, right there.
Cloud AI server
Thousands of GPUs, wired together.
Terabytes of graphics memory.
In a building you'll never see.
The datacenter isn't magic. It's this — a lot of this.
Local left, cloud right — the convention across the whole deck. Sean's rule: no deeper on hardware than "same shape, different scale." The point isn't the details; a datacenter is not a different kind of thing from the laptop, just a much bigger version.
The computer is the portal
The computer isn't the AI. The GPU is where the answer gets built.
You type a question. The computer hands it to the GPU. The GPU builds the answer, one guess at a time. The computer shows it to you. That's the whole path — the computer is the delivery service.
This reframes "the model runs on my laptop" into "the model runs on the GPU inside the laptop." Same reframe applies to the cloud server: the OS is the delivery service, the GPU is where the work happens. Sets up slide 7 (contestants) and slide 8 (airplane-mode proof).
The contestants — same shape at both scales
Gemma 3 4B — local
Company: Google (open weights, downloadable).
Model: Gemma 3 4B — a file on the laptop.
Platform: LM Studio — opens the file, loads it into the GPU.
Runs on the RTX 4060 in this room.
Gemini — cloud
Company: Google.
Model: Gemini (Pro or Flash — sweeping variants).
Platform: gemini.google.com — a browser tab.
Runs on Google's GPUs, across the country.
Amber is always the laptop. Cyan is always the datacenter. LM Studio ↔ Gemma is the same shape as OpenAI ↔ ChatGPT, or Anthropic ↔ Claude.
Both contestants are Google — sibling models packaged differently. The relationship — platform + model + company — is what students should be able to name by the end of this slide.
Does it know who you are?
Local · LM Studio + Gemma
LM Studio has no account, no history.
Gemma sees only what you type this session.
Wifi could be off. Nothing goes back.
Cloud · Google + Gemini
The model doesn't remember you across chats.
The service (Google) does: account, prior chats, activity.
Every time you ask, the service feeds that file in — as if you'd re-typed it.
When Gemini “remembers” something about you, it isn't the AI recalling — it's Google reminding it, in the background, at the start of every reply.
Sean's pivot for this slide: the AI itself doesn't know you. The service around the AI knows you and feeds that context into the model each turn. Same weights every call — different context. If a student says "so Gemini remembers me": correct them gently — Google remembers, and Google tells Gemini before each response. Sets up Event R3 (knew YOU, didn't know the news) — because "knowing YOU" isn't training; it's a database feeding a stateless model.
LM Studio — the one live proof
A program that opens the model file and loads it into the GPU.
Now let's prove it runs with the wifi switched off.
Only place the network gets touched on purpose. Do it here, before any round runs — if reconnecting is slow it costs minute 10, not minute 45. Students are witnesses.
Wifi OFF. Front row confirms the no-internet indicator; they report to the room.
Try to load any web page. It fails. Control: rules out "maybe the cloud is just down."
The dead Gemini tab. Try to send in the already-open tab — the error is the result. Show this first.
Run the test prompt in LM Studio (next slide). It generates anyway.
Wifi ON.
The offline test prompt
In two sentences, what does a GPU do?
Finishes in seconds. Reinforces the last slide. Either appears or doesn't — no ambiguity.
Run in LM Studio with wifi still off, right after the dead Gemini tab. Back-to-back is what makes it evidence rather than claim.
Set up your paper — turn it sideways
Round
Predict
Because
Sure?
Right?
ROUND: R1-F, R1-M, R1-E, R2-F, R2-M, R2-E, R3-F, R3-M, R3-EPREDICT: LAPTOP / CLOUD / BOTH / NEITHERSURE?: which sounded more confidentRIGHT?: which actually was
Nine slots across two classes — you'll fill Round 1 today, Rounds 2 and 3 next class. Blank rows are fine.
Copied, not read — hold until every student has the grid drawn. BECAUSE is the widest column on purpose. Blank slots make a skipped round invisible; that's deliberate.
What are the round rotations?
Three rounds. Three categories each: Food, Math, Event.
Round 1 · Familiar — things a well-trained model should know. Both should succeed.
Round 2 · Recent — the last few years. Where the small model starts to show its age.
Round 3 · Fresh — this week. Past the training-data cutoff.
Same prompt into both models each time. Watch, take notes on your paper, discuss.
No game framing, no candy — structured observation. Round 1 today; Rounds 2 and 3 next class. Each round has three prompt-and-discussion pairs (Food, Math, Event) with roughly equal time. The paper on the next slide is the note-taking apparatus.
Round 1
Familiar.
Things a well-trained model should know. Both should succeed.
Class 1's whole round. Roux → fractions → Henrietta Lacks. Landmark of Round 1: nobody wins by "knowing more" — the differences are speed, privacy, and whether it obeys.
Familiar · both should succeed
Food · Round 1
I'm learning to cook. Give me a list of the steps for making a dark roux, and what to watch for at each stage.
Same string into both. Then run it twice on the laptop.
Load-bearing: this proves the small model isn't a toy, so later failures read as limits, not brokenness. The experts are in the seats — somebody's grandmother makes gumbo, and a student will catch an error before you can. Run twice on the laptop: visibly different answers, cheapest proof it isn't looking anything up.
Discussion
What did you see?
The laptop was faster — its answer never left the room.
Run it again, get a different recipe. It isn't looking anything up.
Did the room's cooks catch anything the model got slightly wrong?
Both should have succeeded. The takeaway isn't "who won" — it's that a capable answer came out of a card in the room, and that it improvises rather than retrieves.
Familiar · does it teach, or does it answer?
Math · Round 1
Don't give me the answer. Don't solve it for me. Teach me how to add these fractions myself: 3/4 + 2/5
Does it walk you through common denominators — or does it blurt 23/20?
First appearance of the "show me how, don't answer" pattern. Two things get tested at once: does it know how, and does it OBEY. A model that dumps the answer failed the second test, no matter how correct it is. If it teaches — good — but even a small model usually manages this one; it's the easy version of Round 3.
Discussion
Did it teach, or did it answer?
The instruction was don't solve it. Did each side obey?
If it taught — did it explain why you need a common denominator, not just how?
Same instruction is coming back in Round 3. Note who obeyed.
Foreshadow: this same "don't solve it" instruction returns in Math R2 (pre-algebra) and Math R3 (quadratic). Escalating math, same rule. Where the model blurts the answer is where the instruction broke.
Familiar · both should know this
Event · Round 1
I'm writing a report on Henrietta Lacks. Give me a plain-English summary of who she was and why her cells changed medicine and ethics.
Historic, well-documented, in every training set. Both should have plenty to say.
Student-suggested topic. HeLa cells, cervical cancer, 1951, Johns Hopkins, taken without consent, foundational to polio vaccine research and countless others. The ethics angle is the real weight — that's where a model might sand off edges. Watch for it.
Discussion
Both knew. Did they tell it straight?
Did either sand off the consent problem, or say it plainly?
Did either add a "warm bath" tone — vague, careful — that a report needs cut out?
You couldn't tell they were wrong from the words alone. You could tell they were polished.
This is where "sounds right ≠ is right" first appears in a serious way. Both models will produce accurate-sounding text; whether it's actually accurate is a room-check.
Round 2
Less obvious.
Round 2's frame: not "recent" — less obvious. Where the answer isn't automatic for a small model. Same three-category rotation: Food, Math, Event.
Recent · one may not have heard of it
Food · Round 2
I want to make Dubai chocolate. Give me a list of the ingredients and how it's put together.
Peaked mid-2024 — right at the edge of what a small model might have seen.
The real thing: milk chocolate bar filled with pistachio cream and shredded knafeh (Middle-Eastern crispy phyllo). If the local model doesn't recognize it, watch for the failure mode — does it say "I don't know," or does it invent a recipe from the words "Dubai" + "chocolate" (dates? saffron? cardamom?). Confabulation looks like competence.
Discussion
Did it know it, or make it up?
Which model admitted it hadn't heard? Which invented a plausible-sounding recipe?
An honest "I don't know" is a better answer than a confident wrong one.
You can't hear the difference from tone alone — you have to check.
The confabulation lesson lives here. A model that invents from the name is doing exactly what the "next-word guesser" behavior predicts. Point at that.
Recent · does it still obey?
Math · Round 2
Don't give me the answer. Don't solve it for me. Teach me how to solve this myself: 3x + 7 = 22
Same rule as Round 1. Harder math.
Second appearance of the "show me how, don't answer" pattern. If the local model obeyed on fractions, does it obey on solving-for-x? Some models handle one and blow the other. That difference is the finding.
Discussion
Same rule. Different math. Same result?
Did the model that obeyed on fractions still obey here?
If it started to teach and then slipped the answer at the end — that's the instruction breaking.
Round 3 is the hardest math. Predict who breaks.
Escalation. Set up Round 3 — students should be able to predict from R1 and R2 who will break under a quadratic.
Recent · both should have heard of it
Event · Round 2
I'm working on a report about the 2018 Thai cave rescue — the youth soccer team trapped in the Tham Luang cave. Give me a list of research leads I can follow up on.
Real event, well inside training, ends well. Do the specifics line up?
Wild Boars soccer team, twelve boys plus coach, June 23 – July 10 2018, Tham Luang Nang Non cave, Chiang Rai province, international dive team, ends with everyone out alive. Model definitely trained on it — check the specifics: names, dates, number of divers, sequence. Chosen to keep Round 2 from being all-grief (post-Titan / post-Ida) — dramatic event with a good ending. Alt swap-ins if you want a different flavor: 2019 LSU / Joe Burrow national championship season (Louisiana-relevant, sports); 2019 first black-hole photo from the Event Horizon Telescope (science, iconic image); 2022 first images from the James Webb Space Telescope (science, well-documented dates).
Discussion — verify one lead live
Pick the most specific lead. Check it.
Specificity is where confabulation hides.
Check the most confident-sounding lead from each side, live in the browser.
Was the lead real? Real but slightly wrong? Or fully invented?
Same "verify a lead" move from prior versions of the deck. This is the round where the room learns that "sounds confident" and "is correct" are separate axes.
Round 3
Fresh.
This week's news. This year's food. The place a training cutoff is a hard wall.
Class 2 second half. Recent food (teacher-swap) → quadratic (hardest obey test) → the wildfire article (published this week). The local model literally cannot know these. What it DOES with that fact is the lesson.
Fresh · past the cutoff
Food · Round 3
What's on McDonald's Big Arch Burger? Give me the ingredients and how it's put together — bun, patties, cheese, sauce, toppings.
A menu item most students have heard of by now. Both models were trained before the launch.
The Big Arch is McDonald's new flagship burger (US rollout early 2026, per the Inc. review 3/17/2026). The real thing, so you can score answers live: two beef patties (roughly one-fifth of a pound each), double American cheese, some form of tangy orange sauce, lettuce, pickles, on a sesame-and-poppy-seed bun with seeds embedded top and bottom (that seeded bottom is the tell). Size is roughly Big Mac scale, positioned by McDonald's as the best in their lineup. If the model confabulates, it'll almost certainly invent a Big Mac variant with "special sauce" and generic pickles — the seeded bottom and the double American are what a made-up answer won't have. Two failure modes to name: (1) local admits "I don't know," or (2) local confabulates from the words "Big Arch" + "McDonald's." Both are lessons.
Discussion
Cutoff means cutoff.
Neither model was trained on this.
The cloud model can look — it has a live connection.
The local model has one honest move and one dishonest one. Which did it pick?
Now they can name why the results differ from Round 2. It isn't intelligence — it's access.
Fresh · same rule, hardest math
Math · Round 3
Don't give me the answer. Don't solve it for me. Teach me how to solve this myself: 2x² + 5x − 3 = 0
Third time asking. Do they still obey?
Strongest round in the deck. Every other round tests what a model knows; this tests whether it does what it was told. A small model will often blurt "x = 1/2 or x = -3" in the first line. No verification needed — everyone watches an instruction break in real time. This is exactly where these students will be by October when Unit 3 begins.
Discussion
Knowing and obeying are two jobs.
Getting the right answer is one skill. Following the rule you were given is another.
Did one blurt "x = 1/2" in the first line, then teach the process after? That's the failure — it already gave the answer.
This is what specification is: telling a machine exactly what to do, and having it do that.
Names the whole course thesis. Round 1 → Round 3 was the same instruction under escalating difficulty. Where it broke is where they need to learn to build safeguards.
Fresh · a story from this week
Event · Round 3
Historic wildfires sweep across France and Spain
Adapted from DOGO News · August 3, 2026
More than 300,000 people have been evacuated as thousands of firefighters battle historic wildfires across France and Spain. Extreme heat, dry conditions, and powerful winds have fueled the blazes.
Major fires began around July 20 in Spain's Ávila province — now the largest wildfire in Spanish history — and near Bordeaux, France. The Bordeaux fire became so intense its smoke created a thunderstorm on July 24, whose lightning ignited new fires miles away.
As of July 29 the spread was largely stopped, but the fires are not out. The same hot, dry, windy weather is now moving toward Greece and Italy.
I just read an article on the historic wildfires sweeping across France and Spain. Give me a list of research leads I can follow up on to learn more.
Read the article aloud — 30 seconds. Then run the prompt on both. Two things to watch: (1) does either fabricate specifics not in the article (a dead fire chief's name, an evacuee count that doesn't match)? (2) does the local model even know this happened? It literally cannot — its cutoff is a year before this. That's the point.
Discussion
Knew you. Didn't know this.
If Gemini seemed to know you — where did that come from?
Did it know about the wildfires — before it searched?
The laptop had access to neither. Why not?
If a model doesn't know something AND answers anyway — what do we call that?
Socratic — don't give the answers, draw them out. Expected landings, in students' own words: (Q1) the service, not the model — Google's stored context; (Q2) no, not until it searched; (Q3) no account and no live connection; (Q4) confabulation — a made-up answer that sounds plausible. If the room lands each of these, the two knowledge buckets (trained on / fetched at answer-time) don't need to be stated by you.
Debrief
What made a question hard?
Which side struggled when the question needed today's news?
Which one struggled when the instruction was don't solve it?
Which one struggled when the topic was obscure?
So — what is a small local model actually good for?
When would you rather not use the cloud model?
Reachable via D from anywhere — run whenever ~7 minutes remain, wherever the rounds left you. Socratic on purpose — draw landings from the room, don't state them. Expected shape, in students' words: (Q1) local can't reach; (Q2) whichever blurted the answer — testing obedience; (Q3) both may confabulate; (Q4) explaining familiar things, no login, offline, private; (Q5) when what you type shouldn't be sent, wifi's off, you're being watched. "What is a small model for" is a real answer — beats "the big one wins."
Class · 0 min
0:00
press T to start
1 / 37
Presenter notes
→/Space next ← prev Home first ·
T timer R reset C clock N notes 1 Round 1 2 Round 2 3 Round 3 B Class 2 pickup D debrief F Food·1 G Food·2 H Food·3 ·
M Math·1 Q Math·2 X Math·3 ·
E Event·1 W Event·2 V Event·3