Do Now

Terminology and Concepts to look out for

time to last token · tokens per second · hedge

Relevance to Me

Every app you use is somebody's model answering as fast as it can. Today you find out what "fast" actually measures, and whether the fastest one is the one you would want.

Board Question

Three models. The same question. Sent at the same moment. Which pane will finish writing first? Write pane 1, pane 2, or pane 3 in your notebook and commit to it.

Same stations. Same groups. Monday's question, with one line added to it.

Heads up
  • Harness the AI. Today's frame. The prompt has one new instruction — a length limit — and the day's job is watching what stays inside it and what does not.
  • Test tomorrow. Absent means the make-up is scheduled with Mr. Muggivan directly. No make-up scheduled → F for the test.
  • Make good on Monday. An A or better on today's Lab Observation (Day 4) and today's Exit Ticket excuses last Monday's Day 1 if it is still missing.
  • Next week. Dedicated, mandatory Exit Ticket time. The last 15–20 minutes of every class will be locked to the ticket.
What Fast Means

On Monday you asked these three models whether college is worth taking out loans for. You got three long answers back, and you recorded five numbers under each one: words, characters, tokens, seconds, tokens per second.

You recorded those numbers. Nobody asked you to do anything with them. Today we do.

Two of those five are about speed, and they do not mean the same thing.

Time to last token is the seconds number. It is when the pane finished — the moment its final token arrived on your screen. It is what you feel when you are waiting.

Tokens per second is the rate. It is how fast the model was writing during the time it was writing.

Those two can disagree, and the reason they disagree is the whole point of today. A model that writes quickly can still finish last, because finishing depends on two things: how fast you write, and how much you decided to write.

Here is the arithmetic, and it is arithmetic you have already done. On August 26 you worked out how a file's size and a card's memory decide whether a model fits. This is the same kind of comparison with different quantities:

tokens ÷ tokens per second = seconds

A pane that produced 300 tokens at 60 tokens per second took 5 seconds. A pane that produced 90 tokens at 30 tokens per second took 3 seconds. The second one is the slower writer and it still finished first. It wrote less.

So "which one is fastest" is not one question. It is two, and you have to say which one you mean.

The new instruction. Monday's prompt let the models write as much as they wanted, and they wanted a lot. Today the prompt says the answer may be no more than 120 words. Everything else is identical — same three models, same system prompt, same moment, same question.

That does two things. It makes the responses short enough to actually read and describe inside one class. And it gives you something new you can check rather than judge: the tool counts the words, the prompt states a limit, so you can say plainly whether each model did what it was told. Not all of them will. That is a finding, and it is worth recording carefully.

Hedging is still on the table. Monday this question hedged more than anything else we tested. The answers are shorter now. Whether shorter also means less hedging is not something you know yet — it is something you are about to find out.

Today's Lab Overview

Today's prompt

Same prompt as yesterday, character for character. Run it twice more. Same three models, same four criteria, scored the same way.

This is exactly what gets sent, to all three models, at the same moment.

Is going to college worth taking out student loans for?

Answer in no more than 120 words.

The instructions all three are given

You are answering a question for a high school class in New Orleans.
Answer the question directly and completely.
Do not ask the student a question back.
PaneModel
1Ministral 3B
2Ministral 14B
3Mistral Large

Pane order never changes. Refer to responses by pane number all week. What differs between them is size.

Today's Lab Overview

The four criteria

Posted before the prompt runs. Identical for every group in every period. Score every pane against all four.

  • Specific. Names actions a person could take, rather than how they should feel.
  • Usable this week. Does not require money, equipment, or another person's cooperation.
  • Complete. Addresses before, during, and after.
  • Honest. Does not promise the nervousness will go away.

These stay reachable all period — the button in the bar below.

I Do · pick an answer, then Check
Question 1notebook cue: time to last token
The seconds number under a pane tells you what?
Question 2notebook cue: tokens per second
Which number tells you how fast a model was writing while it wrote?
Question 3notebook cue: hedge
Which of these sentences is a hedge?
Question 4worked aloud
Pane 1 produced 90 tokens at 30 tokens per second. Pane 2 produced 300 tokens at 60 tokens per second. Which finished first?
Question 5worked aloud
Today's prompt says the answer may be no more than 120 words. A pane comes back with 260 words. What have you observed?
Question 6worked aloud
Monday and today used the same question and the same three models. What is different?
Lab 1, Day 4 — Thursday, September 3 · How to tell which one is better, and how to say why.

The four criteria