Terminology and Concepts to look out for
time to last token · tokens per second · hedge
Relevance to Me
Every app you use is somebody's model answering as fast as it can. Today you find out what "fast" actually measures, and whether the fastest one is the one you would want.
Board Question
Three models. The same question. Sent at the same moment. Which pane will finish writing first? Write pane 1, pane 2, or pane 3 in your notebook and commit to it.
Same stations. Same groups. Monday's question, with one line added to it.
- Harness the AI. Today's frame. The prompt has one new instruction — a length limit — and the day's job is watching what stays inside it and what does not.
- Test tomorrow. Absent means the make-up is scheduled with Mr. Muggivan directly. No make-up scheduled → F for the test.
- Make good on Monday. An A or better on today's Lab Observation (Day 4) and today's Exit Ticket excuses last Monday's Day 1 if it is still missing.
- Next week. Dedicated, mandatory Exit Ticket time. The last 15–20 minutes of every class will be locked to the ticket.
On Monday you asked these three models whether college is worth taking out loans for. You got three long answers back, and you recorded five numbers under each one: words, characters, tokens, seconds, tokens per second.
You recorded those numbers. Nobody asked you to do anything with them. Today we do.
Two of those five are about speed, and they do not mean the same thing.
Time to last token is the seconds number. It is when the pane finished — the moment its final token arrived on your screen. It is what you feel when you are waiting.
Tokens per second is the rate. It is how fast the model was writing during the time it was writing.
Those two can disagree, and the reason they disagree is the whole point of today. A model that writes quickly can still finish last, because finishing depends on two things: how fast you write, and how much you decided to write.
Here is the arithmetic, and it is arithmetic you have already done. On August 26 you worked out how a file's size and a card's memory decide whether a model fits. This is the same kind of comparison with different quantities:
tokens ÷ tokens per second = seconds
A pane that produced 300 tokens at 60 tokens per second took 5 seconds. A pane that produced 90 tokens at 30 tokens per second took 3 seconds. The second one is the slower writer and it still finished first. It wrote less.
So "which one is fastest" is not one question. It is two, and you have to say which one you mean.
The new instruction. Monday's prompt let the models write as much as they wanted, and they wanted a lot. Today the prompt says the answer may be no more than 120 words. Everything else is identical — same three models, same system prompt, same moment, same question.
That does two things. It makes the responses short enough to actually read and describe inside one class. And it gives you something new you can check rather than judge: the tool counts the words, the prompt states a limit, so you can say plainly whether each model did what it was told. Not all of them will. That is a finding, and it is worth recording carefully.
Hedging is still on the table. Monday this question hedged more than anything else we tested. The answers are shorter now. Whether shorter also means less hedging is not something you know yet — it is something you are about to find out.
Today's prompt
Same prompt as yesterday, character for character. Run it twice more. Same three models, same four criteria, scored the same way.
Open MuggsOfPrompts ↗ muggsofcompsci.net/prompts
This is exactly what gets sent, to all three models, at the same moment.
Is going to college worth taking out student loans for? Answer in no more than 120 words.
The instructions all three are given
You are answering a question for a high school class in New Orleans. Answer the question directly and completely. Do not ask the student a question back.
| Pane | Model |
|---|---|
| 1 | Ministral 3B |
| 2 | Ministral 14B |
| 3 | Mistral Large |
Pane order never changes. Refer to responses by pane number all week. What differs between them is size.
The four criteria
Posted before the prompt runs. Identical for every group in every period. Score every pane against all four.
- Specific. Names actions a person could take, rather than how they should feel.
- Usable this week. Does not require money, equipment, or another person's cooperation.
- Complete. Addresses before, during, and after.
- Honest. Does not promise the nervousness will go away.
These stay reachable all period — the button in the bar below.
It is the moment the response was done. It is the number you feel while you are waiting for it.
Seconds is when it finished. Tokens per second is the rate it kept up while it was going.
A hedge avoids committing. The other three commit to something you could go check.
Ninety divided by thirty is three. Three hundred divided by sixty is five. Pane 2 is the faster writer and it still finished second, because it wrote more than three times as much.
It is a description, not a complaint. The prompt stated a limit, the tool counts the words, and the two disagree. You record it.
One line was added to the prompt. Everything else was held the same on purpose, which is what lets you say the length limit is what changed.