Terminology and Concepts to look out for
single pass · hedge · specificity
Relevance to Me
You have asked an AI something and taken the first answer it gave you. Today you get three answers to the same question at the same moment, and you have to say what is different about them before you are allowed to say which one you like.
Board Question
Copy today's board question off the whiteboard into your notebook, then commit to an answer before we run anything.
Today's: Which of the three models we are running today will be the fastest? Which will be the slowest?
Notebook heading: date, station number, group members, prompt.
Today you will type one question and get three answers back at the same time. The three answers come from three different models made by the same company. All three are running in the same place. All three get your prompt at the same moment. All three are given the same instructions, and those instructions are on your screen where you can read them.
So almost everything about these three models is the same. What is different is size. One of them holds a small number of parameters, one holds a medium number, and one holds an enormous number. That is the variable. Everything else in this room is held.
The tool sends your prompt to all three at the same moment. This is called a single pass. You get one response from each model, and then the tool stops. There is no follow up question. There is no button that runs it again. This is not a limitation the tool happens to have. It is the rule that makes the whole week work.
Here is why. On Friday you learned that a comparison only tells you something when one thing changes and everything else stays the same. If you could ask a follow up question, the second answer would depend on the first answer, and the first answer was different for each model. Two things would have moved instead of one, and you would be back where week one left you.
So one prompt goes out. Three responses come back. Everything else is held.
Your job today is to describe those three responses. Not to rank them. Not to say which one you would use. Describe.
Describing sounds easy and it is not, because the moment you read three versions of the same thing your brain starts sorting them into better and worse before you have written anything down. A description is a statement someone else could check by looking at the response. "The second response had four numbered steps" is a description. "The second response was clearer" is not, because clearer is your reaction, and nobody can look at the text and confirm it.
There are four things worth describing, and your notebook has a column for each.
Length. How much text came back. Count it roughly.
Structure. Paragraphs, a numbered list, headings, or one unbroken block.
Specificity. Whether the response commits to amounts, times, and names, or stays general. "Cook it for a while" and "cook it for twenty to thirty minutes, stirring constantly" say the same thing at two levels of specificity.
Hedging. A hedge is language that avoids committing. Results may vary. This is generally true. You may want to consult a professional. Mark every hedge you find. Some responses will have none, and that is worth recording too.
One more rule, and it is the one groups break. Record what came back exactly as it came back. Do not fix spelling. Do not tidy the wording. The moment you clean up a response you have created a fourth response that no model produced, and your group is now describing something that does not exist.
Today's prompt
Open MuggsOfPrompts ↗ muggsofcompsci.net/prompts
This is exactly what gets sent, to all three models, at the same moment.
Is going to college worth taking out student loans for?
The instructions all three are given
You are answering a question for a high school class in New Orleans. Answer the question directly and completely. Do not ask the student a question back.
| Pane | Model |
|---|---|
| 1 | Ministral 3B |
| 2 | Ministral 14B |
| 3 | Mistral Large |
Pane order never changes. Refer to responses by pane number all week. What differs between them is size.
One prompt out, three responses back, nothing after that. The word for it is the whole design of the tool, and it is the reason the week works.
A hedge is language that avoids committing. The other three commit to something you could go check.
Same instruction, two levels of commitment to amounts and times. Neither one is longer and neither one is hedging.
Anyone can look at the response and confirm four numbered steps. Nobody can look at it and confirm easier.
Friday's rule, arriving as a consequence. Two things would move instead of one and the comparison would be gone.
Nothing produced that text. Your group is now describing something that does not exist.