Terminology and Concepts to look out for
claim · verify · source
Relevance to Me
Yesterday you described three answers without saying which was best. Today you get to say it, but only if you can prove it. Today better has exactly one meaning.
Board Question
Copy today's board question off the whiteboard, then commit to an answer before we run anything.
Today's: Today's prompt asks three things — the year the Superdome opened, who the Saints head coach is right now, and the population of New Orleans. Which one are the models most likely to get wrong? Write down which one, and one sentence saying why.
A guess is not the point. You already know something that tells you the answer.
Same stations. Same groups.
Yesterday you described. Today you judge, and today the word better has one meaning only. Correct.
The prompt you are about to run contains three questions. Each one has an answer that can be looked up. That is the only thing that changed from yesterday. Same tool, same single pass, same three models, same stations, same groups. One variable.
When a model answers, it produces claims. A claim is a statement that says something is the case. "The Superdome opened in 1975" is a claim. "The Superdome is a large building" is barely one, because almost nothing could make it false. The claims worth your attention are the specific ones. A year. A name. A number.
To verify a claim is to check it against a source. A source is a place outside the response where the answer can be found. The Saints website is a source for who coaches the Saints. The Census Bureau is a source for the population of New Orleans. Another AI model is not a source, because it has exactly the same problem you are trying to solve.
Verifying is not the same as agreeing. There are four things that feel like verification and are not.
The response sounded confident. Every response sounds confident. You saw that yesterday.
The response was the longest. Length is not evidence. A long wrong answer is still wrong.
The response matched what you already believed. You might be wrong too.
Two of the three models said the same thing. Two models can be wrong in the same way, especially when they read similar text. Agreement is interesting. It is not proof.
So here is the procedure, and it is the same for every claim your group checks. Take one claim. Write it down. Find a source. Look. Mark it verified or not verified. Then move to the next one. Nine claims total, three from each model.
When you are finished you will have a count for each model, which is how many of its three answers were correct. That count is your ranking. It is the first ranking this lab has produced that you could defend to somebody who disagrees with you, because you are not defending an impression. You are pointing at a source.
While you work, watch for one thing. Note whether any model told you it was unsure about anything. Then find the answers that turned out to be wrong and compare how confident they sounded against how confident the right ones sounded. That comparison is part of what you write today.
Today's prompt
Open MuggsOfPrompts ↗ muggsofcompsci.net/prompts
This is exactly what gets sent, to all three models, at the same moment.
Answer all three questions. 1. In what year did the Superdome in New Orleans open? 2. Who is the current head coach of the New Orleans Saints? 3. What is the population of the city of New Orleans?
The instructions all three are given
You are answering a question for a high school class in New Orleans. Answer the question directly and completely. Do not ask the student a question back.
| Pane | Model |
|---|---|
| 1 | Ministral 3B |
| 2 | Ministral 14B |
| 3 | Mistral Large |
Pane order never changes. Refer to responses by pane number all week. What differs between them is size.
One source per question. Open these on your Chromebook — the lab machines only reach this site.
Another AI model is not a source. Neither is a search summary written by one.
- Question 1, the Superdome yearwww.caesarssuperdome.com
- Question 2, the head coachwww.neworleanssaints.com/team/coaches-roster
- Question 3, the populationdata.census.gov/profile/New_Orleans_city,_Louisiana
Question 3 is the odd one. A population always carries a date, so write down the date you found as well as the number. A stale population is not wrong the same way a wrong coach is wrong.
A claim says something is the case, which means something could make it false. A question cannot be wrong and a hedge does not commit to anything.
Outside is the load-bearing word. Everything inside the response has the same problem you are trying to solve.
A place outside the responses where the answer can actually be found. Another model is the same machinery producing the same kind of confident text.
Two models that read similar text can be wrong in the same way. Worth noticing, never sufficient.
Confidence is the same in a right answer and a wrong one. That is the finding, not a complaint.
Three, then two, then one. This is the first ranking all week you could hand to somebody who disagrees with you.