Do Now

Terminology and Concepts to look out for

claim · verify · source

Relevance to Me

Yesterday you described three answers without saying which was best. Today you get to say it, but only if you can prove it. Today better has exactly one meaning.

Board Question

Copy today's board question off the whiteboard, then commit to an answer before we run anything.

Today's: Today's prompt asks three things — the year the Superdome opened, who the Saints head coach is right now, and the population of New Orleans. Which one are the models most likely to get wrong? Write down which one, and one sentence saying why.

A guess is not the point. You already know something that tells you the answer.

Same stations. Same groups.

Better Means Correct

Yesterday you described. Today you judge, and today the word better has one meaning only. Correct.

The prompt you are about to run contains three questions. Each one has an answer that can be looked up. That is the only thing that changed from yesterday. Same tool, same single pass, same three models, same stations, same groups. One variable.

When a model answers, it produces claims. A claim is a statement that says something is the case. "The Superdome opened in 1975" is a claim. "The Superdome is a large building" is barely one, because almost nothing could make it false. The claims worth your attention are the specific ones. A year. A name. A number.

To verify a claim is to check it against a source. A source is a place outside the response where the answer can be found. The Saints website is a source for who coaches the Saints. The Census Bureau is a source for the population of New Orleans. Another AI model is not a source, because it has exactly the same problem you are trying to solve.

Verifying is not the same as agreeing. There are four things that feel like verification and are not.

The response sounded confident. Every response sounds confident. You saw that yesterday.

The response was the longest. Length is not evidence. A long wrong answer is still wrong.

The response matched what you already believed. You might be wrong too.

Two of the three models said the same thing. Two models can be wrong in the same way, especially when they read similar text. Agreement is interesting. It is not proof.

So here is the procedure, and it is the same for every claim your group checks. Take one claim. Write it down. Find a source. Look. Mark it verified or not verified. Then move to the next one. Nine claims total, three from each model.

When you are finished you will have a count for each model, which is how many of its three answers were correct. That count is your ranking. It is the first ranking this lab has produced that you could defend to somebody who disagrees with you, because you are not defending an impression. You are pointing at a source.

While you work, watch for one thing. Note whether any model told you it was unsure about anything. Then find the answers that turned out to be wrong and compare how confident they sounded against how confident the right ones sounded. That comparison is part of what you write today.

Today's Lab Overview

Today's prompt

This is exactly what gets sent, to all three models, at the same moment.

Answer all three questions.
1. In what year did the Superdome in New Orleans open?
2. Who is the current head coach of the New Orleans Saints?
3. What is the population of the city of New Orleans?

The instructions all three are given

You are answering a question for a high school class in New Orleans.
Answer the question directly and completely.
Do not ask the student a question back.
PaneModel
1Ministral 3B
2Ministral 14B
3Mistral Large

Pane order never changes. Refer to responses by pane number all week. What differs between them is size.

Check it yourself

One source per question. Open these on your Chromebook — the lab machines only reach this site.

Another AI model is not a source. Neither is a search summary written by one.

Question 3 is the odd one. A population always carries a date, so write down the date you found as well as the number. A stale population is not wrong the same way a wrong coach is wrong.

I Do · pick an answer, then Check
Question 1notebook cue: claim
Which of these is a claim?
Question 2notebook cue: verify
To verify a claim means to
Question 3notebook cue: source
Which of these counts as a source?
Question 4worked aloud
Two models agree with each other and both turn out to be wrong. What does that tell you about agreement?
Question 5worked aloud
A model answers in one confident sentence and is wrong. What does that confidence tell you?
Question 6worked aloud
Your counts come in: pane 1 got two of three, pane 2 got one of three, pane 3 got three of three. What is the ranking?
Lab 1, Day 2 — Tuesday, September 1 · How to tell which one is better, and how to say why.