You already know the rule for a fair test. Change one thing. Hold everything else still. If you change two things and the result moves, you cannot say which change moved it.
The rule is easy to state and hard to follow, and it gets harder when the thing you are changing is words.
Start with the machine. A model does not receive your sentence. It receives tokens, and the split into tokens happens before any prediction is made. So when you swap a word in your prompt, you are not making one small edit to a sentence. You may be changing how many tokens go in, where they break, and therefore what the model has in front of it at every single step of the loop that follows. A change that looks tiny on your screen is not necessarily tiny inside the machine.
Now add the second problem. Even with the prompt held perfectly still, the answer can move on its own, because the picking has chance in it. So a researcher comparing two models is trying to see a difference through two layers of noise at once: the noise of wording and the noise of chance.
This is not a problem that machines invented. It is the oldest problem in survey research, and the people who do that work for a living are careful about it in a way that is worth copying.
Consider how the Pew Research Center handles it. Pew has tracked American use of AI chatbots for several years running, surveying 5,119 adults in February 2026 for its most recent report (Pew Research Center, 2026). Tracking something across years means asking the same question the same way every time, so that a change in the answers means a change in the country and not a change in the questionnaire.
But Pew needed to change a question. Through 2025 they had asked people whether they had ever used ChatGPT. By 2026, ChatGPT was no longer the only chatbot worth asking about, so the question was widened to cover others as well (Pew Research Center, 2026). That is a better question. It is also a different question, which means the number it produces cannot be lined up against the older numbers as though nothing happened.
Pew’s response is the part to notice. They did not quietly swap the wording and let the line on the graph keep going. They broke the line. On the published chart, the segment covering the wording change is drawn as a dotted line rather than a solid one, and the note under the chart states plainly what changed and when (Pew Research Center, 2026). The reader is told, at the exact point where comparison stops being safe, that comparison has stopped being safe.
That is the standard. Not never change anything. Change what you must, then mark the break so nobody reads across it by accident.
Pew applies the same care when the population changes. Its survey of teenagers reached 1,458 respondents ages 13 to 17, recruited through their parents, and the report is explicit that these findings describe teenagers and not adults (Pew Research Center, 2025). Two numbers from two different populations are not a trend. They are two numbers.
Bring this back to your station. Next week, seven groups will send one prompt to three models. Held constant: the words, the order, the moment. Changing: which model receives them. That is the design, and it is only worth anything if the constant part really does stay constant.
Which raises the question you have to answer in a minute. If your group decides on Thursday to change one thing about the prompt, what counts as one thing? A single word? A whole sentence? The order of two sentences? There is no automatic answer here. There is only the answer your group writes down in advance and then holds itself to, and the honesty to mark the break when you cross it.
Pew Research Center. (2025, December 9). Teens, social media and AI chatbots 2025. https://www.pewresearch.org/internet/2025/12/09/teens-social-media-and-ai-chatbots-2025/
Pew Research Center. (2026, June 17). Americans and AI 2026: Chatbots, smart devices and views on impact. https://www.pewresearch.org/internet/2026/06/17/americans-and-ai-2026-chatbots-smart-devices-and-views-on-impact/
Change one thing. Hold everything else still.
If two things move at once and the result changes, you have no way to tell which one moved it. That is not a technicality; it is the whole reason the comparison exists.
This is the oldest rule in survey research, drug trials, physics experiments, and cooking shows that swap one ingredient. You already use a version of it every time you argue that your team lost because of one specific play. The trick is doing it on purpose — before the answer comes in, not after.
Words look easy to swap. They are not.
A prompt does not enter the model as a sentence. It enters as tokens, and the split into tokens depends on the exact letters. Swap unbelievable for amazing and you have not only changed one word — you have changed how many tokens go in, where they break, and therefore what the model has in front of it at every prediction step that follows.
A tiny edit on your screen is not necessarily tiny inside the machine. This is the first noise layer, and it is invisible from the outside.
Even if you hold the wording perfectly still, the answer can move on its own — the picking has chance in it. So you are trying to see a real difference through two layers at once. The wording noise, and the picking noise.
Any comparison worth trusting has to survive both.
The Pew Research Center surveyed 5,119 adults in February 2026 about their use of AI chatbots. They have been asking that question for several years running.
Tracking something across years is not a trivial data-collection choice. It is a discipline. To claim that a number changed, you have to be sure the question stayed the same. If the answers move, that could mean the country changed — or it could mean you changed the questionnaire and everything else is noise.
Same question, every year, on purpose. That is what makes the line on the chart mean anything.
By 2026, the old question was not good enough anymore.
Through 2025: had you ever used ChatGPT.
2026: widened to cover other chatbots as well, because ChatGPT was no longer the only one worth asking about.
This is a better question. Anyone reading the report gets a truer picture of what people are actually using. It is also a different question, which means the 2026 number and the 2023 number are not measuring the same thing.
Two problems, one moment. Better answer. Broken comparison.
Pew did not quietly swap the wording and let the line on the graph keep going. They broke the line. Solid line through 2025. Dotted line at the wording change. A note under the chart states what changed and when.
Anyone looking at the chart sees the discontinuity before they read across it. That is the point.
The rule is not: never change anything.
Sometimes you have to change something. A question stops being useful. A model gets discontinued. A tool changes its interface. Real research does not freeze in amber; it adapts.
The rule is: change what you must, then mark the break so nobody reads across it by accident. Draw the line dotted. Say what changed. Say when. That way the discontinuity travels with the data, and the next person reading it inherits your honesty instead of your convenience.
Same care, different problem.
Pew's teen survey reached 1,458 respondents ages 13 to 17, recruited through their parents. The report is explicit: these findings describe teenagers, not adults.
If you cite "13 percent of Americans" from the adult survey and "56 percent" from the teen survey and treat them like a trend, you are not doing math — you are comparing two different rooms. Two numbers from two different populations are not a trend. They are two numbers.
Bring this back to your station.
Held constant: the exact words of the prompt, the order they are in, the moment they are sent. Every group in the room hits send inside the same window with the same characters typed. That is the constant part, and it is only worth anything if it really does stay constant.
Changing: which model receives the prompt. Three models, three answers, everything else equal.
The difference you see across the three answers is what this whole design exists to measure.
Your group gets to change one thing about the prompt later in the week. Before you do, you have to decide what counts as one thing.
A single word? A whole sentence? The order of two sentences? Adding a comma? Splitting one sentence into two?
There is no automatic answer here. Different groups will draw the line differently, and different lines will produce different results. The only wrong answer is the one you decide after you see the output.
Write it down in advance. Hold yourself to it. Mark the break if you cross it. That is the whole discipline.
- A1
- B2
- C4
- A1
- B2
- C3
- D4
- E5
- undefined6
- undefined7
The first item is your name. Answer it — that is what puts you on the work now that email collection is off.