Do Now · Lab Notebook · 10 minutes

Set up your notebook. You will need it at the end of class.

Terminology and Concepts to look out for

Copy these. Leave space under each one for a definition.

  • respondent
  • item
  • topline
  • self-report
  • variability
Relevance to Me

Think about these. Nothing to write.

  • How much do you think the person sitting next to you uses AI?
  • Have you ever said you do something less than you actually do?
  • If someone asked how much time you spend on your phone, would your answer be right?
Question of the Day

Copy this into your Lab Notebook and leave room under it. You will answer it at the end of class.

What does a survey answer actually tell you?

What a Survey Actually Gives You

A survey looks simple from the outside. Someone asks a group of people a set of questions, collects the answers, and reports what the group said. The reporting is where the difficulty starts, because a pile of answers is not yet information. Turning one into the other requires several small decisions, and every one of them can be made badly.

Start with the pieces. A respondent is one person who answered. An item is one question on the survey. If forty people answer a survey with twenty questions, there are forty respondents and twenty items, and the survey has collected eight hundred separate answers. Those two words get confused constantly, including by adults who should know better, because both of them can be counted and the counts sound similar.

The first thing a researcher builds from those answers is a topline. A topline lists every item, every answer a respondent could have picked, and how many respondents picked each one. It contains no conclusions and no opinions. It is the raw shape of what the group said, and it is deliberately boring. A researcher who skips the topline and jumps straight to conclusions has no way to show anyone else where those conclusions came from.

Every item on a topline carries a number written as n. The n is how many respondents the item is based on. This matters more than it looks like it should. If a survey has forty respondents but six of them skipped a particular question, then that item rests on thirty-four people, not forty. A careful topline shows the skipped answers rather than hiding them, because a question that many people refused to answer is telling you something. Sometimes the refusal is the finding.

Now the harder problem. Almost everything a survey collects is self-report, meaning the only source for the answer is the person giving it. If a survey asks how many hours you slept last night, nobody checked. If it asks how often you argue with your brother, nobody was watching. The answer goes into the topline looking exactly like a measurement, but it is not a measurement. It is a statement.

Self-report is not worthless. For many questions there is no better instrument available, and people are often roughly right about themselves. But self-report has a specific and predictable weakness: people are inconsistent in the direction of their errors. On questions about behavior that carries any social weight, respondents shade their answers toward whatever seems acceptable. They round their exercise up and their screen time down. They do it without deciding to, which is why simply asking them to be honest does not fix it.

There is a second kind of question that looks like self-report and is not. When a survey asks a respondent to estimate what other people do, the answer is a guess about a group. Guesses about groups fail differently than statements about yourself. People tend to assume that whatever they notice most is what happens most, so a respondent who has seen one dramatic example will estimate high. Neither kind of answer is more trustworthy than the other. They are simply wrong in different directions, and knowing which kind you are holding changes how much weight it can carry.

That brings up the property that separates an interesting item from a dull one. Consider two questions asked of the same group. The first asks whether respondents have ever ridden in a car. Nearly all of them say yes. The second asks how many hours they slept last night, and the answers run from four to eleven. The second item has variability, meaning the answers spread out instead of clustering. The first item has almost none.

An item with no variability can still be true and still be useful. It just cannot explain any difference between the people in the group, because on that item there is no difference to explain. If everyone answers the same way, the item cannot tell you why one respondent does something and another does not. Researchers care about variability for this reason. It is where the differences live, and differences are what most questions are actually about.

None of this requires advanced mathematics. It requires reading a topline carefully, noticing which items rest on how many people, keeping track of whether an answer is a statement about the respondent or a guess about somebody else, and paying attention to whether the answers spread out or pile up. Those four habits do most of the work. The mistakes that follow from skipping them are not subtle ones, and they are made constantly by people who never learned that a survey answer is a claim someone made rather than a fact someone verified.

Work these together
10 questions on the passage. Pick an answer, then check it. Nothing here is collected.
Question 1 of 10the passage
A survey has 40 people answer 20 questions. How many respondents does it have?
Question 2 of 10the passage
According to the passage, why does a careful topline show the answers people skipped?
Question 3 of 10the passage
A survey asks how many hours you slept last night. What makes that answer a self-report?
Question 4 of 10the passage
The passage describes two questions asked of the same group. One asks whether people have ever ridden in a car and nearly everyone says yes. The other asks how many hours they slept and answers run from four to eleven. Why does the passage call the second one more useful?
Question 5 of 10the passage
A survey has 40 respondents. On one item, 6 of them skipped. According to the passage, what is n for that item?
Question 6 of 10the passage
Which of these does the passage say a topline contains?
Question 7 of 10the passage
The passage says that asking respondents to be honest does not fix the weakness in self-report. Why not?
Question 8 of 10the passage
According to the passage, what is predictable about the errors in self-report?
Question 9 of 10the passage
The passage says a respondent who has seen one dramatic example will estimate high when asked about other people. What explains that?
Question 10 of 10the passage
According to the passage, what is true of an item where nearly everyone gives the same answer?
From Our Class Survey
Your answers, collected August 11. 79 respondents. Selected items from the full topline.
A10 · Paying For ItComplete
Do you pay for any AI tool?
  • Yes6
  • No64
  • Not sure8
n = 79
6 + 64 + 8 + 1 = 79. One respondent skipped this item.
A17 · Peer EstimateFor comparison
Out of 10 students in this class, how many use AI without permission?
  • 04
  • 10
  • 22
  • 31
  • 41
  • 56
  • 61
  • 76
  • 89
  • 98
  • 1040
n = 79
All 29 survey items →
The work
Answer the checkable ones together. The discussion questions have no answer today.
Does A10 answer our question?
Our question is how much does this class use AI? Does A10 answer it?
Variability
A10: 64 of 79 gave the same answer. A17: answers landed on 0, 2, 3, 4, 5, 6, 7, 8, 9, and 10.
Which one shows more difference between the people in this room?
The three buckets
About me
The item asks about you.
About others
The item asks about everybody else.
Doesn't answer
The item is about AI, but it doesn't tell us how much this class uses it.
Sort it
Where does A10 go?
The rest of the block

Partner work first. Then quiet work time.

The three rules
  • No cellphones
  • No sleeping
  • No non-academic use of Chromebooks
What you get
  • A quiet room and music
  • No interruptions from the teacher
  • If you are caught up on this course's work, you may work on anything legitimate for another class
If you are picking, pick reading or writing.
Everything we build in this class, we build with words. The tools take writing as input. How well you write is the ceiling on what you can make them do.
Open this now: muggsofcompsci.net/decks/cs1-miniq-day1-wedo