Computer Science I · Fall 2026

The Roadmap

How Computer Science I is built and which CSTA 2026 standards it targets. It shows the five assessment buckets, the cull, and — new this year — the objective codes each claimed standard resolves to. The unit sections show how those standards are distributed across the semester and its assessments.

This is a map drawn from the standards, not a curriculum taught to them. Objectives are ordered by criticality within each unit and taught in that order; the unit runs until its end date and stops wherever it stops. No per-week day budget, no pacing guide.

Pilot semester. A course that exists nowhere else, in its first semester, claiming conservatively on purpose. Year two assigns tiers from the coverage record — dated content-day claims made before scores exist — rather than from intent, which is why every demoted standard stays on the page.

Coverage record notes

2026-08-11. U1-E02 (visualize coded results; identify what the visualization cannot show) is scheduled for Aug 13 as a Week 1 content day, and U1-E01 (compare fixed-response and open-ended items as data sources) for Aug 17 as a Week 2 content day. Both sit above the enrichment line under the calendar’s day-by-day plan and are candidates for promotion when the coverage record is read at end of semester. HS-DAT-DI-25, HS-DAT-DI-26, and S1-DSC-VZ-11 all attach to U1-E02 and would move with it. Tiers change from the record after the days run, not from this note.

How to read this

A bucket is not a unit

Buckets are assessment categories, not units. The mapping between buckets and CSTA concepts is many-to-many, and a unit draws on several buckets.

Two bands run through it

The instrument is graduated middle-school → high-school to defeat the floor effect. MS standards are the substrate, not the target.

The cull is deliberate

226 standards exist; this course touches 63. A roadmap that implied coverage of all 226 would be a lie. Naming 63 and saying why is the design. For calibration, a CSTA-validated professional AI curriculum reached 63% alignment across the 46 high-school foundational standards.

The Cull

Every standard in scope gets exactly one tier. Roughly 28% of the framework is touched at all — the honest number for a one-semester foundational course, and a feature, not a gap.

226
In the 2026 set
20
Core
31
Reach
12
Scaffold
63
Total touched
163
Skip
CORE

Carried by at least one objective above the enrichment line. Assessed. The crosswalk shows the objective codes.

Instrument role: Scored set.

REACH

Carried only by enrichment-tail objectives (U*-E**). An honest partial claim — the enrichment tail runs if there is time.

Instrument role: Diagnostic pool.

SCAFFOLD

Middle-school prerequisite. Not a course target and carries no objective code — the floor the course stands on, present to locate students on the progression.

Instrument role: Graduated items, MS band.

SKIP

Explicitly out of scope, with a stated reason. Absent from the instrument — but named on this page. See what we are not covering.

The Buckets

Five assessment categories. Each carries items at both bands; each maps to CSTA concepts many-to-many.

1

Systems

What a machine physically is, and what physically limits it.

Covers

RAM, CPU, storage, memory vs. disk, bandwidth, compute vs. capacity, what physically limits a machine, local inference hardware.

CSTA

SYS·HWSYS·IM
CORE1
  • HS-SYS-HW-30the hardware thread — bandwidth, memory, what limits a machine

    Demonstrate the capabilities and limitations of a physical or simulated computing device.

    Objectives:U1-07U1-08U1-09
REACH2
  • Differentiate an operating system as a special type of software that manages hardware and other software.

    Objectives:U1-10

    touched where the request path crosses it, not taught for itself

  • Investigate how computing systems and infrastructure impact society and the environment, identifying who is affected.

    Objectives:U1-E06

    data centers, local vs. cloud cost — reached only in the enrichment tail this year

SCAFFOLD1
  • Differences between computing systems by user need.

2

Networks & the Web

How machines talk, and where data physically goes.

Covers

Router, DNS, HTTP/HTTPS, client/server, what "the cloud" actually is, where data physically goes, local vs. remote.

CSTA

SYS·NTSYS·SE
CORE2
  • Diagram a network of computing systems, including hardware and software.

    Objectives:U1-10U1-11
  • Identify cybersecurity and physical security measures and the trade-offs they impose on users, data, and devices.

    Objectives:U1-11U1-12U2-E06

    Unit 2 — what leaves the machine on upload

REACH2
  • Analyze how the internet functions as a network of networks.

    Objectives:U2-E05

    reached only in the enrichment tail this year

  • Classify the causes and impacts of security breaches and social engineering attacks.

    Objectives:U2-E07

    reached only in the enrichment tail this year

SCAFFOLD2
  • How data travels as packets across a network.

    packets

  • How the internet's design supports resilience.

    internet resilience

Buckets 1 and 2 share one CSTA concept. The 2017 Computing Systems / Networks split was merged into Systems & Security in 2026. The buckets stay separate for assessment (different item types, separate diagnostic reads). This is not a mapping error.

3

Data

What data is, how it is shaped, and how it misleads.

Covers

Data types, representation, structured vs. unstructured, distributions, reading a chart, what a mean conceals, metadata, data quality.

CSTA

DAT·DCDAT·DIDAT·IMS1-DSC
CORE4
  • Create a data dictionary describing name, type, and allowable values for each attribute.

    Objectives:U1-22

    survey instrument — students' own data

  • Use a computational tool to clean and organize text-based data.

    Objectives:U1-23U1-25
  • Evaluate different approaches to verifying consistency and compliance with expected data types, values, and ranges.

    Objectives:U2-08U2-09

    Unit 2's verification pipeline — gold standards and consensus thresholds are exactly this, weighed against an unweighted scheme

  • Debate the efficacy of a policy or regulation for the responsible use of data.

    Objectives:U1-C39

    the class AI policy is derived from coded student writing — a terminal artifact, not a partial claim

REACH10
  • Create a data visualization of a multivariate dataset to answer a question or make a classification or prediction.

    Objectives:U1-E02

    reached only in the enrichment tail this year

  • Evaluate a data simulation or visualization to answer a data question and identify potential bias.

    Objectives:U1-E02

    what the visualization cannot show

  • Evaluate the societal, environmental, and ethical implications of large-scale data collection and processing.

    Objectives:U1-E06U1-37

    large-scale collection is not reached above the line; U1-11 and U1-12 are single-request scope

  • Interpret metadata when using data collected by others.

    Objectives:U1-22
  • Apply appropriate analytic and visualization techniques for categorical and quantitative data.

    Objectives:U1-19U1-20U1-21

    storage type against the thing measured

  • Apply methods for analyzing unstructured and text-based data.

    Objectives:U1-23

    hand-coding the open responses — descriptor not yet resolved against the CSTA export

  • Interpret the results of a data analysis to explain patterns, anomalies, and trends.

    Objectives:U1-23U1-24
  • Analyze how graphical conventions support accurate interpretation and how breaking conventions misleads.

    Objectives:U1-E02

    what the mean conceals

  • Apply ethical principles to data collection, analysis, and communication.

    Objectives:U1-22U1-37
  • Assess how data collection and use may impact marginalized and underrepresented groups.

    Objectives:U1-22
SCAFFOLD4
  • Distinguish data and metadata.

    data and metadata

  • Sort, filter, group, and summarize data.

    sort, filter, group, summarize

  • How design choices impact the interpretation of data.

    design choices impact interpretation

  • How personal data and metadata are collected.

    personal data and metadata collection

4

AI Systems

How a model works, and who is accountable for it.

Covers

Tokenization, embeddings, inference, RAG, prompting, single-pass vs. chat, local vs. cloud, what leaves your machine, auditing output, disclosure, stewardship, who decided.

CSTA

ALG·PSALG·MLALG·IMPRO·RDDAT·IMSYS·IMSOC (all)S1-AIN
CORE8
  • HS-ALG-PS-05AIaudit pillar

    Evaluate AI-generated output to assess bias, accuracy, and potential harms.

    Objectives:U1-25U2-C25U2-C26U3-09
  • Evaluate training data by examining its source, quality, representativeness, potential biases, and privacy implications.

    Objectives:U1-27U1-E11

    what the training data did to the resulting model

  • Design a computing technology using human-centered design principles.

    Objectives:U3-18U3-19
  • Evaluate the ethical implications, societal impacts, and potential biases of rule-based and data-driven algorithms.

    Objectives:U1-27U1-28U1-33
  • Articulate the values embedded in the design of an algorithmic system.

    Objectives:U1-32U1-33

    who decided

  • HS-PRO-RD-18AISDD audit gate

    Evaluate AI-generated code for accuracy, reliability, and alignment with program requirements.

    Objectives:U3-08U3-09U3-C31
  • HS-SOC-HU-43AIstewardship pillar

    Evaluate how human choices in using, designing, deploying, and regulating computing technologies have risks, benefits, and impacts.

    Objectives:U1-38U1-C39
  • Connect computing knowledge and skills to personal goals and career aspirations.

    Objectives:U3-30U3-C32U3-E06

    the "keep progressing after this course" exit — the repository is the student's own GitHub account, not a Classroom org that vanishes when the roster does

REACH13
  • Describe the differences between deterministic and probabilistic algorithms.

    Objectives:U1-30

    reclaimed at REACH: fixed seed against reseed in MuggsOfPrompts makes the distinction observable in ninety seconds

  • Develop a machine learning model for a task using an appropriate tool and dataset, and test its performance.

    Objectives:U1-E10

    block-based or a simple library, classification or prediction; the mathematics is out of scope. The standard's own boundary statement names block-based platforms as appropriate.

  • Evaluate the fundamental technological differences between an emerging technology and established technologies.

    Objectives:U1-10U1-11

    local stack against commercial cloud

  • Investigate computing pathways and the range of careers that use computing.

    Objectives:U3-E07

    reached only in the enrichment tail — descriptor not yet resolved against the CSTA export

  • Analyze the historical trajectory of a computing technology and the developments that made it possible.

    Objectives:U1-E05

    the two-date timeline: cloud-only launch Nov 2022, open weights Feb 2023, models on personal hardware days later — descriptor not yet resolved against the CSTA export

  • Justify the selection of an AI algorithm for a given task.

    Objectives:U1-E10U1-E11

    justify AI algorithm selection

  • Analyze the impacts of an emerging technology.

    Objectives:U1-E06

    impacts of emerging tech

  • Compare human intelligence and artificial intelligence.

    Objectives:U1-13U1-14

    human vs. artificial intelligence

  • Explain the rationale behind laws and policies governing computing.

    Objectives:U1-32

    law/policy rationale

  • S1-AIN-HR-08stewardship, specialty tier

    Plan safeguards for AI systems that protect human well-being and privacy while ensuring meaningful human involvement.

    Objectives:U1-33U1-C39
  • Analyze the potential biases and limitations of AI systems.

    Objectives:U1-27U1-31
  • Integrate a prebuilt AI agent into an application.

    Objectives:U3-06U3-17

    MuggsOfCode

  • Assess how unauthorized data collection has influenced the practice of training AI models.

    Objectives:U1-E09
SCAFFOLD3
  • Use an AI tool to generate outputs.

    use an AI tool to generate outputs

  • Analyze AI-generated code.

    analyze AI-generated code

  • Judge when it is appropriate to use AI.

    when it is appropriate to use AI

Merge rationale. All 17 Computing & Society standards carry the AI-related tag. SOC is not a sibling of Bucket 4; it is substantially inside it. v0.1 proposed a sixth bucket; the tag data resolved it as a merge.

5

Specification & Build

Saying precisely what you want, and judging what came back against it.

Covers

Specification structure, checkable criteria as test cases, evaluation by inspection, amending the spec rather than the code, attribution, defined collaboration workflows, documentation and libraries, publishing.

CSTA

ALG·PSPRO·PDPRO·TRS1-SWD
CORE5
  • Evaluate algorithms for efficiency, correctness, and clarity, using metrics or test cases.

    Objectives:U3-09U3-26

    the specification's checkable criteria are the test cases

  • HS-PRO-PD-14disclosure pillar

    Apply appropriate attribution of intellectual property when developing a computing technology.

    Objectives:U2-17U2-18
  • Collaborate on a programming project using a defined workflow that includes design documentation and clear task roles.

    Objectives:U2-10U3-23

    the adjudication routing designed in Unit 2 is the defined workflow; the Unit 3 peer-spec review names the roles

  • HS-PRO-TR-19spec-driven development, near-verbatim

    Evaluate a computing technology's alignment with design specifications and responsible design values.

    Objectives:U3-09U3-C33
  • Refine a computing technology based on user feedback, testing results, and responsible design values.

    Objectives:U3-20U3-27
REACH4
  • Use documentation, libraries, APIs, and other tools in program development.

    Objectives:U3-E01

    reached only in the enrichment tail this year

  • Design software that accounts for complexity through abstraction.

    Objectives:U3-25
  • Use AI-assisted IDE tools or features to understand unfamiliar code and identify errors during debugging.

    Objectives:U3-16U3-17

    the handoff package — deciding which criteria are worth a larger model's limited free tier

  • Design a test plan that exercises functionality, including edge cases and error conditions.

    Objectives:U3-26

    SDD mechanical gate — the specification's checkable criteria are the test cases

SCAFFOLD2
  • Represent an algorithm as a flowchart or pseudocode.

    flowchart or pseudocode

  • Use variables of multiple data types in a program.

    variables of multiple data types

This bucket was Program Logic through v0.2. Program logic taught for its own sake — tracing, control flow, procedural abstraction, data structures used inside programs — is out of scope for a course whose premise is that you change the output by amending the specification. HS-ALG-PS-01, HS-ALG-PS-02, HS-PRO-PD-12, and HS-PRO-VD-16 are named in *What we are not covering* rather than held here as REACH claims with no objective behind them.

Worked exhibits

The page above asserts tiers and lists identifiers. Eight exhibits below show how a CORE claim is earned — the standard, the operating-spec objectives that carry it, what students actually do, the artifact that constitutes evidence, and why the claim holds up to a skeptic.

Eight, not twenty — the argument, not a catalogue. If a bucket description and a worked exhibit ever disagree, the exhibit is wrong; the operating spec is the authority.

Evaluate AI-generated output to assess bias, accuracy, and potential harms.

Objectives:U2-C25U2-C26

What students do

After building a verification pipeline on paper — independent attempts, correctness and teachability judged separately, reliability scored from seeded items, consensus weighted, disagreements routed — students run the model through that same pipeline and verify its solutions under the rules they wrote.

What it produces

A published accuracy number for the model on ACT-style items, plus the fraction of consensus-cleared items that still needed substantive correction.

Why the claim holds

The evaluation is measured rather than asserted, by an instrument the students designed, on a task they now understand deeply enough to judge. A student who reports 40% needing correction has learned more about their own thresholds than one who reports 5%.

Evaluate AI-generated code for accuracy, reliability, and alignment with program requirements.

Objectives:U3-08U3-09U3-C31

What students do

Write a four-panel specification, compile it to a single-file page, then walk each checkable criterion against the rendered output and mark it: matches, missing, diverges, unspecified. Where it diverges, students amend the specification — never the page — and recompile.

What it produces

The session specification with tracked changes: a documented misalignment and the amendment that closed it, accumulating across every compile.

Why the claim holds

The standard says evaluate; it does not say by reading the source. Walking a specification's criteria against a rendered artifact is specification-based testing, and it is more rigorous than reading code and deciding it looks right. CSTA files this under Reading & Documenting, so a reviewer may question the placement — the tracked-changes artifact is the answer.

Create a data dictionary describing name, type, and allowable values for each attribute.

Objectives:U1-19U1-20U1-21U1-22

What students do

Assign a storage type to every field in the survey they themselves answered in week one, then find where the stored type misrepresents the thing it measured: a Likert response held as an integer that cannot be meaningfully averaged, an open response as an unbounded string.

What it produces

The data dictionary, plus a short limitations note naming who is not in the sample and what nonresponse did to the result.

Why the claim holds

The schema is their own data, which means the misrepresentation is legible rather than abstract. The CS content is how a computer holds what a person typed; the research limitations are secondary and stay short.

Debate the efficacy of a policy or regulation for the responsible use of data.

Objectives:U1-34U1-35U1-36U1-C39

What students do

Write reflections and forward-looking statements about their own AI use, code the resulting corpus by the method they used on the survey data, then synthesize it under a prompt written to preserve dissent. Each student locates their own position in the synthesis or finds it absent, revises the prompt, and the class adopts the result.

What it produces

The class AI policy, governing the rest of the semester, with the places the class split recorded in it.

Why the claim holds

The policy is derived from coded data, never drafted as institutional prose. Finding your position missing and having to change the prompt to recover it is the debate — conducted on the artifact rather than about it.

Demonstrate the capabilities and limitations of a physical or simulated computing device.

Objectives:U1-07U1-08U1-09

What students do

Distinguish VRAM, dedicated GPU, integrated GPU, and unified memory against each other, then work three machines from their specifications: what can each run, what can it not, and what is the single number that decides it. The worked example is the rig serving the classroom — photographed, with its memory usage on screen while a model is loaded.

What it produces

The three-machine analysis, and an exit ticket answering a friend who claims their 16 GB laptop can run what the classroom rig runs.

Why the claim holds

The limitation is demonstrated on a machine that is running while students look at it, and the number on screen is evidence rather than assertion.

Articulate the values embedded in the design of an algorithmic system.

Objectives:U1-29U1-32U1-33

What students do

Author a system prompt, run it against a fixed battery of user prompts with the seed held constant, and document what changed in the output. Then read published system prompts as primary source documents — including those of the tools this class uses — and separate a constraint stated to a model from one enforced by the architecture around it.

What it produces

An authored system prompt with its battery results, and a written distinction between stated and enforced constraints in a named tool.

Why the claim holds

Students install the values themselves before reading anyone else's, so the reading lands on something they have already done. Holding the seed fixed is what makes the difference attributable to the prompt at all; nobody is deceived and nobody is guessing.

Apply appropriate attribution of intellectual property when developing a computing technology.

Objectives:U2-15U2-17U2-18

What students do

Locate every saved passage in the original PDF and attest that they found it. Mark direct quotations by selecting them in the source rather than typing them. Record what was discarded and why.

What it produces

A source session in which every quotation is traceable to a location in the original, and every discard is justified.

Why the claim holds

Quotation by selection makes misquotation structurally impossible rather than discouraged — the same stated-versus-enforced distinction the course teaches, applied to attribution.

Connect computing knowledge and skills to personal goals and career aspirations.

Objectives:U3-30U3-C32U3-E06

What students do

Set up their own GitHub account, commit their specification and compiled artifact to their own repository, read a diff, and account for why the work now survives the course. Deploy through Cloudflare Pages and verify the live URL loads from outside the school network. Watch a custom domain go from purchase to resolution.

What it produces

A repository the student owns and a working public URL, both of which outlast the semester and the roster.

Why the claim holds

The connection to personal goals is literal rather than reflective: the student leaves with an account, a portfolio, and a deployed site they can keep adding to. Unit 1 traced a request out of a closet, through a tunnel, to a Chromebook; this runs the same infrastructure in the opposite direction, to an address of their own.

Coverage matrix

Buckets against CSTA concepts. Empty cells are left visibly empty — the gap check only works if the gaps are legible.

1Systems

SYS
COREREACHSCAFFOLD

2Networks & the Web

SYS
COREREACHSCAFFOLD

3Data

DAT
COREREACHSCAFFOLD
S1-DSC
REACH

4AI Systems

ALG
COREREACHSCAFFOLD
PRO
CORESCAFFOLD
SYS
REACH
SOC
COREREACHSCAFFOLD
S1-AIN
REACH

5Specification & Build

ALG
CORESCAFFOLD
PRO
COREREACHSCAFFOLD
S1-SWD
REACH

Practices (ESR / IC / CT / HCD) — mapping pending. The spec calls for a second matrix row-group mapping the four CSTA practice families onto each bucket. That mapping is not yet in the standards data, so it is named here rather than shown — a stated gap, not a silent one. The four families are listed under CSTA 2026 at a glance.

How the course is built

The map above is drawn from the standards; this is the frame the course hangs on. The operating model — ordered objectives, no per-week day budget, closing objectives that run at the boundary regardless — is the discipline that makes tiers coverage-derived rather than intent.

SemesterAug 4 – Dec 15, 2026
School days86
Content days68 — 18 / 16 / 34
Q1Aug 4 – Oct 2 (42 school days)
Q2Oct 5 – Dec 15 (44 school days)
Blocks5 × 90 minutes per week
Calendar basisInspireNOLA 2026–2027, verified

86 is what the school counts; 68 is what can hold an objective. Both matter because every cut decision is made against the 68, and coverage-based item selection in December writes into that number.

How the 86 school days are spent

How the CS I semester's 86 school days are allocated
WhenDaysWhat
Aug 4–74no course content
Aug 10–123onboarding and instruments (U1-01 through U1-05)
8 test days8major-grade tests — Aug 28 · Sep 11 · Sep 25 · Oct 8 · Oct 23 · Nov 6 · Nov 20 · Dec 11
Oct 21Q1 benchmark
Dec 14–152final exam, on the school's exam schedule
Content68objective-carrying days, distributed 18 / 16 / 34 across the three units

Day shapes. Three day shapes: TEACHING — 10 arrival/Do Now · 70 I-do–we-do–you-do · 10 exit ticket. TEST — 45 review · 45 test. HALF-DAY TEST — test only. The 70 is normally two segments and is not required to be.

Unit windows & calendar load

1

Aug 10 – Sep 11 · 18 content days

Aug 4–7 carries no course content — one standalone activity for early finishers and staggered field-trip returns. Aug 10–12 is the fixed opening (attitudinal chunk A cold on a locked Google Form, cognitive benchmark Tue–Wed, chunks B–C). Ordered material starts Aug 13. Unit closes Fri Sep 11 on Test 2.

2

Sep 14 – Oct 8 · 16 content days

Sep 7 Labor Day and Sep 8 School-Site PD are both out. Q1 benchmark Oct 2. Closes on a half day, Oct 8, and fall break Oct 9–13 follows.

3

Oct 14 – Dec 15 · 34 content days

Longest and most fragmented. Thanksgiving Nov 23–27, HS Fall LEAP Dec 1–18, Q2 benchmark and post-test Dec 3–10, half-day test Dec 11. Self-paced against student-set milestones precisely because students are pulled unpredictably. Dec 9 deploy, Dec 10 defense, Dec 14–15 final exam on the school's exam schedule.

Windows are verified against the InspireNOLA 2026–2027 academic calendar. Aug 3 is not a student day for this cohort. Oct 8 and Dec 11 are half days and do count as content-adjacent (test-only). Each unit ends on a test or the exam schedule.

The retired fourth unit

A fourth unit — a program-logic on-ramp — was planned for the December stretch and has been retired. Program logic taught for its own sake is not what this course is for: the premise is that you change the output by amending the specification, so reading and tracing source has no job here. Those days returned to Units 2 and 3.

The three units

What each unit does, the major work it produces, and which standards it targets by tier. Identifiers link to the CSTA viewer; the full descriptors appear in The Buckets above. A standard can read REACH for one unit and CORE for another — the tier badges in The Buckets show the highest tier it is claimed at anywhere.

1

Foundations — Data, Machines, and Models

Aug 10 – Sep 11 · 18 content days

How AI systems are built, what from, where they run, who steers them, and how this class will use them. The unit produces the class AI policy that governs the rest of the semester — derived from coded student writing, never drafted as prose.

Ordered objectives

Taught in the order below. Closing objectives (C) run at the boundary regardless. The enrichment tail (E) is real objectives, ordered, that run if there is time.

Opening — Aug 10–12, fixed

  • U1-01Complete survey chunk A, cold, before any instruction.
  • U1-02Establish course routines, syllabus, non-negotiables, and the provisional AI rule.
  • U1-03Complete the baseline cognitive instrument, Part 1 and Part 2.
  • U1-04Complete survey chunks B and C.
  • U1-05State what the surveys are for, that they recur in December, and that these responses become the dataset the class analyzes later in the unit.

Ordered

  • U1-06Account for why the AI in this room does not run on the Chromebook in front of them.
  • U1-07Explain how a computer computes.
  • U1-08Distinguish VRAM, dedicated GPU, iGPU, and unified memory.
  • U1-09Determine what a given machine can and cannot run, and why — the two cards in the teacher's rig as the worked example.
  • U1-10Trace one request twice — same prompt, local stack and commercial cloud — naming each hop on both paths.
  • U1-11Compare what each endpoint is permitted to do with what arrived: local storage against a commercial provider's terms.
  • U1-12Distinguish *nothing persists* from *requests are unlinkable to a person*, and say which claim each system actually makes.
  • U1-13Describe pretraining, post-training, freezing, and inference as distinct phases.
  • U1-14Explain parameters at the level of what the model stores.
  • U1-15Explain tokenization at the level of what the model reads, using live token boundaries in MuggsOfPrompts.
  • U1-16Account for why token count, word count, and character count disagree.
  • U1-17Trace a prompt from text to tokens to output, inspecting candidate tokens and their probabilities at three points.
  • U1-18Account for why the model struggles to count letters, spells unusual words oddly, and treats whitespace as meaningful.
  • U1-19Assign a storage type to every field in the survey schema — integer, float, string, boolean, timestamp.
  • U1-20Explain what the machine actually holds for each kind of question.
  • U1-21Identify where a stored type misrepresents the thing measured: a Likert response stored as an integer that cannot be meaningfully averaged; an open response as an unbounded string.
  • U1-22Write the data dictionary, including a short limitations note — who is not in the sample, what the instrument could not ask, what nonresponse does to the result.
  • U1-23Code the class's open-ended responses by hand: familiarize, code line by line, collate into themes, check themes against extracts.
  • U1-24Locate the response that fits no theme; decide to force it or name it; write the criterion used.
  • U1-25Submit the same responses to a model with no special instruction and compare against their own codebook.
  • U1-26Write a prompt that protects the outlier, using the criterion they authored.
  • U1-27Distinguish bias originating in training data from bias installed after training.
  • U1-28Locate control at the layer where change is cheap: weights are expensive and frozen; system prompts, tools, and wrappers are editable in seconds.
  • U1-29Demonstrate installed bias in MuggsOfPrompts — author a system prompt, run it against a fixed battery, document the difference in output.
  • U1-30Hold the seed fixed and account for why that is what makes a difference attributable to the prompt at all; then reseed and account for what changes.
  • U1-31Test whether an instruction holds across a battery or only fits one case, and name that failure as overfitting.
  • U1-32Read published system prompts as primary source documents.
  • U1-33Read the system prompts of the tools this class uses, and distinguish a constraint that is *stated* to a model from one that is *enforced* by the architecture around it.
  • U1-34Write a reflection on how they have experienced AI so far.
  • U1-35Write a closing statement of how they intend to use AI academically going forward.
  • U1-36Code the resulting corpus by the method from 23–26: theme it, protect the outlier, synthesize under a prompt that preserves dissent.
  • U1-37Identify social desirability in a corpus written to their teacher, about AI use, in the class that grades AI use — and state what that implies about which findings to trust.
  • U1-38Distinguish uses of AI that build capability from uses that substitute for it, and state what performance under no-AI conditions requires.

Closing objective — runs at the boundary regardless

  • U1-C39Locate their own position in the synthesis, or find it absent; revise the prompt; adopt the resulting document as the class AI policy.
Enrichment tail12 objectives, if time
  • U1-E01Compare fixed-response and open-ended items as data sources, and identify a theme no closed item offered.
  • U1-E02Visualize the coded results to answer a question about the class, and identify what the visualization cannot show.
  • U1-E03Explain why training is a one-time capital cost and inference is a permanent per-query cost.
  • U1-E04Connect that asymmetry to why few organizations build models and many deploy them.
  • U1-E05Trace the historical trajectory that made local AI possible: cloud-only launch November 2022, open weights February 2023, personal hardware within days.
  • U1-E06Connect the datacenter buildout that trained these models to the memory supply that determines what hardware this room can afford.
  • U1-E07Regenerate a model's thinking trace with a different seed and account for what that shows about whether the trace is reasoning or output.
  • U1-E08Build a timeline of a documented incident in which an instruction change preceded a behavior change, using public commit history.
  • U1-E09Evaluate the competing explanations offered by the company and by its critics.
  • U1-E10Train a machine learning model for a chosen task using an appropriate tool and dataset, and test its performance. *(Block-based or a simple library; classification or prediction; the mathematics is explicitly out of scope.)*
  • U1-E11Evaluate what the training data did to the resulting model, and justify whether this type of model suited the task.
  • U1-E12Compare the September corpus to their August open responses and account for what changed in their own language.

Major work

The semester dataset

Three instruments in Week 1 — cognitive, attitudinal, behavioral — each pairing fixed-response items with open-ended ones. Students are told on day one what the survey is for, that it recurs in December, and that their own responses become the dataset the class analyzes in Week 4. Identifiers are stripped before distribution.

Data dictionary and storage-type audit

A storage type for every field, plus the place each type misrepresents what it measured — a Likert response stored as an integer that cannot be meaningfully averaged, an open response as an unbounded string. The limitations note stays short; the CS content is how a computer holds what a person typed.

Hand-coding, then machine-coding

Familiarize, code line by line, collate into themes, check themes against extracts. Then the same responses to a model with no special instruction, compared against the student's own codebook — and a prompt written to protect the outlier, using the criterion the student authored.

The system-prompt experiment, in MuggsOfPrompts

Students author a system prompt and run it against a fixed battery of user prompts, with the seed held fixed so the difference is attributable to the prompt at all. A battery rather than one case makes does my instruction hold across cases the visible question — a system prompt is a spec, and a battery is its test suite. Students author the bias themselves and nobody is deceived. Published system prompts are then read as primary source documents, and a documented incident is timelined from public commit history.

Seed, reseed, and the thinking trace

Reproducibility first, then an explicit reseed — because same prompt, different output is one of the more important facts about these systems and this is the only place in the course positioned to teach it. Regenerating a model's thinking trace under a different seed produces different reasoning reaching the same answer, which is how students learn a trace is generated text that looks like reasoning rather than the reasoning itself.

The class AI policy

Week 5's writing — reflections, authored system prompts, revised user prompts, aspirational statements — is the corpus. The policy is coded out of it, not drafted as institutional prose. The adopted document records where the class split.

Standards

Anchors are shown bold. Click any code to open the CSTA viewer.

2

Directing and Verifying — MuggsOfSources

Sep 14 – Oct 8 · 16 content days

Verification is the unit's subject. Students build the apparatus that decides whether an answer can be trusted, running it on paper before it runs anywhere else: redundancy, reliability weighting, adjudication. The substrate is ACT-style reading items on passages the class already reads, so the item bank is a byproduct of the reading rather than a project of its own. MuggsOfSources adds the other side of the pipeline, where students are the apparatus: every passage is theirs to locate, attest to, and defend. The unit closes by running the model through the students' own pipeline and publishing its accuracy, measured on passages they have read closely, by an instrument they built.

Ordered objectives

Taught in the order below. Closing objectives (C) run at the boundary regardless. The enrichment tail (E) is real objectives, ordered, that run if there is time.

Ordered

  • U2-01Restate what the class AI policy permits, what it requires, and where the class disagreed.
  • U2-02Answer an ACT-style reading item independently and in silence, then compare within your group and account for where you converged and where you did not.
  • U2-03Pass your group's answer and its explanation on, and check another group's answer for correctness. A binary judgment, quickly made, checkable against the key.
  • U2-04Evaluate that same group's explanation against a teachability rubric: the trap is named, the reason it is tempting is stated, and the sentence in the passage that settles it is identified.
  • U2-05Distinguish the two judgments, and identify an answer that is correct but unteachable, or wrong but well reasoned.
  • U2-06Name the parameters of the process the class just ran: how many independent attempts per item, what counted as agreement, who reviewed whom.
  • U2-07Explain why agreement on correctness is machine-checkable and free, and why teachability is not, and what follows from that asymmetry.
  • U2-08Compute a reliability score per reviewer from performance on seeded items whose quality was known in advance.
  • U2-09Weight consensus by reliability, and justify the weighting against an unweighted scheme.
  • U2-10Design adjudication routing for items where reviewers disagree: to a second review, then to the instructor.
  • U2-11State a claim about your topic before seeing any evidence, and commit to it.
  • U2-12Vet a scholarly source against a checklist: peer review, author credentials, publication venue, date, presence of methods and references, primary versus secondary.
  • U2-13Select supporting passages from a set of unranked candidates, and justify each selection.
  • U2-14Explain why the tool shows no ranking and no score, and what that would let you skip.
  • U2-15Locate every saved passage in the original PDF and attest that you found it.
  • U2-16Write how each passage relates to your claim, in your own words.
  • U2-17Mark direct quotations by selecting them rather than typing them, and explain why that makes misquotation impossible.
  • U2-18Record what you discarded and why.
  • U2-19Conclude, where it is true, that the article does not support the claim — and treat that as a finding rather than a failure.
  • U2-20Account for why PDF-to-text extraction can produce a fluent, grammatical sentence the author never wrote — column merges, lost ligatures, reassembled hyphenation, footnote markers mid-clause.
  • U2-21State the general principle: a machine can be entirely scrupulous and still hand you something false, because the error entered upstream of the part being careful.
  • U2-22Explain retrieval-augmented generation as a pipeline — chunk, embed, retrieve, ground — and identify which stage this tool deliberately omits.
  • U2-23Account for why a passage that supports a claim may share no vocabulary with it, and what query expansion and keyword search do about that.
  • U2-24Compare where the document goes in this tool against where it goes in a commercial one, and account for what *client-side extraction* means for who holds your article.

Closing objectives — run at the boundary regardless

  • U2-C25Run the model through the pipeline the class built: batch-answer the item bank, verify the answers and explanations under the same rules, and publish the accuracy number.
  • U2-C26Account for how often the model was right on passages you have read closely, and for the fact that you produced that number by building the measurement.
  • U2-C27Defend a source session orally: what you claimed, what you kept, what you discarded, and why.
  • U2-C28Perform an unassisted checkpoint and account for the gap.
Enrichment tail7 objectives, if time
  • U2-E01Report what fraction of consensus-cleared items still needed substantive correction, and what that says about the thresholds the class chose.
  • U2-E02Identify collusion as an attack on the process, predict how it would appear in the data, and propose a countermeasure.
  • U2-E03Specify how the process would scale from one classroom to eighty students and hundreds of items, and what breaks first.
  • U2-E04State what each tool did for you and what it could not do for you.
  • U2-E05Analyze how the internet functions as a network of networks and how it differs from other networks.
  • U2-E06Identify security trade-offs in a computing system for users, data, and devices.
  • U2-E07Classify the causes and impacts of security breaches and social engineering attacks.

Major work

The rotation, on paper

Week 6 runs the verification process before naming it: answer and explain independently and in silence, compare in groups, pass the group's answer on, judge another group's twice. Experience before abstraction, the same order as vibecode before spec panels, and paper before the tool in Unit 3. Independence is the design detail that matters most: a group reading a passage together produces one reading with four names on it, which is review but not redundancy. Five silent minutes first gives four independent attempts per item, and convergence becomes something the teacher watches rather than infers.

Correctness is not teachability

Two separate judgments on the same answer. Correctness is a key match, binary and quickly made. Teachability is a rubric: the trap is named, the reason it is tempting is stated, and the sentence in the passage that settles it is identified. Students find the answer that is correct but unteachable, and the one that is wrong but well reasoned. Every item carries at least one labeled distractor type. Seeding one item with an explanation that lands on the right answer for the wrong reason measures reviewer reliability, costs ten minutes to write, and needs no infrastructure.

Formalizing and scaling it

Week 7 names the parameters of what the class just did, then designs the version that would survive eighty students and hundreds of items: reliability scoring from seeded items, consensus weighted by reliability, adjudication routing when reviewers disagree, a countermeasure for collusion, and a statement of what breaks first. The assessed artifact is the review and the pipeline design, never the worked item.

One article, one claim

Vet a scholarly source against a checklist. Commit to a claim before seeing evidence. Select passages from unranked candidates — no score, no ranking, nothing that would let you skip the reading — locate each in the original PDF, attest to it, and write the relation in your own words. Quotations are marked by selection rather than typed, which makes misquotation impossible. Concluding that the article does not support the claim is a finding, not a failure.

Where the error enters

RAG as a pipeline — chunk, embed, retrieve, ground — and which stage this tool omits. Why a supporting passage may share no vocabulary with the claim. Why PDF extraction can produce a fluent, grammatical sentence the author never wrote: column merges, lost ligatures, reassembled hyphenation, footnote markers mid-clause.

Running the model through it

The unit close. Batch-answer the item bank, verify the answers and explanations under the same rules the class wrote, publish the accuracy number, and report what fraction of consensus-cleared items still needed substantive correction, which is what the chosen thresholds actually cost. The model becomes another contributor whose reliability is measured the same way the students' is. The model is being tested on the course's own subject matter, since the passages are about AI.

The closed loop — they build so the bank can be studied from

The bank the class verifies is built so that it can be studied from later, before an ACT. ACT Reading is a scored section of the same test, so the substrate change does not weaken the argument. That framing is what makes the consensus threshold a live argument rather than a housekeeping detail: a student arguing for a stricter one is arguing on behalf of someone who will use this bank. In Louisiana the ACT score attaches to TOPS eligibility, so unreliable verification is unreliable preparation weeks before a test that is money. The reliability score stops being a classroom metric and becomes a forecast: how much can this bank be trusted?

Standards

Anchors are shown bold. Click any code to open the CSTA viewer.

3

Building and Shipping — MuggsOfCode

Oct 14 – Dec 15 · 34 content days

Spec-driven development. Students write a specification, a local model compiles a single-file page from it, and students evaluate the output against what they specified — then amend the specification, never the code. The durable artifact is the spec with its revision history; the compiled page is disposable.

Ordered objectives

Taught in the order below. Closing objectives (C) run at the boundary regardless. The enrichment tail (E) is real objectives, ordered, that run if there is time.

Ordered

  • U3-01Decompose an observed website from memory into four structural categories: Project, Structure, Style, Behavior.
  • U3-02Write a specification in language and structural notation alone, with no visual representation.
  • U3-03Reconstruct a peer's specification, drawing only from what is written, having never seen the source.
  • U3-04Evaluate the precision of a specification by comparing it against a stranger's reconstruction, and identify the specific words or omissions that produced each divergence.
  • U3-05Name the gap types that appear: contradiction, ambiguity, undefined behavior, gap.
  • U3-06Enter the tool in vibecode: describe something, watch it appear, refine by talking.
  • U3-07Write a four-panel specification in the tool: Project, Structure, Style, Behavior.
  • U3-08Compile, and read the artifact against what was specified.
  • U3-09Evaluate three chosen criteria against the rendered output: matches, missing, diverges, unspecified.
  • U3-10Amend the specification — not the page — and recompile.
  • U3-11Account for what changed between compiles and why.
  • U3-12Run a Markdown QC quick pass and record what it surfaced.
  • U3-13Identify specification lines that cannot be checked against any output, and rewrite them as checkable ones.
  • U3-14Identify anything in the artifact that the specification never asked for, and decide whether to constrain it out or write it in.
  • U3-15Distinguish a specification fault from a model ceiling, and justify the call.
  • U3-16Decompose a criterion the local model cannot execute into simpler ones, and re-evaluate.
  • U3-17Build the handoff package: identify which criteria are worth a larger model's limited free tier.
  • U3-18Design a page using human-centered design principles and name its intended user.
  • U3-19Identify accessibility requirements and specify them as checkable criteria.
  • U3-20Complete a full cycle against that design: specify, compile, evaluate, amend.
  • U3-21Read a partner's exported specification and predict what it will produce.
  • U3-22Compile a partner's specification and compare the result to your prediction.
  • U3-23Give and act on structured feedback naming specific specification lines.
  • U3-24Write a capstone specification with scope, criteria, and constraints.
  • U3-25Decompose the capstone into buildable parts.
  • U3-26Design a test plan covering edge cases and error conditions — the specification's checkable criteria are the test cases.
  • U3-27Build against the specification and refine on what evaluation finds.
  • U3-28Set a project plan: scope, milestones, and the date each milestone is due.
  • U3-29Direct a self-paced project against milestones you set, reporting progress and blockers.
  • U3-30Set up a repository, commit the exported artifact and its specification, and read a diff — and account for why the work now survives the course.

Closing objectives — run at the boundary regardless

  • U3-C31Assemble the portfolio: the session specification with tracked changes, QC findings, and the amendments each produced — legible to a reader with the student not in the room.
  • U3-C32Deploy through Cloudflare Pages and verify the live URL loads from outside the school network. *(Dec 9)*
  • U3-C33Defend the build against its specification. *(Dec 10)*
Enrichment tail7 objectives, if time
  • U3-E01Use documentation and libraries where the specification calls for them.
  • U3-E02Refine based on what the output evaluation found.
  • U3-E03Read and review a peer's exported specification and artifact.
  • U3-E04Collaborate using a defined workflow with task roles and design documentation.
  • U3-E05Evaluate a build's alignment with its specification and with responsible design values.
  • U3-E06Follow a custom domain from purchase to resolution — ten dollars, DNS pointed, watched live. Demonstration only; the free Pages URL is the assignment.
  • U3-E07Connect computing knowledge and skills to personal goals and aspirations, and investigate computing pathways.

Major work

Sketch, Spec, Swap — paper first

Decompose a remembered website into Project, Structure, Style, Behavior. Write the spec in language alone; a stranger reconstructs it having never seen the source. The divergences name the gap types: contradiction, ambiguity, undefined behavior, gap. First tool contact is vibecode only — no spec panels yet.

The spec-amend loop

Four panels, compile, read the artifact against what was specified, a Markdown QC quick pass, three student-chosen criteria evaluated against the rendered output, amend the spec, recompile. The discipline is stated once and enforced all unit.

Peer-spec compilation

The Week 11 swap repeated with the model as the reconstructor. Read a partner's exported specification, predict what it will produce, compile it, and compare the result to the prediction.

The handoff package

The tool exports the spec, the current artifact, and a continue-from-here instruction, meant for a larger model's free tier. Scaffold cheaply on local hardware; spend borrowed tokens only where they matter. This is a budget decision on the student's own project.

The portfolio — a repository review, not a presentation

Self-paced against student-set milestones. The student sets up their own GitHub account and commits their specification, exported artifact, and diffs to a repository they own. Dec 9 deploy to Cloudflare Pages and verify the live URL loads from outside the school network; Dec 10 defend the build against its specification. The portfolio — the session specification with tracked changes, QC findings, and the amendments each produced — is reviewed as a repository without the student present. It has to survive without narration: README stating what the project is and who it is for, the spec with tracked changes, QC findings, amendments.

Standards

Anchors are shown bold. Click any code to open the CSTA viewer.

This is a pilot semester of a course that exists nowhere else. Tiers are conservative on purpose — CORE is now defined as *carried by an objective above the enrichment line*, not as *the teacher intends to reach it*. Every demoted standard stays on the page. Year two will assign tiers from the coverage record — dated content-day claims made before scores exist — rather than from intent.

HS-ALG-PS-05 is Unit 2's strongest claim and it is not made with a worksheet. U2-C25 and U2-C26 run the model through the verification pipeline the class designed and publish its accuracy on a task the students now understand deeply. That is evaluating AI-generated output for accuracy — measured, by an instrument they built, rather than asserted.

HS-DAT-DC-24 and HS-PRO-PD-15 both come from the pipeline. Verifying consistency against expected values is what gold standards and consensus thresholds do; a defined workflow with task roles and routing is what adjudication is. Neither is a programming project in the conventional sense, and PD-15's wording is broad enough to hold — but note it, because a crosswalk reviewer may expect code.

HS-ALG-PS-04 is reclaimed at REACH rather than culled. Holding a seed fixed and then reseeding makes deterministic versus probabilistic observable in ninety seconds, which is a genuine touch — U1-30 carries it.

HS-PRO-RD-18 is claimed on output inspection, not source reading. The standard says evaluate AI-generated code for accuracy, reliability, and alignment with program requirements. It does not say by reading it. A student who walks every criterion in their specification against the rendered artifact and finds a divergence has evaluated the generated code — that is specification-based testing, and it is more rigorous than reading source and deciding it looks right. The evidence is the tracked-changes specification: a documented misalignment and the amendment that closed it. CSTA files RD-18 under Reading & Documenting, whose framing is about reading and interpreting code; a crosswalk reviewer could raise the placement, and the tracked-changes artifact is what answers it.

HS-PRO-RD-17 is not claimed. Analyzing how a segment of code works — parameters, return values, data structures — is code as an object of study, which has no job in a course whose premise is that you change the output by amending the specification.

HS-ALG-PS-03 and S1-SWD-TR-07 are stronger than they look on paper. Under this model the specification's checkable criteria are the test cases, so evaluating with metrics or test cases and designing a test plan are the mechanism of the unit rather than exercises beside it.

HS-SOC-CE-46 is CORE on the strength of U3-30 and U3-C32, and that only holds if the repository is the student's own GitHub account rather than a Classroom org that vanishes when the roster does. The student sets up their own account and commits their own portfolio — the standard's *personal goals* language is literal here, not aspirational.

For calibration: a CSTA-validated, professionally built AI curriculum reached 63% alignment across the 46 high-school foundational standards. Full coverage is not the bar.

One caution from the standards themselves

CSTA states plainly that computer science is not using technology tools like word processors, slide decks, or generative AI tools — those appear in every subject. The prompt-writing strand, the policy document, and Unit 2's tool work are the three places this course could drift toward AI-use-as-content. Each is anchored: prompt writing sits on top of tokenization and system prompts, the policy is derived from coded data, and Unit 2 is evaluative rather than consumptive. Keep the anchors.

How the tools hold the design

Three constraints on the purpose-built tools that are load-bearing curriculum decisions rather than scoping ones. Each is the reason a claim above survives contact with a classroom.

Why reading is the right substrate

Reading is the substrate for the verification pipeline, and the correctness and teachability asymmetry that Unit 2 rests on is cleaner on reading than it was on math. Correctness is a key match: free, binary, machine-checkable. Teachability is naming which trap a distractor is, why someone would fall for it, and what in the passage settles it. That needs a rubric and a human. A student can explain why a distractor is tempting without being able to work the problem the item is drawn from, which matters in a required course where nobody should be locked out of the teachability judgment by their math level.

The course reads sources regardless, so the item bank is a byproduct of the reading rather than a project of its own. By December that is fifteen to eighteen worked passages on tools, capabilities, jobs, and the industry, each of which is simultaneously the reading, the ACT practice, the Unit 2 corpus, and the evidence base for the class AI policy.

Every item carries at least one labeled distractor type. The four are: true in the world but not in the passage; true in the passage but answering a different question; half right, wrong attribution; and extreme wording (always, never, all, none, proves). The label lives in the key, not on the student's screen.

MuggsOfMath is retired for the pilot year. Not deferred, retired. Google Forms with a locked column schema does everything the model-free core was specified to do: item loads, answer captured, string comparison against a key, export to the longitudinal table. The tool was solving a problem the existing pipeline already solves. Item codes MB-01 through MB-10 stay on disk, and year two decides whether they come back.

Paper before software, and why independence matters

It runs on paper before it runs anywhere else. U2-02 through U2-05 is groups answering reading items and rotating; U2-06 through U2-10 names what happened and designs the scaled version. Experience before abstraction, the same order as vibecode before spec panels, and paper before the tool in Unit 3. Nothing in U2-01 through U2-10 requires software, which also means Unit 2 opens with a floor under it. Independence is the design detail that matters most: a group reading a passage together produces one reading with four names on it, which is review but not redundancy. Five silent minutes first gives four independent attempts per item, and convergence becomes something the teacher watches rather than infers.

The closed loop — they study from what they build

The corpus is the durable asset, not the tool. Tools get rewritten. Models change, frameworks change, the interface will be rebuilt at least twice. A verified bank of 160-odd items with distractor mappings survives all of it, which is why the schema deserves more care this fall than the tool does. At three-way redundancy across roughly 160 items with 75–80 students, each student authors somewhere between six and eighteen of them, about 4–11% of the bank, so if any of these students later studies from it, the vast majority is fresh to them. Contamination is smaller than it looks.

On reading it is smaller still. The math version had a real contamination question: a student who authored an item and later studied from the bank had seen it. On reading, every item is attached to a passage. A student who re-reads a passage they worked in September is doing the thing the practice is for. Recognizing your own item is review, not contamination. Authoring an item is stronger preparation than practicing it in either version, and recognizing your own item later is a good moment.

Curation — the class drafts, the instructor is the target pass

Curation is a stage in the design, not a repair of it. The corpus the class produces is a draft, and it is cleaned before any of it reaches the tool. That relationship has a name students meet in Unit 1: speculative decoding. A cheap draft model proposes several tokens at once; a stronger target model verifies them in a single pass, keeps the run that matches, and corrects at the first divergence. Here the students are the draft — three of them independently working the same item, cheaply and in parallel — and the instructor is the target pass that accepts or corrects before anything ships. The draft is not the lesser thing in that arrangement. It is what makes the expensive pass affordable, and a draft that is usually right is the whole reason the scheme is worth building.

What the design requires is that the target pass be visible rather than silent. If the bank is trustworthy because the pipeline the students designed worked, the reliability numbers mean something. If it is trustworthy because everything was quietly rewritten afterward, the pipeline is decoration and the numbers measure something that does not determine the outcome — precisely the failure this course exists to name. Made visible it becomes data: of the items that cleared three-way consensus, what fraction still needed substantive correction? Five percent means they built something good and they will know why. Forty percent means their thresholds were wrong and they will want to know where, which is the best possible thing to hand a class that just spent two periods arguing about consensus rules.

What keeps the correction figure honest

Three rules keep that number honest. Key-gate failures are blocking, not a review queue — an item whose solution does not reach the keyed answer never ships to students, full stop, because the stakes here are a wrong solution studied from weeks before a test tied to TOPS eligibility. The two kinds of edit are separated and only one is counted: normalizing notation and tightening prose are invisible and need no logging, while fixing a wrong intermediate value or restructuring reasoning is substantive, and only that category belongs in any reported figure. And the pass is committed in two steps — the generated version, then the cleaned version — so the git history is the audit trail and the precision analysis generates from the diffs rather than from anyone's memory.

Why a weak model is the right instrument

A frontier model would ruin both MuggsOfPrompts labs. It produces good output from a vague system prompt, and a student who writes something careless and gets something decent concludes prompt engineering is vibes. A small model's weakness is the instrument's sensitivity — vague instructions produce visibly worse output, specific ones visibly better — and flatter probability distributions mean clicking a token shows genuine competition between candidates rather than a column of 0.99s. This has a direct consequence for fallbacks: comparable must mean comparably small, not comparably capable. Falling back to a frontier API does not preserve the lab; it inverts it. The same holds for MuggsOfCode, where the loop depends on a weak spec producing visibly weak output.

Stated versus enforced

Stated versus enforced is the transferable idea, and it comes from the tools themselves. MuggsOfSources has the model emit sentence IDs, so fabrication is structurally impossible rather than discouraged — the guarantee lives in the architecture, not in an instruction the model is asked to follow. MuggsOfCode forbids the AI from writing spec content and enforces it by never routing model output into the spec panels. A student who can ask is this behavior promised, or enforced? has a test that outlasts every tool in this course.

Build deadlines — two tools ship this semester

MuggsOfPrompts

Aug 24

Highest-priority build item. U1-15 through U1-18 (tokenization live, candidate probabilities, why token count disagrees with word count) and U1-29 through U1-31 (the system-prompt battery with a fixed seed) land in the second half of August under the operating-spec ordering. Earlier v18 artifacts listed Aug 19 and Sep 1; the current build-need-by date is Aug 24, aligned to the Week 3 opening under the calendar's day-by-day plan. Minimum viable form: fixed seed, fixed battery, outputs, and the seed displayed. The entire what-happened panel sits behind one toggle and can ship empty.

MuggsOfSources

Sep 21

Not the soonest, but the deepest — sentence IDs, attestation, extraction damage, and quotes-by-selection are U2-11 through U2-19 objectives rather than conveniences. That block has no floor without it.

MuggsOfMath

retired for pilot year

Retired, not deferred. Google Forms with a locked column schema does everything the model-free core was specified to do: item loads, answer captured, string comparison against a key, export to the longitudinal table. The tool was solving a problem the existing pipeline already solves. Unit 2 has moved to ACT-style reading items on passages the class already reads, which turns the item bank into a byproduct of the reading rather than a project of its own. Item codes MB-01 through MB-10 stay on disk; year two decides whether they come back.

Assessment

A pre/post growth instrument spanning two bands, unit assessments for grades, and a portfolio in place of a sit-down final.

Aug 10–12
Cognitive benchmark — 100 items in two sittings of 50 (Aug 11, Aug 12). Separate forms, editing after submit OFF, permanent item codes (item 7 stays item 7 in December). Participation grade per part, completion only. Attitudinal survey in three chunks: A cold on paper Aug 10 before any instruction; B Aug 11; C Aug 12. Editing ON, resumable, partial response is acceptable and is not made up. Not graded.
Baseline. The survey responses become the dataset the class analyzes later in Unit 1 — the data dictionary, the hand-coding, the corpus that becomes the AI policy.
Aug 28 – Dec 11
Eight major-grade tests on fixed dates: Aug 28 · Sep 11 · Sep 25 · Oct 8 (half day, test only) · Oct 23 · Nov 6 · Nov 20 · Dec 11 (half day, test only). Standard shape is 45 review / 45 test. Minor grades from Do Nows, exit tickets, and classwork.
Major/minor grade backbone. Sep 11 closes Unit 1; Oct 8 closes Unit 2.
Oct 2
Q1 benchmark — coverage-selected subset of the August instrument, drawn from item codes marked delivered on the coverage record
Graded. Item-level results captured; the coverage record is what makes the selection defensible before scores exist.
Unit 2 (Sep 14 – Oct 8)
Unassisted checkpoints — U2-C27 (defend a source session orally) and U2-C28 (a solo checkpoint the student accounts for)
The gap between assisted and unassisted performance is the measure, and students account for it themselves. The only oral and only unassisted assessments in the unit.
Dec 3–10
December post-test — coverage-selected subset, expected ≈ 70 items, one sitting. Selection from the coverage record, never from August scores.
Growth measure. Windowed across Fall LEAP because students are pulled unpredictably; offered repeatedly until everyone has taken it.
Dec 9 · Dec 10
U3-C32 deploy through Cloudflare Pages and verify the live URL loads from outside the school network; U3-C33 defend the build against its specification
Terminal performance, staged so a pulled student can rejoin. Runs the Unit 1 request-trace in the opposite direction, to an address of the student's own.
Dec 14–15
Final exam on the school's exam schedule
The school's final.
Unit 3 (throughout)
SDD portfolio (U3-C31) — the session specification with tracked changes, QC findings, and the amendments each produced, in the student's own GitHub repository. Reviewed as a repository without the student present.
Authorship demonstrated through the artifact that survives every recompile — must be legible unnarrated. In place of a sit-down capstone presentation.

COVERAGE RECORD. One line per content day, written into the weekly lesson plan already submitted the Wednesday before delivery — no new document, no new deadline. Format: `DATE | OBJECTIVE CODES DELIVERED | ITEM CODES COVERED`. "Delivered" means taught and practiced, not mentioned; a day lost to an assembly is recorded as not delivered, and that entry is the finding. December's item subset is codes with at least one covered date — a filter rather than a judgment. Dated before any score exists, which is what makes coverage-based selection defensible where score-based selection would not be.

ITEM CODES. Permanent, position-independent. Item 7 stays item 7 in December even if 4, 5, and 6 are not asked. Flat serial as identity; domain is a spreadsheet column. Objectives carry item codes, added when each unit is built out — retrofitting this in December is the weekend the design exists to avoid.

ITEM DISTRIBUTION FOLLOWS THE ORDERING, NOT THE BUCKETS. Objectives near the top of a unit carry more items; enrichment-tail objectives carry few or none. Items written for content that does not run get filtered out in December, so weighting by rank protects the post-test.

The instrument spans two bands (MS and HS) so that a cohort arriving below grade level lands somewhere real rather than at the floor. Score resolves to highest band answered reliably, per bucket — not percent correct.

Matched pre/post pairs are the reported figure. Unmatched students are excluded from the growth number, and that count is stated.

The Unit 1 corpus writing (U1-34 through U1-C39) is not anonymous and students should say so: it is written to the teacher, about AI use, in the class that grades AI use. Social desirability is a finding to surface, not a flaw to apologize for, and it bears on which findings to trust (U1-37).

Volume of revision is never read as effort. A student thrashing toward something that looks right generates many tracked changes; one who wrote a precise spec first generates few and did better.

What we are not covering, and why

163 standards are out of scope. Named omissions with reasons beat silent gaps.

HS-ALG-PS-01, HS-ALG-PS-02, HS-PRO-PD-12, HS-PRO-VD-16

Program logic taught for its own sake — data structures used in algorithms, procedural abstraction and control-structure optimization, modular programs, data structures inside programs — is out of scope for a course whose premise is that you change the output by amending the specification, not by reading and refining source. These sat at REACH through v18 with no CS I objective behind them; the roadmap rework retired the claim rather than continuing to hold it as decoration.

HS-PRO-RD-17

Analyzing how a segment of code works — parameters, return values, data structures — is code as an object of study, which has no job in a course whose premise is that you change the output by amending the specification. HS-PRO-RD-18 is still claimed, on output inspection rather than source reading; the tracked-changes specification is the evidence.

S1-AIN-DD-02/03/04, S1-AIN-DS-06/07, S1-AIN-DD-05

Neural-network internals, supervised-learning applications, and evaluating whether an AI or non-AI solution fits a problem — out of scope for a foundational course. S1-AIN-DD-05 moved here in the roadmap rework: no CS I objective carries it, and the earlier CORE claim on it was intent rather than placement. HS-ALG-ML-08 is not in this row — the block-based training lab in the U1 enrichment tail claims it at REACH, which the standard's own boundary statement names as appropriate.

S2-* (all 62)

Specialty II is the advanced tier; it assumes Specialty I completion.

CYB (27), GMD (17), PHY (16)

Different specialty pathways. Not this course.

XCS (5)

Interdisciplinary integration into other subjects.

E*-* (elementary)

Below band.

Remaining MS foundational (~33)

Prerequisite content not load-bearing for this course's targets. MS-ALG-PS-01/03/04 and MS-PRO-RD-17 sit here alongside the standards demoted when the program-logic module was retired.

HS-SYS-SE-33, HS-DAT-DC-21, HS-SOC-HI-39, HS-SOC-ET-42

Genuine content, no room. Candidates for year two, when the coverage record from this year makes the case for what to add.

CSTA 2026 at a glance

Concepts

  • ALGAlgorithms & Design
  • PROProgramming
  • DATData & Analysis
  • SYSSystems & Security
  • SOCComputing & Society

Practices

  • 1–2ESR — Establishing Supportive Relationships
  • 3–5IC — Inclusive Computing
  • 6–9CT — Computational Thinking
  • 10–12HCD — Human-Centered Design

Identifier convention

BAND-CONCEPT-SUBCONCEPT-## — e.g. HS-ALG-PS-05. Numbering is continuous across concepts within a band, so an identifier is not predictable from its subconcept — always resolve against the source.

Specialty areas take an S1- / S2- prefix: AIN · CYB · DSC · GMD · PHY · SWD · XCS (prefixed S1- / S2-).

What changed from 2017

The 2017 Computing Systems and Networks & the Internet concepts were merged into a single Systems & Security concept in 2026. Buckets 1 and 2 keep them separate for assessment (different item types, separate diagnostic reads) — that is a deliberate choice, not a mapping error.

AI-related tag distribution

73 of 226 standards carry the AI-related tag. The concentration in Computing & Society is why that concept merges into Bucket 4.

AI-related tag distribution across CSTA 2026 areas
AreaTaggedTotal
Artificial Intelligence (S1+S2)2424
Computing & Society (MS+HS)1717
Algorithms & Design (MS+HS)1622
Systems & Security (MS+HS)617
Data & Analysis (MS+HS)417
Software Development (S1+S2)317
Programming (MS+HS)218
Cybersecurity (S1+S2)127