You’ve probably experienced this by now, and probably more than once.
An assignment comes back and it’s fine. The grammar is clean, the structure is sound, the content is (as far as you can tell) correct. And something about it rings hollow in a way you can feel but can’t quite put your finger on. You read it again, looking for the error—for something—that would give you cause to make a useful note in the margin. But the hook isn’t there. There is nothing to respond to. There is nothing there.
You reflexively wonder if the student wrote it. And while that’s an understandable reaction and question, it’s also a dead end. You can’t answer it reliably, with certainty.
And here’s the thing: even if you had that answer, it wouldn’t tell you the thing you actually want to know.
The better question is the one in the title. If a machine can produce a passing response to your assignment in nine seconds, what was that assignment measuring in the first place?
The scale we already had
We’ve described the Levels of Knowledge for decades now. The version we like best lays them out like this:
- Level 5: Research, Inventions, or Performances (Creative Enterprise)
- Level 4: Working Expertise (Problem Solving)
- Level 3: Transferable Knowledge (Generalizing)
- Level 2: Conceptual Understanding (Teaching)
- Level 1: Information (Memorization)
Look closely at the parenthetical for each level, because that’s the test…the thing you can do that proves you’re there. Level 1: you can state it. Level 2: you can teach it. Level 3: you can carry it somewhere new. Level 4: you can solve with it.
Now consider what a language model does. It states facts and definitions. It articulates understanding, describes relationships, sees linkages, seeks underlying principles. It will explain a concept to you at whatever depth you ask, adjust when you say you don’t follow, and produce three to five fresh analogies on request.
That is a Level 2 performance. Our own test for conceptual understanding—can you teach this to someone else?—is now something software does on demand, nicely, for free, and at midnight.
This is a claim about evidence, not about learning. Nothing about how human beings learn has changed in the last three years (since AI became common and accessible). Conceptual understanding is still hard-won, still necessary, and still the floor everything above it stands on. What changed is that a submitted artifact at that level no longer demonstrates the student got there.
The part we would rather not say out loud
If your ability to fairly evaluate student learning performance collapsed the moment a language model showed up, it was very likely sitting at Levels 1 and 2 all along.
That was just as true 10 years ago. We simply couldn’t see it because producing a competent summary cost a student four hours, and four visible hours of effort seemed like they must have produced learning. Sometimes they did. Often they produced a summary—the work product.
AI did not break evaluation. It removed the labor that was concealing what evaluation had been measuring. That is an unpleasant gift, more reminiscent of something from Pandora than a box with a bow under a Christmas tree, but it IS a gift. And we think the institutions that treat it as one, learning its lesson, will come out of this decade in far better shape than the ones buying detection software.
Which word we’re using
The primary distinction Pacific Crest has been insisting on for years matters more now than it ever and it’s worth a quick review because it has a great deal of bearing on where teachers find themselves now.
Evaluation determines the level of quality of a performance. A stakeholder requests it and has skin in the game as far as criteria go. A report describes the level of quality attained and it exists to support a decision. A grade is an evaluation.
Assessment produces feedback to strengthen a future performance. Optimally, the assessee requests it and helps set the criteria. And the report describes what made the performance strong and what would make the next one stronger—and it says nothing about the level of quality at all.
Look again at what’s happened in the last three years:
What generative AI broke is evaluation. Evaluation rests on an inference: that an artifact (work product) is evidence of the capability of the person who submitted it. AI severed that inference, and no rubric repairs it.
What AI cannot break is assessment—not because assessment is harder, but because there is no product to outsource. The object was never a product…it was the next performance.
Nobody can be assessed on your behalf.
Juggling
Here is where the ground shifts, and the clearest way to see it is an example we keep coming back to. (What can we say? It’s an extremely useful example that’s easy to understand.)
Suppose you want to learn to juggle. There are underlying principles: catching means reacting smoothly as the object enters your hand; throwing means accounting for the object’s weight and shape; you watch the object in the air, never the one in your hand; and you have to attend to the throw as much as the catch. Four principles. You have just read all of them.
You cannot juggle.
Now ask a language model. It will give you those four principles, more precisely than described above. Ask it how juggling changes with three potatoes instead of beanbags. Ask about torches, about clubs of unequal weight, about juggling while walking…it will answer every question you ask, fluently and correctly.
You still cannot juggle.
This is a critical distinction useful to anyone redesigning (or even just tweaking) a course this fall: generalizing knowledge is not the same as transferring knowledge. Transferring is applying something in a new context. Generalizing is developing enough working expertise with the underlying principles that you can transfer at will and to contexts nobody hands you.
An LLM can hand a student a transfer. It will produce the application, in the new context, on request, and do it well. What it cannot do is generalize on the student’s behalf, because generalizing is a change in the learner, not a property of the output. There is no artifact to pass across. You cannot ask a machine to juggle for you. It can describe every principle and name every context, and it cannot move your hands.
This isn’t a claim about what the technology will or won’t be able to do—that line keeps moving, and will continue to do so. It’s a claim about where the development happens. Generalizing changes the learner, so the learner has to do it, no matter how capable the tool on the other side becomes. That is where our evaluations need to live now.
What to do on Monday
The Methodology for Generalizing Knowledge gives the sequence, and its middle four steps are the key pattern:
Familiar → Similar → Different → Unfamiliar
Apply the knowledge in the context where it was learned. Then in one that is less familiar but recognizably similar. Then in one with key differences. Then in one well outside the comfort zone. Beanbags, potatoes, juggling while walking, torches.
Three things follow from taking that somewhat amusing progression seriously.
Move the context, not the topic. The common redesign instinct is to pick a harder subject. That doesn’t help because the model isn’t struggling with your subject. What it cannot know is which of your course’s contexts a particular student has and has not been in. “Explain the second law” is answerable by anyone and anything. “Apply it to the compressor failure we analyzed in week four, and say where the analogy breaks” is answerable by someone who was in the room.
Evaluate the choice, not the answer. Require students to justify the approach—what they ruled out, which assumption they accepted, where they were unsure. Reasoning is the thing being judged. A student who cannot defend a choice does not own the knowledge, whatever produced the paragraph. (As a side note, years ago a professor found out that the answer keys to a chemistry activity had been making the rounds among his students. He contacted us and asked if we had a different version of the key questions and exercises available. We did not. What we did have and shared with him was the idea of process and validation. Instead of grading the answers the students arrived at, ask the students to validate the choices they made in their calculations and demonstrate how their answer was correct. It worked well in a case where the students were armed with the correct answers for a single activity. It will work just as well in cases where they can generate correct answers at will.)
Stop treating one hard assignment as the whole design. This is where most AI-era redesigns fail. A single “authentic” capstone is a leap straight to Unfamiliar, and what it often produces is a student who copes. Coping is not generalizing. The progression is the pedagogy—four contexts of deliberately increasing distance, which also means four chances to catch a student who is drifting rather than one autopsy in week fourteen.
One more thing worth noticing
Step 1 of the Methodology is Validate Meaning: confirm the learner is genuinely at high Level 2 before trying to generalize anything, because new generalized knowledge only builds on knowledge already generalized. Skip it and whatever you build is fragile and falls apart on transfer.
Which raises the question this whole article has been circling. If students can now produce Level 2 artifacts without reaching Level 2, how do you validate that they’re standing on anything at all?
You will have to ask them.
In person,
out loud,
early,
and more often than is convenient.
That’s not a workaround or some kind of special academic cludge; it’s ASSESSMENT—the process many of us have described as a something it would be nice to have in the classroom for the llast 20 to 30 years while evaluation quietly did the real work.
Evaluation is now doing considerably less of it. Which means the practice we may have treated as optional is the one we’re going to need going forward.
Where to take this
The International Academy of Process Educators is forming a Special Interest Group on AI in the Classroom, alongside groups on lifelong learning and self-growth. If these questions are live for you this fall—and if you are teaching, they are—that’s where the conversation continues with people working the same problem in their own courses. Reach out to them at www.processeducation.org
This is not something any of us solves alone in a syllabus revision over a long weekend.
