Steven Schwartz had been practicing law for thirty years when he asked ChatGPT to find supporting case law for a personal injury brief. It gave him six cases. Each came with a name, a docket number, a judge, and a holding. Each was cited with the fluent confidence of a system that has never once said I don’t know. He never verified them against a reporter, the published record of real court decisions, before filing. When opposing counsel couldn’t locate the cases, he asked the tool that had invented them whether they were real. It said yes. He filed that answer with the court.
None of the six cases existed. The case is Mata v. Avianca (S.D.N.Y. 2023). Judge Kevin Castel’s sanctions order runs to forty-six pages, and reading it produces a specific kind of discomfort. The discomfort is not because Schwartz was careless in any ordinary sense. It is because at no point does he appear to have done what a lawyer with thirty years of citation-checking in his hands would once have done without thinking about it. That lawyer would have opened a reporter and felt for the shape of a real case. He would have noticed the smoothness where a genuine one would have had grain.
He had the credential. Whatever reflex should have stopped him at the reporter wasn’t running that day.
What matters for what follows isn’t whether Schwartz once had that reflex and let it go quiet. Nor is it whether he never built it as firmly as thirty years might suggest. Either way, the reflex is not information you can look up when you realize you need it. It has to already be running before the need arrives. And it is built by exactly the kind of work the tool had just done for him.
This is not a story about carelessness. It is a story about where the reflex actually comes from, and what happens when it isn’t there when the boundary case arrives. The boundary case is the one that looks almost right and is wrong.
The Three Stages
Technology has always moved the human role through the same sequence, one stage at a time.
First comes the craftsman, who does the work directly. The craftsman’s judgment is embodied. It is built by running the attempt and sorting what comes back into one of three things: confirmation, disconfirmation, or a flaw in the attempt itself. Done long enough, the sorting becomes fast enough to feel automatic.
Then comes the curator. The curator selects, edits, approves, and governs what the technology produces, rather than producing it directly. A curator is competent only to the degree the curator can still recognize the work from the inside.
Then comes the teacher. The teacher no longer evaluates individual outputs. The teacher shapes the system’s future output for everyone downstream. That makes the teacher a corrective signal, and its value depends entirely on the judgment behind it.
Neither of the later stages is a decline from the one before. Each is a genuine form of value, and each depends on something the prior stage built. The next section defines the compound precisely. For now, think of it as the working unit formed when human judgment and AI output operate together. A curator who can evaluate a compound’s output only by comparing it to the compound’s own prior output is not a curator in the sense that matters. That is fluency, not evaluation. A teacher whose only frame of reference is what the system has already produced is training it toward its own average, not toward the work.
The dependency is not sequential coincidence. It is structural. Curation requires knowing what the work actually is, from the inside. The curator has to know it well enough to recognize when an output looks right and is wrong. That knowledge is built by independent, consequence-bearing practice. You generate an output. You diagnose where it went wrong. You recover. And you do all of it without an existing answer available as the standard. Teaching requires curation’s judgment, aggregated and generalized. Each stage is derivatively dependent on the one beneath it. It is not lesser than what it depends on. But it is unreachable without it.
In short: craftsman, curator, and teacher are three different jobs, and each later one is real work. But each stands on the one before it. Nobody can judge the work, or teach a system to do it better, unless that judgment was first built by doing the work.
The Compound Threshold
AI: Endgame 2 established where the technology crosses from tool to something structurally new. Heavy use of AI by a practitioner does not, by itself, cross that line. The line is crossed when the institution reorganizes training, accountability, and expected performance around output the practitioner is no longer expected to produce alone. One way to tell: would removing the AI component merely reduce efficiency, or would it dismantle the productive architecture built around the work? If it would dismantle it, the threshold has been crossed. The reorganized human-plus-AI unit on the far side of that line is what this series calls the compound.
The clearest indicator is the training program itself. Below the threshold, an institution trains independent practitioners who later learn to use AI. Above it, the institution trains people to operate the compound from day one.
That threshold is also, usually without being named as such, where an institution stops training craftsmen by default.
Crossing it does not guarantee the loss. An institution can cross it and still deliberately preserve independent practice elsewhere in its pipeline. This piece returns to that possibility. Absent that deliberate preservation, though, once an institution trains for the compound from day one, it has stopped training craftsmen.
Put plainly: the threshold is the point where an institution rebuilds its training around the human-plus-AI unit. Unless it deliberately keeps independent practice alive somewhere, that is also the point where it stops producing craftsmen.
What the Craftsman Built
The judgment built by skilled work is tacit. It is carried and used, but never fully put into words. That is the single, patient argument Richard Sennett spent The Craftsman making. The judgment develops through the resistance of materials. Think of wood that splits along a grain you have to learn to feel. Think of a joint that looks sound and isn’t, or a citation that reads plausibly and is invented. And it is not fully teachable, because it is not fully articulable. You cannot hand someone the finished judgment. You can only hand them the conditions under which they build their own. That means handing them the struggle rather than sparing them from it.
A writer on AI tutors converges on the same structural point from a completely different direction. Bill Gates, writing in August 2026, describes a system that preserves what he calls productive struggle. The domains don’t match. A single tutoring session is not a career of embodied practice. But the mechanism holds at both scales. The system gives a student the full explanation when a concept is new. Then, at the moment it checks whether the student has actually understood it, it withholds the answer and lets the student work it out. Same tool, same student, same material. The entire difference between whether learning happens turns on when the answer is released. A tutor that hands over the answer at every stage is not a lesser tutor. It is a different instrument doing a different job: answering, not teaching.
This is what the craftsman stage is actually for. It is not production for its own sake. It is not inefficiency to be optimized away. What it builds is not tied to any single historical form of production, such as the workshop, the apprenticeship, or the manually flown approach.
What is load-bearing is narrower and more portable. It is independent, consequence-bearing practice. The learner generates an output. The learner meets what the world actually returns. Then the learner has to sort the result into one of three things: confirmation, disconfirmation, or a flaw in the attempt itself. And the compound’s own verdict is not available to do the sorting for them. This section will give that condition a name in a moment.
Sennett’s workshop supplies one version of that condition. Other processes might supply it too. One is a simulator run graded against a genuinely blind failure. Another is an adversarial review with no answer key. Another is a certification exercise that withholds the compound’s verdict until the trainee has committed to their own. None of these are proven substitutes. No case here shows one producing judgment equivalent to what independent production built for the practitioners described below. But each meets the specification on its face: a process in which the compound’s output is not available as the standard being learned toward. On the same logic, what doesn’t meet it is any process, however effortful, in which it is.
Call this condition craftsman development, regardless of its specific form. The word names the structural role, not the historical guild.
Independent calibration is the sorting described above, done where the compound’s verdict is not the standard. Remove independent calibration from every available channel, and the same judgment is not acquired more cheaply.
It is not acquired.
To bank the result: the craftsman stage exists to build a judgment that cannot be handed over, only grown. What grows it is practice where the learner’s own call meets a real result, and the tool’s answer is not the yardstick. Take that away everywhere, and the judgment does not come cheaper. It does not come at all.
The Curator without the Craftsman
On the night of June 1, 2009, Air France 447 was flying above the Atlantic. Ice crystals obstructed its pitot probes, the sensors that feed the airspeed reading. The result was unreliable and then lost airspeed indications. The autopilot did what it was built to do when it can no longer trust its own inputs. It disconnected and handed control back to the pilots.
The co-pilot at the controls pulled back on the stick. The aircraft’s nose rose. The correct response to an impending stall is the opposite: nose down, rebuild airspeed, recover lift. It is taught in the first weeks of flight training. At one point the stall warning sounded continuously for fifty-four seconds. The captain was resting in the cabin. Neither pilot working the controls at 35,000 feet appeared to register what the aircraft was telling them. The crew never understood they were in a stall, and never attempted recovery. That was the conclusion of the French accident investigators in their final report (BEA, 2012). All 228 people aboard died.
The investigators’ language is precise on the point that matters here. The pilots lacked training in manual handling at high altitude. That is not a lack of competence in the ordinary sense. Both were qualified and experienced within the automated envelope they normally flew.
AF447 is one data point. It is not proof of what happened in that cockpit down to the mechanism. What makes it more than an anecdote is that two later reviews examined the same pattern industry-wide. One was a 2013 FAA-commissioned panel. The other was a 2016 report by the Department of Transportation’s Office of Inspector General, the OIG.
Here is what those reviews describe. Airline pilots typically flew manually only during takeoff and landing. FAA officials estimated that automation held the aircraft roughly ninety percent of the time. The OIG noted that no industry-wide analysis had actually validated that estimate. Most carriers had no system at all for tracking how much manual flying time their pilots actually got. What manual practice did happen was concentrated in the routine minutes, not in the high-altitude boundary conditions where AF447 went wrong. Safety reviews following AF447, Colgan Air 3407, and Asiana 214 all fed into the mounting concern about automation reliance that produced both reports.
Long-haul aviation had, reasonably, optimized direct manual experience at the edge of the flight envelope almost entirely out of a career. The automation handled the routine majority of every flight so well that the remaining edge cases arrived to pilots who had not been given the chance to build the reflex the moment required. Those edge cases are the moments when automation hands back control precisely because it no longer trusts itself. That is the panels’ own account.
Aviation and legal practice don’t share a timescale, a physical substrate, or a feedback loop. A stall announces itself in seconds. A bad citation might not surface for months. What they share is the structural position of the evaluator. In both, judgment about a boundary case is exercised by someone whose calibration was built entirely inside the system now producing the boundary case. The mechanism doesn’t depend on a domain’s tempo. It depends on where the evaluator’s reference point was built.
This is the specific shape of the failure the curator stage is poorly positioned to see from inside itself. Picture a curator who has learned to evaluate a compound system’s output only by comparing it to what the compound usually produces. That curator will be fluent under ordinary conditions. That curator will also be blind at exactly the boundary. The boundary is the case that looks almost right. It is the citation that reads plausibly, the diagnosis that fits the pattern but not the patient, the airspeed reading that should not be trusted.
Ordinary conditions are, by construction, the conditions the curator has practice recognizing. The boundary case is, by construction, the one kind of failure the curator’s training never covered. Covering it would have required doing the work directly, under conditions where the result carried consequences. It would have required doing that long enough to build the reflex before it was needed.
In short: a curator whose sense of right was built only inside the compound will be fluent on ordinary days and blind at the boundary, the one case that training never covered. AF447 shows the pattern in a single cockpit. The later reviews show it across an industry.
What the Typewriter Did Not Kill
The obvious objection is already waiting. The typewriter killed penmanship. Every major production tool has atrophied a prior skill. The next generation never acquired it, and the work continued. From a distance, dependence on AI looks like that same sequence. First the skilled stop using the hand. Then the cohort that follows never builds the hand.
The sequence is right. Mata, the Schwartz case, and AF447 are not a different mechanism from the generation that never gets the runway, the years of independent practice before the tool arrives. They are the same trap at two points in time. Skill goes quiet inside the tool’s envelope. Then it is never built. What the typewriter did to penmanship, dependence does to whatever the tool has taken over. That is one unfolding risk, not two.
What is lost is not the same.
Penmanship was a method of producing text. The typewriter ate a motor skill and left the evaluator intact. The evaluator, here, is the capacity that knows a wrong answer when it sees one. We still read. We still know a bad sentence when we see one. The person who cannot form a copperplate Q can still tell whether a paragraph is doing its job. The tool occupied the hand. It did not occupy the apparatus that would have noticed.
The calculator is the middle case. It reduced the cognitive load of calculation. With that load went the practice that kept estimation running: the habit of asking, before you trusted the display, whether the result was the right size. The evaluator was not occupied. It went quiet from disuse. Students could still produce an answer. They were less and less likely to know if it had come out the wrong size.
Stall recovery is not penmanship. Neither is the reflex that should have stopped Schwartz at the reporter. Those are not methods of producing the output. They are the evaluator’s reference point: the grain of a real case, the feel of an aircraft that has stopped flying. That reference point is built by doing the work under consequence. It is available only if it is already running when the boundary arrives.
A tool that writes the brief and also confirms that the cases exist is not occupying the hand. It is occupying the apparatus that should have noticed.
That is the inheritance from the first paper in this series, AI: Endgame. Prior tools changed the environment in which cognition worked. This one performs parts of the cognitive work through which evidence is assessed and error is recognized.
The typewriter sequence still describes how the skill disappears: atrophy, then absence. It does not describe what has disappeared. File this under skills we always lose, and the filing closes the wrong question. The history is real. The analogy survives only as a sequence, not as a verdict on what the sequence costs.
What this section establishes: the typewriter objection gets the order of events right. Skill fades, and then it stops being built. It gets the cost wrong. Earlier tools took over a way of producing the work and left the noticing behind. This one can take over the noticing too.
The Teacher Stage
The teacher stage is where this compounds rather than merely repeats. A curator’s blind spot affects one output at a time. A teacher’s corrective signal shapes what the system produces for every downstream user of it.
Suppose the teacher, too, entered as a curator and never independently produced the work. Then the correction is calibrated by comparison to the compound’s own prior output. The system is being corrected toward its own average, not toward the work itself. It can be made more consistent. It cannot, by this route, be made more correct in the domain where it is currently wrong.
The reason lies in where the correcting signal came from. It was built entirely from comparison to the compound’s own output. So it has no independent access to what correct looks like in exactly the domain where the compound is wrong.
This is not a failure of diligence. A conscientious teacher who has only ever curated compound output will do the job conscientiously. That teacher will still train the system to converge on itself. The corrective signal is only as good as the judgment behind it. That judgment has to have been built somewhere the compound’s own output was not the standard. Building it takes years of independent calibration, and nobody in the pipeline has had reason to require them.
Put plainly: a teacher who learned the work only by reviewing the tool’s output can make the tool more consistent, but not more correct where it is wrong. Effort is not the missing piece. Judgment built outside the tool is.
The Generational Compounding
None of this is acute in the first generation. The first cohort of curators trained as craftsmen before the compound threshold was crossed. Their evaluation is grounded in years of independent production, and it shows.
The problem belongs to the generation after that. That generation enters directly as curators. Its baseline for what the work should look like is the compound’s own output rather than the work itself. The problem also belongs to the generation of teachers after that. They correct a system using judgment calibrated by the system they are correcting. By the third generation, craftsman knowledge risks becoming institutional history rather than embodied practice. It is known to have existed. It is described in training documents. It is no longer running in anyone’s hands.
The mechanism is slow enough to be invisible while it is happening. That is what makes it dangerous rather than merely difficult. Nothing about the second generation of curators looks obviously deficient. They are fluent and fast. They are comfortable inside the compound in ways the craftsman generation sometimes was not. The gap does not show up in ordinary performance.
It shows up once, at the boundary, when the compound is subtly wrong and there is no one in the room whose hands remember what right feels like. By the time that failure is legible enough to name, the craftsman generation has typically already retired. And the knowledge is not recoverable from any record. Tacit knowledge, the kind carried in practice rather than in words, was never the kind of thing records could hold.
That is the typewriter sequence applied to the evaluator. First the reflex goes quiet in people who once had it. Then there is no one left who ever did. The first movement looks like convenience. The second looks like fluency. Neither looks like loss until the boundary case arrives. That is exactly when the missing capacity cannot be looked up.
None of this is a property of the technology in the abstract. It is a property of a pipeline that has not been deliberately kept independent of the compound’s own output. The mechanism describes a default, not a certainty. Still, nothing in the cases above shows what interrupting it actually looks like in practice. What interrupting it actually requires is the subject of the piece that follows this one.
In short: the loss is slow and generational. The first curators still carry craftsman judgment. The curators after them learn the work from the tool, and the teachers after them correct the tool with judgment the tool shaped. Nothing looks wrong until the boundary case arrives, and by then the people who could have caught it have retired. That is the default, not a certainty.
This is the same observer constraint that has run through this series from the start. The constraint is that the compound entity cannot evaluate its own output from outside its own frame. This piece applies it to the pipeline that produces observers, rather than to the systems they observe.
What this piece adds is where the external evaluator was supposed to come from. It also adds what happens when the conditions that built that evaluator are exactly what the compound’s efficiency removes.
The human component was never external to the compound by virtue of being human. It was external by virtue of having built its judgment somewhere the compound couldn’t reach. That place was the years of running the attempt and sorting the result directly, before anything was curating for it.
Shorten that runway and the evaluator is no longer outside the frame. It is inside. It runs on instruments the compound itself helped calibrate. It checks the compound’s work against a standard the compound had a hand in setting.
The failure is not that the compound gets things wrong. Every productive system does. The failure is that the capacity to notice, at the boundary, is being quietly retired along with the generation that built it. This is not happening through neglect. It is not happening through any decision anyone made on purpose. It is the ordinary byproduct of a system optimizing itself for exactly the conditions under which that capacity is never tested.
The compound entity requires craftsman judgment to evaluate its output and is, by its own efficiency, eliminating the conditions under which craftsman judgment develops. The typewriter did that to a hand. AI does it to the instrument that was supposed to be measuring the result.
References
- American Bar Association, Formal Opinion 512: Generative Artificial Intelligence Tools (ABA, 2024)
- Bureau d’Enquêtes et d’Analyses, Final Report on the Accident on 1st June 2009 to the Airbus A330-203 Registered F-GZCP, Air France Flight 447 (BEA, 2012)
- Carr, Nicholas, The Glass Cage: How Our Computers Are Changing Us (W.W. Norton, 2014)
- Gates, Bill, “The Turbulent AI Era Is Here. The Choices We Make Now Are Critical” (GatesNotes, August 2026)
- Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023)
- Sennett, Richard, The Craftsman (Yale University Press, 2008)
- U.S. Department of Transportation, Office of Inspector General, Enhanced FAA Oversight Could Reduce Hazards Associated with Increased Use of Flight Deck Automation (2016)
Copyright © 2026 – Lloyd W. Taylor – https://lloydwtaylor.com – ltaylor@netelder.com