Arc Consolidation | Episodes 263–269
The case for putting AI into institutional work was always partly a memory argument. People rotate out and take the reasons with them. They leave the file and remove the account of why one threshold sat where it did, which exception had been negotiated, and what the terse note in the margin was protecting. A system that could search the archive, retrieve the precedent, and hold the working context through a staff turnover looked like a repair to a problem the field has carried for decades. The Interim Protocol names that use as one of the honest cases.
Then the memory-holder acquired a tenure of its own, and it turned out to be shorter than the tenure of the people it was brought in to outlive.
Nobody Remembers the Method opened on Chekhov's old servant. Firs remembers when the orchard's cherries were dried and sent off by the cartload, that there was a recipe for it, and that the recipe has gone. The trees are still there. Everyone can see them. What has vanished is the method that made them worth anything, and the only person still carrying a fragment of it is the person nobody is listening to. Institutions produce this loss constantly, and they have learned to call it turnover.
The version that arrives now comes with a date on it. A model deprecation notice reads as routine supplier communication: a model named, a replacement suggested, an end-of-support date given. Anthropic's public deprecation table is unusually explicit about the arrangement. Publicly released models carry at least 60 days' notice. Opus 4.1 was notified on 5 June 2026 and stopped answering on 5 August. The page makes support into a dated commitment and settles, in passing, the question of who owns time in the relationship.
What the institution built around that model is not a configuration line. Reviewers learned its characteristic errors and compensated for them. Thresholds were calibrated against its outputs. Its phrasing settled into templates, and its habits settled into the review culture. A retirement date reaches all of that at once.
Which reading is under review
Validated Against a System That Is Gone put the problem in the smallest case that carries it. A determination is made in March. It reaches appeal in October. The policy has not changed and neither has the evidence. The model that read both has been withdrawn.
Re-running the case on the successor produces a clean answer to the wrong question. The October reviewer is meant to be examining the determination the institution actually made, and a fresh answer is a different event wearing the same file number. The historical record has to explain the act under appeal: the model version, the prompt, the retrieved evidence, the configuration, and the human who accepted the result. The comparative record is a separate obligation, testing whether the institution still stands behind the method now.
The tempting shortcut is to validate the successor by asking it to reproduce the predecessor's answers, and to treat agreement as continuity. The Container Is the Tell already showed where that ends, in a citation checker that passed invented references because the strings looked right. A successor can clear the same bar by learning the surface of the evaluation set. Resemblance gets the certificate while the judgment underneath stays brittle.
Perturbation asks something harder. Change a harmless detail, reword two clauses without touching the substance, and watch whether the determination flips. The PromptRobust benchmark tests exactly this sensitivity across character, word, sentence, and semantic variation, and more recent work on benchmark reliability under paraphrase found that systematic rewording moves measured performance and that familiar formulations can inflate it. A method that survives only the wording it was tuned on has preserved a test surface.
None of this is new to the newsletter. Capability Regression Does Not Look Like Failure described an upgrade that improves on the measured surface while narrowing on an unmeasured one. What succession adds is a calendar. The regression arrives whether or not anyone was watching for it, because the supplier has scheduled it.
The rational delay and the eviction
Given that, staying put is the sensible local move, and Pinned follows what that sensible move accumulates into. The present system has known faults with known workarounds. The successor promises better benchmarks and carries unfamiliar failure modes. Every month of postponement avoids a real cost that would land this quarter.
The cost compounds quietly instead. Documentation drifts toward what experienced operators remember rather than what is written. Exceptions settle into prompts where nobody reviews them. The test set preserves the cases that once caused trouble and misses the ones the old model handled so naturally that nobody thought to check. The institution's method becomes fitted to a service it does not control, and the fitting is invisible until the notice arrives.
A retirement date converts accumulated dependence into an eviction. Work that never won internal priority now has to fit a supplier's notice window, alongside recovering the undocumented method, testing a successor, running both in parallel where that is possible at all, retraining reviewers, and deciding what happens to cases already in flight. The invoice for inference captured none of this. A procurement process that priced only tokens priced the smaller half.
Four questions belong in the contract before the first request is sent. How much notice accompanies retirement. Which version stays callable during parallel testing. What configuration and usage records the institution can export. Who absorbs the cost of re-validation when the supplier initiates the move. Those terms will not remove migration cost. They make it visible at the moment dependence is being chosen, which is the only moment the choice is still open.
Policy written by infrastructure
Somebody Set the Router went after the quietest part of the stack. Most requests go to the cheapest model likely to cope, and a smaller share goes to the expensive one held for work where being wrong costs more. Vercel's AI Gateway Production Index showed the shape in its June data: open-weight models running 29 percent of gateway tokens on under 4 percent of spend, with Anthropic at the other end taking 61 percent of spend on 32 percent of tokens.
That split is a theory of consequence, executed thousands of times a day by code. Rule 3 of the Protocol already names the judgments that deserve the better reader: eligibility, significance of impact, adequacy of mitigation, compensation, consent, and the scope of consultation. A cost router almost never receives that list. Its categories come from what the implementation can cheaply detect, which is usually length, task label, customer tier, or latency budget.
Succession moves the line without touching the code. Swap a model and the relationship between cost and capability changes underneath a routing rule that still reads the same. The cheap tier may become adequate for work that used to go upward. The premium tier may lose the specific capacity a high-impact workflow depended on. In the constructed appeal, the October request meets the March routing logic, and if the input looks like ordinary document review it goes to a cheap reader before any human sees it. The log records which endpoint answered and nothing about what was decided.
The commitments that cannot be rerouted
The Plant Outlives the Buyer followed the same asymmetry down into concrete, and carried the argument back to The Ground Address. Power purchase agreements run 15 to 20 years in the US EPA's description and 10 to 15 in the UK's 2026 call for evidence. The workload justifying the substation can move on a quarterly product decision.
A developer left holding a built asset, a debt schedule written against contracted revenue, and no counterparty is the infrastructure form of a pinned model. The reverse case is worth naming too, because a gas asset can outlive the forecast used to justify it and still find a buyer, which transfers the capacity and its emissions to an owner with a fresh incentive to keep it running. Irreversibility is distributed unevenly, and the party with the shortest commitment keeps the most flexibility.
What the institution actually owns
The Harness Is What Stays found the constructive turn in the research. Yang, Zhao, Wu, and Kästner tested whether smaller models could recover performance through adapted harnesses. Across seven business-oriented tasks and three small-model families, optimized harnesses improved 16 of 21 task-model pairings, closed the gap entirely in seven, and at the strongest pairing recovered 89.7 percent of the larger model's performance at 4 percent of its cost. The authors describe task difficulty being "lifted from the model into the harness" through instructions, tools, and orchestration.
The limits are real: routine workflows adapt better, and the small model still needs sufficient base capability. Within them, a substantial share of what institutions buy as model capability turns out to live in arrangements they already own. The system prompt, the tool and permission map, the retrieval sources, the retry logic, the routing table, the character profile used to stabilize behaviour, the versioned evaluation cases, the perturbation tests, the register of consequential judgments, and the reversibility classes proposed in Before the Commit. Assembled, those are a handover file, and a handover file survives a model.
The succession clause
End of Support closed by refusing two comforts. The first is preservation as sufficiency. Anthropic has committed to preserve the weights of every publicly released model for at least the lifetime of the company, which matters for research and for accountability, and which cannot preserve the institution that surrounded the model. Tools change, retrieval sources go dark, prompts and permissions drift. A callable model is one component of a historical system.
The second comfort is interchangeability. The regression findings and the routing findings both answer it. An upgrade can narrow where nobody was measuring, and a swap can redraw the boundary between ordinary and consequential work while the code that draws it stays untouched.
What replaces both is a clause, sitting in procurement beside the reversibility class. It names who controls the end-of-support date and what notice follows. It lists the artifacts that must survive the swap: configurations, evaluation sets, decision records, routing logic, and evidence held somewhere the supplier cannot withdraw. It identifies which determinations must stay reviewable after migration, and it accepts that reviewability sometimes means a preserved environment and sometimes only a documented reconstruction, or a plain statement that exact reproduction is no longer available. Any of those beats letting a successor answer as though it had made the original decision. And it assigns authority to redraw the routing line, because a price change is not a mandate to decide which claims get the more capable reader.
Four earlier inquiries now stand along a time axis. The record has to survive its reader. The building has to survive its tenant. The signature has to remain attributable after the tool that made it is gone. The remedy has to reach determinations made under a system that no longer answers.
Firs knows only that the household once knew how to make the orchard pay, and that nobody knows now. He is the last witness to the loss and he cannot repair it. The institution has one advantage over him, and it expires on a date somebody else has already chosen: it can still write the method down.
End of Support ran from 20 to 26 September 2026. What the next arc inherits is the question the succession clause leaves open, of who is owed an account when the system that made the decision is no longer available to give one.
