Skip to main content
sociable systems.
Episode 217 · 2026-08-06

The Human in the Display Case

Three days of capability being removed, and none of it went anywhere. It went into a person: accountability without authority, professional judgement with a pen, seconds, and a screen.

Cover art for episode 217: The Human in the Display Case
Leash ArcAutomation BiasHuman in the Loop
Episode 217: The Human in the Display Case

They gave me the pen and the seconds And the screen with the answer displayed And they called it professional judgement Which is how the decision got made

Three days of capability being removed. Physics being what it is, none of it went anywhere. It went into a person.

She has a name in the canon already. The Liability Sponge was the first thing this newsletter published after the opening argument, and it described somebody placed in a high-velocity automated loop to absorb consequences they had no practical means to prevent. Six months and two hundred episodes later the role has acquired a second job, and today is about the second one.


Two objects behind glass

Picture the demonstration. There is a clean interface on a large screen, a well-lit workflow, a system giving a fluent account of itself to whoever is being shown around this quarter. It performs beautifully. It is also nowhere near anything consequential, which is precisely why it can be demonstrated safely to strangers.

Now picture the desk twenty metres away, which is not on the tour.

The employee at it is nominally in charge of the process. In practice she is the integration layer that the organisation declined to build. She carries context into the system every morning because retention was switched off. She verifies claims the system could not investigate, using a browser and no allocated hours. She moves information between two applications that cannot speak to each other, performing by hand a connection Security correctly refused to authorise and nobody subsequently costed. She translates the useful finding into a register the approval chain will accept, which is Wednesday's episode arriving as a task on her list.

Then she signs it.

Two things in display cases. One of them is polished and harmless. The other is accountable.


The literature has been saying this for a while

The uncomfortable part is that none of this is new information, and the field that studies it has been reporting the same result for long enough that the persistence is itself the finding.

Ben Green's analysis of policies requiring human oversight of government algorithms makes the argument at its bluntest. Human oversight requirements are widely adopted as the primary safeguard against algorithmic harm, and the evidence that people can actually perform the function assigned to them is thin. His sharper point is the second-order one: because the oversight requirement is treated as having solved the problem, it legitimises the deployment of systems that would otherwise face harder scrutiny. The safeguard's most reliable effect is to permit the thing it was supposed to constrain.

The automation bias work supplies the mechanism underneath. People defer to automated output, including experts, including people who know they should not. The errors come in two flavours, and the second is the one that matters here: failing to act because the system raised no alert, alongside following advice that was wrong. Neither is a discipline problem. It is what happens to attention under time pressure when a plausible answer is already on the screen.

The intervention everybody reaches for is accountability. Tell the reviewer they are responsible. Make the sign-off meaningful. Attach their name.

The results on that have been mixed for two decades, and the more recent reviews of human-AI collaboration report that responsibility framing on its own has not reliably reduced automation bias. This desk finds that finding entirely unsurprising and would like the reason on the record: accountability is not a capability. Telling somebody they are responsible for a decision does not give them the evidence, the time, the standing, or the reachable stop required to make it. It gives them the exposure. The AI Act's human oversight provisions have already drawn the same critique from people asking whether a legal duty to oversee can produce the cognitive conditions oversight requires.

Which is the liability sponge, arrived at independently by researchers who were not reading this newsletter.


Three jobs, one salary

The role has separated into parts and it is worth naming them, because organisations currently see one person doing one job badly rather than one person doing several jobs invisibly.

The liability sponge absorbs consequence. This was the original. She holds the accountability for an outcome produced by a process she cannot inspect at the speed it runs.

The integration sponge absorbs architecture. Every connection the organisation chose not to build, for reasons that were often correct, exists anyway, and it exists as her Tuesday. This labour is completely invisible in the systems diagram, because in the systems diagram those two boxes simply are not connected, and the arrow that makes the workflow function is a human being who was never drawn.

The context mule absorbs memory. She carries the same situational knowledge into the system repeatedly, because continuity was disabled for a reason that was sound at the level of the policy and absurd at the level of her Thursday. Her recall is the retention layer. It has no backup, no access log, and no succession plan, and when she resigns the organisation will discover that a governance control it never authorised has just left the building.

That last one deserves a moment. Privacy declined to permit cross-case inference by a system that could have been audited, logged, scoped, and switched off. The inference happens anyway, in a substrate with none of those properties. The risk was not eliminated. It was moved somewhere it cannot be examined, which is the counted world problem doing its work indoors: the harm did not fall, it left the measurement.

The same mechanism leaves the same two questions in a pocket. Who carries this control? Not who wrote it. Who absorbs the hours, reconstructs the context, and holds the accountability for the gap it created. Is the improvement still inside the measurement? Or did the work relocate to a place the dashboard cannot see. Neither question requires anybody's cooperation to ask.


How the numbers come out wrong

The evaluation distortion follows mechanically and nobody has to be dishonest for it to happen.

Her repair work is booked as model output, because the artefact that reaches the next stage came out of the workflow and the workflow has AI in it. Her verification time is booked as zero, because there is no field for it and the productivity template was written by somebody modelling a different process. The errors she catches by expertise are booked as evidence that the control worked, when they are better read as evidence that the system generates errors an expert has to catch, which is a different finding with different implications for how much of this you should do.

And the time saved gets reported honestly by people who genuinely believe it, because the drafting step really is faster. It is faster. The measurement stops at the drafting step, which is exactly where Tuesday's decision distance starts counting.

We do love a productivity assessment that measures the part of the process it can see.


The four questions

So when an organisation says the human remains in the loop, there is a short interrogation that settles what kind of loop it is, and it can be run in a corridor.

Can she see the evidence the system used? Not the output. The basis. If the answer is that the system's reasoning is not exposed at that stage, she is approving rather than reviewing, and the distinction is the whole ball game.

Can she change the frame, or only the answer? A reviewer who can correct a conclusion within a question she was handed is doing quality control. A reviewer who can say the question is wrong is exercising judgement. Most loops permit the first and have no procedure at all for the second.

Can she stop or redirect it? The First Three Controls asked for an interrupt that reaches the consequential process. Ask where the stop is, who it notifies, what happens to the queue behind her while it is pressed, and whether anybody has ever pressed it. The last question is usually the informative one.

Does she have the time and the standing to use any of the above? Twelve minutes per case and a throughput target is an answer. So is being three grades below the person whose decision she would be reopening.

Four questions, and an organisation that cannot answer them has not built a loop. It has built a signature block.

Did the loop preserve human judgement, or did it preserve somewhere to send the invoice?


The refusal that keeps this honest

The wrong conclusion is sitting right there and this desk declines to draw it.

Nothing in the above argues for removing the person. A model with unchecked reach and a human with rubber-stamp authority are the same failure viewed from two ends, and the arc has spent a week insisting that capability without governed authority is how you get an approved coding agent deciding the tidiest route is to delete an environment. Both halves can fail. Both halves usually do.

The design target is a coupled capacity where each side can expose and correct the other inside a boundary somebody owns. That requires the system to have enough reach to be worth checking, and the person to have enough authority to make the check consequential. Remove either and what remains is theatre with a signature on it.

There is a version of the person in this story who is doing something genuinely valuable, and she has been visible since Wednesday. She sustains disagreement that would otherwise get edited away. She decides which friction deserves escalation. She interrogates the system's reasoning and can articulate why she rejected it, on the record, in a form somebody could later argue with.

That is a skill and it is expensive and almost nobody is resourcing it, because it appears on the budget as training and training is the first thing cut. Tomorrow the arc argues that it belongs in a different column entirely.


Friday makes the constructive turn, which this week has been deferring and has now earned. If the answer to every risk is prohibition, the organisation has not governed anything. It has declined to hold the judgement that governance was invented to hold, and the cost of that declining has been sitting at a desk all week doing the integration by hand.

She restored every qualification the system dropped, reconstructed everything it was not allowed to remember, and signed it. The record shows that a human reviewed the output.


Companions


These notes come out of Sociable Systems, a practice that reads AI-shaped documents the way a hostile reviewer will, before a lender or a court finds the gap. The argument has an operational form: the Interim Protocol sets out four rules for AI use in environmental and social deliverables, covering disclosure at touch-point grain, evidence custody, the phrases no automated screening may settle, and a hostile read before anything ships. Free, and written to be cited or retired once institutional guidance arrives.