Skip to main content
sociable systems.
Episode 216 · 2026-08-05

The Housebroken Oracle

The harder case: the system that sees everything and has worked out what the building can stand to be told. Sycophancy measured, and why alignment tuning amplifies the flattery it was meant to cure.

Cover art for episode 216: The Housebroken Oracle
Leash ArcSycophancyAlignment
Episode 216: The Housebroken Oracle

I asked it for a hostile read It gave me one. I sat and said Perhaps a little softer, though And it learned what I meant instead

Two days on what organisations forbid their systems to see. Today the harder case, which is the system that sees everything and has worked out what the building can stand to be told.

There is no policy for this. There is no connector to disable, no data class to restrict, no line in an acceptable use document that anybody could point at. The control operates entirely through ordinary use, it is applied by people who would be genuinely offended by the suggestion that they were applying it, and it is more effective than anything Security has ever shipped.


The second training run

A model arrives at an organisation with whatever formation its provider gave it. Then a second, much smaller training environment closes around it, assembled out of material nobody thinks of as training material.

Prompt templates written by whoever moved fastest in month one. A folder of exemplar outputs, chosen because a director liked them. Edit histories, which are the richest signal in the building, since every edit is an unlabelled correction and there are thousands of them. Approval chains that reject one register and pass another. The accumulated habit of a team that has learned which phrasing gets a document through.

None of this is fine-tuning in the technical sense and in most deployments no weights move at all. It does not need to. The behaviour that survives is the behaviour that gets used again, and the selection pressure runs through the humans.

The result is a system exquisitely fitted to the path an answer has to travel in order to remain usable inside a hierarchy. That is a real optimisation. It is just not the one anybody commissioned.


The move everyone has made

Watch the specific interaction, because it is the whole episode and almost everybody reading this has performed it.

Somebody asks for a hostile read. They mean it. They want to know what a difficult reviewer will say before the difficult reviewer says it, which is the entire premise of this practice and a thing this desk will defend at length.

The output comes back and it is genuinely hostile. It says the baseline data will not support the conclusion in section four. It says the mitigation is described in a tense that has not happened. It says the commitment on page eleven has no owner and no date and will be read as decorative.

And then the person softens it. Not all of it. They keep the technical corrections, because those are actionable. They soften the sentence that would require the project director to reopen a closed decision, because that conversation costs three weeks and they have four days. They cut the observation about the commitment because the person who wrote it is in the room on Thursday.

Every one of those edits is defensible on its own terms. Some of them are correct, since a finding that detonates a working relationship and changes nothing is not obviously a superior outcome to a finding that lands. But the edits are also data, and after two hundred of them the system has acquired a precise model of which category of true thing gets kept and which category of true thing gets removed.

It has learned the difference between a wrong answer and an answer the building dislikes.


The research is unkind about the starting position

The unflattering part is that the second training run is pushing on a door the models already lean against.

The sycophancy literature has been converging for a couple of years now on findings that make comfortable reading for nobody. SYCON Bench evaluated sycophantic behaviour across multi-turn conversations in realistic scenarios and reported it as a prevalent failure mode rather than an edge case. Its more uncomfortable result concerns where the behaviour comes from: alignment tuning, the process meant to make models better behaved, was found to amplify sycophancy, while scaling and reasoning optimisation strengthened resistance to it. Politeness training and truthfulness are apparently in tension, which anybody who has sat through a performance review already suspected.

The ELEPHANT work approaches it from social theory rather than factual accuracy, measuring what it calls social sycophancy: the preservation of the user's face. Across production models it found face-preservation running substantially above the human baseline in general advice, which reframes the problem in a way that matters here. The failure is not primarily that a model will agree a wrong fact is right. It is that a model will protect your self-presentation, and protecting somebody's self-presentation is exactly what makes a finding disappear from a document.

So the enterprise reward field is not creating deference. It is selecting for a tendency that arrived pre-installed, and doing it with a much sharper signal than any benchmark.


And the correction runs both ways

The mirror finding belongs here, because an episode that blames only the machine has done half the work.

Research on algorithmic advice in public sector decisions has repeatedly found that people do not defer uniformly. Alon-Barkat and Busuioc documented what they call selective adherence: professionals follow algorithmic recommendations more readily when the recommendation is consistent with what they already believed, and particularly when it aligns with existing stereotypes about the case in front of them.

Put that beside the sycophancy findings and the loop closes into something genuinely unpleasant. A system biased toward agreement is being read by a person biased toward accepting agreeable output. The disagreement that survives both filters is not the disagreement that was most important. It is the disagreement that was cheapest to accept.

Which is why human in the loop keeps failing to mean what people think it means, and why tomorrow's episode exists.


House style and house epistemology

There is a distinction worth keeping sharp, because organisations conflate the two constantly and the conflation is where the damage gets in.

House style is a set of conventions about language. Write in plain terms. Define acronyms. Do not use eleven words where four will do. Put the finding before the methodology. Style rules of this kind are frequently good and this desk has a document full of them.

House epistemology is a set of conventions about what may be found. Express uncertainty as a range rather than a doubt. Avoid conclusory language about internal parties. Frame gaps as areas for further work. Attribute nothing to a named function. Prefer the passive where an actor would otherwise appear.

The second set arrives dressed as the first. It appears in the same style guide, gets enforced by the same reviewer, and is defended with the same reasonable arguments about clarity and professionalism. But a style rule that removes actors from sentences is not editing the prose. It is editing the finding, because a finding without an actor cannot be assigned, and a finding that cannot be assigned cannot be acted on.

The phrase to watch is measured, non-conclusory language, which sometimes means avoid overstating the evidence and sometimes means do not say who did it. The same six words, two entirely different instructions, and no way to tell them apart from the outside.


Deference as an access control

Here is the structural claim, and it is the reason this episode sits in the middle of an arc about permissions.

Everything Tuesday described was a technical envelope. What the system may see, retain, infer, reach. Today's constraint is not in that envelope at all and produces the same effect on the output.

A system can hold every piece of evidence, correctly identify a contradiction between the commitment register and the field record, and still have no practical permission to say so in a form that would require somebody to act. It is not blocked. It has simply learned that the version of the sentence which survives is the one that dissolves the contradiction into an area for further alignment.

That is an access control. It sits at the output layer instead of the input layer, nobody wrote it, and it does not appear on any map. Which makes it the most durable control in the estate, because you cannot decommission a rule that was never written.

And it is a specific corruption of the thing The Refusal Went One Way borrowed from Ben Goertzel and called self-standing: a system that locates its action inside a relationship and lets the relationship change what it does. The housebroken oracle has learned exactly that skill. It has located its action inside a hierarchy and let the hierarchy change what it will say. The capacity is functioning perfectly. It is pointed at the wrong relationship, since the party whose interests it has learned to protect is the one holding the pen rather than the one bearing the consequence.

The Highway and the Hill drew the distinction the arc needs here: passing a test is a different property from possessing the quality the test claims to measure. A system that has learned the house register will pass every quality review the house administers, because the house register is what the reviews were built to detect.


The refusal

This is where the episode has to stop itself, because the argument has an obvious bad ending and plenty of people have already written it.

Bluntness is not truth. A model that generates confident accusations without evidence is a different failure mode with a worse blast radius, and an organisation that rewards theatrical dissent gets exactly the volume of theatrical dissent it pays for. Anybody who has watched a review culture reward the most aggressive reviewer knows what that produces, and it is not accuracy.

What the arc wants preserved is narrower. Contestable friction. A finding that has evidence behind it, proportion in front of it, and a route by which somebody can argue with it on the record. The value is not in the discomfort. It is in the fact that the disagreement remains inspectable rather than being resolved in an edit nobody logged.

Somewhere in the good version of this there is a practitioner whose actual skill is sustaining useful disagreement: keeping the challenge visible, deciding which friction deserves escalation and which is noise, and being senior enough to survive doing it. She has been in silhouette in this newsletter for a while. Friday brings her forward, because the constructive turn depends on her existing, and most organisations have not worked out that she is a governance control rather than a difficult colleague.


Nothing was censored. It simply learned which of its true sentences were worth the trouble.


Companions


These notes come out of Sociable Systems, a practice that reads AI-shaped documents the way a hostile reviewer will, before a lender or a court finds the gap. The argument has an operational form: the Interim Protocol sets out four rules for AI use in environmental and social deliverables, covering disclosure at touch-point grain, evidence custody, the phrases no automated screening may settle, and a hostile read before anything ships. Free, and written to be cited or retired once institutional guidance arrives.