They built the thing with three eyes where one was plenty for the job. It did not see the same thing twice, and that was how it learned to see.
Disagreement among accounts has been treated across this arc as a condition to be held. It is also a resource, and a branch of the field already treats it as one.
Sampling one model several times and taking the majority answer, self-consistency, is standard practice precisely because independent reads of the same question are expected to differ, and the vote uses the spread. The vote is the averaging move, and the empirical record on when it fails is sharp. On hard science problems, small models' sampling distributions concentrate on the wrong answer, so the vote locks the error in: majority voting reduced per-problem accuracy on most GPQA Diamond problems for the two models tested, 56.6% of problems for Qwen2.5-7B and 65.7% for Llama-3-8B (Bahuguna, "When Self-Consistency Backfires", 2026). A companion analysis decomposes the wrong-consensus case and lands on a sentence worth quoting in full: "Agreement is graded evidence, not certification" (Zhang et al., "Decomposing Wrong-Consensus Agreement in LLM Self-Consistency", 2026). On the cases that matter most, the vote can be less accurate than the individual reads it flattened.
The panel designs are the interesting half. Where several differently configured models read the same evidence and the outputs are held side by side instead of voted, divergence becomes a detector: ensemble hallucination-detection frameworks treat disagreement among independent checks as the alarm signal, and multi-agent work uses adversarial debate and explicit voting to surface disagreements that a single model would have averaged away (Yang et al., "Minimizing Hallucinations and Communication Costs: Adversarial Debate and Voting Mechanisms in LLM-Based Multi-Agents", Applied Sciences, 2025). The Humane Loop arc's neighborhood of models is the same design read socially: plurality kept on purpose, because the disagreement is where the information lives.
Where disagreement is signal
The distinction that matters is between disagreement that carries information and disagreement that is noise, and it has to be made procedurally rather than by intuition. The routing rule can be stated in one line: disagreement between differently situated instruments routes; disagreement between identical instruments votes. When two different readers of the same record diverge, the divergence marks ambiguity in the record, an instability in one of the readers, or a genuine bifurcation in the evidence, and each of those deserves a routing decision rather than a vote. When the same instrument reads unambiguous evidence twice and jitters, averaging is correct. The failing systems are the ones that cannot tell the two cases apart, so they vote on everything, and the vote converts the signal cases into confident single answers.
The Interim Protocol's third rule, the never-machine-settled list, names the domain where this matters most. Eligibility, impact significance, mitigation adequacy, compensation, consent and consultation scope are the judgements where the vote is most tempting, because the case volumes are highest, and least permissible, because they are the judgements where the losing account is a person's account of their own life. The Protocol's list reads along this arc's axis as the list of places where plurality must survive the pipeline.
A panel on a grievance file
Take the forty submissions from the March meeting again. A panel design gives the stack to several differently situated readers: one model prompted to extract the company's procedural record, one prompted to extract each complainant's account in their own terms, one prompted to list every point where the two disagree, and a community monitor reading the same stack with none of the prompts. Their outputs go side by side into a single review sheet. Where they converge, the reviewer can move quickly. Where they diverge, the divergence is the agenda.
The design works only if the difference among the readers is real. Four copies of one model with the same prompt and different random seeds are identical instruments, and their disagreement measures jitter. Different prompts on one model buy some independence. Different models buy more, and a reader from the affected community buys a kind no configuration of models can supply: a frame built by living through the meeting. Nothing here demotes the models. They read forty submissions in a minute and can be asked, repeatedly and without fatigue, to look for the account they missed. The human and the models are differently situated readers, and the panel is valuable because they are.
Divergence has a cost. A panel that flags everything flags nothing, and a reviewer facing three hundred divergences will start voting in their head. The routing rule needs a threshold, written down, and the threshold belongs in the Protocol's disclosure register with the other choices about where machines touched the file. Below it, divergence is logged and the file moves. Above it, a person reads the original testimony. The record keeps both the log and the reading.
Companions
- The neighborhood of models, plurality as design: The Pantry and the Neighborhood.
- The vote uncalled on purpose: The Highway and the Hill.
- The judgements no screening may settle: the Interim Protocol.
- Bahuguna, U., "When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs" (arXiv:2608.11403, 2026, preprint).
- Zhang, L. et al., "Decomposing Wrong-Consensus Agreement in LLM Self-Consistency" (arXiv:2608.18795, 2026, preprint).
- The fiction: Akutagawa, "In a Grove", on Wikisource.
These notes come out of Sociable Systems, a practice that reads AI-shaped documents the way a hostile reviewer will, before a lender or a court finds the gap. The argument has an operational form: the Interim Protocol sets out four rules for AI use in environmental and social deliverables, covering disclosure at touch-point grain, evidence custody, the phrases no automated screening may settle, and a hostile read before anything ships. Free, and written to be cited or retired once institutional guidance arrives.
