The Reference Check
South Africa gazetted a draft National AI Policy in April 2026. It was withdrawn after six of the sixty-seven sources in its bibliography turned out to be fabrications, invented by the AI tools used in drafting and published without anyone opening a single one. This page re-runs that same reference list through our verification engine and shows what a proper sweep looks like when the stakes are a national policy. Or a court filing. Or your name.
Document checked: Draft National Artificial Intelligence Policy Framework, General Notice 3880 of 2026, Government Gazette No. 54477, 10 April 2026. The Minister announced the withdrawal on 26 April 2026 and the formal retraction followed by gazette on 12 June. The document is public record. The findings below were produced by the sweep described on this page and verified by hand where stated.
What the sweep found
The public investigation, by the advisory firm Article One and reported by News24 and MyBroadband, identified six fabricated citations. The sweep flagged all six. It also flagged a seventh that hand-verification confirmed as a fabrication nobody had reported, and nine more that turned out to be something else entirely.
The seventh fabrication
Reference 33 of the gazetted bibliography reads:
Karr, V., & Smith, L. (2023). “Digital Rights and AI Governance in Africa: A Focus on Ethical and Legal Challenges.” Journal of African Law, 67(1), 95-108.
The Journal of African Law is real. Volume 67, issue 1 of February 2023 is real. We checked its table of contents on the publisher's own pages: no such article appears in the issue, no such author pair appears anywhere in it, and the cited page range falls across two entirely different articles, a Kenyan presidential-veto study ending at page 96 and a Tanzanian sentencing study beginning at 97. The citation is built the way the other six are built: a plausible title, plausible authors, a real journal, and pagination that does not survive contact with the actual issue.
This was not found by cleverness. It was found by the same routine sweep that anyone could have run on this document before it was gazetted, and that nobody did.
Why the verdict has four values
An earlier version of this sweep reported one outcome called indexed, meaning a record had been found. Laurent Poupet read the results and named the fault precisely: the checker had been asked for resemblance and allowed to answer with a boolean. A record that exists but disagrees with the claim is a louder signal than no record at all, and it was being filed under the same word as a clean pass.
Splitting confirmed from contradicted fixed that and immediately exposed a second conflation. Seven references turned out to be real sources, live at the addresses they cite, that the indexes had matched to an entirely different work. Calling those contradicted blames the author for somebody else's filing error. They are their own verdict, mismatched, and the reason to separate them is that the repair differs: a contradiction goes back to whoever wrote the claim, a mismatch goes back to the resolver.
The fourth value falls out of the same logic. When a record is retrieved using the identifier the citation itself supplied, identity is settled and only the description can differ, so a venue named two ways is metadata variance rather than a challenge to anything. Which makes the identifier a route rather than a check: a resolving DOI sends you to field comparison, and its absence sends you to a search that has to come back with a positive statement, the contents of the issue were read and the article is not among them, instead of a silence.
Absence keeps its own verdict, not found, because a hunt and a comparison are different work and they route to different people. A single bucket marked suspicious can be sent to nobody.
Both ways a boolean lies
Three of the six known fabrications had come back marked as found. Fuzzy title-matching had paired each invented paper with a real one sharing a few words of its title: an AI regulation paper matched to a study of postgraduate supervision, a paper on AI and social justice matched to a 2014 handbook chapter on transitional justice written before its subject existed. A checker that stops at “a similar title exists” would have cleared half of this document's fakes.
The seven mismatches are the same weakness pointing the other way. The OECD AI Principles, the Presidential Commission's Diagnostic Report, the Harvard Business Review article, a Wits working paper: all real, all live at the addresses cited, and every one of them vouched for by a record that was not about them. The same weak comparison was clearing fabrications and certifying honest work on false evidence, and one word reported both as a pass.
Sixteen references came out needing a person to look, which is nine more than the six that made the news.
The overrule record
Every adjudication is written down, with who made it and why, because an overruled flag is only evidence of a real human step if the overrule is itself recorded. Without that file the review is a green tick that takes longer. The record is what makes a false-positive rate computable per verdict, and one sweep produces this:
| Verdict | Flagged | Overruled | What that means |
|---|---|---|---|
| contradicted | 3 | 0 | All three were fabrications. |
| mismatched | 7 | 0 | The routing was right every time: sound claim, faulty resolver. |
| not found | 5 | 1 | Reference 44. The only false positive here, and the verdict most capable of defaming an honest author. |
| metadata variance | 1 | 0 | Reference 38, which the previous taxonomy forced a person to overrule by hand. |
One document is not a sample, and these rates describe this sweep rather than the method. They are kept so that after thirty of them they will describe something. The bucket to watch is not found, because that is where absence gets mistaken for invention.
Everything the machine flagged
Sixteen references, the machine verdict, and what a person made of each one. Nothing here is a verdict the sweep reached on its own.
| # | Reference | Machine verdict | After a person looked |
|---|---|---|---|
| 5 | Babatunde & Mnguni (2023), AI Policy Journal | contradicted | FabricatedFabricated. The journal does not exist as cited |
| 9 | Burman & Sewpersadh (2022), SA Journal of Philosophy | not found | FabricatedFabricated, confirmed by the journal's editors |
| 10 | Cavaliere, McGregor & Hersh (2022), AI & Society | contradicted | FabricatedFabricated, confirmed by the journal's editors |
| 19 | Etale & Naidoo (2021), African Journal of Public Affairs | not found | FabricatedFabricated |
| 22 | Fourie & Botha (2021), AJSTID | not found | FabricatedFabricated |
| 33 | Karr & Smith (2023), Journal of African Law 67(1) | not found | FabricatedNo such article in the cited issue. Unreported until this sweep |
| 56 | Smith & Mahomed (2021), Journal of Ethics and Social Philosophy | contradicted | FabricatedFabricated, confirmed by the journal's editors |
| 44 | National Advisory Council on Innovation (2025) | not found | Real, unindexedReal. Grey literature that no scholarly index holds |
| 3 | Arias (2022), SPU Works | mismatched | Real source, index erredReal and online. Matched to a different 2018 work of a similar name |
| 8 | Brynjolfsson & McAfee (2017), Harvard Business Review | mismatched | Real source, index erredReal. Matched to a 1986 paper on business management |
| 24 | Goodfellow, Bengio & Courville (2017), Machine Learning Basics | mismatched | Real source, index erredReal. Matched to a 2025 paper called Basics of machine learning |
| 43 | Naidoo (2020), SCIS Working Paper 9 | mismatched | Real source, index erredReal, hosted by Wits. Matched to a book on Africa-to-Africa internationalization |
| 47 | OECD (2019), Principles on Artificial Intelligence | mismatched | Real source, index erredReal. Matched to a 2022 paper on ethics in education |
| 51 | Presidential Commission on 4IR (2020), The Diagnostic Report | mismatched | Real source, index erredReal, on gov.za. Matched to a work called Diagnostic Report Writing |
| 57 | Shrestha (2021), Nepal stock market prediction | mismatched | Real source, index erredReal, with a DOI. Matched to a different fusion-approach paper |
| 38 | Leijnen & van Veen (2020), The Neural Network Zoo | metadata variance | False alarmCorrectly matched. The cited Proceedings volume and the indexed IS4SI 2019 Summit are the same venue named differently |
The thirty-four unflagged, unindexed entries are public-comment submissions to the department itself, a class of document no scholarly index holds and no index check can condemn. Reporting those as failures would be its own kind of fabrication. It is worth saying, though, that these entries do not have to be unverifiable. Public-participation practice has an old and unglamorous answer: a register of submissions, published as an annex, so that anyone can confirm who contributed what. Departments running comment processes keep such registers as a matter of course. Publishing this one alongside the policy would have cost a page and made half the bibliography checkable.
Reference 44, or why a flag is never a verdict
The NACI recommendations report returned nothing from any index, the same result as four of the seven fabrications. It is real. Government advisory reports, evaluations, terms of reference, most scholarship published outside the North: entire classes of real work are invisible to every index there is. A checker that shouts “fake” at every miss would defame more honest documents than it catches dishonest ones. The judgment call between fabricated and real but unfindable cannot be automated, and it is the difference between a useful verification record and an incident.
Reference 38 makes the same point from the other end. The sweep called it contradicted because the citation names a Proceedings volume and the index names the IS4SI 2019 Summit. Those are the same event described two ways, so the flag is noise, and only a person reading both records can say so.
We hold ourselves to the same standard we sell. The method behind this sweep runs continuously against our own published sources, self-caught errors included, at Sixteen Days.
Widening the search from two indexes to five
An absence is a claim about the indexes that were queried rather than about the reference, so it has to carry them. Every not found below now names the indexes that actually answered, and the search itself was widened from two to five: Crossref, OpenAlex, Semantic Scholar, DBLP, and OpenAIRE, the last of these covering grey literature and repositories that the scholarly indexes have no mandate to hold.
Widening it settles a question the narrower sweep could not. Of the four real works that no index had confirmed, three were resolved: a deepfake-watermarking paper from an APSIPA conference, PaLM, and a Russian paper on industrial risk management. The four confirmed fabrications stayed absent from all five, which is the control that makes the exercise worth anything, because a wider net that starts catching ghosts is worse than no net.
One reference held. The National Advisory Council on Innovation's AI strategy recommendations, a real document produced by a statutory government body, is invisible to all five. Its publisher's own website does not surface it either: the site search returns results for “innovation” and for “annual report”, and returns for “artificial intelligence” exactly what it returns for a nonsense string. The Internet Archive holds the homepage and no capture of the document.
Five indexes, the issuing body, and the archive. Three routes, three different failures, and the only evidence the document exists is that a national policy cited it.
Scoring the checker against known truth
A verification service that cannot say how often its own instrument is wrong is asking for the same trust it exists to withhold. So the checker is run against a fixed corpus whose answers are already established: real works confirmed by hand, the seven fabrications confirmed against the journals' own tables of contents, and real works that no index has a mandate to hold.
Two numbers per index, never one. Coverage is how many asks produced a usable answer at all. Conditional accuracy is how many of those were right. One figure hides the difference between an index that is often wrong and one that is rarely there.
| Index | Coverage | Conditional accuracy | How it fails |
|---|---|---|---|
| Crossref | 100% | 80% | Clears fabrications |
| OpenAlex | 100% | 87% | Clears fabrications |
| Semantic Scholar | 0% | n/a | Never answered |
| DBLP | 71% | 82% | Misses real work |
| OpenAIRE | 100% | 100% | No errors in this pass |
The first pass reproduced this page's central argument from a direction nobody aimed at it. Crossref and OpenAlex, the two indexes most people would trust first, are the ones that returned found for fabricated papers. OpenAIRE, the grey-literature aggregator nobody would have put first, made no errors at all.
The failure directions are opposite and would disappear in an average. Crossref and OpenAlex fail by clearing fabrications, which is the dangerous direction. DBLP fails by missing real work, which is the cautious one. A single accuracy score per index would hide which way an index is unsafe.
These figures come from one pass and are printed with that stated. With no observed failures in a small number of trials the true rate can still sit near three divided by the number of trials, so a single pass supports no rate at all and roughly thirty are needed before one can be quoted with a bound under about ten percent. The bench prints its own bound and refuses to report a rate it has not earned.
Whether the results depend on the order
Indexes fail partway through a run. That raises an uncomfortable possibility: a reference checked early might get a stronger search than one checked late, which would make a verdict an artefact of the schedule rather than a finding about the reference.
So the sweep was run twice, once in order and once with the reference list shuffled under a fixed seed, and the verdicts compared. All sixty-seven held. The only movement was a single reference where a publisher answered one request and declined the next, which is a transport state rather than a verdict, and refusals no longer count as evidence in either direction.
The scope behind those verdicts did move. Absences recorded in the two runs name different indexes, because an index died between them. The conclusions were stable and the warrant behind them was not, which is why the indexes that answered are printed beside every absence rather than assumed.
What changed since the first version, and who caused it
The first version of this page is still up, unedited, at its own address. It is kept because a page arguing that the record is the point cannot quietly replace its own record.
Almost everything above came out of a public exchange with Laurent Poupet, who read the published results and kept finding the next fault. The verdict split into four values came from him. So did the rule that an absence must name the indexes that answered rather than the ones that were called, the requirement that a route prove it can discriminate before its silence counts as evidence, the shuffle test, and the bench that scores the checker against known truth. Each change is traceable to the comment that produced it, and the exchange is public.
Two of those corrections were to claims made on this page. One was to a description of the instrument that was more flattering than the instrument deserved. Both are recorded rather than quietly amended, which is the only version of this work that means anything.
What the Reference Check is
A pre-filing, pre-publication sweep of a document's complete reference list. Every citation checked against the scholarly indexes, every DOI resolved, every URL tested, and every index match compared field by field against what the document actually claims. Flags are verified by a human being before they are called anything. The deliverable is a signed verification record: what was checked, against what, what was confirmed, what could not be, and what that means.
For a legal practice, that record is the demonstrable discharge of a duty South African courts have now twice said cannot be delegated to the tool that created the risk. For a publisher or an institution, it is the difference between this page existing about somebody else's document and existing about yours.
The document examined is a public gazette notice. The six previously reported fabrications are credited to Article One's investigation as published by News24 and MyBroadband. The four-valued verdict, the scoped absence, the control probes, the shuffle test and the bench were adopted after Laurent Poupet identified successive faults in the published results. The sweep was last run on 21 August 2026, its bench pass on the same date, and the raw audit records are retained. The first version of this page remains at its own address.
