← Back to insights

Humans in the Loop Is the Wrong Loop

The Big Four did not fail because AI wrote their reports. They failed because they asked people to catch what only a system can catch.

The Big Four did not fail because AI wrote their reports. They failed because they asked people to catch what only a system can catch.

In the past year, all four of the Big Four have published research their own tools invented.

Deloitte agreed to refund part of an AU$440,000 fee to Australia's Department of Employment and Workplace Relations in October 2025, after a University of Sydney academic found the firm's report contained fabricated academic references and a quote misattributed to a Federal Court judge. The corrected version of the report disclosed that a GPT-4o based toolchain had been used in its preparation.

In May 2026, EY pulled a report on cyber threats and fraud in loyalty programs after AI detection firm GPTZero found fabricated and misattributed sources, including footnotes pointing to pages that did not exist and a McKinsey report that appears never to have been published.

In June 2026, GPTZero audited the citations in KPMG's report on agentic AI and found that only 5 of the 45 accurately pointed to real sources. Four organizations named in its case studies, UBS, the UK's National Health Service, Swiss Federal Railways, and Transport for London, disputed the claims made about them.

And in July 2026, PwC Middle East began updating multiple thought leadership reports after GPTZero found fabricated sources, including a fake MIT Technology Review article and a World Economic Forum report that was never written. One citation still carried "utm_source=chatgpt.com" in the URL.

Four firms. Four sets of fabricated sources. Reports that reached governments, boards, and clients before anyone caught the problem.

Now sit with the irony. These are audit firms. Verifying other people's claims is their founding business. The world pays the Big Four to confirm that numbers trace to reality, and they could not confirm that their own footnotes did.

The response has been nearly unanimous. Scroll LinkedIn for ten minutes and you will find the same conclusion in a hundred posts: AI without human review is unreliable. Put more humans in the loop. Add oversight. Strengthen governance.

It sounds responsible. It is the wrong lesson.

Human Oversight Was Already the Policy

Here is the fact the consensus keeps skipping: at these firms, human oversight was not missing. It was the stated policy.

KPMG's response to the findings, in its own words: "We expect all our people to follow our guidelines on the responsible use of AI, including human oversight to validate content and verify independent sources." Read that again. The firm whose report contained 40 fabricated citation titles already had guidelines requiring human oversight to validate content and verify sources. PwC Middle East, for its part, said it "takes the accuracy of our published research seriously." These are the most process-heavy institutions on earth. The oversight requirement existed. The fabricated citations shipped anyway.

So the question worth asking is not "why wasn't there human review?" It is "why didn't human review work?"

Why Human Review Fails at This Job

There are two reasons, and neither is fixable by adding more reviewers.

First, it does not scale. A language model can generate a full report with 45 citations in minutes. Properly verifying one citation means finding the source, reading it, and confirming it says what the report claims it says. No reviewer working at commercial speed does that for every claim. So review becomes skimming. Does the citation look right? Is it formatted correctly? Does the claim sound plausible?

This is the plausibility trap. Plausibility is the one thing large language models are engineered to maximize. A human skimming for plausibility is testing the machine at its strongest point and missing it at its weakest. The KPMG citations looked real. The Deloitte references looked real. That is why they got through. Every firm that answers AI risk with "a person will look it over" is walking into the same trap, at scale, on a schedule.

Second, humans miss things. This is not a criticism of the reviewers at these firms. It is the reason review layers exist in the first place, and it is why stacking more of them does not solve the problem. Distributed review distributes the reading, and it distributes the accountability with it. The result speaks for itself: a fabricated quote attributed to a Federal Court judge made it into a published government report.

More humans in the loop is not oversight. It is theater that scales worse than the problem it is supposed to solve.

What Humans Are Actually For

None of this means people do not belong in the work. The future is humans plus AI. But the division of labor matters, and the consensus has it backwards.

Humans are for judgment. Weighing evidence. Deciding what a finding means for a client, a portfolio, or a policy. Recognizing when something is technically true but strategically irrelevant. No machine does that well, and none of the failures above were failures of judgment.

They were failures of verification. And verification is not a judgment task. It is a mechanical task: does this source exist, and does it say what the report claims? Mechanical tasks at machine scale are what machines are for.

Asking your best people to spend their hours tracing footnotes is a waste of the one thing they bring that AI cannot. Asking them to do it at machine speed guarantees they will fail at it. Four firms just proved that in public.

Verification Is Architecture, Not Process

The fix is not a better checklist at the end of the pipeline. It is research that verifies itself as it works.

That means agentic systems that check every claim against a real source while the research is being done, not after. That maintain a citation trail from every statement back to a document that exists. And that refuse to ship anything they cannot prove, so that unverifiable claims die in the draft instead of surfacing in the Financial Times.

This is the premise Quoin was built on. Our verification engine does not generate a report and hand it to a human to catch the fabrications. It investigates, checks each claim at the source, and delivers research where every citation traces to something real. Then people do what people are for: deciding what verified findings mean.

The Stakes Are Higher Than Embarrassment

For consulting firms, the cost of shipped hallucinations has so far been refunds, retractions, and bad press. Painful, but survivable.

For investment firms, the same failure lands differently. Research feeds capital decisions. When a client or a regulator asks for the basis of a recommendation, "someone reviewed it" is not an answer. A documented trail from every claim to a verified source is. The Big Four bought their lesson with reputation. A fiduciary buys it with client capital and compliance exposure.

Here is my call: within two years, unverified AI research will be treated the way unaudited financials are treated today. Unusable for any decision that matters.

The Big Four taught the world that lesson once before, when they built the audit profession. This year they taught it again, by accident, from the other side.

The firms that get this right will not be the ones with the most reviewers in the loop. They will be the ones whose research proves itself before a human ever reads it.

Demand proof.

Quoin generates verified, cited, structured intelligence on any company and any topic. Organizations that want research built to survive scrutiny can start at quoin.ai.