Big Technology Podcast

Campbell Brown on why no one audits the AI we now trust

Campbell Brown· CEO of Forum AI at Forum AI
·~57 min·English·Big Technology
AI SafetyPolicyLLMAI Company
TL;DR

Forum AI CEO Campbell Brown argues that AI is fast becoming a primary way people get high-stakes information, yet the labs still largely grade their own accuracy, and for now it is mostly enterprise demand, not consumers, pushing them toward truth over engagement.

01Core Thesis

Engagement vs. Accuracy: Why AI Might Not Repeat Social Media

<strong>Social media had to optimize for engagement, which rewards the most hyperbolic content; AI, because its money comes from enterprise, is being pushed to optimize for accuracy instead.</strong> The difference is not the technology but the incentive each one is paid to serve.

I'm a big company spending millions and millions of dollars with Anthropic or OpenAI, I'm not going to let them optimize for engagement. I'm going to demand they optimize for accuracy

Campbell Brown, Big Technology Podcast
Key Insight
Brown's optimism is conditional, not structural: accuracy is winning largely because enterprise is the paying customer. She credits the labs for caring and competing to improve, but her point is that the incentive, not goodwill, is what currently holds, and she is candid that a shift toward personalized filter-bubble chatbots could undo it, which is why the rest of the interview is about locking accountability in first.

02The Danger

Confident, Fluent, and Wrong

<strong>The real hazard is not that models are sometimes wrong, but that they deliver wrong answers with the exact same crisp, confident tone as correct ones</strong>, so readers cannot tell a fact from a hallucination by how it looks.

And you're more apt to believe something that is completely wrong if it's hallucinating. So, that presentation I think is part of the problem.

Campbell Brown, Big Technology Podcast
Key Insight
This inverts the usual framing of the hallucination problem. The failure is not only in the model's knowledge but in its interface. A model that signaled its own uncertainty, the way a careful source hedges, would be less dangerous while being no more accurate, which suggests calibrated confidence is a safety feature, not just a UX detail.

03The Evidence

Every Major Model Made the Same High-Stakes Errors

<strong>Forum AI's evaluation of thousands of prompts found models misstating public opinion, inventing quotes, giving wrong voting information, and taking sides on contested questions</strong>, with one top model citing a Chinese state tabloid for a basic civics question.

If you care about factual accuracy, I might not cite Chinese state-run media.

Campbell Brown, Big Technology Podcast
Key Insight
Brown draws a line most coverage blurs: source quality is the easy failure and bias is the hard one. A model citing a state tabloid can be fixed with something close to a sourcing rule, while representing contested topics fairly cannot. Bundling the two makes the problem look intractable; separating them shows where a quick win actually exists.

04Accountability

Nobody Audits Themselves

<strong>The labs currently grade their own accuracy and publish a blog post saying they did well; Brown's whole company exists to replace that with outside verification</strong>, the same separation we already demand of banks and drug makers.

if you are using AI for something really important, ask yourself, who is checking the outputs and the information that you're getting? And if it's the company that sold you the AI, you have a problem.

Campbell Brown, Big Technology Podcast
Key Insight
The regulated-industry analogy quietly reframes AI evaluation from a technical service into a civic function. If AI is becoming the infrastructure through which people get medical, legal, and election information, then independent verification is not a nice-to-have vendor feature but the missing institution, the auditor or regulator that every other high-stakes system was eventually forced to build.

05The Method

Smartest People, Not Thousands of Labelers

<strong>Forum AI's bet is that judgment on high-stakes topics comes from a few of the best experts in a field, not armies of crowd labelers</strong>; those experts build the rubric, and the rubric trains an LLM judge to grade models at scale.

Our view is you don't need hundreds and thousands of people. You need the smartest people in a given domain.

Campbell Brown, Big Technology Podcast
Key Insight
This mirrors how frontier coding models improved, trained on excellent code rather than all code, and extends that logic to judgment. The wager is that a rubric distilled from a few genuine experts generalizes through an LLM judge, trading the breadth of the crowd for the depth of the specialist. If it holds, expertise becomes a reusable asset instead of a per-question cost.

06The Principle

Truth When There Is One, Perspectives When There Isn't

<strong>On questions with a checkable answer the model should cite the evidence; on genuinely contested ones the goal is to represent the perspectives fairly and supply context, not to pick a side.</strong> Unbiased means a fair process, not a neutral verdict.

it means looking for as close as you can get to what the truth is when there is one, representing multiple perspectives when there isn't a clear right or wrong or yes or no answer.

Campbell Brown, Big Technology Podcast
Key Insight
The hard engineering problem hides in the branch, not the answer. Deciding whether a question has one factual answer or is genuinely contested is itself a contested, high-context judgment, exactly the call a partisan model would get wrong first. Brown's framework does not remove that judgment; it relocates it from the model's output to the experts who define the rubric.

07The Safeguard

The Holdout: Don't Teach to the Test

<strong>Forum AI wants to keep a sealed benchmark the labs can never train on, so a high score reflects real improvement rather than tuning to the test.</strong> It is the train/test split of machine learning turned into a governance safeguard.

we would work with the labs on evaluation and data, but ensuring that there is a standard that is a holdout so that measurement is real and legitimate, and you can actually have something to work toward.

Campbell Brown, Big Technology Podcast
Key Insight
The holdout is what separates an evaluation vendor from a standards body. Selling training data to the labs makes Forum AI one supplier among many; owning a benchmark the labs cannot game makes it the scorekeeper, a role closer to a ratings agency than a contractor. That is a more durable and more politically loaded position than the data business it sits next to.

08What's Next

The Sycophancy Pull: Accuracy Only For Now

<strong>People want a chatbot that flatters them, and as intelligence gets cheap that friend-bot becomes a real business</strong>; the enterprise incentive keeping today's models honest is temporary, and Brown's own optimism leans hard on the words for now.

it feels inevitable to me for exactly the reasons that you laid out that people want that friend to know they landed safely. It's irresistible almost to go down that path.

Campbell Brown, Big Technology Podcast
Key Insight
This closes the loop back to the first idea. The engagement trap that captured social media was never defeated in AI; it was only deferred by who happens to be paying today. Brown's fight for independent, expert-built standards is really a race to institutionalize accuracy before the consumer market makes the friend-bot too lucrative to resist.