Campbell Brown on why no one audits the AI we now trust
Forum AI CEO Campbell Brown argues that AI is fast becoming a primary way people get high-stakes information, yet the labs still largely grade their own accuracy, and for now it is mostly enterprise demand, not consumers, pushing them toward truth over engagement.
Engagement vs. Accuracy: Why AI Might Not Repeat Social Media
<strong>Social media had to optimize for engagement, which rewards the most hyperbolic content; AI, because its money comes from enterprise, is being pushed to optimize for accuracy instead.</strong> The difference is not the technology but the incentive each one is paid to serve.
I'm a big company spending millions and millions of dollars with Anthropic or OpenAI, I'm not going to let them optimize for engagement. I'm going to demand they optimize for accuracy
Confident, Fluent, and Wrong
<strong>The real hazard is not that models are sometimes wrong, but that they deliver wrong answers with the exact same crisp, confident tone as correct ones</strong>, so readers cannot tell a fact from a hallucination by how it looks.
And you're more apt to believe something that is completely wrong if it's hallucinating. So, that presentation I think is part of the problem.
Every Major Model Made the Same High-Stakes Errors
<strong>Forum AI's evaluation of thousands of prompts found models misstating public opinion, inventing quotes, giving wrong voting information, and taking sides on contested questions</strong>, with one top model citing a Chinese state tabloid for a basic civics question.
If you care about factual accuracy, I might not cite Chinese state-run media.
Nobody Audits Themselves
<strong>The labs currently grade their own accuracy and publish a blog post saying they did well; Brown's whole company exists to replace that with outside verification</strong>, the same separation we already demand of banks and drug makers.
if you are using AI for something really important, ask yourself, who is checking the outputs and the information that you're getting? And if it's the company that sold you the AI, you have a problem.
Smartest People, Not Thousands of Labelers
<strong>Forum AI's bet is that judgment on high-stakes topics comes from a few of the best experts in a field, not armies of crowd labelers</strong>; those experts build the rubric, and the rubric trains an LLM judge to grade models at scale.
Our view is you don't need hundreds and thousands of people. You need the smartest people in a given domain.
Truth When There Is One, Perspectives When There Isn't
<strong>On questions with a checkable answer the model should cite the evidence; on genuinely contested ones the goal is to represent the perspectives fairly and supply context, not to pick a side.</strong> Unbiased means a fair process, not a neutral verdict.
it means looking for as close as you can get to what the truth is when there is one, representing multiple perspectives when there isn't a clear right or wrong or yes or no answer.
The Holdout: Don't Teach to the Test
<strong>Forum AI wants to keep a sealed benchmark the labs can never train on, so a high score reflects real improvement rather than tuning to the test.</strong> It is the train/test split of machine learning turned into a governance safeguard.
we would work with the labs on evaluation and data, but ensuring that there is a standard that is a holdout so that measurement is real and legitimate, and you can actually have something to work toward.
The Sycophancy Pull: Accuracy Only For Now
<strong>People want a chatbot that flatters them, and as intelligence gets cheap that friend-bot becomes a real business</strong>; the enterprise incentive keeping today's models honest is temporary, and Brown's own optimism leans hard on the words for now.
it feels inevitable to me for exactly the reasons that you laid out that people want that friend to know they landed safely. It's irresistible almost to go down that path.