Watch who is actually stress-testing frontier models right now. It's engineers — building games in one prompt, running physics sims, racing benchmarks, making the model play Pokémon. It's wonderful, and I read all of it. But something is missing from the picture.
Where are the business people?
I mean the people who own a P&L problem. The controller staring at unapplied cash. The owner of a bakery who can't afford a marketing team. The operations director with a reporting process held together by three heroes and a macro. Almost none of them are experimenting — not at the level engineers are, often not at all. They're waiting: for the vendor demo, for the enterprise rollout, for someone to tell them what AI means for their function.
This is backwards, because the most interesting tests of these models are P&L-shaped, and only business people can run them.
What a business-shaped test looks like
This year I built an accounts-receivable agent for a Fortune 500 media & streaming company. The problem is classic: customers pay large lump sums decoupled from invoicing, and someone has to figure out what the money is for. In its first two days in production, the agent applied roughly $100 million of cash.
Here is the thing — the interesting parts of that build were not engineering parts. They were business judgments. Deciding that the model must never do arithmetic (scripts return exact numbers; the model judges). Deciding that confidence about which deal a payment belongs to must be reported separately from confidence about which invoices inside it — because a finance team's trust dies the first time the system is confidently wrong. Deciding the agent drafts every email and sends none, so the analysts keep the decision and the byline. That last one is why it was adopted instead of resisted, and no benchmark measures it.
Early in that project I kept pushing for a proper data layer and losing the argument. My stakeholder kept feeding gigantic spreadsheet exports straight into the model. The way I finally explained it: you hired the smartest intern in the world, and you're asking him to find matches by reading line by line and taking notes in a notebook. That sentence did more than any architecture diagram. It's also, I think, the job description of the business experimenter: translate what the technology needs into what the organization can hear.
The lab
So I run experiments the way engineers run them, but on businesses. A real artisan bakery — its whole commercial engine, website to payments to an agent-supported marketing calendar to an Airtable-wired operational backbone — built and operated as a test of how far one person plus agents can carry a small company. A tool that reconstructs WhatsApp conversations into legally defensible evidence, built because a real case needed it. A database of my own career with provenance on every claim, which writes my CVs and is about to become an agent you can interrogate instead of reading one.
None of these are demos. They have customers, deadlines, and consequences. That's the point: a model that looks brilliant in a chat window meets a very different judge when actual money moves on its output.
Why so few of us
I have three guesses. First, business people were trained to buy software, not build it — experimentation feels like it belongs to engineering, so it gets delegated, and the learning goes with it. Second, there's no benchmark culture for business problems: engineers have leaderboards; controllers have month-end. Nobody publishes "we pointed a model at our reconciliation backlog and here is exactly where it broke." Third — honestly — fear of looking naive in front of technical colleagues, which keeps the people with the best problems away from the tools that want those problems.
But the barrier that used to be real — you needed engineers to even start — is gone. I build these systems in conversation with the models themselves. What's left is the part that was always ours: knowing which problem matters, what "correct" means in context, and what an organization will actually adopt.
If you own a P&L problem, you are sitting on a better eval than anything on a leaderboard. Pick one messy, unglamorous process. Give it to a frontier model seriously — with your judgment wrapped around it. Then write down where it broke, because that's the most valuable AI research almost nobody is publishing.
The models are ready. The organizations are the experiment.