Ten years leading AI and digital transformation across the enterprise โ now building governed, multi-agent AI systems hands-on, and researching what actually scales: not the model, but the verification, observability, and trust around it.
I've led enterprise AI and digital transformation at some of the world's largest companies, I build governed multi-agent systems hands-on, and I write research grounded in the data those systems produce. The rare combination: I've run it at scale, I ship it myself, and I reason rigorously about why it does or doesn't scale.
My current work centres on one thesis: capability is abundant and its returns are bounded; advantage comes from the architecture that converts a capable model into verified, observable, trustworthy value.
Short, plain-language briefings for executives โ the research translated into cost, risk, and what to do about it. No jargon, no equations.
How autonomous production collapses while verification, observability, and trust do not.
The flagship thesis: AI collapses the cost of production but not of verification, observability, or trust โ and those, in an Amdahl-law sense, bound how far it scales. Then the architecture that bends them. Grounded in a 52-million-record natural experiment.
Can an AI-operating system observe itself?
The ceiling on autonomous operation is not the agent's intelligence but whether the system can answer questions about itself. A defect taxonomy, a field study, and a family of failures root-caused to a system scanning its own substrate.
North-star adjudication for multi-agent systems.
How to guarantee a swarm of agents actually finished the job when the agents grade their own homework. Reify work as a durable packet; adjudicate completion against evidence, not self-declaration.
An economic and trust architecture for multi-tenant AI.
Why a platform that never buys a token โ and never proxies a customer's key โ is a stronger trust and safety design. Decoupling payment, custody, and provenance.
A mathematical framework for AI decision routing.
Six borrowed formalisms as one control vocabulary โ audited self-critically into two load-bearing 'engines' and four design 'compasses'. Rigor through honesty about what's proven versus proposed.
If you're working on how AI actually scales โ cost, latency, safety, the systems around the model โ I'd like to hear about it. Direct is best: