MEMO · JULY 2026
Why most AI pilots produce nothing.
In 2025, researchers at MIT’s Project NANDA examined hundreds of corporate generative-AI initiatives and reached a number that made headlines: 95% of pilots showed no measurable impact on the P&L. The finding is usually quoted as an indictment of the technology. It is better read as an indictment of the pilots.
The study’s more useful finding got less attention. Deployments built with an outside partner, customized to the organization’s actual work, succeeded roughly twice as often as internal builds. The failures cluster where the tool was generic and the workflow was not — demonstrations run on clean sample documents, then quietly abandoned when the real contracts, statements, and correspondence turn out to be long, messy, and particular.
Our reading: the pilot fails at the selection stage, before any software is installed. A pilot chosen because a tool was available tests the tool. A pilot chosen because a specific piece of work eats hours every week — a reporting cycle, a review queue, a first draft that is always the same draft — tests something the firm actually wants. The second kind produces numbers; the first produces demos.
This is why an engagement here begins with a written assessment rather than an installation. Deciding what deserves to be built is most of the work. The building is the easy part.