Business Ideas
Enterprise AI Use Cases: Choose a Pilot With a Baseline
Most enterprise AI pilots never show a measurable return. The difference is rarely the model. It is whether the use case had a buyer who owns the outcome, a baseline measured before the build, and a success metric agreed in advance.
Choose enterprise AI use cases that have three things before any build starts: a named business owner who is accountable for the outcome, a baseline measured from today's process, and a success metric with a threshold everyone agrees on. Prefer workflows that are high-volume, repetitive and already measured, and plan how the system fits into daily work.
The failure rate is high. MIT's Project NANDA report The GenAI Divide: State of AI in Business 2025 found that "just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact." Its research covered over 300 publicly disclosed AI initiatives, 52 interviews and 153 survey responses from senior leaders (Virtualization Review, reporting on MIT NANDA).
Key takeaways
- Most pilots show no measurable return. Pick use cases that can.
- Name a buyer who owns the outcome and the budget.
- Measure the baseline first, from today's process.
- Agree the success metric and threshold before building.
- Plan the workflow fit, not just the model.
Why do most enterprise AI pilots fail to show value?
According to the MIT NANDA report, mainly because systems do not fit and learn from the way work is actually done, not because the models are weak.
The report describes the core barrier as learning capability: "Most GenAI systems do not retain feedback, adapt to context, or improve over time." It also cites "brittle workflows, weak contextual learning, and misalignment with day-to-day operations" (Virtualization Review).
Two further causes are within a buyer's control from day one: pilots started without a baseline to compare against, and pilots without an owner who is accountable for the result.
What does a good enterprise AI use case need before the build?
A named buyer, a measured baseline and an agreed success metric. Without these, a pilot can run indefinitely without anyone being able to say whether it worked.
| Requirement | What it looks like | Why it matters |
|---|---|---|
| Buyer | A named leader who owns the outcome and the budget | Someone decides whether to scale it |
| Baseline | Today's time, cost, error rate or volume, measured | You can show the change |
| Success metric | One number and a threshold, agreed in writing | The pilot ends with a decision |
| Workflow owner | The team whose daily work changes | Adoption, not just deployment |
| Data access | The data exists, is accessible and is allowed to be used | Avoids a stall in week two |
Which enterprise AI use cases are most likely to succeed?
High-volume, repetitive workflows with clear rules and outputs that can be checked quickly, where the current cost is already measured.
| Criterion | Stronger candidate | Weaker candidate |
|---|---|---|
| Volume | Thousands of similar items a month | Occasional, varied tasks |
| Rules | Clear, documented steps | Heavy judgement, little precedent |
| Checkability | Output can be checked quickly | Errors hard to detect |
| Baseline | Already measured | Nobody knows today's cost |
| Risk of error | Low or easily caught | High legal or safety impact |
Examples that often fit, depending on the organisation: document intake and extraction, clinical or case documentation, invoice matching, compliance deadline tracking, and first-line support triage. NELL's forward deployment page maps use cases like these across 11 industries.
How should an enterprise AI pilot be designed?
Small scope, fixed dates, one metric against the baseline, the real users doing real work, and a decision meeting booked for the end date.
- Scope: one team, one workflow.
- Dates: a start, a midpoint review and an end.
- Metric: one number compared with the measured baseline.
- Users: the people who do the work, using it in their normal day.
- Decision: scale, change or stop, decided by the named buyer at the end date.
The same structure works for a startup selling into an enterprise: see founder-led sales for how to agree pilot terms before starting.
Should an enterprise build or buy AI for a use case?
Decide per use case. Buying fits common workflows where vendors already have depth; building fits workflows unique to your business. Either way, the buyer, baseline and metric come first.
Whichever route you take, apply the same pre-build checklist. A purchased tool without a baseline and an owner is as likely to stall as a custom one.
How does NELL work with enterprises on AI use cases?
Through forward-deployed engagements, which put NELL engineers alongside your team on a production AI system.
The forward deployment page describes where the work is across 11 industries, and is explicit that it is "the map NELL operators work from with founders - not a record of results we are claiming." To discuss a specific use case, contact NELL.
Frequently asked questions
Why do most enterprise AI pilots fail?
Research from MIT's Project NANDA points to systems that do not adapt to how work is actually done, plus brittle workflows. Missing baselines and unclear ownership also leave pilots unable to show results.
How do you choose an AI use case in an enterprise?
Look for high-volume, repetitive workflows with clear rules and checkable outputs, where today's cost is already measured, and where a named leader owns the outcome.
What metrics should an AI pilot use?
One primary metric compared with a measured baseline, such as time per item, cost per case, error rate or volume handled, with a threshold agreed before the pilot starts.
How long should an enterprise AI pilot run?
Long enough to measure the metric against the baseline with real users in their normal work, and no longer. Fix the end date and the decision meeting at the start.
Who should own an enterprise AI pilot?
A business leader who is accountable for the outcome and controls the budget, supported by the team whose daily work will change.
Start where you are
Buyer, baseline and metric first. Then the model.
Sources
- Virtualization Review, MIT Report Finds Most AI Business Investments Fail, Reveals 'GenAI Divide' (August 2025)
- NELL AI Labs, AI Opportunities Across 11 Industries
The example use cases are patterns, not recommendations for any specific organisation. NELL AI Labs offers the forward-deployed engagements described in this guide.
