How to pick the right first AI project (without betting the company on it)
Most AI projects fail before anyone writes a line of code. They fail at selection. Here is the two-axis method we use to choose — and the one question it cannot answer.
Most AI projects fail before anyone writes a line of code. They fail at selection. Here is the two-axis method we use to choose — and the one question it cannot answer.
The failure is usually not technical. The model works. The vendor was competent. The pilot ran.
The failure is a choice failure. Teams pick the project the vendor demoed best, or the one the loudest department asked for, or the one that sounded most like what a competitor announced. None of those is the same as the project your own economics can afford to run first.
The question is not "what can AI do?" That question has too many answers and none of them are about your business. The better question is narrower and much harder to dodge: which of our processes, automated first, pays for the rest?
Answering it requires a way to compare candidates that does not collapse into whoever argues hardest in the room. Enthusiasm is not a ranking method, and neither is seniority. What you need is a scoring pass that both the operations lead and the finance lead can look at and agree reflects reality, even when they disagree about the result. That is what a value map is for.
The CCAIE Value Map is a structured way to get from "we should do something with AI" to "these three specific activities, in this order, for these reasons."
It works by drilling from the broad to the specific. Start with your functions — where work actually happens. Within each function, identify the value driver: what actually creates or protects value there. Then drill to a key activity: a concrete, repeatable task, not a department and not an ambition.
Every candidate should trace to at least one of six drivers. If it traces to none of them, it is an interesting idea rather than a project:
| Driver | The question it answers |
|---|---|
| Revenue | Does it help win, grow, or retain customers? |
| Cost | Does it remove manual effort or waste? |
| Speed | Does it shorten a cycle time or a wait? |
| Quality | Does it reduce errors, defects, or rework? |
| Risk | Does it prevent loss, non-compliance, or failure? |
| Experience | Does it make life better for a customer or employee? |
Only then do you score. Each candidate activity gets two numbers, 1 to 5:
Plot those two axes and you get four quadrants. High value and high feasibility is the quick win — that is where your first project lives. High value and low feasibility is a big bet: real, but it needs research support or a funding route before it is safe to start. Low value and high feasibility is a fill-in: do it when there is slack, not instead of the quick win. Low on both, park it and say so out loud.
Drill from function to a named activity, score it on two axes, then read the quadrant. The first project should come from quick wins.
Here is what that looks like filled in. The scores are illustrative — the point is the shape of the output, not these particular numbers:
| Function | Driver | Key activity | V | F | Bucket |
|---|---|---|---|---|---|
| Operations | Quality | Visual defect check on the packaging line | 4 | 4 | Quick win |
| Supply chain | Cost / Speed | Demand forecasting for perishable inputs | 5 | 3 | Big bet |
| Customer service | Experience | Drafting replies to repeat customer emails | 3 | 5 | Fill-in |
| Finance | Risk | Fraud detection on low accounts-receivable volume | 3 | 2 | Park |
Notice what the table settles. The demand-forecasting idea is the most valuable thing on the list and it is still not the first project — it scores a 3 on feasibility, which means it needs better data or outside research support before it is safe to start. The defect check is less exciting and goes first, because it will actually finish. That reordering is the entire value of doing this on paper rather than in a meeting.
One thing happens before any of this scoring, and skipping it wastes the whole exercise. Every candidate has to clear two screens first: data readiness — is the data available, and can you actually get at it — and responsible-use fit, which for most Canadian organisations means privacy obligations under PIPEDA and, in Quebec, Law 25. A candidate that fails either screen does not get a low score. It comes off the list until the screen is passed.
Scoring gives you an order. These five questions tell you whether the top candidate is actually ready to be a project. They are the same questions we use as the assessment rubric in the Managers course, deliberately — the article and the service should share a vocabulary.
Questions three and four come straight out of our governance work; if you want the architecture underneath them, that is what the Octopus Protocol specifies.
The map scores feasibility. It cannot tell you what the process actually does today.
Only a measurement can, and the gap between those two things is wider than almost anyone expects. In an executive programme we delivered in August 2026, five cross-functional teams each built a documented AI project for their own organisation. The work was good. The business cases were structured, the arithmetic was sound, and the teams knew their processes well.
Not one of the five had measured how long their process step actually takes. Every duration was estimated from how the process is described — from the procedure, not the stopwatch.
One team reported eight hours for a step. On inspection, eight hours turned out to be how long a document takes to travel between desks, not how long a person works on it. The two numbers are not close, and only one of them belongs in a business case.
Across the five pilots, the portfolio identified roughly 160 hours a month of recoverable effort. That figure was estimated, not measured, and we said so in the report we gave the owners. It is an illustration of scale, not a result.
This is the gate in front of every ROI claim — including ours. No payback number survives contact with an unmeasured baseline. If you take one thing from this article, take that: the map orders your candidates, but a baseline is what makes the first one defensible.
The worksheet is one page and deliberately blank. It takes a working session, not a project.
That is the whole method. It is not complicated, and it is not meant to be — the difficulty was never the framework, it was being specific enough to use one.
One page, free, no email required. A blank quadrant grid and the five-question checklist — enough to run the session described above.
The assessment is the full version: we run the map across your functions, screen the candidates, and measure the baseline behind the first one.