Almost every company we meet has a list of AI ideas. The difficulty is not generating ideas but choosing the one to do first — the one that will show a real result within weeks and earn the organisation the confidence to do the next. Over time we have settled on six criteria. We score each candidate from 1 to 5 on each, in a workshop with the people who own the process.
The six criteria
1. Frequency
How often does the task happen? Something done two hundred times a day by ten people is a better first project than something done once a quarter by a director, however painful the quarterly task is. Frequency is what turns a small improvement into a visible one.
2. Pain
Does anyone actually mind? The best candidates are tasks people complain about — retyping shipping documents, answering the same twenty questions, chasing month-end numbers. If the current process is tolerable, the new one will not be adopted.
3. Data availability
Is the information the AI needs already in a system, in a usable form, and can we get access within the pilot? "It is in people's heads" or "it is in WhatsApp" scores low. "It is in the ERP and a SharePoint library" scores high. This criterion kills more candidates than any other, which is why we ask it early.
4. Measurability
Can we state, before starting, the number that should move — minutes per document, first-response time, answer accuracy on a test set, forecast error? If the benefit can only be described as "better", the project will be hard to defend and impossible to scale.
5. Risk if wrong
What happens when the AI makes a mistake — and it will, sometimes? A draft that a person reviews before sending is low risk. A decision applied automatically to a customer's account is high risk. First projects should sit at the low-risk end, with a human in the loop, so that errors are caught and corrected rather than feared.
6. Ownership
Is there a manager who wants this, will make their team use it during the pilot, and will say so afterwards? A project without a business owner becomes an IT experiment. We would rather do a slightly less valuable project that has an enthusiastic owner.
What usually wins
Scored honestly, the winners are often unglamorous: a customer-service assistant for the most common enquiries, document extraction for a high-volume form, a forecast for one product family, a bilingual draft generator for marketing. These are also the projects that build the foundation — the data connections, the access model, the measurement habit — that the more ambitious ideas will need.
The pilot that follows
With a use case chosen, we run a four-to-six-week pilot on the client's own Azure subscription:
- Week 1 — Scope and sources. Confirm the measure, gather the data or documents, write the test set, settle access and privacy questions.
- Weeks 2–3 — Build. Connect the data, build the assistant or model, run it against the test set, iterate with the business owner.
- Weeks 4–5 — Use it for real. A named group uses it in daily work alongside the old process. We watch the measure, collect corrections and fix what is wrong.
- Week 6 — Decide. Present the measured result against the baseline, the cost to run it, and a plan to scale or to stop. Either answer is a good outcome if it is based on evidence.
If you would like to run this scoring exercise on your own list, our whitepaper includes the sheet, and we are happy to facilitate a workshop.