Messy is fine. Missing is not
Teams usually apologise for their data before showing it: inconsistent names, duplicates, gaps, years of spreadsheet convention. None of that blocks a forecast. It is normal, it is workable, and cleaning it is ordinary engineering.
What blocks a forecast is data that was never recorded. If nobody ever wrote down why a deal was lost, no model will tell you why deals are lost. This distinction — messy versus missing — decides more AI projects than any choice of model.
Check one: is the outcome recorded anywhere?
Every prediction needs a history of the thing being predicted. To forecast churn you need past churn, labelled, with dates. To predict admissions you need enrolments by cycle, by programme, and the applications that did not convert.
Ask where the outcome lives and who records it. If the answer is that people know it but nobody writes it down, your first project is to start recording it, and the forecast comes a cycle or two later. That is a real and useful finding, and it costs nothing to discover now.
Check two: how much history, in cycles rather than rows?
Volume matters less than coverage. A business with millions of transactions over four months has less forecasting power than one with three years of monthly figures, because the second has seen a full seasonal cycle several times.
For anything seasonal — education intake, festival demand, quarter-end buying — count cycles, not rows. Two full cycles is thin. Three or more is workable. Less than one, and any confident forecast is arithmetic dressed as insight.
Check three: is the past comparable to the present?
History only helps if the world it describes still exists. A pricing change, a new distribution channel, a merged product line or a change in how a field is filled in can all make older data describe a different business.
This is where domain knowledge beats technique. Ask the person who has been there longest what changed and when. Their list of dates is usually the most valuable input to the whole project.
Check four: can the signal be seen at all, in one place?
Most organisations hold the answer across several systems: demand in the CRM, credit terms in finance, commitments in operations. Each is right and none can see the others.
Before modelling, check that the join is even possible: that there is a shared identifier for the customer, the student or the order, or that one can be built. If there is not, the first piece of work is integration, and pretending otherwise produces a model trained on one third of the picture.
Check five: is there a decision waiting for the answer?
The least discussed readiness test is organisational. Who will act on the forecast, what would they do differently, and does that action have to happen before the data currently arrives?
If nobody owns the decision, the forecast becomes another report. We would rather find that out in a thirty-minute conversation than deliver a technically sound model that changes nothing.
What a readiness check should produce
A short written answer to one question: can this data support this decision, and if not, what would have to be true?
Sometimes the honest output is that the history needed does not exist yet. That finding has value: it stops a year of work, and it tells you exactly what to start recording today so the answer exists next year.
Next
AI opportunity assessment: 30 minutes, written answer either way