"Our data isn't good enough" is the most common reason given for not starting, and it is almost always said about the wrong problem.
It usually means the categories are inconsistent. Three people class the same fault three different ways. Half the work orders sit under a general heading nobody has bothered to refine. The terminology drifted when the last shift supervisor left. If you are the person who would have to defend the data, that is what you are picturing when you say it is not good enough.
That kind of mess is survivable now. Something else, which usually gets less attention, is not. Knowing which is which is the difference between a project that can start and one that cannot.
What it forgives: structure
The requirement that dropped is the requirement for structure.
The old approach needed every user to categorise everything correctly at the point of entry — right failure class, right component, right cause, from a dropdown, at the end of a shift. Leonard Lin, Product Manager for Maintmaster Manufacturing Intelligence, is blunt about how well that worked: "nobody does that; that's way too much effort."
Mandatory category fields get populated. They do not get populated accurately. The first plausible option gets picked, and the taxonomy you designed becomes a record of what was at the top of the list.
That is precisely why weakly structured data can now be analysed. Free text can be read. Your classification does not have to be complete, consistent or even particularly good, because the analysis no longer depends on it to carry the meaning.
So: inconsistent categories, drifting terminology, a general bucket with far too much in it. Forgiven.
What it does not forgive: absence
Here is the boundary, and it is not about tidiness. It is about whether anything was written down.
Consider four closures for the same fault, which is roughly the spread you will find in any system that has been running a few years:
| The closure | What it supports |
|---|---|
| Nothing. There is no record that this happened. | |
| "Fixed." | A count, and nothing else. |
| "Replaced bearing." | What was done. Not what was wrong, so not whether it worked. |
| Bearing seized on drive end, replaced and re-greased | What was wrong and what was done. Usable. |
Only the last one can be learned from, and it took eight words.
The third is the interesting case, because it looks like a real record and is the one most likely to pass an internal audit. It tells you an action was taken. It does not tell you what problem the action was for, so it cannot tell you six months later whether replacing the bearing was the right call or whether the same misalignment is quietly eating a bearing a quarter.
What the empty ones cost is specific. Leonard again:
"Then you know the problems are there, but you don't know how to solve them, and then you can't learn from it."
Note what survives and what does not. The count survives. You can still see that this asset failed eleven times, still build the failure-rate KPI, still put it on a dashboard. What you cannot do is learn: not what went wrong, not what worked, not whether the same fix has been applied eleven times to a fault nobody has diagnosed.
This is, however, a great starting point. A great deal of maintenance history is in this state, and it already gives you visibility.
The events are in a system and somebody is closing the jobs — that is the expensive part, and it is already done. With this visibility, you can start improvement discussions in a data-driven way. You have data to justify why you want to focus on topic X. The next step, "add your fixes", is a logical step that makes sense for the team to follow. Your data improvement project is on its way automatically.
The floor, and the direction it runs
When it comes to data comment quality, AI is very capable nowadays, but there is also a simple floor. There is a test for this you can run in an afternoon, with no tooling. Leonard's version:
"If a human can't understand it, then the software will not be better."
Read a sample of last month's closures. If you cannot tell what happened, neither can anything else.
Leonard was careful enough to write down the limit of his own rule afterwards, and it is worth stating precisely because it is the part that gets misread. The rule runs one way only. If a person cannot read a note, the AI has no chance with it — but there are still cases where a person can read a note and the AI cannot.
So human-readable is a floor, not a guarantee. It reliably tells you what is unusable. It does not certify that everything above it will work. Reading the rule as our notes are legible, so we are fine runs it backwards, and it sets an expectation your data may not meet.
Leonard's other limit is older and has not moved: "if you have inaccuracies in your reports, it will not be good analysis." A lower bar for structure is not a lower bar for truth. A record that says the wrong thing is worse than a thin one, because a thin record announces itself and a wrong one does not.
Where you actually stand
The useful question is not whether your data is good. It is where it is bad — and that is answerable without an audit.
Manufacturing Intelligence measures what proportion of notes are of usable quality and breaks it down by department and by shift, so you get a map rather than a verdict: which areas, which teams, where the recording thins out. That turns a data-quality programme into a conversation with three named supervisors.
One honest limit. It shows you where recording is weak; whether teams then improve after seeing it is not something we have tracked, so we are not going to present it as a result. The map tells you where to look. Someone still has to go and change how the recording is done.
And when the analysis is working from records that are thin, the system tells you it is uncertain — it signals low confidence rather than presenting a thin answer as a solid one. That is not a guarantee of correctness.
Nobody has perfect data
One objection survives all of that, and it is the one usually said out loud: we don't have perfect data.
Nobody does, and Leonard is specific about why:
"Nobody has perfect data, because nobody ever used it, so naturally it's not good."
That is a causal claim, and it runs the opposite way round to how the problem normally gets described. Records do not become good and then get used. They get used, and that is what makes anyone bother to write them properly. Data nobody has ever read is data nobody has ever had a reason to improve.
Which is why you can start on middle-good data. Say 50% of your closures carry both the problem and the solution. Or say all of them name the problem and none of them record the solution. Neither is a set anyone would call good, and in both the issues that matter are still sampled and visible — because the issues that matter are the ones that keep happening, and anything that keeps happening lands in a sample that size. What comes out is usable and directional: it points at the right assets and the right repeated faults rather than measuring them precisely, which is what you need when the decision in front of you is where to look first.
It changes how the recording gets better, too. A team that can see its notes being used has a reason to write them; a team filing notes into a system nobody reads does not. Whether that happens is not something we have tracked, as above — it is the argument for starting early, not a result we can show you.
So start using the data when you are mid-way there, and let Manufacturing Intelligence boost the recording from there rather than waiting for it to be right first.
The three questions
About a sample of your own closed work orders, in this order:
- Is there a record at all? No record is not a data-quality problem. It is a data-existence problem, and it is a different project.
- Can a person tell what happened and what was done? If not, nothing downstream will do better. This is the floor.
- Is what it says correct? Worth a look, but it is rarely where a team is stuck.
That is not a survey finding, and we are not going to dress it up as one. It is what we hear in demo conversations: teams arrive assuming they have a question three problem, and what they go on to describe is question one. Not a wrong record but no record — paper, a spreadsheet, a shift book, a conversation at handover.
What that means you fix first
Each question points at a different first move, and none of them is a clean-up you have to finish before you start.
Fails question one — get the event into a system. Not a better system: a system. Leonard's framing of the first step sets a deliberately low bar — it "doesn't need to be ours." What matters is that the event is captured at the time it happens rather than reconstructed later. For maintenance that is a CMMS instead of a spreadsheet; for line stops it is production counters capturing events automatically instead of a shift book written up at the end of the day. Neither is an AI project. Both are what makes one possible.
Fails question two — change the habit, not the schema. This is where the note-quality map earns its keep: you are asking three supervisors for two short phrases instead of one word, not redesigning a taxonomy.
Fails question three — try it and see where you are. You do not have to settle this before you begin: Manufacturing Intelligence gives you guidance on where the weak spots are.
Worth naming the order this reverses. Most organisations arrive at all this from the analytics end: someone wants a dashboard or a prediction, and the data appears as an obstacle in front of it. Leonard's distinction is between "someone who wants to analyse the data versus digitising your data. And then you can analyse later." Those are two projects, and they get confused constantly.
If you have already tracked key maintenance KPIs for a while, you have a head start on telling the three apart — the KPI gives you the count, and these questions tell you whether anything sits behind it.
Pull twenty closures and read them. Whatever proportion fails question two is your real starting point, and you now have it without buying anything. If most of them fail, the fix is not a data programme. It is somewhere to type two short phrases at the point of work, which is what Maintmaster CMMS is for.
