Artificial intelligence is becoming increasingly useful in engineering, inspection, training and assessment. It can review large amounts of information quickly, identify patterns, organise evidence and produce remarkably professional reports.
That capability is precisely why we need to understand its limitations.
At Electrical Testing Ltd, we had been using AI to help review video evidence from highway electrical assessments for several months. The results had been consistently useful. It could recognise technically demanding activities, identify relevant evidence and organise its findings in a way that made the review process considerably easier.
Then, without warning, it identified a serious safety incident. It failed the operative. It criticised the assessor for not intervening. On another analysis, it even questioned the integrity of the video evidence.
There was only one problem: the incident never happened.
A dropped lantern that was never dropped
The assessment footage showed routine highway electrical work being carried out from a mobile elevating work platform.
During the video there was a sudden, clearly audible bang.
The AI subsequently reported a critical safety failure, stating that an old lantern had been dropped from the MEWP basket to the ground. Based on that finding, it judged the operative “Not Yet Competent”.
On another run, the finding escalated. The AI said that the assessor should have stopped the work. It questioned the assessor’s supervision and later described the video as “heavily edited”.
The account had therefore developed from one supposed observation into four significant conclusions:
• a lantern had been dropped from height;
• the operative was not competent;
• the assessor had failed to intervene appropriately; and
• the integrity of the evidence itself was questionable.
It was detailed, structured, timestamped and written with confidence.
And it was wrong.
When the original footage was reviewed by a competent person, the cause of the noise was obvious. A gust of wind had blown over a plastic pedestrian barrier. The barrier hit the pavement. The lantern remained on the lighting column throughout.
That distinction matters enormously. The AI had not simply used the wrong word. It had created an event which had not occurred and then used that false event to make consequential judgements about real people.

The central problem: a real sound was linked to the wrong object, and the false event then drove a chain of further conclusions.
How could AI make such a convincing mistake?
One of the difficulties with modern AI is that its output can look remarkably authoritative.
Formatting, technical terminology, timestamps and confident language can make a conclusion feel as though it has been established through a conventional evidential process.
But generating a plausible answer is not the same thing as proving that answer is true.
A simplified way of thinking about a large language model is as an extraordinarily sophisticated prediction system. Text is broken into pieces called tokens. Other information, including images, audio and video, must similarly be represented in forms that the system can process. The model then generates its response piece by piece based on the input, the surrounding context and patterns learned during training.
A simple analogy
“I made a cup of…” Tea would be an entirely plausible continuation. So would coffee. But the fact that tea is statistically or contextually plausible does not establish what was actually made.
The same distinction appears to have mattered in our assessment. The AI had highway electrical work taking place at height. It had a lantern being worked on. It had a sudden loud impact. “Dropped lantern” was a plausible explanation.
But plausible is not evidence.
The information the report failed to account for was the plastic barrier. The sound was real. The object was wrong.
We cannot claim to know the precise internal mechanism that caused this particular model to make the mistake. What we can establish from the evidence is much simpler: the report attributed a real sound to an object that demonstrably remained attached to the column.
And importantly, a competent person did not need access to the AI’s internal workings to disprove the finding. They simply needed to go back to the original evidence.
A better prompt is not the whole answer
A reasonable reaction might be that the AI was simply given poor instructions. It wasn’t.
The prompt used for the assessment was extensive. Among other things, it instructed the system to record only what could be seen or heard, avoid inference or assumption, provide timestamps, identify safety concerns, assess evidence against defined competence requirements and review the behaviour of the assessor.
Those are sensible controls. They still did not prevent the error.
More interestingly, once the AI had wrongly accepted the dropped lantern as an observed fact, those same instructions helped give the mistake additional authority.
The timestamp made it look precise. The safety review categorised it as a serious failure. The competence assessment converted the false observation into a failed operative. And the mandatory assessor review then created a second person apparently responsible for an event which had never happened.
Good prompting matters. Better prompts can undoubtedly improve AI performance. But prompt engineering cannot be the sole safety control where the consequence of an incorrect answer matters.
Repeating an answer does not make it true
There was another particularly interesting feature of the case. The video was analysed three times. All three analyses identified the supposed dropped lantern.
At first sight, that feels reassuring: three analyses; three matching conclusions.
But they were not three independent witnesses. They were repeated runs of the same AI system over the same evidence.
Consistency demonstrates that an answer is repeatable. It does not establish that the answer is correct.
That is an important distinction as organisations increasingly use AI for quality assurance and verification. A wrong answer does not become evidence simply because the machine can reproduce it.

None of this means AI is useless
There would be an equally serious mistake in drawing the opposite conclusion and dismissing AI altogether.
In the same assessment footage, the AI correctly recognised several technically demanding activities, including the prove-test-prove safe isolation sequence, appropriate PPE and LV glove discipline, use of a torque wrench and appropriate handling of WEEE waste.
That is significant capability.
AI can help competent people process evidence faster. It can organise information, identify areas deserving attention, highlight inconsistencies and produce a useful first-pass analysis.
The objective should therefore not be to choose between AI or people.
A more useful question is: Which parts of the process can AI improve, and which decisions still require competent human judgement?
For safety-critical or consequential decisions, we think that distinction is essential.
The competence paradox
There is also a longer-term issue that engineering organisations need to consider.
We frequently say that important AI output should be checked by an experienced, competent person. That is sensible. But where do experienced people come from?
Engineers, technicians, inspectors and assessors do not become experts simply by reaching a particular age or job title. Expertise is built by doing the work.
Junior people inspect installations, check drawings, read specifications, write reports, make decisions, make mistakes and have those mistakes corrected. Repetition and feedback gradually develop the judgement that allows somebody to become an experienced reviewer.
Yet many of those routine activities are precisely the tasks AI is becoming capable of performing.
The competence paradox: if AI removes too much entry-level work, organisations may eventually need experienced people to supervise AI while producing fewer people with the experience required to do so.
Human oversight only works if we continue developing humans capable of providing it.
AI strategy therefore needs to include competence development. Sometimes allowing a developing engineer or technician to undertake a task that AI could complete more quickly may still have value, because undertaking that task is how the individual develops the judgement required later in their career.

The strongest model is not “AI instead of people”, but AI used within a competent, evidence-based workflow.
Six practical guardrails for using AI

The witness and the storyteller
A plastic barrier fell over. AI transformed that into a dropped lantern, a failed operative and a negligent assessor.
Fortunately, the workflow included an experienced person who returned to the original evidence and challenged the conclusion.
That is perhaps the most useful lesson from the whole exercise.
AI is becoming extraordinarily capable and it has a valuable role to play in engineering, inspection, training and assessment. We should use it.
But when the conclusion matters, someone still needs to understand the work well enough to say: “Show me the evidence.”
Because ultimately, the raw footage was the witness.
The AI was the storyteller.
Website upload notes
This final section is for the website developer and is not intended to appear as part of the article body.

Editorial note: Article adapted from Simon Hobbs’ presentation “When the Machine Sees a Crime That Never Happened”.