A missed comma, a fake citation: the small AI mistakes professionals were never taught to catch

Most people met their workplace AI assistant in the same way. Someone from IT or the vendor shared a screen, typed in a request, and a neat summary of a 40-page contract appeared a few seconds later. There was a slide listing everything the tool could do, from drafting emails to translating documents and writing code, and then everyone was given a login. Hardly anyone walked out of that session knowing when the assistant was likely to get something wrong, what a wrong answer would look like, or what they were supposed to check before sending its work on under their own name.

A global survey of more than 48,000 people in 47 countries, carried out by KPMG and the University of Melbourne in 2025, shows how wide that gap has become. Fewer than half of employees, 47%, said they had received any AI training at all, and only 40% said their employer had a policy or guidance on using it. Two-thirds, 66%, admitted relying on what the AI produced without checking whether it was accurate, and 56% said they had made mistakes in their work as a result.

One of the most public examples came in October 2025, when Deloitte agreed to repay part of a A$440,000 fee to Australia's Department of Employment and Workplace Relations. A report the firm had written for the department cited academic papers that did not exist and quoted a court judgment that had never said what the report claimed. The corrected version admitted that an AI tool had been used in preparing it. On the page, every one of those false references looked just like a real one.

Kurt Muehmel, Head of AI Strategy at Dataiku, said these are the mistakes that should worry companies most. “The hardest errors to catch are the ones that do not look like errors. Certain errors are obvious. Others are more insidious, in the sense that they are more difficult to detect,” he said.

Hassan Abu Sheikh, Co-Founder of CNTXT AI, said the risk is highest in everyday work, where people have stopped paying attention. “The error that looks trivial is the error nobody is checking for, because everybody has decided the system is good at the hard problems and has stopped asking whether it is good at the easy ones,” he said.

Catherine Bozhenko, Product Manager at DataRobot, said she would be wary of any company that handed out these tools without first finding out where they break. “If you did not do this before launch, I would be very scared to use it in your organisation,” she said.

Companies talk about what worked and keep quiet about what did not

Muehmel compared the way companies talk about AI with the way people post on social media. “It works a little like what we see in people's personal lives when they are posting on Instagram. They are not showing all the failed outfits they tried on before they found the one that looked really great. They are not showing all the terrible restaurants they went to. They are only showing the good ones. There is some filtering going on when organisations present their work. Every organisation wants to be seen as extremely successful and savvy,” he said.

He said a lot of failure is to be expected at this early stage, and the harm comes from hiding it. “The reality is not that, and I think that is very understandable. We are at a phase with this technology where there is a great deal of learning, and with learning comes a great deal of failure, rather like a young child learning to walk. I recently met a new nephew of mine who is in exactly that phase, and he spends far more time falling than he does walking. We know he will be walking very soon,” he said.

So the employee in the training room sees the tool at its best and rarely at its worst. Abu Sheikh said one reason is that companies are adopting AI in two very different ways at the same time. Some move fast and clean up afterwards, while others go slowly and fix problems as they appear.

“You either grow at massive scale and fix afterwards, or you grow step by step and fix step by step. With AI, in some industries, we are seeing these two approaches happen side by side in a way I think is unseen. In some areas we are going back and fixing all the mistakes that were made along the way. In some areas we are going step by step,” he said.

He said the problem starts when both are sold with the same demonstration. “The areas where mistakes are acceptable and the areas where they are not have been treated as though they are the same thing. They are not the same thing, and the conversation about failure gets lost because everybody is looking at the same demonstration regardless of what the system is actually being asked to do once it goes into a business,” he said.

In fields such as healthcare, banking and government, where a wrong answer can hurt someone, Abu Sheikh said the pace should be set by what happens when the tool is wrong. Those fields can also learn from the companies that moved first. “

There are sensitive areas where a high chance of a mistake cannot be allowed to exist. That is where the pace of adoption has to answer to the consequences of the system being wrong. The industry that moved quickly has already discovered where things go wrong, has already gone back and fixed them, and that knowledge exists. A sensitive sector that moves carefully does not have to rediscover every one of those mistakes for itself,” he said.

The same question can get a different answer every time you ask it

Staff have been trained on software for decades on the understanding that it behaves the same way every time. A spreadsheet formula that works on Monday still works on Friday. AI assistants are different. Ask the same question twice and you may get two different answers, one right and one slightly wrong, and someone who has only ever seen the right one has no reason to expect the other.

Muehmel said this changes what it means for software to be faulty.

“With software it used to be the case that when you did the same thing, only one thing should happen. With generative AI, the same input does not always give the same output, and that is by design. The idea of what used to be faulty software, what used to be a bug, is very different now, and that same unpredictability is what gives it all the abilities traditional software does not have. There are a lot of benefits that come along with it,” he said.

The benefits are one reason the failures get less attention. “We have seen very clearly that generative AI is able to deliver productivity gains, especially in software development. That is where we have very clear evidence of significant gains, of 30% or more depending on the organisation,” Muehmel said. Outside software, he added, the gains are harder to pin down. “The difficulty is that it does not do just one thing, and organisations use it in very different ways, which means its impact runs across the whole business and becomes difficult to separate from everything else a company is changing at the same time,” he said.

Nobody is yet sure who pays when something goes wrong. “We are finding that out right now, and mostly it is being settled in the courts. Organisations are in somewhat uncharted territory when it comes to liability. Insurance companies are very interested in the question as well, about who is actually carrying the risk when these AI models are used for important business work,” Muehmel said.

He said companies should settle this before they sign anything. “I would encourage organisations to have a very close look at the terms and conditions of the AI providers they choose to work with, and to make sure that when they sign those contracts, it is clear who is liable and what counts as an error by the system,” he said. Until then, he added, early use should be limited to work where a mistake can be absorbed.

“We are in a phase of learning, and these tools must be used on real business problems, though in areas where a mistake is more tolerable than it would be with traditional software.”

The missed comma and the wrong sign cause more harm than the dramatic failure

When people picture AI going wrong, they tend to imagine something dramatic. Abu Sheikh said the mistakes that actually reach clients are usually small and boring. “The biggest problem is that we forget we trained it. The mistakes we made before are trained into it. A lot of memes say that AI cannot do simple maths, and yet it can do complex physics. That is true, because sometimes we make mistakes, and those mistakes we do not catch,” he said.

He said the tools have picked up the slips of the people whose work they learned from. “We go away, we sleep on it, we come back to the desk, and we see that we missed a column, or we missed a comma, or we forgot that this was a deduction and not an addition. That is where the simplest mistakes become the problem. On complex work it still makes mistakes, though day-to-day tasks are where I see it happen most,” he said.

Think of an accountant who asks an assistant to match up two sets of figures, or a lawyer who asks it for cases that support an argument. What comes back looks finished and tidy, and the one wrong number or made-up case sits among material that is otherwise correct. Only someone who reads every line will find it.

“This is where we need to be careful and actually look at what comes out, instead of saying generate a document, send it, and moving on to the next thing. The habit of reading what came back is the whole safeguard,” Abu Sheikh said.

That habit, he said, fades fastest when the tool has been getting things right. “The moment the person stops reading the output because the last 50 were correct, all the testing that came before has been wasted, because the system has been given a level of trust that no testing was ever meant to justify,” he said.

Muehmel said one person reading carefully matters, though some problems only show up when you look at thousands of answers together. His answer is to keep a set of answers that are known to be correct and to check the tool's responses against them regularly, so that a slow rise in mistakes gets noticed. “It becomes critically important to have a clear record of what a good response is supposed to look like. You need that to compare against, so that you can check performance across everything the system produces, instead of one answer at a time with a thumbs up and a thumbs down, and see whether its answers are starting to get worse. It is by looking at them all together that you find the real errors,” he said.

He added that the person who signs off on the work still has a job to do. “On an individual basis anyone needs to keep a clear eye on what is going out, especially if they are the person responsible for checking it, and to use their judgement to spot problems. The real solution in the long term is to test the answers against what is expected and confirm whether they are accurate,” he said.

Staff need to hear that the work is still theirs

If the usual training session covers what the tool can do, Abu Sheikh said the more useful message is about who still owns the work. “The most important thing we tell an employee is that the task will still be done by them. Before AI, you did the task. After AI, you need to do the task ten times better, sometimes 100 times better, and faster. The value of AI does not come from waiting at the chat window for an answer. It is there to make your job better. It does not exist to do your job for you,” he said.

Once people see it that way, he said, checking the output becomes their responsibility. “Human and AI need to work together, instead of a person looking at a screen waiting for the AI to produce the answer. You were already doing the task before any of this existed. The idea is to multiply what you can do, and the moment somebody understands their role that way, whether the output can be trusted becomes their question, and stops being a question they assume somebody else has answered,” he said.

He told a story from his own office. “The other day my co-founder was looking at my screen and asked whether that was the answer the AI had given or the instruction I was giving the AI. I said it was the instruction I was giving it. That is how long my instructions are. The length of that instruction is the work. It is the subject knowledge, the context and the judgement going in before anything comes back out, and it is the difference between using the tool and waiting on it,” he said.

The person who puts that much in, he said, is also the one best placed to notice a bad answer. “The person who types one line and takes whatever comes back has no way of judging the answer. The person who has put everything they know into the request already knows what a correct answer should look like before they read it,” Abu Sheikh said.

Muehmel said honest training depends on companies being honest internally about what has already gone wrong. Plenty of employees who were promised a revolution have gone back to their managers saying the assistant makes as much work as it saves. “This is a really important question about company culture. Even if organisations are not sharing their failures publicly, they absolutely need to be sharing them internally so that people can learn from them, and so that you do not repeat the same failures over and over again within the organisation,” he said.

Keeping a record of how each tool is actually being used, he said, is how a company finds out whether the tool is badly set up or should never have been used for that job. “Finding the projects that are not delivering what employees expected lets an organisation dig in and understand what is going on. Perhaps there is a problem with the way it was set up that can be fixed. Perhaps it has been applied to work that does not need AI at all and ought to be used elsewhere,” he said. “Without this kind of tracking, and without a culture where people give quick feedback, organisations are going to go round in circles, rolling out new technology, spending more, and unable to say what is working and what is not. They will not be able to help the employee who is frustrated that the assistant is not helping them the way it should.”

The way a tool will fail is decided long before anyone uses it

By the time an employee opens an AI tool, most of the choices that decide how it will fail have already been made. Bozhenko said the work of understanding those failures has to begin when the tool is first being designed. “If it starts at the moment you are launching, it is too late, and that is something you should never do. It should start at design, when you begin to think about what data you will use, how you will test it, and what you need to do to make sure it is safe before anyone uses it. That is exactly the right point to start,” she said.

She said companies too often treat this as something to deal with later. “People treat this as though it is a nice extra, something you attend to once everything else is working, and it is really the opposite. It is much more practical and much more basic than people think, and far more concrete than philosophy or a set of principles written down somewhere,” she said.

Bozhenko described four things any system should be able to show before it is released, each of which is a job someone can actually check. “Fair means you have tested whether it treats certain groups of people unfairly and you know where the gaps are. Explainable means somebody can account for why the system gave the answer it gave. Transparent means there are records that let you look back at what the system did after the event. Accountable means there is a person, and a written process, standing behind the decision to release it,” she said.

How much of this is needed depends on what the tool will be used for. “It starts with what you are building it for, because that is so different each time. Who is it for, what data does it use, how much does it deal with people, and how many decisions about people will it make? The more it touches people's rights, or decisions about people, or personal data, the higher the risk, and the more carefully it needs to be handled,” Bozhenko said. “Two systems can run on the same AI model and carry completely different levels of risk because of who they affect and what they decide. Testing them the same way is how organisations end up with a system that was checked as though nothing was at stake.”

Muehmel gave an example of how the same technology carries very different stakes. “A retailer and an insurance company will deal with AI differently. For the insurer, it is crit’t take certain things about a customer's identity into account when setting prices or deciding who can buy a policy. We do not want to discriminate against people because of their background. A retailer might be able to use data a little more freely, for something with less at stake than granting an insurance policy, such as targeting advertising,” he said.

Most companies use AI models built by the big technology labs, and Muehmel said their main duty is to be honest with the people on the other end. “The real question is how we use a technology we did not build ourselves in an ethical way. A lot of that is about disclosure, making sure we are not pretending to be human in any way, that people know clearly when they are dealing with an AI and when they are not, and that we are open about the whole process,” he said.

Every AI model fails in its own places, and general testing will not find them

Bozhenko said companies often assume that the safety checks they built for one AI model will work on the next. “This is the part people underestimate. They assume that what worked on one model will carry over to the next one, and it does not, because the weak spots are in different places in every model, and your safety checks are only as good as your knowledge of where that particular model goes wrong,” she said.

Her answer is to test each model on the kind of material and the kind of people it will actually deal with. She described making up realistic sample data, based on the real thing, and running it through the system to see where it stumbles.

“You create sample data based on the data you intend to use, you try it, and you see where it fails. That is the most practical way to stop this being guesswork. You can be sure that you tested it on the groups of people who actually live in your region, and you can see the specific problems, whether that is bias against certain nationalities or certain religions or anything else that shows up. Without this, you are working blind,” she said.

A tool built for a bank in Frankfurt, a hospital in Singapore or a government office in Abu Dhabi will serve very different people, and Bozhenko said no software company knows those people as well as the staff who serve them. “You are the expert in what you are building, and you know the details. The AI models are general. They are not built for each particular job, and they cannot be. Knowledge of your area, your market and your culture is something that will never be automated away. It has to come from a person who knows what a wrong answer looks like in that setting. A tool can run the test. It cannot tell you what to test for in a market it knows nothing about,” she said.

Unfair treatment of certain groups is only one of the dangers. Bozhenko said each of the others needs testing of its own, including personal data leaking out, the model making things up, and people hiding instructions in a document or email to trick the system into doing something it should not. The list gets longer once AI is allowed to take actions by itself, such as sending messages or changing records. “Now the system can act, and it does not only answer, so you need to make sure it never does anything it was not permitted to do. An action nobody authorised is a different kind of failure from an answer that was wrong. This is the part that gets skipped, because you can skip it without anything visibly going wrong on the day you launch,” she said.

At CNTXT AI, Abu Sheikh said every product is put through thousands of trial runs before a customer touches it. “Before it reaches a customer, it is tested rigorously. Thousands and thousands of test runs go through the system, so that if a problem does happen with a customer, we have already seen it coming. We make sure no half-cooked product goes out to the market. The main thing we look for is the AI making things up, and the whole point of the testing is that we find it before the customer does,” he said.

He said machine testing still has limits, so a small group of chosen users tries each product next. “Having AI test the system is not enough on its own, however many runs you put through it. Testing tells you how the system behaves in the situations you thought of. A person using it in their real work tells you about the situations you did not think of, and those are the ones that matter, because those are the ones that will actually happen. That small group of testers exists so that the wrong answers show up in front of people who were told to expect them and know what to do with them, instead of in front of a customer who was shown a demo and told the system works,” he said.

Companies that rushed in can still catch up, though what went unrecorded is lost

Plenty of companies moved fast and are now running tools that were never checked this way. Bozhenko said that is common and can be put right. “It is possible, even if you were one of the early adopters who started without thinking about any of this, which definitely happens, because when you are trying to get something exciting out quickly, safety rules are not the first thing on your mind. It usually becomes an afterthought, until someone from compliance comes and raises it. The way to deal with it is to accept that the system is running, and that thankfully it is running well, and then go back and add the things that were never built in, instead of taking the fact that nothing has gone wrong yet as proof that nothing will,” she said.

Most protections can be added afterwards, she said, with one exception. “If you did not test it enough and there are gaps around bias, you can still add checks on top to protect against it. You can change the instructions the system runs on, you can add extra checks, and then relaunch it. It can be fixed. If you did not keep records of what the system was doing, you can start keeping them now. You will lose what came before, because that history does not exist and cannot be recovered, though at least you will be able to track everything from that point on,” she said.

She added that a system which passed every test on the day it launched will still change once people start using it. “Its behaviour will change over time, and new risks will appear once people start using what you have given them. Even if you launched it properly, if you start to see it slipping, or people start trying to attack it, you will need to update it. That is normal work. It is far better to start now than much later, because every week the system runs without anyone watching is a week in which you have no record of how its behaviour has changed,” Bozhenko said.

The law is starting to push companies in this direction, though not equally everywhere. “Under the EU AI Act it is mandatory, whereas in the US it is more of a recommendation, and as far as I am aware, in the Middle East it is also more of a recommendation. Data protection, bias, transparency and human feedback are all things you need to work through before anything goes live. These rules came after some people already had systems running, so for some organisations they will apply to what is already there,” she said.

Muehmel said the European law works because it starts with what the tool is being used for. “The EU AI Act is a very good example. At its core it looks at a specific use and decides which level of risk it falls into, and based on that there are different requirements and limits on how the technology can be used. For the most part those risks focus on people's liberties and human rights and the ways they could be abused, which I think is a pretty good line to draw,” he said.

The blame should sit with whoever decided how the tool would be used

When an employee uses an AI tool in good faith, and it gets something wrong, many companies look first at the person who pressed send. Muehmel said the answer depends on how the tool was built and who decided where to use it. “Responsibility may not be the same for every use. In some cases it rests with the person using the tool to make sure they are using it properly. In other cases it is higher up. In many cases organisations need to be building systems that are safe for their employees to use, where the employee essentially cannot misuse them, and if they are not doing that, then the responsibility lies with the leadership,” he said.

He said the people who build the tools carry their share. “Employees have to use a technology as intended, and the leaders and engineers building those systems need to design them so that, while nobody can promise they will never make a mistake, they are hard to misuse. There is also a responsibility for making sure the technology is not used for things it should never be used for at all,” he said.

That responsibility, he added, has to be written down. “Who is responsible for what may change from one use to another and from one stage to another, but it absolutely needs to be assigned. It needs to be named and it needs to be tracked.”

That same record is what persuades a board to keep paying for AI. “The organisations that do this will be able to go to their CFO, their CEO and their board and explain very clearly what they spent on the technology, where they spent it, what they used it for, and what it did for revenue or profit. Organisations that have some success but cannot trace it back will find their leaders, their boards and their shareholders unwilling to invest any further, and they are going to fall behind their competitors. Without that record there is no proof that this technology is anything more than a fun science project,” Muehmel said.

Abu Sheikh said a company's attitude to failure is set at the top, and it starts with senior people using the product themselves. “If you are someone with power, or someone senior, you should be the first to test your products, instead of building another hyped-up product that you launch into the market while hoping for the best. The second thing is being with the engineers, in the war zone, actually building the product with them, actually testing it with them and actually checking it with them, instead of looking down from up high and leaving it at that,” he said.

The employee who was given a login and a demo is the last person in that chain, and Abu Sheikh said the first person in it is the one who decides how seriously failure gets taken. “A leader should be responsible for whatever the company puts out, from A to Z, because the employee is part of your organisation and part of your culture, and you are the one who spreads that culture. You should be involved in the systems you are building. Nobody below you will take the ways a product can fail seriously if the person at the top has only ever seen the demonstration,” he said.

Sindhu V Kashyap

Global Technology Journalist & Multimedia Storyteller | Covering Founders, Investors & Leaders Reshaping Tech | Writer · Interviewer · Moderator · Editor

Previous
Previous

A growing workforce now signs off on AI output it cannot question, change or refuse

Next
Next

Companies taught staff what AI can do and skipped the part where it gets things wrong