Companies taught staff what AI can do and skipped the part where it gets things wrong

Most people who use AI at work learned it the same way. Someone ran a training session, showed what the assistant could write, summarise and analyse, and sent everyone back to their desks. The session covered what the tool could do. It said very little about when the tool gets things wrong. For a lawyer, an analyst or a civil servant who has to sign off on the result, that is the most useful thing to know.

Kurt Muehmel, Head of AI Strategy at Dataiku, said the most dangerous mistakes look perfectly reasonable. "The hardest errors to catch are the ones that do not look like errors," he said. "Certain errors are obvious. Others are more insidious, in the sense that they are more difficult to detect."

Hassan Abu Sheikh, Co-Founder of CNTXT AI, has seen the same problem in everyday work. "The error that looks trivial is the error nobody is checking for," he said. "Everybody has decided the system is good at the hard problems and has stopped asking whether it is good at the easy ones."

Many employees have already found this out. They were told the assistant would change how they work, and they came back saying they still had to redo much of what it produced. Catherine Bozhenko, Product Manager at DataRobot, said companies that put these systems in front of staff without first finding out how they fail have taken a risk she would not take. "If you did not do this before deployment, I would be very scared to use it in your organisation," she said.

Companies share their wins and keep the failures to themselves

Muehmel said companies talk about AI the way people post on Instagram. "They are not showing all the failed outfits they tried on before they found the one that looked really great. They are not showing all the terrible restaurants they went to. They are only showing the good ones," he said. "Every organisation wants to be perceived as extremely successful and savvy with its deployments."

He sees nothing strange in a new technology failing often. The trouble starts when a company pretends it doesn't. "With learning comes a great deal of failure, rather like a young child learning to walk," he said. "I recently met a new nephew of mine who is in exactly that phase, and he spends far more time falling than he does walking. We know he will be walking very soon."

Abu Sheikh said the market is moving at two speeds. Some industries are racing ahead and cleaning up their mistakes later. Others are moving carefully and fixing things as they go. The trouble, he said, is that both get sold the same way. "The areas where mistakes are acceptable and the areas where they are not have been treated as though they are the same thing," he said. "The failure conversation gets lost because everybody is looking at the same demonstration, regardless of what the system is actually being asked to do once it goes into a business."

He added that sectors such as healthcare or government can learn from the fast movers without repeating their mistakes. "The industry that scaled quickly has already discovered the failure modes, has already gone back and fixed them, and that knowledge exists," he said. "A sensitive sector that moves carefully does not have to rediscover every one of those mistakes for itself."

The same question can get two different answers

Most workplace training still treats AI like ordinary office software, which does the same thing every time you press the same button. Generative AI does not work that way. Ask it the same question twice and the answers can differ. "It used to be the case that software was always deterministic, meaning that when you did the same thing, only one thing should happen," Muehmel said. "The notion of what used to be faulty software, what used to be bugs in software, is very different with generative AI."

The same unpredictability makes the tools useful, and the gains are real. "We have seen very clearly that generative AI is able to deliver productivity gains, especially in software development," Muehmel said. "That is where we have very clear evidence of significant gains, on the order of 30% or more depending on the organisation."

Nobody yet agrees on who pays when an AI system causes harm. "We are finding that out right now, and primarily it is being handled through the court system," Muehmel said. "Insurance companies are very interested in the question as well, about who is actually carrying risk when using these new AI models for critical business processes."

His advice is to settle that before signing anything, and to start with work where a mistake will not do serious damage. "Have a very close look at the terms and conditions of the AI providers you choose to work with, and make sure that when you sign those contracts, liability is clearly assigned and what constitutes an error is clearly defined," he said. "Then apply it in areas where there is a certain tolerance for mistakes."

The small mistakes are the ones that get through

Abu Sheikh said the errors that slip past people are usually small and dull. "A lot of memes come out saying that AI cannot do simple maths, and yet it can do complex physics," he said. "That is true, because it learned from us, and sometimes we make mistakes that we do not catch."

He described the kind of slip every office worker knows. "We go away, we sleep on it, we come back to the desk, and we see that we missed a column, or we missed a comma, or we forgot that this was a deduction and not an addition," he said. "On day-to-day tasks is where I see the highest probability."

His answer is simple, and many people skip it: read what comes back. "We need to be meticulous and actually look at the output, instead of saying generate a document, send it, and moving on to the next thing," he said. "The moment the person stops reading it because the last 50 outputs were correct, the testing that came before has been wasted."

Muehmel agreed that the person checking the work matters. He added that one person reading one answer at a time will miss a system that is slowly getting worse. Companies need a clear picture of what a good answer looks like and need to check the tool against it regularly, across hundreds of answers at once. "It is important for any one person to keep a clear eye on what is going out, and to use their judgement," he said. "The real solution long term is to test the responses against what is expected and to confirm whether they are accurate."

Bozhenko said most of the decisions that determine how a system fails are made before anyone uses it. "If it starts at the moment you are deploying, it is too late," she said. "People treat this as something you attend to once everything else is working, and it is really the opposite. It is much more practical and much more basic than people think."

What that means depends on the job the tool is doing. A system that helps write marketing copy carries far less risk than one that screens job applicants or approves loans. "The more it makes decisions about people, or handles any data about people, the more it falls into the high-risk category," she said. "Two systems can use the same underlying model and belong in completely different risk categories because of who they touch and what they decide."

Each AI model gets different things wrong

Bozhenko said many companies assume that checks which worked for one AI model will work for the next. They won't. "The gaps sit in different places in every model," she said. "The safeguards you design are only as good as your knowledge of where that particular model fails."

The way to find out is to test the tool with the kind of information it will really handle, and against the people it will really serve. "You can see the particular problems, whether that is bias towards nationalities or bias towards religions or anything else that shows up," she said. "Without this, you are doing things blind."

Software can run those tests. Deciding what to test for takes someone who knows the market and the people in it. "The models are generic. They are not built for each particular case, and they cannot be," Bozhenko said. "A tool can run the evaluation. It cannot tell you what to evaluate for in a market it knows nothing about."

Unfair treatment of certain groups is only one risk. Bozhenko also listed private data leaking out, outsiders tricking the system into misbehaving, and the tool making things up. Newer AI agents can act on their own, sending emails or changing records, so there is another risk: they may do something nobody approved. "An action that was never authorised is a different category of failure from an answer that was wrong," she said. "This is the part that gets skipped, because it can be skipped without anything visibly going wrong on the day you launch."

At CNTXT AI, Abu Sheikh said every product is put through thousands of test runs before a customer sees it. "The word we look for when a problem happens is hallucination," he said, using the industry term for AI making things up. "The testing exists so that it is something we have already seen ourselves and not something the customer discovers on our behalf."

Even so, he does not trust machine testing alone, so a small group of chosen users tries each product first. "Simulation tells you how the system behaves in the situations you thought to give it," he said. "A human being using it in their actual work tells you about the situations you did not think of, and those are the ones that will happen for real."

Employees need to be told the work is still theirs

Abu Sheikh said the first thing he tells staff about AI is who is responsible for the result. "The most important thing we tell an employee is that the task will still be done by them," he said. "The work of AI is to enhance your job. It does not exist to do your job for you."

He told a story about his co-founder looking over his shoulder. "He asked whether that was the answer that came out or the prompt I was giving the AI. I said it was the prompt I was giving the AI. That is how long my prompt is," he said. "The person who types one line and takes what comes back has no way to judge the answer. The person who has put everything they know into the request already knows what a correct answer should look like."

Muehmel said companies that keep their AI failures out of public view still have to talk about them inside the business. "They absolutely need to be sharing those failures internally so that there can be shared learnings from them," he said. "That way you do not repeat those same failures over and over again within the organisation."

When an employee complains that the assistant isn't helping, that complaint is useful. It can show that the tool was set up badly, or that it was put onto a job that didn't need AI at all. Companies that don't collect these reports, Muehmel said, end up "going in circles, deploying new technology, spending more, and not being able to report back on what is working and what is not." They also lose the argument with their boards. "Without that record there is no proof that this technology is anything other than a fun science project," he said.

Tools already in use can still be fixed

Plenty of companies rushed AI tools out to staff with none of this in place. Bozhenko said that is common and can be put right. "When you are trying to get something out fast and exciting, governance is not what you think about first," she said. "It usually comes up when someone from compliance raises it."

Most of the missing checks can be added later, she said, though a company can never recover a record of what the tool did before anyone was keeping one. "You can still add guards on top, you can add additional checkers, and redeploy it. It is fixable," she said. "You will lose what came before, because that history does not exist and cannot be recovered. You will at least be able to start tracking from that point onwards."

A tool that worked well at launch will not stay that way on its own, because people start using it in new ways and some will try to break it. "It will drift, and there will be new risks once users start to interact with what you have given them," she said. "Every week the system runs unmonitored is a week in which you have no record of how its behaviour has changed."

Regulators are starting to push, though unevenly. "With the EU's AI Act it is very mandatory, whereas in the States it is more of a recommendation, and as far as I am aware, in the Middle East it is also more of a recommendation," Bozhenko said. For early adopters, she added, the rules now reach back to systems they put in place before those rules existed.

Someone has to own the mistake

When an employee uses an AI tool in good faith, and it gets something wrong, Muehmel said the blame depends on how the tool was set up. "In some cases it may rest with the end user to make sure they are appropriately using the technology," he said. "In many cases organisations need to be designing systems that are safe for their employees to use, where the employee essentially cannot use it irresponsibly. If they are not doing that, then the accountability is on leadership."

He said responsibility has to be spelt out for every use and every stage. "It needs to be named and it needs to be tracked," he said. "Organisations need to hold their employees at every level of seniority to account."

Abu Sheikh said it starts with the person at the top. "If you are someone with power, you should be the first to test your products, instead of launching another hype product into the market and hoping for the best," he said. "That means being with the engineers, in the war zone, actually building and testing the product with them, instead of looking down from up high."

That, he said, is what decides whether staff take the tool's weak points seriously. "You are the one who spreads the culture," he said. "Nobody below you will treat the failure modes of a product as serious if the person at the top has only ever seen the demonstration."

Sindhu V Kashyap

Global Technology Journalist & Multimedia Storyteller | Covering Founders, Investors & Leaders Reshaping Tech | Writer · Interviewer · Moderator · Editor

Next
Next

Seagate and SK hynix cut AI reply times by 95% by moving chatbot memory onto hard drives