Who Is Accountable When the Tooling Decides: AIOps and the Shrinking Team

Somewhere this year, in a network operations centre that will never be named in a case study, an engineer with three decades of service handed back a laptop and went home. The systems he looked after are still running. The runbooks he wrote are still in the wiki.

What went out of the door with him was the reason he always checked one particular switch before anyone raised an alert about it, a habit built from an outage in 2011 that nobody documented and everyone remembered. His replacement is a dashboard.

Multiply that morning across the industry and the shape of the problem becomes clear. Roles have been cut, retirements have gone unreplaced, and the volume of data those teams once interpreted has multiplied several times over. Observability was built to tell operators what was happening inside their systems, and it is now being asked to supply the judgement that used to sit in the heads of the people reading the screens.

Phil Lenton, Head of Product for SaaS and AIOps at Riverbed Technology, said the drain is running from two directions at once. “You are seeing this brain drain, either self-inflicted because companies are making cuts, or because people are ageing out,” he said.

The exposure is sharpest in disciplines where experience takes decades to build and the underlying technology has been stable for just as long. “Twenty, 30, 40 plus years of experience has really been a necessary part of getting the job done, and networks have not changed that much for many decades in many ways,” Lenton said.

AI cannot rebuild knowledge that was never written down

The window for capturing expertise closes when the expert leaves, and most organisations have reached for the tooling after that has already happened. “AI cannot magically reconstruct institutional knowledge that was never captured,” Lenton said.

What the technology can do, in his account, is record that knowledge while the people who hold it are still on the payroll and still solving problems in view of the system. Watching an engineer work also produces a better record than asking them afterwards.

“It can observe someone in action and probably record their actions better than if you ask them themselves how they solved that problem last week,” Lenton said. An engineer might recall checking DNS records and finding a conflict, when the record shows three other checks before it, in an order learned through years of incidents.

Rizwan Kareem, Business Unit Manager for Cognitive Engineering and Industries at Omnix International, drew the division in a single line. “AI understands patterns. Engineers understand consequences,” he said.

The architectural mistake Riverbed sees most often, and one the company tested on itself, is connecting the reasoning system straight to the data. “It is an incredibly fast way to spend an awful lot of tokens, and you end up with a mess, because you have no consistency,” Lenton said.

The system then works through every incident from first principles and produces a different answer each time it meets the same failure. Riverbed sits a skills library in between, carrying the weight a departing engineer used to carry.

“They capture how you should go about measuring this telemetry and looking for indicators, and they also really usefully capture the institutional knowledge of how that customer solves problems,” Lenton said. Jeff Stewart, Group Vice President of Product Management at SolarWinds, described enterprises reaching the same conclusion by another route, with internal knowledge bases and workflows wired into AI tooling through MCP servers so the context travels with the automation.

Accountability moves to whoever wrote the rule

When tooling detects, correlates and remediates without a human touching the incident, the accountability question stops being about the operator on shift. “The question for most businesses now is who created the automated policy,” Kareem said.

Responsibility sits with the engineering team that defines the governance model, the business rules and the operational limits within which the system is allowed to act. Stewart said this increases the human burden, because governing a recommendation engine is harder work than answering an alert.

“We as humans become even more accountable for the recommendations that these systems will be making,” he said. The vendor cannot supply the part that matters most, which is the organisation’s own definition of what is critical.

“The skills really need to come from the institution. We provide a base set, but that does not know how you as a bank or as a healthcare organisation operate,” Lenton said. Kareem made a similar point about design, saying every recommendation should be explainable, measurable and aligned with business objectives, and that the organisations gaining most from AIOps will be the ones practising the most responsible automation.

The expertise problem also has a supply side that the current cuts are cutting through. Junior engineers learned the discipline by doing the repetitive work that AI now absorbs, so the training ground for the next set of senior operators is disappearing at the same time as the current set retires.

“What are junior folks going to cut their teeth on? Where are they going to learn the ropes, and where are they going to learn the hard way?” Lenton said. Omnix sees the same risk in industrial deployments where model behaviour is understood by a handful of specialists.

“If only a small group of specialists understand how AI models generate recommendations, organisations risk replacing one dependency with another,” Kareem said. The organisation then depends on systems that few of its people can confidently interpret or maintain, and the concentration of knowledge has simply moved somewhere else.

Partial coverage protects only the failures a company has already seen

Many enterprises still deploy observability selectively, covering the systems where they have been burned before. Lenton considers the approach flawed on its own terms.

“It is flawed thinking to use observability systems on just the gaps, because what you really mean is just the gaps you know about,” he said. He compared it to a driver who uses indicators only when other vehicles are visible, at the exact moment the signal matters least.

Selective deployment also weakens the machine learning it is meant to support. Riverbed’s systems generate petabytes of raw telemetry a day for the largest clients, nearly all of it normal or within tolerance, and that normal behaviour is what the models learn from.

“If you deprive an AI observability system of that huge extent of situation normal data, you just make it less accurate for when the spikes occur,” Lenton said. Stewart described the human limit that makes the volume unmanageable without it, saying pattern recognition and anomaly detection at that scale sit beyond anything people can process in a meaningful way.

“People are still essential, especially for understanding business context, which AI generally does not understand, assessing risk and deciding when exceptions should override automated recommendations,” Stewart said.

The next thing enterprises will have to watch is the agents

The estate under observation is now growing faster than the teams responsible for it, as enterprises deploy agents and agent-to-agent loops across their operations. “What is watching those agents? Who are those agents talking to? Is the data clean? Is there governance around that data set?” Stewart said.

He expects observability to extend into agent behaviour, performance and token efficiency, and expects that extension to be substantial. He also sees a second-half acceleration in security posture work as newer models expose vulnerabilities at a pace defenders have not previously faced, with the pressure carrying into 2027.

He cautions that deployment is running ahead of governance. “The people, the process, the governance and the training are all things that I would like to see come before a lot of the investment in immediate deployment into some of these systems,” Stewart said.

Kareem described what resilience actually requires, calling it a continuous learning process in which every operational event strengthens future decision-making by embedding engineering knowledge into AI-assisted workflows. The constraint is the same in Germany, Japan, the United States and the Gulf, where heavy investment in automation and intelligent manufacturing meets a shortage of experienced people, and the answer everywhere starts with capturing what the leavers know before the tooling is asked to stand in for them.


Sindhu V Kashyap

Global Technology Journalist & Multimedia Storyteller | Covering Founders, Investors & Leaders Reshaping Tech | Writer · Interviewer · Moderator · Editor

Next
Next

Veeam adds six hypervisors and an archive tier as backup repositories become the ransomware target of choice