Executing a known task on a mainframe is a solved problem. Diagnosing an incident that spans five environments is not. And no amount of hiring solves it, because the challenge isn’t a shortage of hands.
Why faster diagnostics may be more important than faster task execution
Ask any IT leader about the mainframe skills gap, and you’ll hear a familiar story. Experienced engineers are retiring. The talent pipeline is thin. Institutional knowledge is walking out the door. It’s a real problem: 81% of IT leaders described their mainframe skills gap as very or extremely significant in a June 2026 study by Hanover Research for Rocket Software.
But that framing misses something important. The crisis isn’t only about who’s left to do the work. It’s about how much harder the work itself has become. And it reframes the whole conversation about AI in enterprise IT.
Consider what happens during a major incident today. A single business transaction no longer lives in one place. It crosses the mainframe, distributed systems, cloud services, APIs, security tools, and third-party services before it completes.
When something breaks, the signals that explain why are scattered across environments that don’t naturally communicate. Even a room full of seasoned experts can’t hold that entire picture in view at once. That challenge shows up in Rocket Software research: 51% of IT leaders reported difficulty bridging cloud and on-premises environments, highlighting the complexity of operating across increasingly hybrid technology estates.
That’s the real shift. Executing a known task on a mainframe is a solved problem. Diagnosing an incident that spans five environments is not. And no amount of hiring solves it, because the challenge isn’t a shortage of hands. It’s the difficulty of seeing the whole system clearly enough to know what to do next.
The bottleneck is knowing what the work is.
When an incident hits, teams race to answer the same set of questions:
- What happened?
- What’s impacted?
- What changed?
- Who needs to be involved?
- What’s the likely root cause?
None of these are execution questions. They’re diagnostic ones. And in a fragmented, multi-platform shop, they’re getting harder to answer, not easier.
“Faster execution doesn’t move the needle.
You can’t act on an answer you don’t have yet.”
This reframes the whole conversation about AI in enterprise IT. Most of the excitement has centered on speed and helping people complete tasks faster. But if diagnosis is the real bottleneck, then faster execution doesn’t move the needle. You can’t act on an answer you don’t have yet.
The Hanover Research study suggests leaders already understand this. When asked where they’re most willing to deploy agentic AI, the top-cited use case wasn’t automated action; it was detection and correlation.
- 50% want agentic AI to detect anomalies and correlate signals across systems.
- 48% wanted proactive identification of security risks
- 42% wanted proactive identification of operational risks
The priority is understanding before action.
Modern transactions make diagnosis harder
The mainframe itself isn’t the obstacle here. The study is direct on this point: 52% of leaders view mainframes as an AI advantage, citing reliability, data depth, and transaction integrity. Only 22% see them as a limitation. The difficulty comes from the connections between systems.
A transaction that starts as a mainframe batch job, moves through a distributed application, touches a cloud service, and passes through an integration layer leaves a trail of evidence spread across tools that were never designed to be read together.
Correlating those signals during a live incident, under time pressure and regulatory scrutiny, is genuinely hard work, regardless of how experienced your team is.
This is also why conversational AI has largely fallen short in operational settings. The Hanover study found that 43% of leaders say it depends too heavily on prompt quality and user expertise, 40% call it a poor fit for regulated environments, and 38% say it doesn’t integrate into incident workflows.
A chatbot can answer a question if you already know what to ask. It can’t survey your environment and surface what you should be asking.
Real diagnosis requires context. It requires linking applications, jobs, transactions, infrastructure, and services into a coherent picture, and then tracing an issue from symptom to root cause. That’s a fundamentally different capability than responding to prompts or running predefined tasks.
Instrumentation Layer
One of the most practical steps toward that context is happening at the instrumentation layer. IBM, Rocket Software, and others are working to bring OpenTelemetry instrumentation to the mainframe, so that z/OS subsystems, transactions, and jobs can emit traces, metrics, and logs in the same open, vendor-neutral format already used across distributed and cloud environments. That matters because it attacks the correlation problem at its source.
Today, much of the effort in a cross-platform incident is spent normalizing evidence, reconciling different timestamps, identifiers, and formats, before diagnosis can even begin. With common instrumentation, a transaction that starts on the mainframe can
carry the same trace context as it moves through a distributed application, a cloud service, and an integration layer, making the mainframe a first-class participant in the observability tooling enterprises already run rather than a separate island that requires specialized interpretation.
It also makes AI-driven diagnosis more dependable, because agentic systems reason far better over consistent, well-structured signals than over fragmented ones. This work is still maturing, and adoption will be incremental; existing monitoring investments don’t disappear, but the direction of travel is clear: the less human effort required to stitch environments together by hand, the less the quality of diagnosis depends on scarce expertise.
Don't miss these other great articles
Better intelligence before faster response
There’s a parallel lesson on the security side that reinforces the same argument. As threats grow more sophisticated and cross-environment in nature, the instinct is to automate response as quickly as possible. But you can’t safely automate a response to a threat you don’t fully understand.
“Moving from signal to diagnosis must come before moving from diagnosis to action.”
Operational teams facing cross-environment threats need clear intelligence first, a confident read on what’s happening and where, before automated action becomes wise. Rushing to respond without that understanding introduces risk, not resilience. In regulated environments, where a single misstep carries real financial and compliance consequences, moving from signal to diagnosis must come before moving from diagnosis to action.
The Hanover study makes clear that most leaders agree. Nearly half (48%) permit AI to take limited actions while humans remain accountable overall. Another 34% allow AI to provide insights while humans stay fully in charge of decisions. The appetite isn’t for autonomy. It’s for trusted intelligence that supports faster, better-informed human decisions.
From signal to diagnosis to action
The future of enterprise AI in mainframe environments isn’t simply about helping people work faster. It’s about compressing the time between signal, diagnosis, and action with confidence and auditability at every step.
Leaders already find this credible. In the Hanover study, 91% consider AI-driven diagnostics credible for reducing mean-time-to-resolution, with most expecting a reduction of 20% or more when implemented well. And 87% believe AI will help address the mainframe skills gap within two years, not by replacing expertise, but by making complex environments legible enough for both senior and junior staff to act on.
Rocket® EVA™ is built for exactly this: a governed, agentic AI platform that delivers precise, end-to-end diagnosis across core systems to link applications, jobs, transactions, infrastructure, and services so an issue can be traced from symptom toward root cause across the full estate—mainframe, distributed, and cloud.
It surfaces patterns and connections that no single team could consistently see on their own, within a framework where every action is tied to identity and role, high-risk steps route to human approval, and a full record exists of what happened, when, why, and who approved it.
That governance matters. In environments where 78% of leaders say regulatory considerations significantly limit their ability to deploy AI, faster diagnostics only count if you can prove how you got there.
Address operational complexity, not succession planning
The mainframe skills gap is real, and it deserves serious attention. But treating it purely as a staffing problem leads to the wrong solutions. You can’t hire your way out of operational complexity. You can’t train someone to manually correlate signals across environments faster than those environments generate them.
What you can do is invest in diagnostic capability that sees the whole picture, turning scattered signals into clear answers, keeping humans in control, and leaving a defensible audit trail at every step.
The organizations that handle major incidents best won’t be the ones with the fastest task automation. They’ll be the ones that can answer what happened, what’s impacted, and why, before the incident becomes a crisis.
Learn more about how Rocket EVA and governed agentic AI can help your organization move from signal to diagnosis to action faster.








0 Comments