How to Reduce ‘Mean Time to Innocence’ in Mainframe Incident Management

Sep 30, 2026

Biao Hao is a Principal Architect at IBM, where he helps some of the largest U.S. financial institutions advance digital transformation and innovation through IBM Z and AI. He is the technical lead for IBM’s participation in the Banking Industry Architecture Network (BIAN) and contributes to industry and community initiatives. With more than 20 years of financial services experience, he has led the strategy, architecture, and delivery of banking and insurance solutions and advised clients worldwide. He holds an M.Sc. in Physics and Computer Science from McGill University.

A mobile banking app begins timing out. The cloud operations team sees API failures. The IBM Z team sees Db2 lock contention. Both teams launch separate investigations, even though they are troubleshooting the same outage.

When a customer transaction spans distributed services and IBM Z, each team can see a valid symptom without seeing the full cause. Effective mainframe incident management connects a prioritized Z event group to the enterprise incident and application topology. That gives responders shared context and less time spent proving which team is innocent. 

Why ‘Mean Time to Identify’ Becomes ‘Mean Time to Innocence’

A payment authorization, insurance quote, claims update, or account lookup may begin in a mobile channel, pass through an API gateway and several services, use messaging, and complete against systems of record on IBM Z. Customers experience one transaction. Operations teams may see a chain that spans Kubernetes or OpenShift, MQ, CICS, IMS, Db2, batch, storage, network, and z/OS infrastructure.

When teams monitor those environments separately, each sees valid but incomplete evidence. The platform SRE team may see API latency, 5xx errors, pod restarts, or failed service calls. The Z SRE or system programmer may see CICS, Db2, WLM, JES, storage, or LPAR symptoms in a Z-focused console. Neither view alone explains the customer-facing failure.

“Teams can spend less time proving their innocence and more time fixing the problem.”

That is how MTTI (mean time to identify) becomes “mean time to innocence.” Teams spend the first part of an outage proving that their domain is not responsible. MTTR then depends more on handoffs, conference bridges, and manual timeline comparisons than on the technical fix itself.

In the opening example, the question is whether the API failures and Db2 lock contention belong to the same incident. Answering it requires a high-signal Z alert and a reliable link between the affected Db2 resource and the banking service.

A Practical Model for Mainframe Incident Management 

The model combines two complementary functions. First, a Z-domain operations capability collects Z metrics, logs, and events; correlates related conditions; reduces noise; and creates a prioritized event group for deeper diagnosis. Second, an enterprise operations capability correlates alerts across application, cloud, network, middleware, and mainframe environments, enriches them with topology, and connects the resulting incident to workflow and automation.

Concert for Z provides Z-specific operational intelligence. It uses Z data providers and OpenTelemetry collectors to consolidate and normalize events, logs, and metrics from integrated products, associate them with resource topology, and compress related events into event groups. [1]

Concert Operate provides enterprise-wide unified operations and incident management. It brings together alerts and operational data from distributed, cloud, network, middleware, and Z environments; correlates them into incidents; enriches them with topology; and connects them to workflows and automation. [2]

The design principle: promote the prioritized Concert for Z event group as an alert in Concert Operate, then attach it to the right resource in the shared topology. Distributed and Z symptoms can then converge in one incident without forwarding every raw Z event.

Figure 1. Event and topology integration across IBM Z and distributed environments

Figure 1. Event and topology integration across IBM Z and distributed environments

The figure shows two key connections: the event group carries the Z diagnosis into the enterprise workflow, and topology links the alert to the application and its distributed symptoms.

Event Integration: Promote High-Signal Z Events

First, make the Z event group a first-class alert in the enterprise incident system. IBM documentation identifies Cloud Pak for AIOps, the former name of Concert Operate, as the example webhook target. The receiving platform uses JSONata to map the incoming JSON payload into its event schema. [3, 4]

Preserve the alert summary, severity, classification, source event-group ID, a deep link back to Concert for Z, and a stable resource identifier. Those details let responders recognize the issue, return to the Z evidence, and match the alert to a resource in enterprise topology.

For implementers: The payload can include resource.name, resource.sourceId, type, sender, and relevant sysplex, LPAR, subsystem, job, transaction, CICS region, or Db2 subsystem identifiers. The exact mapping depends on the resources and incident policies in use.

Don't miss these other great articles

Topology Integration: Connect the Alert to the Application

Event integration tells the enterprise platform that something happened on Z. Topology explains why it matters. A Db2 alert should attach to the correct Db2 subsystem and, where the model supports it, to the CICS region, LPAR, sysplex, application service, and business service that depend on it.

In the banking example, the API alert and Db2 contention alert become meaningfully related when their resources map to the same service dependency path. Concert Operate can match resource.sourceId and resource.name against topology resource matchTokens; stable identifiers make that match possible. [5]

For implementers: IBM Z Resource Discovery can aggregate sources such as the z/OS Discovery Library Adapter, ADDI, CMCI, and HMC to populate enterprise topology. File or REST SDK observers can also model a critical path, such as mobile app to API to MQ to CICS to Db2 to LPAR and sysplex. [6, 7, 8]

The practical formula is simple: high-signal event group + stable resource identity + topology match = one cross-domain incident instead of two disconnected investigations.

Close the Loop Through the Incident Workflow

The primary event flow is from Z-domain correlation into enterprise incident management, but Z operators still need to know when their alert becomes part of a broader incident. Workflow integrations can notify the Z team and link the enterprise incident to the original Z event group. ServiceNow, Slack, Microsoft Teams, runbooks, and automation are possible routes. [9]

Clients with an established Netcool/OMNIbus environment can preserve existing enrichment, deduplication, maintenance suppression, and NOC procedures while adding enterprise correlation and topology. This is an environment-specific bridge, not a prerequisite for the overall model. [10]

What This Looks Like in Practice

Db2 Lock Contention and API Latency

An API gateway reports elevated latency for OrderService while OMEGAMON for Db2 detects lock chains and thread contention. Z-domain correlation produces one prioritized event group. The enterprise platform groups it with the API alert because both resources map to OrderService. Responders can see the customer-facing symptom and the likely Db2 cause in a shared incident, then follow the link to the Z evidence for diagnosis.

Batch SLA Protection

A delayed settlement or reconciliation job may not initially look like an online outage, but it can threaten downstream reporting, payments, or regulatory deadlines. Linking the Z workload event to the affected business service makes the priority and downstream impact visible.

How to Get Started

Start with one critical service and a Z event group that matters to its operation. Validate the event connection, resource match, and incident workflow before expanding coverage.

Phase 1: Surface the Z Event Group

Configure the outbound event path, map the payload, retain severity and classification, and include the source event-group ID and deep link. Success means the enterprise platform receives a recognizable, high-signal Z alert with a route back to its source.

Phase 2: Align Resource Identifiers

Standardize application, business-service, resource, subsystem, sysplex, and LPAR identifiers. Confirm that the incoming Z alert names the same resource used in the enterprise topology. This identity match supports meaningful cross-domain correlation.

Phase 3: Add Topology and Correlate Across Domains

Load discovered or lightweight topology for the most important services first. Validate that a Z-originated alert opens on the correct resource and can correlate along dependency paths to distributed components. Success means the Z and API symptoms appear in a shared incident when they affect the same service.

Phase 4: Operationalize the Workflow

Add ITSM integration, ChatOps notifications, runbooks, automation, ownership rules, and operating procedures. Ensure Z specialists receive enterprise incident context while retaining direct access to their domain tools.

From Parallel Investigations to One Coordinated Response

The point of integration is faster coordination across the systems that deliver one customer transaction. Correlate within the Z domain, promote the resulting event group, attach it to shared topology, and bring the right teams into the enterprise incident workflow.

Then a Db2 lock chain and the API latency it produces become one incident with shared context and a probable cause. Teams can spend less time proving their innocence and more time fixing the problem.

0 Comments

Submit a Comment

Your email address will not be published. Required fields are marked *

Sign up to receive the latest mainframe information

This field is for validation purposes and should be left unchanged.

Read More

AI-Powered WLM Cuts Batch Queue Time by Up to 70%

AI-Powered WLM Cuts Batch Queue Time by Up to 70%

What happens when Workload Manager can anticipate batch demand instead of simply responding to it? The latest issue of Cheryl Watson’s Tuning Letter leads off with findings from testing IBM AI-powered Workload Manager (WLM). The test used an AI model trained on 30...

Linux, Unix, Windows, and OS/2: How Demand Beat Design

Linux, Unix, Windows, and OS/2: How Demand Beat Design

The Rise of UNIX and Portable Computing  “OS/2” There. I said it. And I’ll say it again… but not right away so I don’t seem too fixated on something that inexplicably became an “also ran.” You know what it proved? A bunch of things, beginning with the fact that IBM is...