Keeping in-memory data consistent across multiple LPARs introduces coordination challenges. This design separates local reads from Sysplex-wide updates, using native z/OS services to synchronize changes.
Synchronizing In-Memory Data Across a Parallel Sysplex
Parallel Sysplex in-memory data synchronization presents a critical challenge: keeping local data copies consistent across LPARs while preserving the fast reads that transaction processing demands.”
Mainframe applications have long relied on in-memory data structures to deliver the sub-millisecond response times that modern transaction processing demands. CICS regions, batch programs, and other online workloads load configuration parameters, lookup tables, translation maps, and cached reference datasets into memory at startup and access them thousands of times per second.
On a standalone LPAR, in-memory data management is straightforward. Applications read from a local copy, updates are applied in place, and consistency is guaranteed by default. There is simply one source of truth and nothing to coordinate.
A Parallel Sysplex changes the equation. The same subsystem now runs independently on multiple LPARs, each with its own address space and copy of the data. These copies start identically, but they aren’t inherently linked. The first update applied on any single member causes the copies to diverge—and from that point forward, reads on non-updated members return stale data.
The real challenge is not performing the update. It is coordinating the update across every member of the Sysplex while maintaining data integrity and availability.
Two known established alternatives exist, each with trade-offs. The first involves placing the data in a shared Coupling Facility structure, which eliminates synchronization concerns but sacrifices read performance – since every access must now traverse the Coupling Facility link.
The other involves moving changed pages across the Plex—the approach DB2 for z/OS employs in Data Sharing. This works well for database workloads but introduces disproportionate complexity for simpler read-intensive data structures.
A New Model
There is a third design approach. One that is deliberately generic and can be adopted by any product or subsystem that manages in-memory data requiring Sysplex-wide consistency.
This approach is also optimized for read-intensive in-memory data—the kind that is loaded once and queried constantly but updated infrequently—and leverages native z/OS cross-system services to propagate changes consistently across the entire Sysplex.
It preserves the full speed of local reads while coordinating updates across the Sysplex only when they occur—using XCF Signaling for inter-member communication and the Coupling Facility’s IXLUSYNC service for coordinated two-phase processing.
The diagram below illustrates the high-level topology. Each LPAR in the Sysplex hosts application workloads (CICS, batch, IMS, or any other subsystem), a local copy of the in-memory data, and a Plex Server that manages update coordination. All Plex Servers connect to the Coupling Facility, which provides the shared infrastructure for cross-member event management.
Figure 1: Sysplex Architecture for In-Memory Data Synchronization
The key is the separation of read and update paths. Read requests are satisfied entirely from local memory—no network traversal, no Coupling Facility access, no coordination overhead. Update requests are coordinated centrally — routed to the local Plex Server, which manages the Sysplex-wide commit protocol through the Coupling Facility. This bifurcation ensures that most operations execute at full local speed.
Just as importantly, everything is validated first, broadcast second. The Primary member executes the update locally before sending anything to the rest of the Sysplex. Only proven-successful updates are propagated, eliminating most cross-system error scenarios.
These principles yield a system where the common case—high-frequency reads—is entirely unaffected by the synchronization machinery, while the rare case—an update—is handled through a controlled, well-defined protocol.
Don't miss these other great articles
How Updates Propagate Across the Sysplex
When an update arrives, the coordination protocol unfolds in two broad phases.
Phase 1: Lock and Update
The Plex Server on the receiving member writes the work to a LIST structure in the coupling facility. Primary members pick it up and then apply the update to its own local data. If the update fails locally, the process stops immediately and no other member is contacted. If it succeeds, the Primary broadcasts the update to all members via XCF Signaling, and the Coupling Facility coordinates the same update on every other LPAR. Each member applies the change to its own local copy.
Phase 2: Unlock
Once every member has successfully applied the update, the Coupling Facility triggers the second phase: all members release their write locks simultaneously, and the originating application is notified of successful completion. The Sysplex is now consistent—every member holds an identical copy of the updated data.
Figure 2: Two-Phase Update Flow Across the Sysplex
The elegance of this model lies in its delegation of coordination complexity to the infrastructure. The Coupling Facility tracks member responses, manages event sequencing, and triggers phase transitions automatically. The Plex Servers simply respond to events as they arrive.
Sample Applications
As described, this model is designed for applications making use of in-memory techniques and technologies to deliver better system performance.
Scenario 1: Financial Reference Data
Foreign exchange rates, interest rate tables, fee schedules, and more are loaded into memory at startup for sub-millisecond access during transaction processing.
On a single LPAR, this works well—but it creates a hard ceiling on throughput. All transaction channels (ATMs, online banking, wire transfers, trading desks, and batch processing) must compete for resources on that one member.
When transaction volumes spike during peak hours, the single LPAR becomes a bottleneck. Queues build, latency rises, and SLAs are breached. Meanwhile, other LPARs in the Sysplex sit idle because they do not have the reference data loaded—or if they do, their copies may be stale.
With synchronized in-memory data, all LPARs hold identical, current copies of the reference data. A workload distributor spreads incoming transactions across all members. Each LPAR contributes its own full TPS capacity, yielding a threefold increase with zero additional hardware.
There is no restart, no batch sync delay, and no risk of one LPAR processing at a stale rate while another uses the new rate. Every transaction, regardless of which LPAR processes it, uses the same reference data.
Scenario 2: Real-Time Fraud Detection “Hot Lists”
The critical weakness in the existing architecture is that this process runs on a fixed schedule. When an account is flagged as fraudulent, the hot list update sits in a queue until the next batch cycle completes. During that interval, every LPAR in the Sysplex continues to operate with a stale hot list, and the flagged account is not blocked anywhere. Transactions from the compromised account are approved on whichever LPAR they happen to reach until updated – resulting in potentially significant losses.
With this Sysplex-aware approach, the batch synchronization dependency is eliminated entirely. As soon as the fraud analytics engine flags an account, it broadcasts the hot list update to all members via XCF Signaling and commits it through IXL USYNC. Within sub-second, every LPAR’s in-memory hot list includes the flagged account. All fraudulent transactions that would have been approved during the batch window are now denied instantly—on every LPAR, from every payment channel.
Additionally, the fraud screening workload itself is distributed across the Sysplex. Instead of funneling all transactions through a single fraud engine, each LPAR independently screens its own traffic against the same current hot list. This triples fraud-screening capacity and reduces per-transaction latency, enabling more sophisticated real-time detection rules without impacting throughput.
Challenges and Considerations
Two areas demand particular attention during implementation. The Coupling Facility services are powerful but carry a learning curve—teams must invest in understanding event lifecycle management, structure allocation, and the behavioral nuances under failure conditions.
Recovery logic, meanwhile, is the most intricate aspect of any Sysplex-aware design, handling:
- Mid-update member failures
- Coupling Facility structure failures
- Member reintegration after recovery
All require careful implementation and testing.
Start Simple, Enhance Later
This design provides a clear path to achieving in-memory data consistency across a Parallel Sysplex without sacrificing read performance. By building on native z/OS infrastructure—XCF Signaling, IXLUSYNC, and the Coupling Facility—it avoids custom coordination logic and instead leverages services that IBM has engineered and supported for decades.
The recommended approach is pragmatic: start with a single-Primary model and a straightforward two-phase commit, prove correctness in production, and enhance incrementally. Future iterations can introduce more granular locking, parallel update pipelines, and automated recovery flows as operational experience accumulates.
Start simple. Build the product first and enhance it later. The foundation must be solid before the architecture can grow.
For any team managing in-memory data on z/OS and facing the challenge of Sysplex-wide consistency, this approach offers a robust starting point—one that balances engineering rigor with practical simplicity and is designed to evolve from the outset.










0 Comments