Enterprise data gets moved around a lot. An application needs it, an analytics team wants a copy, or an AI project needs training data. Each request may make sense on its own. Add them together, though, and there are more pipelines to maintain, more copies to protect, and more questions about which data is current.
That was the starting point for Planet Mainframe’s October 6 IMS Virtual User Group session, “Data Gravity and IBM Z: Bringing Processing to the Data.” IBM’s Vishal Gupta and Cüneyt Göksu discussed what happens when organizations bring processing closer to their mainframe data, with examples covering virtualization, analytics, and AI.
Gupta is a Senior Solution Architect with IBM’s Z Ecosystem team in India, with more than 18 years of experience connecting mainframe data and applications with hybrid cloud environments. Göksu is an Executive IT Specialist at IBM’s Germany Development Lab, with more than 30 years of experience working with Db2 for z/OS and IBM Z. Together, they brought perspectives on data integration, architecture, and the day-to-day requirements of enterprise workloads.
Here are five things we learned.
1. Data gravity depends on how you use the data
Gupta started with a familiar idea from physics: More mass means a stronger gravitational pull. As data accumulates, it creates a similar pull on the applications and services that need it.
Database size is part of the picture. So are application complexity, the frequency of data access, and how much latency a workload can tolerate. An application that needs current transaction data throughout the day has different requirements from a report that can use yesterday’s numbers.
For an architecture decision, that means looking beyond data volume. How often will the application request data? How closely is it tied to other systems? How long can it wait for an answer? Those questions help determine where the processing belongs.
2. A copy comes with more than a storage bill
Gupta identified four considerations when moving data: security, consistency, latency, and cost.
Every additional copy creates another place where sensitive information needs protection. It also creates another version to keep consistent with the source. If a nightly ETL job supplies the data, the application is working with a snapshot. Continuous replication can provide more current data, but the pipeline still needs maintenance and monitoring.
Then there is the cost of building the pipeline, keeping it running, and storing what it produces. A decision to copy data is also a decision to take on those responsibilities. The question is whether the application’s requirements justify them.
3. SQL can provide a route to IMS data
For IMS users, one useful example was IBM Data Virtualization Manager for z/OS, or DVM. It gives applications a way to access different data sources through interfaces such as SQL.
Gupta walked through representing IMS segments as virtual tables. The setup uses information about the database, including the DBD, PSB, and COBOL copybook layout. DVM preserves the parent-child relationships between the virtualized segments.
Once that layer is in place, a developer can query the data without handling the underlying IMS access details. Virtual views can also bring together information from different sources. During the Q&A, Göksu pointed to joining a Db2 table and an IMS segment in one SQL statement as an advantage of using DVM.
The data remains in its source, while the application gets a relational way to work with it.
Don't miss these other great articles
4. Where analytics runs and how current its data is are separate decisions
Göksu described how IBM Db2 Analytics Accelerator supports complex analytical queries alongside Db2’s transactional work. Applications submit queries to Db2, and eligible queries can be routed to the Accelerator.
For IMS data, he discussed two options: periodic snapshots loaded through Accelerator Loader, and near-real-time replication through Classic CDC.
The choice comes back to what the application needs. A snapshot may be sufficient for one report. Another use case may require changes to be reflected throughout the day.
Keeping analytics within the IBM Z ecosystem still involves decisions about loading, replication, and synchronization. Location and data currency both matter when choosing an approach.
5. AI near the data can mean different things
The session covered two approaches to AI, each addressing a different need.
SQL Data Insights and SQL Data Insights Pro support semantic analysis: looking for similarities, patterns, and relationships beyond conventional SQL predicates. The speakers also described using DVM to bring IMS data into the SQL Data Insights Pro infrastructure for model training.
Machine Learning for IBM z/OS focuses on deploying models near the transactions that need their predictions. Gupta discussed importing models developed outside the mainframe, using formats such as PMML and ONNX, and integrating inference with mainframe applications.
He summed up that approach with a clear instruction: “Train your model anywhere, build your model anywhere, but execute it where your transaction happens.”
The thread running through the session was a useful question to ask before creating another copy: What does this application need, and can the work happen closer to the data already available?








0 Comments