[00:00:09] – Amanda Hendley
We’re here today for a virtual user group series. My name is Amanda Hendley. I am with Planet Mainframe. I’m happy to host you. This is one of 3 user group series that we run. We have our IMS, we do CICS, and we have Db2. And I’m really excited because this month is Security Month at Planet Mainframe. So we’re covering, um, all 3 of our user groups have a session this month and we’ll be talking about security, data privacy, DR, all of that. So we’ve got some great programming, and if you read Planet Mainframe this month, you’ll find, uh, on the site, uh, we’ll be publishing a lot of articles and content on those topics as well. So we’re here for Data Gravity and IBM Z, bringing processing to the data. You can— love for you to drop your location and job title in chat so we can see where everyone is from. It’s always fun to find out where people are joining from.
[00:01:16] – Amanda Hendley
Oh, let me go back real quick.
[00:01:25] – Amanda Hendley
Just wanted to check my audio for you.
[00:01:30] – Amanda Hendley
Okay, I think my audio might be a little bit better now. So, um, a couple of other things is, uh, we want to do a call for sessions. You’ll see an email about this out to the Planet Mainframe community this week, but we are looking for, uh, writers, presenters, thought leaders If you want to publish with Planet Mainframe, we’ve got a lot of great topics coming up over the next couple of months that we would love to be able to feature you and your voice on Planet Mainframe. You can contact me. I’m at ahenley@planetmainframe.com, and I’d love to talk to you about an idea or something you want to publish. Before we get going today, our Q&A, we’re going to let you drop it into chat. And we’ll answer it during the transitions. And I’m sure if there’s something more complicated that you would like to have answered as we go through the session today, we can take care of that. We can take care of that during, during the end, maybe even get you off mute if it’s something that you need to talk about a little bit more in depth. All right. Now to tell you about something else coming up, Arcati.
[00:02:47] – Amanda Hendley
We will be releasing the Arcati Mainframe Navigator survey soon, probably in the next 2 weeks. This is our annual research report. So we collect data from you, the end users, and the people who work on mainframe in order to put together our report. In February, we’ll release the whole thing. The opportunities to get involved include participating in the survey. The vendor directory, which is a free inclusion in our vendor directory, or you can sponsor the report and have an opportunity to present your own data, have it included in their advertising within the report and more. And then I would love to see folks in the next couple of events and programs coming up. I am going to be at IDUG North America and Tech Exchange in Atlanta where I’m based. October 27th through 29th. If you are an influential mainframer and haven’t gotten your pin this year, I will be there with pins. And then we are also exhibiting at GS UK Forge November 2nd through 5th. So some great programs. These are exceptional events, and I hope to see you at one of them. And now for our session, let me hop into my Zoom workspace and stop my screen share and introduce you to our speakers.
[00:04:28] – Amanda Hendley
I hope you can see the screen. Our speakers today are Vishal Gupta. He is a Senior Solution Architect with IBM’s z Ecosystem team in India. He’s got 18 years of experience, and he helps clients connect mainframe data and applications with hybrid cloud environments. And then we have Mehmet Göksu, an executive IT specialist at IBM’s Germany Development Lab. He’s got 30 years of experience with Db2 z/OS for IBM Z and helps clients with architecture, deployments, migrations, and performance. And with that, I’m gonna turn it over to you and let you take it away.
[00:05:07] – Vishal Gupta
Hi, everyone. Good morning. Good evening. Hope am I audible and my screen is visible?
[00:05:13] – Amanda Hendley
It looks great.
[00:05:15] – Vishal Gupta
Sure. Thanks. So thanks, Amanda, for introducing. So everyone, today we will be talking about a topic called data gravity and IBM Z and how it matters when we live in a hybrid cloud world. So before talking about the content, let me quickly talk about what we have planned to cover today. So we’ll talk about what is data gravity and why it matters. We’ll talk about what are the different parameters on which data gravity is dependent on and what are the risk factors when someone tries to fight against data gravity. And we’ll talk about how data gravity and IBM Z are related to each other and how it matters most for the application running on IBM Z. And then we’ll talk about that industry standard approaches which supports data gravity and the corresponding IBM solutions which supports the data gravity. So this is what we have a plan to cover today. And starting with, let me just start with the very simplest concept from physics called gravity, right? So as we know, gravity is directly related to the mass and we say The greater the mass, the stronger the gravitational pull, right? So the data behaves in a similar manner.
[00:06:37] – Vishal Gupta
So as data starts accumulated in a system, it makes something which is called data mass, right? So, and then that data mass attracts applications, processing services closer to it. What does that mean is when— this application and processes runs closer to the data, it gives us some benefits such as the latency benefit, it gives us throughput benefits and the cost optimization factor. So data gravity matters a lot when we think of running applications closer to the data. And this concept brings in an architectural shift. And that architectural shift which I’m referring here is What we say is instead of moving data closer to the processing where it happens, let’s bring processing closer to the data, right? The other way of thinking of it is let all the processing flow into the data where it originates and it resides instead of other way around, right? So this concept was introduced by, I, IT researcher Dave MacLaurin in 2010, which we are referring here, right? So now we’ll talk about why data gravity matters. So what are the different factors where it matters a lot, right? We know data is getting increased day by day. And now when we think of core systems like IBM Z, creation of data is not every hour or minute or second, right?
[00:08:15] – Vishal Gupta
It is like every millisecond. We are creating new data. And when that data starts accumulating and continuously it gets generated, it becomes very difficult for an organization to move the data from one platform to another because it becomes very complex. It’s very expensive to move the data, right? So the data size matters a lot when we think of data gravity. Talking about hybrid cloud platforms, None of the organization right now is just having their applications running on one single platform. They all are living in a hybrid cloud world. And when they live in a hybrid cloud world, their data spread across the system. And then it becomes very challenging for them to access the data or move the data from one platform to another between the platforms. So hybrid cloud also brings in the complexity. And the third, if I’ll talk about the business user expectations, right? The performance expectation which business users are looking for. Now they are not looking for after-the-fact information. So what business user needs is real-time insights from their data, real-time insights from their AI applications. And for that, a low latency is the must. And to achieve the low latency, If we have a processing running away to the data where we are introducing time lags because we are spending a lot of time while moving data from one platform to another, right?
[00:09:49] – Vishal Gupta
Which brings in the higher latency. So meeting the business expectation, moving data from one platform to another and then processing it is no longer a solution for the organizations. And if we talk about the cost factor and the risk factor, that also matters in the data gravity because movement of data is And when we make multiple copies, it brings in the risk factors. And we’ll talk about these 2 aspects, cost and risk, later in our next slides as well. Okay, so in short, we can say, right, moving data everywhere is no longer a scalable solution or an efficient and sustainable solution. Hence, the data gravity matters a lot in today’s hybrid cloud world. So I was talking about in my previous slide about the data size. So the question comes in our mind is, is data size is the only factor which drives data gravity? The answer is no. So there are other factors as well which drives data gravity pull, right? And those other factors are application mass. The complex and tightly coupled applications increases the data gravity. And if we think of an acceptable latency by a process or by an application, right, that also plays a major role.
[00:11:09] – Vishal Gupta
And to achieve the lower latency, processing must stay close to the data. So latency also plays a role. And if you think of a frequency of accessing the data, that is also very important, right? For example, we are looking for data to be accessed again and again, and if our processing is running away from the data then every time we are introducing that complexity. So if we think of mass, the data mass, higher the mass, stronger the gravity. And when we talk about the latency, higher the latency is acceptable, the weaker the gravity. So these are the other factors along with data size, which also depends in the data gravity concept, right? So we should worry about all these factors when we are thinking of should we go ahead with the data gravity concept or should we go against data gravity? And now in this slide, so I am trying to depict what an organization may face if they go against data gravity, right? So what are the risk factors they may see? Topmost concern for an organization should be data security. So when we make multiple copies of data, not within a system, even across the system as well, right, that gives more exposure to that sensitive data.
[00:12:31] – Vishal Gupta
And when we are giving more exposure to the sensitive data, right, it brings in the security concern for an organization. So making multiple copies brings in security concerns. And if we talk about the consistency, when we have multiple copies of data, making all the copies exactly same how we have it in our system of origin, right, it’s very challenging for an organization to make consistent data. Because when we are trying to make data consistent, it depends on multiple processing in between. It could be a replication process which is making or which is trying to data consistent, right? So data consistency is another issue when we think of making multiple copies of data across the platform. And the next in the list is a latency factor. So when we copy data from one system to another, It depends on what kind of flow we have established to make a copy of data. It could be my nightly ETL process, right? It means the copy which I’m getting in my target system is 24 hours stale. And even if you think of our latest replication processes, right, still we have a latency in between. So it all depends on our consuming applications and business users.
[00:13:57] – Vishal Gupta
how old, how stale data they are okay with. So latency is another risk factor which organization may face. And next we are talking about cost. So cost is 2 aspects of cost one has to bear if they think of making a copy of data. The one is they have to establish the pipelines, they have to spend money on building the pipelines of replicating data from one environment to another. then they also have to spend to manage those pipelines. And it’s not about those applications which we are building across the platforms where we wanted to replicate the data, it’s about the storage cost as well. So maintaining replication pipelines and introducing more storage is very expensive for an organization and they have to bear this cost when they think of fighting against data gravity. So moving data gravity always creates more problems than it solves, right? So we need to think of what is needed by our business, what kind of risk factors we may face when we think of fighting data gravity. So here we would like to relate it with IBM Z. And Suneet, I would like to hand it over to you to cover this piece that how it is related to IBM Z.
[00:15:21] – Mehmet Göksu
Thank you. So, so far you have seen that the data gravity, particularly z data gravity, is essential for both customers and of course for IBM as well. So again, the essence is instead of moving data off the platform for analytics, for AI, for streaming use cases, bring those applications back to the source. Okay. And again, why? Because the transactional data, which is Db2 and IMS, is created on the source, and which is the essential data for enterprise. So the question is how we will bring this actually capabilities into the source. So this is what we are going to discuss From now on. Next page.
[00:16:21] – Amanda Hendley
Yeah.
[00:16:22] – Mehmet Göksu
So this is the current big picture of IBM Z data and application integration architecture. So this picture has actually 3 pillars. Okay, on the left-hand side is IBM Z and the software stack But please focus, this is a picture focused on data and application integration architecture. And in the middle, on-prem platforms such as non-z data sources, IBM or non-IBM, x86, other hardware platforms, etc. And on the right-hand side, mostly it is public cloud and hyperscalers such as Databricks, Salesforce, Snowflake, and so forth. And these 2 lines, okay, blue line and green line, have different representations. So blue line is a direct access to the data sources, and green line can also be used as streaming use cases or for replication requirements. So let’s make a simple example. For instance, your system of record, okay, sitting on either IMS or Db2, all right, and your legacy transactional workload is running on IBM Z. But assume you are building a new application on a public cloud, and this application requires Direct access to Db2 or IMS. Okay. And let’s say the assumption is the enterprise decided to have a direct access instead of replicating data. So for Db2, for instance, you can use legacy JDBC, ODBC interfaces.
[00:18:27] – Mehmet Göksu
You can use REST API. And you can directly access to Db2 z/OS. Okay, and with that, no replication, single truth, direct access, zero latency, and very simplified architecture. You don’t need any other component, okay, in the middle as a middleman to land the data in a different location. And for IMS, Again, you can use virtualization layers, for instance, DVM or IBM Z Digital Integration Hub, ZDIH. And you can access directly into IMS and Db2 in a transactional fashion or LTP. And again, the zData gravity fully satisfied. And if the application on the right-hand side is an analytic application, okay, not a transactional application, so in that case you can attach IDAA into Db2 and suddenly you have a hybrid engine, okay. Db2 is your transactional engine, IDAA is your analytical engine, But you still have a single copy of the data and you can still directly access for both OLTP and OLAP purposes. And if you have IMS into this picture, you can load IMS data into IDAA directly. And again, you can access IMS data with analytical patterns via SQL, for instance. And if your use case is AI, on the right-hand side, you can use SQLDI Pro, okay, and which is attached to Db2.
[00:20:29] – Mehmet Göksu
And now you can run AI, basically semantic queries against the Db2. And again, you can satisfy AI requirements without moving data off the platform. And if the use case is streaming, event publishing, event processing, and if the use case requires Kafka, again, you can use the latest Confluent platform. Now Confluent is an IBM company, and you can use DataGate for Confluent for Db2-related streaming use cases. And this middle box, this Confluent Platform, runs on LinuxONE or Linux on z/OS LPAR. And this is still considered as an IBM z ecosystem. And with that, you can still use 4 legacy use cases, OLTP, OLAP, AI, and streaming. All these works run on IBM Z or LinuxONE platform. Okay. So eventually, if those use cases requires IBM Z data, but if the enterprise doesn’t want to access mainframe, then the only thing that you can do is the replication. Okay, then you can start considering replication from IBM Z to a different data store. So this is where we are from the IBM Z data perspective. Next page, please. So to bring data or to bring application or processing closer to data, There are techniques, okay, such as data virtualization technique.
[00:22:32] – Mehmet Göksu
This is not a replication scenario. This is a virtualization layer running on Z, and with that layer, you do not access the data source directly. You just access a virtualization layer, and this layer accesses this data source. And the value here is data still stays on Z. But you can provide a single unified application layer to any application on the Z, off the Z, and access to any data source on IBM Z platform. So analytics on platform, again, like IDAA, we will go into a little bit detail. So you can run analytics basically all up. Patterns directly on the source without moving data off the platform. AI near data means actually you can run real-time AI with the source systems such as SQLDI and SQLDI Pro. You can run training, the Db2 data, and then you can run semantic queries against this data. Okay, so all in all, these 3 patterns, okay, can be executed within the same platform without moving data off the platform. Next page. So with that, now we will go into details with Vishal. So we will go through 4 different products and solutions. So we will go through data virtualization as a virtualization layer.
[00:24:19] – Mehmet Göksu
We will go into details of Accelerator, which is the analytics platform running on Z. We will go to machine learning, MLZ, AI near real data. And last but not the least, SQL Data Insights Pro, which is an embedded AI in your database engine. OK, so let’s proceed. Yeah, please go ahead, Vishal.
[00:24:47] – Vishal Gupta
Sure. So as we have seen, there are different approaches to achieve data gravity when we talk about IBM Z. And one of the approach, or standard industry-wide approach, is data virtualization. And on IBM Z, with IBM Z, with IBM solution, we can achieve it via IBM Data Virtualization Manager for z/OS. And this is also known as DVM, so I refer it DVM from here onwards. So what DVM is? So DVM is a data service layer which gives us a way to access the data real time without making a copy of that data. And it also gives us a flexibility to integrate disparate data sources, right? Let’s understand it in a simpler way within scenario, right? So I wanted to access my z data. I wanted to access it from outside world, and outside world could be my x86 distributed servers or could be my hyperscalers and in cloud from where I wanted to access my z data. So instead of thinking of how to access that data, this virtualization layer gives us a flexibility to access that data in a simpler form.
[00:26:03] – Amanda Hendley
Right.
[00:26:04] – Vishal Gupta
So it also gives us an abstraction layer to shield our developers. And when we are talking about shielding our developers here, we are trying to emphasize on an industry standard way of accessing the data. So what we are trying to say is the developers who are looking for data from different data sources, they no need to worry about how to access the data from different data sources. They can deal with all different data sources with industry standard data access mechanisms such as SQL way of accessing the data, API way of accessing the data. And to give an example, if I wanted to access my IMS database in an application running on cloud, a developer no need to worry about how to access the IMS data. They can run a SQL query against DBM and DBM will take care of pulling the information from IMS and giving it back to the user. So that’s how we are saying it shields our developer to have specific skill sets for each data source. It has a metadata catalog, which is behind the scene, not related to users, but it’s good to know that it keeps a metadata catalog where it keeps track of all the virtual objects which we create in DVM and the corresponding original data source so that DVM knows when someone comes for a request to a DVM virtual object to which original data source request has to go.
[00:27:49] – Vishal Gupta
And if we talk about what all ways of interaction or transaction DVM supports, it’s not only limited to read-only transaction that I wanted to read the data from different data sources then DVM will help. DVM will definitely help, but beyond read, it also supports write back to the original data source. And to think of with an example, we can run an insert query, insert SQL query to insert data into VSAM file residing on C, right? So it support that kind of write back to original data sources. And if we think about a latest data solutions which are coming in the market. So from IBM, we have Watson xDotData, which is an open lakehouse solution. So DBM also gives us a way to integrate z/data with IBM Watson xDotData. So that kind of capability it also has. Now if we talk one level deeper to know how it works, So DVM has its own server which runs as a started task in its own address space on z/OS, and it has its own SQL engine, right? That’s what you can see in the middle of this chart. That’s a representation of DVM server which is running as a started task in its own address space in z/OS, right?
[00:29:15] – Vishal Gupta
And if you think about its key components, a metadata repository, It is the same one which I was talking in the previous slide, that metadata catalog where it keeps a track of all the original data sources and a virtual layer. It has a MapReduce parallel I/O capability to give us an efficient way of accessing the data to return the data way faster than someone is expecting. So it has an in-memory caching layer. So this caching layer is not for all the Scenarios. But if we think of a smaller file where we just have rule-based information, right, which doesn’t get dynamically changed, we can keep that kind of small files right into the DVM cache. And when consumer ask and look for that data, request doesn’t go to original data source again and again. And it’s being running on z/OS. It takes all the advantages of z/OS hardware and our z/OS operating system, and which makes it more optimized. On the right-hand side, we have a representation of multiple data sources which it supports. So if I talk about in the numbers, DVM supports more than 30+ different data sources which can be virtualized via DVM.
[00:30:31] – Vishal Gupta
And if I’ll name few, if I talk about the data sources on mainframe, it supports IMS, ADBAS, VSAM. From the operational side, it does support SMF files, log streams. So all these mainframe data sources can be virtualized via DVM. If we think of IBM non-mainframe data sources, Db2 Warehouse, the services running on IBM Cloud Pak for Data, so these can also be virtualized via DVM. And it’s not only limited to IBM data sources, it’s beyond that. The third-party data sources are also supported, which can be virtualized via DVM, such as MySQL, Oracle, Postgres. So all these data sources can also be virtualized via DVM. And on the left-hand side, we have a representation of consuming applications or consumers who would be interested to access that data via DVM. And we have a representation of how they can access data. On the top, you can see we have a representation of SQL way of accessing the data via DVM. So JDBC, ODBC drivers supports this SQL way of interacting the data. We can also expose our virtual layer, our virtual objects which we have created in DVM with the help of z Cyber Vault. Z/OS Connect in an API form.
[00:31:47] – Vishal Gupta
So consumer can access those endpoints of APIs to access the data. Our standard SOAP-based data services can be implemented on virtual objects. And if you think of Z/OS native applications who would also be interested to access the data, thinking of an example, my COBOL program wanted to access the Oracle data. Oracle data can be virtualized via DVM and then My COBOL application running on mainframe can access that Oracle data via DVM. So these are the different way of accessing the data via DVM. So here I will touch base on one example and I will try to walk you through how this process works, right? So here I’m taking an example of my IMS database, which I wanted to virtualize. And I wanted to give a relational way of accessing my IMS database to my users. So what that needs is, so even before talking about it, let me talk about 2 virtual objects which DVM supports. So DVM has one of the virtual object called virtual table, which is one-to-one relation to the data source, and it also supports virtual view. As an example, if we have 2 different data sources virtualized with the help of virtual tables in DVM, then I can create a virtual view on top of those 2 virtual tables.
[00:33:19] – Vishal Gupta
If you visualize in a way that one VSAM file has been virtualized into virtual table 1, one of my IMS segment is virtualized in virtual table 2, and my user is interested to fetch 2 columns from a VSAM file and 2 columns from IMS segment. They no need to run 2 different select queries against DVM. They can come to view which has been created on top of those 2 virtual tables and they can access it. So now if we think of an IMS database and we wanted to virtualize it in DVM, so DVM comes up with its UI, and it also comes with a batch process to virtualize a data source. And when we think of a virtualization process of any data source, we need to supply some input information. So in the case of IMS database, that data virtualization process of creating a virtual table, what it needs is it wanted to know the path where the DBD resides for my IMS database. We need to supply the PSB information of my IMS database. We have to supply the COBOL copybook layout of my IMS segment, which I wanted to virtualize. So in that process, when I supply this information, it creates a virtual table for us on top of an IMS segment.
[00:34:43] – Vishal Gupta
And once we have that IMS segment virtualized in a virtual table, then my user can run a select query against that virtual table to pull the information. And another question arises in the case of IMS databases. We do have a concept of root segment and child segment. How do we maintain that relation? So when we create virtual tables against multiple segments of an IMS database, DVM introduces a parent-child relationship between the tables so that relationship between the tables which we have between the segments stays intact, right? So in this way, we can transform our IMS database into a relational form via DVM, and then user can access it. They can select the data from any IMS segment which they wanted to. They can access the data from multiple segments in a one go, in a one SQL query. They no need to worry about what should be the SSA they have to use and how to access that IMS database. Everything will be taken care by DVM. behind the scenes. So that’s how we can virtualize IMS database via DBM for users to access. So next we will talk about 2 other solutions in the same space of data gravity.
[00:36:00] – Vishal Gupta
And Zunaid, I would like to hand it over to you to cover these.
[00:36:05] – Mehmet Göksu
Thank you, Vishal. So the next product and solution is IBM Db2 Analytics Accelerator. So Db2, in short term, it is called Accelerator. And it is actually a 10-plus-year-old solution. And it is actually a logical extension of Db2 z/Engine to make Db2 as a hybrid engine for OLAP analytic workloads. So it is actually designed for running complex analytic queries much faster than Db2 z/OS. And it addresses mostly reporting, business-critical applications, depends on complex queries. And with that, Db2 and IDEA provides a transactional and analytic engine, okay, a database system, okay, without moving data off the platform. So with that picture, actually, the transactions, okay, transactional queries, short-running queries running in Db2, and analytic queries are offloaded to accelerator. So in the next page, I’m going to show how that works. So on the left-hand side, there is a user, and this user actually represents an application, and this application could be anywhere. So it could be a CICS IMS transaction manager application, it could be batch, it could be any application accessed to Db2 via DDF. And this application sends a query, and Db2 Optimizer checks the access path of this query. If it’s a short-running, well-optimized query, For Db2 engine, it runs in Db2 engine and results sent back to application.
[00:38:07] – Mehmet Göksu
So it’s business as usual. So if there’s an attached accelerator, so the optimizer checks the query. If it’s a complex query, this query is offloaded to accelerator. It is executed outside of Db2 in the accelerator and results sent back to the user. So when you look at this picture, you need to visualize 2 things. One, this is fully transparent to the application. So basically, the application sends just a query, and then optimizer actually selects the best access path for this query, either in Db2 or in the accelerator. So it is transparent. Second, the data sitting in the accelerator is continuously updated and replicated via Integrated Synchronization Protocol. So the data actually also resides in the accelerator in a columnar format. Okay, so this is how it works. So let’s look at— yeah, you can switch— and let’s look at from the IMS perspective, right? So if you want to use IMS data with the accelerator, you have 2 options. There’s a separate product called IDA Loader, Accelerator Loader, and you can load IMS data to the IDA like a relational table. And remember the representation from Vishal in the DVM case, so this virtual table context.
[00:39:45] – Mehmet Göksu
So IDA Loader also virtualizes IMS segments and use this virtualization layer, and then it loads these segments into a relational table structure in the accelerator. So the advantage is you can use this option loading for both Db2 and accelerator. So your Db2 data, your IMS data can also be loaded into Db2. With HA load capability, you can load multiple accelerators at the same time. But with this approach, actually, it’s a snapshot of your IMS data, right? So if you load this every hour, every day, so the data currency is just related with this mini-batch cycle. On the right-hand side, you can use IBM Change Data Capture, which is called Classic CDC. And you can also use near real-time replication from IMS into the accelerator. So the advantage here is now you have almost a continuous replication in the accelerator for real-time analytics. Next page, please. Yeah, so in a nutshell, the loader is a separate product outside of the accelerator. And it is used to load external data, z or non-z data, into the accelerator. And IMS is just one of them. So it’s running on IBM Z platform. It reads IMS data, again, segments.
[00:41:33] – Mehmet Göksu
And these segments are virtualized as a relational table. And then basically the loader reads from IMS and writes this data into the accelerator. Okay, next page. So this is interesting use case. If you need real-time analytics with your IMS data with IDA, then loader may not satisfy you, okay, because eventually you need to load periodically. as a snapshot. With that, on the right-hand side, this is your Db2 system and your accelerator. So Db2 and accelerator are paired on the right-hand side, and the data is flowing from Db2 and accelerator. So this is a typical Db2 and IDAA use case. On the left-hand side, you have IMS, okay? And classic CDC reads from IMS, and replicates data, okay, to the accelerator. So with that picture, okay, on the right-hand side, the green line is integrated synchronization data coming from Db2 to accelerator. And on the left-hand side, the blue line is IMS data is flowing from IMS into the accelerator directly. So this picture tells you, okay, your Transactional system, okay, could be either one of them in Db2 and IMS. So different applications, different use cases, but you can create an enterprise data warehouse, okay, or enterprise analytics platform within the accelerator, okay, and your transactional systems still running in their best environment, Db2 and/or IMS.
[00:43:25] – Mehmet Göksu
But your enterprise analytic use cases runs in accelerator. And remember, the entry point of the accelerator is always Db2, which is SQL. Okay, so that means suddenly you have an SQL-based analytic platform with every data source. Okay, and please remember, this is a continuous replication in either case. Next page, please. And this was an analytics story. So now if we switch to AI story, okay, and AI within the data engine, now we start talking about the DI Pro. So SQL Data Insights Pro actually is a continuation of the SQL Data Insights, okay, which is part of the Db2 13. And with Db2 13, we brought AI into the database engine. And what does that mean? Okay, so Db2 data, okay, is trained, okay? And all these AI capabilities, okay, training semantic queries, okay, running directly in the source. Okay, this is from 10,000 feet SQL Data Insights provided to us. So this is all unsupervised training. Okay, and it is not large language model. It is, it is large database model, which is maybe the first one in the market. Okay, and why? Because the modeling Okay, the training, okay, happens in the source, happens on IBM Z, and happens in the Db2 z/OS ecosystem.
[00:45:19] – Mehmet Göksu
So suddenly you integrate your applications with Db2 with semantic queries. They are basically regular SQL-based queries. Okay, so you can run analytic queries with SQL. You can run transactional queries against Db2. Now you can run AI queries, semantic queries against Db2 data. And what is the IMS story here? So next page, please. Oh, before that, so what is the business value here? Again, in a short term, Remember the z/data gravity story? So the insights from the transactional data, okay, can be received, okay, without moving data off the platform. Please remember, if you replicate data, it is a cost. It’s a product cost. It’s a CPU cost of your z system. It brings additional complexity. Sometimes it’s a single point of failure because if the replication breaks, your lower environment can be impacted. So more relevant insights means without moving your transactional data off the platform for AI, you get insights directly from the source. And since this is single truth, remember, this is your single source of truth, which is Db2 and/or IMS, the decisions will be faster. Low cost means, remember, no data movement. There is no data duplication. There is no other system of records or system of insights outside of the z.
[00:47:13] – Mehmet Göksu
That’s why it is lower cost. Reduced risk. This is very important. If you replicate data, You expose your valuable z data outside of the platform. And all the access to the z data actually done through the RACF or other non-IBM security packages. So you lose this security, and it is open to exposure. So next page, please. So AI, again, with Queries, the right-hand side, there are 4 semantic queries. Basically, this is provided by SQLDI and SQLDI Pro. With that, now with regular queries, okay, you can use queries for data discovery, but with semantic queries, now you generate insight. Again, let’s be very specific. With a regular SQL semantic, actually, you can ask certain strict questions to the database, okay, with regular predicates such as WHERE clause, BETWEEN, greater or equal than, etc., etc. But you can’t ask fuzzy questions to the data. Basically, fuzzy questions means bring me the similar data, okay, bring me the data which is similar to this pattern. For instance, this is a fuzzy question. So this can’t be done with regular SQL semantics. That’s why AI, basically similarity queries or semantic queries, provide this capability. So similarity, semantic clusters, commonality, analogy— these are 4 different purposes, but essentially they are AI capabilities and executed through SQL.
[00:49:12] – Mehmet Göksu
Next, please. Yeah, so for the VSAM and IMS data, SQL DI Pro runs in the same infrastructure with IDAA, which is a Linux run box. So imagine this box is hosting your analytic plus pro capabilities. So with that, as Vishal mentioned, if you use DVM, Data Virtualization Manager, so you can access IMS data with DVM. And you can use and access IMS data for model training. Because remember, your primary data source or one of the primary data sources could be in IMS. And if you want to run, or if you want to train your models with the IMS data, you can use DVM. You can connect to DVM, and DVM pulls the IMS data into SQL DI Pro infrastructure, and you can train this data. So with this composite use case, okay, remember DI Pro DVM and Db2 and IMS all together. Now you are bringing any data source, including IMS, into the SQL DI Pro infrastructure for model training and semantic capabilities with Db2. So this is a very powerful use case, building AI on top of Db2 data with IMS. Yeah, machine learning. Yeah, Vishal, please continue. Sure.
[00:51:03] – Vishal Gupta
Thanks, Vineet. So all, like, as we have— if I’ll give a quick recap, right? So we have talked about how we can access the data with the help of virtualized layer. We have talked about how to have analytics application running right onto the platform on our mainframe. We have also talked about how we can have our data smart enough to answer those all fuzzy questions, right? So it means we are bringing AI right embedded into my data. But we have not talked about how to bring AI applications right onto the platform on my IBM mainframe, right, on my IBM Z. So when I’m talking about AI application is if I’ll say it in an more simpler understandable words, right? So I am referring to how I can deploy my AI models— could be machine learning model, deep learning model— right onto my IBM mainframe and running them closer to my transactions, running them closer to my data, right? Can we do that? So yes, that is doable, and Machine Learning for IBM z/OS helps us in this scenario where we can have our machine learning, deep learning models right available next to my transaction and data.
[00:52:30] – Vishal Gupta
And if I’ll talk about core benefits and features of machine learning for IBM z/OS, which it provides, right? The first thing is, right, if I’ll know on a high level, right, it is a full-featured machine learning platform for z/OS. That helps us bringing in AI right into my mission-critical applications which are running in my transactional systems, which are running in my batch application on mainframe. And for that, Machine Learning for z/OS gives us an end-to-end platform. It’s not just limited to one part of that SDLC of AI model lifecycle management. It’s complete, starting from developing to training to deploy. It helps us, or it gives us an end-to-end platform for an AI model lifecycle. And if we think of models which has been developed outside my mainframe, outside my Z platform, it gives us a way to import those models which has been developed by my data scientist outside Z, could be on in cloud or on-premise environment. It gives us that way to import those models and deploy them right onto the mainframe itself. And the next key benefit is scalability, right? Now with that latest hardware, right, the Z17, which has come up with Telem2 and a provision to have Inspire accelerator, Z16 with Telem1, right?
[00:54:05] – Vishal Gupta
We should have a way to take advantage of those accelerated chips on those AI accelerators which we have in our hardware. So machine learning for Z helps us leveraging those hardware accelerators when we deploy models on Z. Last but not the least, because these are the key features, trustworthy features it also provides us, right? It gives us an explainability. As an example, why my model has given me this result, right? It gives me a way to detect the drifts upfront instead of just using the model and getting stale results after some time. And it tells us now our model needs retraining because it has started drifting, right? So it gives us that capability with regards to the trustworthy. Now if I talk about from a different angle, the same process, right? We know for an AI model inferencing, what we need in step 1 is we need to have training data which we wanted to use to train our AI models. So once we have that data handy with us, then what our data scientist does is they create a machine learning, deep learning model for my applications based on their expertise, which they are comfortable with.
[00:55:36] – Vishal Gupta
Plus, they try to use the tech stack which is more suitable to their use case scenario. So what we say as a part of this solution is we no need to worry about which tech stack is used to create that model, right? You build your own model and that too anywhere. So once you have your model ready, we in MLZ gives a way to transfer or to convert that model into PMML or ONNX industry standard forms. It gives us a way to convert my model into ONNX and to take an advantage of my Z hardware. It uses the Z Deep Learning Compiler for ONNX model to get that efficiency. And once I have my model available in my ONNX farm, Machine Learning for z/OS helps us deploying that model right onto my mainframe, closer to my data and my transactions. And once I have my model deployed closer to my application, it gives me an endpoint. So Machine Learning for z/OS gives me an endpoint for my data along with an input and output copybook which that model is expecting. And then with that URL, with that endpoint, I can put that endpoint, I can call that endpoint in my COBOL applications, on my mainframe applications, and send the input to my model running just next to my application instead of depending on a model which is running in some other platform, which is dependent on a streaming platform, as an example.
[00:57:15] – Vishal Gupta
I have my running application. In that application, I need some inferencing result from one of the model. One way of doing it is my mainframe transaction is putting some information onto a streaming application. And on the cloud side, one of my application is reading that information, passing it into an AI model, getting an inferencing result, sending it back. It comes into my mainframe program, and then I take a decision based on that. So instead of depending on that long pipeline, when we have a way to deploy the model right next to my application, I can use that and I can get rid of all that long pipeline which introduces latency because my transactions which are running on mainframe cannot deal with this kind of latency. So with Machine Learning for Z, we can deploy our processing Right next to my data and my transactions on Z itself. So in a one-liner, if I’ll say, train your model anywhere, build your model anywhere, but execute that where your transaction happens. And for that, Machine Learning for Z helps us. And with that, it closes our core solution offerings which we wanted to cover, which supports data gravity.
[00:58:39] – Vishal Gupta
So Amanda, back to you. And if there are any questions, you can read out for us.
[00:58:44] – Amanda Hendley
Yes, we just got several in. So Sri has said, how does a DVM recognize the indexes, AIXs, if we are providing only copybooks? Do we also share the info of VSAMs?
[00:59:16] – Vishal Gupta
Okay, uh, Sunit, would you like to take this question?
[00:59:25] – Mehmet Göksu
Well, I think this is referring to VSAM.
[00:59:38] – Amanda Hendley
Yeah.
[00:59:39] – Mehmet Göksu
So the indexes, again, we need to double-check. Okay, I don’t want to speculate, but if the indexes doesn’t taking care about the virtual table, so the additional indexes could be created with your virtual table definition. But let’s take it as homework and we will get back to you.
[00:59:58] – Vishal Gupta
Same, I was also thinking. Thanks for pitching it.
[00:00:02] – Mehmet Göksu
Well, the second one, what is the accuracy percentage for IBM SQLDI? Yeah, again, it’s a probability. Again, yeah, basically it generates a percentage of the accuracy, but it’s related with the data size. So the more data you provide, the better the percentage will be. The constraints, the third point, are you referring to, for instance, the constraints defined in Db2 table? I’m not sure. But in general, so the data accessed for the training is the full table. Okay. And if Again, I assume— I’m just making an assumption. If you are referring to foreign key kind of constraints in the table definition, this is not actually considered because the table is fully accessed. But during the training, you can use predicates. For instance, you can use the WHERE predicate. So just subset of the data is considered for the training. So there’s another question from Yan. Yeah, you’re welcome. What are the main advantages of using DVM versus a direct call to Db2 or IMS? Yeah, good question. So you can still directly access individually for both data sources, but with DVM, basically you can write a single SQL statement, okay, And with that single SQL statement, you can join a Db2 table and IMS segment in a virtualized way.
[00:01:52] – Mehmet Göksu
So that’s the advantage. And application actually see both data sources like a relational table. And DVM can also be used to expose PostgreSQL data.
[00:02:10] – Vishal Gupta
Yes, please. So Postgres is supported. I just double-checked.
[00:02:13] – Mehmet Göksu
Thank you. Thank you so much, Vishal. So yeah. So PostgreSQL is supported. Yeah. Thank you.
[00:02:23] – Amanda Hendley
Okay, great. Well, before we go, I don’t have another meeting date to share with you, but I did want to let you know that we will have our Recording of this session that we will share on the site, and you’ll get an email about it. And I’m hoping that we will be able to get the slides from today. Is that possible, Michelle? And—
[00:02:51] – Vishal Gupta
Yes, I will share with you.
[00:02:52] – Amanda Hendley
Oh, thank you so much. So we’ll get that out. So you’ll be able to reference back to it, and I’m sure people could reach out if they have any additional technical questions that they need help with. As I said, I’m, I’m shooting for December for our next session. But of course, December is a month, despite 31 days, that feels like it’s about a week long. So it might be the new year before we are able to meet again. So with that, I want to thank you both for presenting today. It was such a great session. And I want to thank everyone for joining us, and we’ll be in touch really soon. Thank you. Thank you.