Dev48
Language
  • About
  • Services
  • Industries
  • Technologies
  • Articles
  • Contacts
Book a call
    Home/Articles/The critical role of a streaming first data architecture
Dev48

© 2026 · All rights reserved.

The Critical Role of a Streaming First Data Architecture

Источник: Striim

The Critical Role of a Streaming First Data Architecture

Source: Striim

Learn about the critical role of a streaming first data architecture and see a demo of Striim's streaming integration and streaming analytics platform.

September 25, 2026

Steve Wilkes, Striim Co-founder and CTO, discusses the need for a “streaming first” data architecture and walks through a demo of Striim’s enterprise-grade streaming integration and streaming analytics platform.

To learn more about the Striim platform, go here.

Unedited Transcript:

Today we’re going to talk about what is a streaming first architecture and why is it important to your data modernization efforts. We’ll talk about the Striim platform and we’ll give you some examples of what customers are doing on the solutions they are building using our platform, but take some time to give you a demonstration of how it all works and then open things up at the end for questions and answers.s

So most of you are probably aware of this, the wealth news fast and it’s not just the world is the data that moves fast. Data is being generated largely by machines now. And so businesses need to run at machine speed. You need to be able to understand what’s happening right now and react with immediately and also that it’s not too late. Yeah. At the same time, customers and employees expectations are rapidly rising and positive reason for this is the smartphone boom and people’s access to realtime information, the ability to see what’s happening to your friends, what’s happening in the world. Uh, instant access to news and some access to, uh, messages, communication, um, because the consumer world is instant. Those consumers are also employees and executives and the, they expect instant responses and insight into what’s happening within the enterprise. And the quality of applications used to deal with is also driving desire for similar quality business applications.

The other side of the coin is that businesses needs to compete. Um, technology has always been a source of that competition. And as technology is rapidly changing and we’re getting more and more data. Businesses that are more and more data-driven have a competitive edge and this is now have almost all departments ranging from engineering or manufacturing all the way through to marketing being incredibly. And so the survival of most businesses depends on the ability to innovate and utilize new technologies and data itself is also massively increasing. Almost everything can generate data, it could be trucks or your refrigerator. Um, TV is a wearable devices that you have, health care devices that oh becoming more and more portable. Even things like Istomin making its way from being caught to a restaurant. It’s tracked and produces large amounts of data and this data is growing exponentially.

IDC did this study like a couple of months ago and they’re estimating that today the 16 Zettabytes of data by 2025 is going to increase 10 fold and all of that data around 5% of it now is real time increasing to 25% of it in 2025 and by real time they mean is produced and needs to be processed and analyzed in real time. So that’s 40 zettabytes 21 zeros of data that will need to be processed in real time. And of that 95% it will be generated by devices. The kick that really is the only a small percentage of the state it can ever be stored there physically, not enough hard drives being produced to store all this data. So if you can’t store it, what can you do with it? Well, the only logical conclusion is that you need to process [inaudible] analyze this data in memory in a streaming fashion close to where the data’s generated. It may be you’re turning the raw data, you know, thousand data points a second into aggregated data that is less frequent but still contains the same information content. Yeah. So that kind of thing is what people talk about is age processing was really trained to handle these huge volumes of data that people see coming down the line.

And it’s not just IoT data that the rise in streams, you know, okay, everything, every piece of data is generated because something happened, some kind of event, someone was working on the enterprise application, someone was doing stuff on a website or using a web application and machines were generating logs based on what they were doing. Applications generation, those databases as generating logs and network devices, everything generating logs. But they’re all based on what’s happening based on events. So if the data is created based on events in a streaming fashion, then it needs to be processed and analyzed in a streaming fashion. If you collecting things in batches, then there’s no way you’ll ever get to a real time architecture and a real time insights into what’s happening. But if you collect things as streams, then you can do these other things. You can do batch processing on the streaming data, you can deliver it somewhere else, but at least the data needs to be streaming.

So your stream pro so thing as he merged as a major infrastructure requirements and is helping drive enterprise modernization. So moving to a streaming fest, data architecture readings that you’re transitioning at these data collection to real time collection. You’re not doing batch collection of data. You’re doing real time collection of data, whether it’s from devices, files, databases, uh, wherever it’s originating and you’re doing this increments that you’re not trying to boil the ocean and replace everything in one go. You’re doing it. Use case by use cases, proud of your data modernization projects. And this means that things that have high priority to become real time to give you real time insights and a potential business competitive edge or better support for your customers or reduce the uh, amount of money you’re spending on manufacturing and by improving product quality, any of these things. Um, can we drive as today to modernization and your doing it use case by use case, you’ve placing pieces of it, bridging the old a new worlds of data.

Now some of the things that our customers are telling us, uh, and we have these legacy systems and by legacy that can mean anything that was installed, you know, over a year ago and they can’t keep up or they don’t predict that they’ll be able to keep up with the large amounts of volumes of data that they’re expecting to see with the requirements for low latency, kind of real time insights into data and with the ability to rapidly innovate and rapidly produce new types of analysis, new types of processing give you new types of insights into what’s happening in your business. Okay. They’re also telling us that we can’t just rip and replace these systems and the need to have the new systems and the old systems work together, uh, with potentially fail over from one to the other. While you’re doing this replacement. It’s a Striim has been around for around five years now.

Uh, we are the providers of a platform, the Striim platform that does streaming integration on analytics. The platform is mature. It’s been in production with customers for more than three years now. There’s customers all in a range of industries from financial services, Telco, healthcare, retail. We’re seeing a lot of activity in Iot. Striim is a complete end to end platform that does streaming integration on, on the mystics across the enterprise, cloud and IoT. We have an architecture that’s very flexible that allows you to deploy data flows to bridge enterprise, cloud and Iot. You can deploy pieces of an application at the edge close to where the data is generated. And that doesn’t have to just be IoT data. It can be any data that’s generated but close to that data. Other pieces that are running on premise, doing some processing and other pieces are in the cloud.

Also doing processing our analytics. It’s very flexible how you can actually deploy applications using our platform because applications consists of continuous realtime data collection. And that can be from things that you may think of as kind of real time, uh, sensors. Sending events, a message cues, et cetera. Are there things like files that you may think are AAP is batch processing? Um, we could read at the end of the file and as new records are written to the file stream that has immediately turning files, the rollover, et Cetera, into a source of streaming data. Suddenly with databases, um, most people think of databases as a historical record what’s happened in the past. But by using a technology called change data capture, you can see the inserts, updates, deletes, everything that’s happening in that database in real time. So you can collect that nonintrusive Lee from the database on stream.

That’s it. So now you have a stream of all the changes happening in the database. Okay. So all of the applications built with that platform use some form of continuous data collection. On top of that, you can then do real time stream processing and this is through a SQL based queries. There’s no programming involved in the Java, no c sharp, no Java script. You can build everything using SQL and this allows you to do filtering transformation aggregation of the data. Yeah. By utilizing data windows. So you can say what’s happened in the last minute, uh, you can look for a change in data and only send that out and et cetera. And then enrichment of data, which is also very important. And that is the ability to load large amounts of reference data into memory across the distributed cluster and join that in real time with streaming data to add additional context to it.

So an example would be if you have device data coming in and it’s device x, Y, z value one, two, three. Okay, that doesn’t mean much to things downstream that might be trying to analyze it. But if you join that with some context and you say, well, device Xyz is this sensor on this particular motor, on this particular machine, now you have more context. And if you include that data, you can do much better on top of the stream processing. You can actually do streaming analytics that can be correlating data together, joining data from multiple different data streams and looking for things that match in some way. So maybe using a web blogs and network logs and you’re trying to join by IP address. Um, and you’re looking for things that have happened on either side in the last 30 seconds. That kind of correlation to complex event, a processing which is looking for seek because of events over time, the mass, some kind of pattern.

So if this happens, followed by this, followed by this and it’s important you can do a statistical analysis on anomaly and integrate with third party machine learning. Yeah. We can also generate alerts and trigger external systems and build these really rich streaming dashboards or later visualized results of your analytics. Yeah. And any of the data that’s initially collected, the results of processing, the results of analytics that can all be delivered somewhere and you can deliver to lots of different targets in a single application. So you can push stuff to enterprise and cloud databases, files or do, uh, Kafka, et cetera. Okay. As a new breed of middleware that supports streaming integration analytics, it’s very important that we integrate with your existing software choices. So we have lots of data collectors and data delivery that work with systems you may already have. It wasn’t the big data systems, enterprise databases, open source, um, pieces we can integrate with it and do all of this in a enterprise create fashion that is inherently clustered, distributed, scalable, reliable and secure as a general purpose piece of middleware.

We support lots of different types of use cases from real-time data integration, uh, analytics and being able to build dashboards and monitor things. And these use cases across all different industries and they can range from a building your data lake and preparing the data before you land it. I’m doing tag migrations, I’m doing iot edge processing. And then on the other lytics and patterns side, there’s things like fraud detection, uh, predictive maintenance, uh, anti money laundering is some of the things that we received from customers, uh, around that. And then if you want to build these dashboards and monitor things in real time and look to see whether things are, for example, meeting SLAs or meeting expectations and yeah, we’ve done things like call center quality monitoring, SLA monitoring. Yeah. Looking at the, that worked from a customer perspective and being able to alert when things aren’t running normally. And I use cases across many different industries. There’s a lot of texts on here. Yeah. The takeaway is that we have used cases in a lot of different industries.

Yeah. One of the examples is, uh, using Striim for hybrid cloud integration. And that’s really where you have a database on premise and you want to move it or copy it to the cloud. And it’s one thing to just like take the database and put it in the cloud, but that will miss anything that’s happening while you’re doing it or miss things have happened since you’ve done it. So it’s really important that you include a change data capture in this to continually feed your hybrid cloud database with new information. And so by using a set of wizards, you can build this really quickly that allows you to join a on premise, for example, oracle database and deliver real time data from that into, uh, for example, Azure SQL DB. So you now have an exact copy that is always up to date of the on premise database.

Another totally different example is uh, using us for security monitoring, which is where you have lots of different logs being produced by VPNs firewalls, network routers, individual machines, essentially, uh, microcontrollers. Anything that can produce a log and you recognize a unusual behavior, um, is most often seen by affecting multiple systems, no security unless they get a lot of alerts from all these logs and all these systems all the time. But a lotamz of those are false positives. So the goal of this was to identify things that were really high priority for them to look up first by seeing what’s the activity happening that was affecting multiple things. So for example, if you have a port scan from a network row to the same, this guy’s looking at other stuff. Is there any activity on the other machines that he’s looked at? Okay. Are they doing port scans?

Are they connecting to external sites and downloading malware? So by doing this correlation in memory in real time, you can spot threats at a higher priority. And also by pre correlating all the data together and providing that to the analysts, they can see immediately the information they need rather than having to go and manually look for this across a whole bunch of different bugs. And this really increases the ominous productivity. So a couple of other examples from our customers. One is a very simple, uh, realtime data movements where data from, uh, HP nonstop and SQL server databases is being pushed out into, uh, multiple targets, uh, whether it’s Hadoop, HDFS, Kafka, HBase, and they’re using a as a analytics hub for their communities. So basically ensuring that wherever they want to put the data, that’s always up to date and that is always containing real time information from there and they can see on the other databases.

And then the glucose monitoring company, uh, are using us to see, uh, events coming in from, uh, devices on these implantable devices, uh, real time monitoring of glucose. And it’s really important that these things work. So they are looking at the, whether the device is having any errors, whether it’s suddenly going offline and being able to see in real time any of these devices not working properly. And this is really important to their, their patients because their patients rely on these devices to check their glucose glucose levels. So this has really reduced the, uh, times detect that there’s an issue going on and has improved patient safety massively. Okay. We have recognized generally by a lot of the analysts in both the in memory computing and the streaming analytic landscapes. And we’re also getting a lot of recognition from various publications and [inaudible] a trade show organizers and then also very importantly, one that best places to work, uh, which is really vindication of, you know, us being a, a really great company, a key differentiation.

Striim’s end to end platform does everything from collection and processing, Oh, lytics delivery visualization, the streaming data that is easy to use, uh, with the SQL language for building a processing and analysis that allows you to build and deploy applications in days. Um, and we’re enterprise grade, which means that we are inherently a scalable in a distributed architecture, reliable and secure. Okay. And that you can integrate us. We’re easy to integrate with, uh, your existing technology choices. So those are the kind of key things to remember about why we’re different. So with that we’re going to go into a demonstration. Sothe first part of this, um, basically going to show you how to do the integration rather than to type a lot of things. Uh, we’re just going to go through, uh, how to build a change data capture a into Kafka and do some processing on that and then do some delivery into other things.

So this is pure integration play. You start off by doing a change data capture from SQL, in this case, my SQL and okay, build the initial application and then configure how you get data from the source so we can figure the information to connect into my sequel. When you do this, we’ll check and make sure everything is going to work, that you already have. Change data, capture, configure properly. And if it wasn’t with how you had to fix it and how to do it, you don’t select the tables that you’re interested in. We’ve got to collect the change data from, and this is going to create a data stream, that data stream. Yeah. But then go to two different to Kafka. So we’re going to configure how we want to write into Kafka. Um, and that’s basically setting up what the broker configuration is, what the topic is and how we want to format the data.

In this case we’ve got the right to add as JSON, when we save this, this is going to create a data flow and the data flow is very simple. In this case it’s two components. We’re going from my SQL CDC source into a Kafka writer. We can test this by deploying the application and it’s a two stage process. You deploy first, um, which we’ll put all the components out over the cluster and then you run it and now we can see the data that’s flowing in between. So if I click on this, I can actually see the real time data. And you see there’s a data and there’s it before. That’s basically the four updates. You get the before image as well, so you can see what’s actually changed. So is real time data flooding through [inaudible], um, um, my sequel application. Okay. But it doesn’t usually end there.

Uh, the raw data may not be that useful. And one of the pieces of data in here is um, a product id. Uh, and that probably is, it doesn’t contain enough information. So what we’re going to do first is we’re gonna extract the various fields from, from this and those various fields include the location id, product Id, how much stock there is, et cetera. This is a inventory monitoring table and we’ve just turned that from kind of a rural, a format into a set of name fields. So I’ll make it easier to work with later on. You can see the structure is very different. Now what we’re actually seeing in that data stream. If we then, uh, once add additional context to this, what we’ll be able to do is join that day. There was something else. So, first of all, we’ll just configure this so that instead of writing the raw data at Cafca, we’ll write that process data ad and you can see all we have to do is change the input stream. So that will change the data flow. Now to right that uh, process data at into Kafka.

But now we’re going to add a cache and this is a distributed in memory data grid that’s going to contain additional information that we want to join with a raw data. And so this is product information. So every product ID is a description and price and some other stuff. So first of all we’ll just create a, a data type that corresponds to our database table. Yeah. And configure what the key is. And the key in this case is the product Id. Then we specify how we are going to get the data. And it could be from files, it could be from acfs. Yeah. We’re going to use a database reader to load it from my SQL table. So especially specify all the connections and the query we’re going to use. And we now have a cash of products information. So use this, we modify as sequel to just join in the cache.

So anyone that’s ever written any secret before knows what a join looks like. We’re just joining, uh, on the product Id. So now instead of just the raw data, we now have these additional fields that we’re pulling in in real time from the product information. So if we start this and look at the data again, you’ll actually be able to see the additional fields like description, um, and brand and category and price that came from that other type that’s all joined in memory. There’s no database lookups going on is actually really, really fast. So that’s where I seem to Kafka. If you already have data on Kafka or another message bus or anywhere else for that matter is new files. Um, you may want to kind of read it and push at some of the targets. So, well we’re going to do now is we’re going to take that data that we just wrote to Kafka.

We’re going to use a Kafka, a Rita in this case. So we’ll just search for that and tracking the capital sauce. And then we can figure that with the properties connected to the broker that we just used. So the uh, and because we noticed JSON data, we’re going gonna use it Jason Pasta. I was going to break it up into a adjacent object structure and then create this data stream. Okay. When we, uh, deploy this and uh, start this application, it’ll start reading from that Kafka a topic and we can look at that data and we can see, uh, this is the data that we were writing previously with all the information in it and it’s adjacent full Max. You can see the adjacent structure though. So the other targets that we go into right to, uh, the Jason Structure might not work. So what were you going to do now?

Is We got after in the query that’s going to pull, uh, the various fields edit that Jason’s structure and creates a well-defined, a data stream that has various, um, individual fields in it. So we’ll write a query to do that. That’s directly accessing the JSON dSata and save that. And now instead of original data stream that we had with the JSON in it, when we deploy this on, uh, start it up and look at the data. And this is incidentally how you would build applications, looking at the data all the time, um, as you’re building and adding additional components into it. Um, if we’re, uh, look at the data stream now, then you’d be able to see that, uh, we have those individual fields, which is what we had before on the other side of Costco, but doesn’t forget that, um, it may not be stream rights into Catholic. It could be anything else. And if it, you were doing something like we just did with CDC into Kafka than Kafka into additional targets, you don’t have to have Kafka between, um, you can just take the CDC and push it out to the targets directly.

So, uh, what are we gonna do now is going to add a simple target, which is going to write to a file. And, uh, we do this by choosing the file. So the file writer and especially finding the formats we want. So we are going to write this. I’ve seen the CSV format. Um, we actually call it DSV because it’s delimiter separated, right? Um, and the, the limits can be anything. It doesn’t have to be a coma and save that. And now we have something that’s going to rotate to the file. So if we deploy this and start this up, then we’ll be creating a file with the real time data.

← All articles