This video can't play due to privacy settings
To change your settings, select the "Cookie Preferences" link in the footer and opt in to "Advertising Cookies."
Can agentic AI patch the internet? ft. Mike McGrath
When a security vulnerability hits an open source application dependency, enterprise IT teams are often forced to choose between breaking application upgrades and running vulnerable code in production. Join Red Hat CTO Chris Wright and VP of Lightwell Engineering Mike McGrath as they reveal how Project Lightwell leverages agentic AI and expert human review to surgically patch vulnerabilities without disrupting runtime stability.
Transcription
Transcript
00:00 - Chris WrightHere's a problem every CTO and CSO grapples with. Your team writes clean, audited code, but under the surface, your application relies on thousands of open source packages your team didn't write. When a CVE hits one of those dependencies, you may be trapped between breaking your application with an unvalidated upgrade or running vulnerable code in production. So today we have Mike McGrath, Vice President of Lightwell Engineering here at Red Hat to talk about how we're working towards eliminating that choice entirely. Welcome to Technically Speaking, where we explore how open source is shaping the future of technology. I'm your host, Chris Wright. Mike, great to have you here today. I thought before we got totally into the details of Lightwell and patching and security vulnerabilities. What's your background? We've been working together for a while. Give us a taste of what you've been up to before this.
01:01 - Mike McGrath
Yeah. Going back to the beginning, I started in the '90s with the Slackware Bible that I probably got at Sam Goody or who knows what. And then I started volunteering in the Fedora project right around Fedora Core 3. And Red Hat hired me about 19 years ago. And I've had various jobs working on RHEL or Fedora or OpenShift. And prior to Lightwell, I was leading core platforms, which included RHEL and OpenStack and Satellite and a lot of those support communities. And I think by any measure, I've had a great and storied career at Red Hat. I've had a good time here.
01:37 - Chris Wright
Yeah. I'm now remembering downloading the entire alphabet of floppies off the internet for Slackware. Thank you for that. So the context here for Lightwell starts, I'd say back in the spring, we had what's often referred to as the "Mythos Moment." Anthropic announced Mythos and established a program around that called Glasswing and put a bunch of content out there describing a shocking set of vulnerabilities and capabilities of the models and harnesses had to find vulnerabilities in software. That was many months ago at this point. And of course, there's many programs out there. It's not just a single frontier lab like Anthropic. There's Daybreak with OpenAI and there's open source projects in this space like Akrites. Walk us through that initial phase of recognizing the reality of something like massive vulnerability scanning and trying to translate that into a response.
02:45 - Mike McGrath
I think my main focus on this back then was on RHEL and we had seen this slow but steady increase in vulnerabilities that were coming in from the community. And it was not even. Some communities were getting hit very hard. Some communities weren't. And as we looked at it, we realized a lot of these vulnerabilities, even pre-Mythos, were coming from AI. And so for us, originally it was less about trying to find vulnerabilities and it was more about how do we respond to this? Because RHEL is the enterprise Linux as far as we're concerned and people expect it to be secure. And so for us, it was kind of about managing that increase. And then earlier this year, we decided, well, let's just start taking a look at this ourselves and be more proactive about it. That was maybe a month before you started to hear sort of leaks about this Mythos thing that Anthropic was working on. Instead of just waiting for it to happen, we had some conversations about scanning proactively using our own harnesses. And at the time, I think Opus was the top dog for that kind of thing. We just started looking and it did not take us long to find unknown vulnerabilities. And I think the big aha moment for my team was that we did not have security researchers doing this. These were people that were skilled in AI. They were software developers, but they were not inherently security people. And in finding these, it was also a matter of where to scan and how to scan. And that kind of led into then Mythos being released. And I think that was, even for us, I had access to Mythos, trying to figure out what it did and how it did it. It took only a few weeks and then it's, now we're in a different world.
04:40 - Chris Wright
For the longest time, we've had from a Red Hat point of view, a lot of experience working with open source communities and doing security vulnerability responses and disclosures responsibly and focusing predominantly on that platform tier where we sit. Linux kernel being a prominent example of areas where we pay close attention to security issues. The RHEL world has an application runtime, a language Java, Python, et cetera. And we had already started bringing newer versions of this and maintaining patched versions of these language runtimes. But the Mythos scanning and the associated cybersecurity vulnerability discovery showed a whole different world up the application layer. So how have we shifted our focus from that very platform-centric view of the world to starting the application down and coming through those dependencies?
05:41 - Mike McGrath
Yeah, as we've started to look at this in actual enterprise environments where this code is actually running, it is not at all uncommon for more than half of the code that is being run to be in that sort of middle of the sandwich. They've got underlying enterprise support at the operating system layer, usually something like RHEL, and then they'll have some platform support higher than that. Maybe that's a Jboss or something like that. But more than half of the code is kind of running as dependencies in between those stacks or on those stacks. And for us, the real question was understanding and talking with customers that unless that full stack is secured somehow, they're always going to have problems and vulnerabilities. And in the past, we've shied away from going higher in the stack because it is an expensive proposition to do that. But this is one of those things that, while AI scanning kind of frightened us into looking at this space, the capabilities of those large language models have evolved enough that we can also seriously look at evaluating, fixing that space. And I think that's where all of the, this culmination of all of these conversations came together to create Lightwell.
06:55 - Chris Wright
It's an insane goal, really. If you go back in time, scaling with people to do all of this work is a global scale problem. Internally, we've kind of referred to it as patching the internet. That's all. No big deal.
07:09 - Mike McGrath
Just patch the internet.
07:10 - Chris Wright
And yet we need automation. We need tools. We need AI, LLMs to help us through this process of not just discovery, but remediation. I think a really interesting aspect here is we take an average enterprise, which is built of thousands of applications, each of which has thousands of dependencies. So you get this sort of order of magnitude thing with that sandwich that you're describing, on the order of hundreds of thousands of different small language specific modules. If you're trying to manage all of that as an enterprise and think about the change associated with all of that, I think that's by itself overwhelming just this year volume. In this Lightwell context, the versions that are being deployed into production are typically quite older than the versions that are being worked on actively in the upstream community. I think part of the critical focus of Lightwell is patching those versions that are in production while sort of this parallel life of bringing it into the upstream so we're making the overall internet and open source a safer place.
08:25 - Mike McGrath
I think looking at community releases where they're always focused on securing the most recent code is a fine goal for communities to work on. But the code that is actually running in these production environments can be much older. In fact, the oldest vulnerability that we've discovered that is in production in a customer's environment was a vulnerability found in 2001. And so at that point, you have to assume that whoever put that together has probably moved on to another job. It's not their full-time job looking at that. And that goes true for the hundreds or thousands of applications that these customers are running. And so a focus for us with Lightwell has really been on securing the content that they're actually running in their environment. And that's the last thing most customers want to do is completely upgrade some application that on its own is running just fine. But because the communities have moved forward and the technologies have moved forward, sometimes it can be weeks, months, maybe even years of work to fully upgrade that application. And as a business, you want to focus on the things that matter today. And so by being able to focus just on the very surgical issues related to security, it allows them to keep running application or upgrade it on their terms when they're ready to do it. This is that sort of peace of mind that Lightwell, I think, brings to the table.
09:55 - Chris Wright
The ripple effect of simply bumping the version number of a single dependency through all the transitive dependencies in your application could effectively touch the entire dependency tree and then totally destabilize the application. So I love that surgical view. It's a really great way to think about it. What I've seen is there's a back to basics here. Not all businesses have a robust awareness or understanding of what they have deployed where. So sort of like the table stakes is, what is your dependency tree? What are your applications? Where are they deployed? Because being public facing versus multiple firewalls into your data center has a very different risk profile. And then what are those application dependencies? As we produce something, you can make an intelligent choice about how quickly you respond in one context and where you might batch it in another context. The other piece is just good old fashioned CI/CD. How quickly can you automate an update to an application? So again, this back to basics, what do you have? Where is it? And how quickly could you have a push button system that can update an application and the production? All right, we got a pretty good view for the problem domain. Let's talk a little bit about under the hood, what is Lightwell ? We're building a lot of content. We're doing the scanning. We're doing the remediation. What does that look like inside the Lightwell factory, if you will?
11:33 - Mike McGrath
Yeah, I think most engineers that have worked with Claude or cursor can kind of work through in their mind how finding and fixing vulnerability to work on their machine. I think Lightwell has taken an additional step of building fully agentic harnesses to do this sort of headless, if you will. And so we are sort of taking that step from hunter gatherer society type work with security into fully industrialized mass production fixes and remediations. What that looks like is it starts with an assessment of what comes in. One of the big things I think I've enjoyed about Lightwell is instead of focused just on how bad a security vulnerability is, we also have an assessment of just how complex it is because some very serious vulnerabilities could be a simple one line change. And even some very minor vulnerabilities can require an entire application rewrite. Anybody from my team that's about to listen to this will roll their eyes. But there's a famous quote that I think sometimes attributed to Lincoln, sometimes Washington, which is, "If you give me five hours to chop down a tree, I'll spend the first four sharpening my axe." And I think that is very true in Lightwell where we spend quite a bit of time prepping for the remediation so that we can do it. And so that assessment at the early stages of our intake process kind of informs everything else we do with that vulnerability, how much human oversight it will need. And certainly the checking and rechecking of the packages that are being built because we know that if we just because we fix something does not mean that it has been fixed well. And the last thing we want is for us to ship a package out that looks like somebody took Thor's hammer to the source code, making it unrecognizable.
13:26 - Chris Wright
Yeah. Yeah. So, we have a lesson in there that goes beyond security. Just it's easy to imagine how agents will solve all of our engineering problems. But it turns out there's still a lot of noise in the system and getting that noise out of the system so we can focus on the real problem. Key. Key here. Now, let's talk a little bit about the scale. We are just patching the internet. What does that look like in the Lightwell context? Like GA Lightwell , we're in the early stages of building out all of this content and our content repository. What does that look like and what are the longer term goals and vision for Lightwell ?
14:05 - Mike McGrath
Yeah. We're mostly focused on Java and Python, which is great because we have a lot of expertise in that area. We've got a middleware team that has been building Java packages for decades. And just within Red Hat and some of our internal tooling, and if you've used Rel, you've probably noticed quite a lot is written in Python. So we have a lot of expertise in this area already. When it gets to actually trying to bring this flywheel up to full speed, quite a lot of it is built into what we're calling the validated packages. And it is just very simply trying to take what upstream did and get a build out of it. We haven't fixed any vulnerabilities in it. It's just about making sure that we know how to build that package, that we can test and make sure that that package built correctly and that we didn't break anything in the process. That's really the funnel of everything that we're doing. After that, it's a matter of applying patches to it and then making sure those patches didn't break anything. And I think one of the surprising things for me has been, you've got a package that's got known vulnerabilities in it, maybe it's got some unknown vulnerabilities in it. Going through and fixing each of those one at a time, pretty easy for automation and AI to do. One of the more complicated steps has been, how do you smash all of those vulnerabilities into a single build, which is what the customer expects. They want all the vulnerabilities fixed, not just a couple. And that turns out to be a fairly error-prone process, even agentically. And so, I would say the other end of the funnel is once we've got everything built and we know that it can build, how do we then prepare the final deliverable for customers? It's another area where we have quite a bit of a genetic workflow built into it, but also a pretty smart subject matter expert review process so that human eyes are involved when they're really needed. And that's really where we're spending our time on speed. I think the other part of that too is just the error rate. The higher error rate you have in any system, the more human eyes you're going to need. And so, that's just one of those internal metrics that we really keep an eye on.
16:15 - Chris Wright
It's a pretty sophisticated system and obviously still being refined and being built. That's generating the content, recognizing the inputs, generating the content, delivering that content. So we've talked a lot about the patching process, discovery, remediation, delivering content that's version-specific to what's in production. What's in production typically trailing pretty far what's in the open source community development trees. And so, another key aspect of our work is getting those patches that aren't simple backboards from what's already been fixed upstream into the upstream communities. What is that looking like? That's a monumental task. We're talking hundreds of thousands of different open source projects.
17:04 - Mike McGrath
Yeah. And I think this is one of the areas where I'm really happy with the business model that Red Hat has come up with. Because it really does find a way to sort of funnel this enterprise business into bolstering open source communities. And I think one of the things, Red Hat's always been very involved and committed to our upstream communities. And this is another good example of that, especially where we've been able to partner with IBM just to find people to help take these things upstream. While we have this great agentic factory to find and fix things, that does not do us a lot of good in terms of sending things upstream. Many upstream communities simply would not be ready for an onslaught of content. And I think anybody that's been involved in open source knows how poorly a sort of drive-by bug report, how poorly that can go. And so, we've taken a very human approach to this. We will be sending people upstream to work with those communities. If they have a embargo policy, we will follow that. If they don't, then we know to reach out to somebody and say, "Hey, I've got a vulnerability. I know better than to just open a bug and leave it out there." And I think there's a sort of rigor and care required when taking things upstream. Especially because not all upstreams are the same. Some of them are very vibrant and will say economically sound, where they will be able to respond to this very quickly. Some of them are somebody's hobby still, even though they're running a critical piece of library infrastructure for Java or Python or whatever. And we also understand that just because we found this vulnerability doesn't mean that upstream is going to drop everything that they're doing to fix it. And one other policy we have is when we aren't just going to go upstream with a vulnerability, we'll go upstream with a fix in mind and we'll work with them on that fix. And for the actually our very first vulnerability that we sent upstream, they decided not to take the fix that we provided. The approach was the same, but the code was different. And our policy is to then go back and use whatever upstream did. And we'll get rid of our fix and we'll rebuild based off of whatever upstream did. And then the other thing we're looking at is other communities like Akrites, which is coming up. And I think we're interested in seeing how that goes. You know, that has a lot of potential for helping both the future of Lightwell and open source security everywhere, which at the end of the day is the goal here.
19:43 - Chris Wright
Yeah, I think it's really important to recognize the human nature of the trust relationships in communities. So, some communities have turned off pull requests from non-identified members of their community. So, the drive by patcher is just simply rejected at the door. Others are going to want to see some due diligence put into the process so that you're not just getting garbage.
20:09 - Mike McGrath
Yeah.
20:10 - Chris Wright
There's plenty of conversations around AI slop and this kind of thing that really can slow down anybody's ability to even look at a potential issue. And then is it closed, not a bug? System works as designed or is it a real issue that needs to be fixed and how do you fix it? All these things I think are really important. Even though these dependencies are in production and they're older, there is a future state where that application could easily bump a version dependency to a newer version. So, part of this effort is broadening just the security stance of all of open source software and as you go forward to a newer release, the security fixes are baked in. I think this is also a really important aspect of how we sort of slowly shift the tide from a bulk of vulnerability discovery to secure development practices and rapid remediation. And at the end of the day, what do you think? Open source comes out more secure?
21:10 - Mike McGrath
Yeah, I think it has to. And while I'm not too worried about job security at the moment, I do think that the future of open source is not just more secure, but more secure than the proprietary options. And I think a big part of that is because large language models will always know more about open source communities and projects because a lot of times they've been trained on those projects. And so I think that there's a really great, I happen to think this is a really great marriage of large language models, harnesses and open source.
21:43 - Chris Wright
Yeah, I love that. So we got a more secure internet and better rested CISOs and security teams.
21:50 - Mike McGrath
That's right.
21:51 - Chris Wright
That's awesome.
21:52 - Mike McGrath
That's right.
21:53 - Chris Wright
Thank you so much, Mike. What a great conversation.
21:55 - Mike McGrath
Yeah. Great to be here.
21:57 - Chris Wright
Software maintenance shouldn't force us to choose between innovation speed and security. By decoupling vulnerability patches from forced upgrades, Project Lightwell fixes security flaws without breaking runtime stability. Just as open governance and enterprise Linux brought trust and standardization to operating systems, automated patch factories are now building that same trusted baseline for the era of AI. Thanks for joining the conversation. I'm Chris Wright and I can't wait to see what we explore next on Technically Speaking.
About the show
Technically Speaking
What’s next for enterprise IT? No one has all the answers—But CTO Chris Wright knows the tech experts and industry leaders who are working on them.