Loading summary
Patrick Gray
Foreign and welcome to this soapbox edition of the Risky Business podcast. My name is Patrick Gray. For those of you who don't know, these soapbox editions of the show are wholly sponsored. And that means everyone you hear in one of these editions of the show paid to be here. The idea is it's like a keynote interview, right? So someone from one of the sponsors gets to come along, talk to us about what they're doing, talk to us about how they see their problem space, so on and so forth. And joining me today is someone you would have heard on the show many times before. It's for us, Abu Khadija, who is the founder and CEO of Socket. Now, Socket is a software supply chain security company which really started out making a product slash platform that would help software developers discover when some of their dependencies or packages that they were relying on were malicious. But Socket has also grown these days, I guess, like what Socket does is not just that anymore. So joining me to have a chat about a few things now is for awesome Booker dj And a big feature of this interview which I think is going to be really interesting to people, not just in software development actually is we're going to be talking about reachability analysis, which is the idea that there might be a bug in a package, but you know, is it reachable and you know, how do you go about testing that? So that's going to be a big part of this conversation. But for us, let's kick it off with a bit of an update. Know you're not just doing the malicious package tracing stuff anymore, are you?
Abu Khadija
No, that's right. And Pat, it's great to be here, I think. We originally launched Socket on a Snake Oilers back in April, in 2023. So we've been doing this for a little while. Yeah, yeah. And so, yeah, the pitch has definitely changed since the early days. When we initially launched, we were very focused on solving this problem of how do we help companies safely use open source software. We looked around and we saw all these supply chain attacks happening in open source and we were frankly like a bit mystified as, like, why is no one taking this problem seriously? All the vulnerability scanner vendors, you know, the SCA tools were just, you know, purely focused on CVEs to the exclusion of, you know, what we felt as a bigger risk than vulnerability. Right. Like a supply chain attack is something that's going to compromise you 100% of the time. Whereas a vulnerability, you may have some time to fix that you might have some days, weeks, months before somebody finds that and so we were just surprised. No one is really tackling this problem and it's going to be a growing problem, we believed. And so that's kind of where we started as a company building the first kind of ecosystem wide scanner, going out and really analyzing every open source package that exists, looking for signs of supply chain attacks. And so we built what I think is a best in class tool in that space. And we've had a lot of success with folks bringing it in using it. And we're now finding about 1,000 supply chain attacks per week in the open source communities and ecosystems that we scan. We cover all the top ones. So that's kind of where we started and then now we had customers. We're fortunate to actually work with some of the biggest companies in the world. Now we've had success getting the this into more and more hands. And what they told us is like, hey, why aren't you guys also handling the CVE problem? It seems like it's a related problem. Just tell me if this open source dependency is safe, help me understand the full risk profile of it, not just is it malicious? And so then we kind of, we listened to our customers and so we were like, this seems like a reasonable request, so let's kind of take a look at the vulnerability space. And then that's kind of what led us now to doing. We didn't want to just do it like everyone else. And vulnerability scanning is a commodity. Right. There's open source tools, there's a million vendors that can tell you if you have CVEs in your dependencies. So we wanted to take a socket like approach and kind of approach it with fresh eyes and really try to do something different.
Patrick Gray
Yeah. So when we first met and first started talking about your product, I always thought that a big risk for you was going to be that some of these software composition analysis companies like Snyk and White Source and whatever, that they were going to move into this space. Right. And that would be a bit of a challenge for you. What I didn't necessarily expect was for you to actually go and sort of enter their space and start challenging them there. I mean, I'm guessing that's where you're moving, right? You're trying to become like a bit of a challenger to those companies now.
Abu Khadija
That's right, yeah. And there's actually been many instances already where we've had customers switching off of those tools for various numbers of reasons.
Patrick Gray
I mean, is it just the case that like with any sort of software in infosec I mean, I see this pattern over and over because I've been doing this for 25, five years. Is new vendor comes along, builds a cool tool. It's amazing. Everybody uses it to like, for like five to 10 years. Then it sort of gets a bit stale, then the new one comes out and everyone sort of switches to that. Is like, is that kind of like what you think is happening here or is it like what you want to happen here?
Abu Khadija
I guess, I mean, there's a bit of that. I think people do want to use the hot new thing. I do think, though, that the existing vendors have stumbled in a couple of key ways. I think one is, you know, they are inundating folks with too many alerts and people feel that these tools are failing them in a key way. Like, if everything is an urgent priority, then nothing is a priority. And so I think if you look at those, some of the folks you mentioned, they've been really slow to adopt things like reachability analysis to try to cut down on the noise and help you understand which of the vulnerabilities are actually real. And I think on the supply chain attack side, they're relying on literally the CVE system to tell you if a package is malicious, which is just not what that. Like not what the NVD is for. It's only tracking vulnerabilities. And so a lot of folks have realized they have a gap in that area and they just haven't.
Patrick Gray
But they are, they are catching up to you on that, right?
Abu Khadija
I mean, not really. So they do have like a, you know, they do have some research teams that are kind of sporadically looking for things here and there, but it's not systematic. Like, Socket is literally looking at every single open source package. And from the moment it's published, we're crawling all these ecosystems. We get updates when a new one is published.
Patrick Gray
Well, I mean, I'm regularly seeing Socket cross my desk, just through Catalan's news bulletins and whatever. It seems that you are absolutely the number one source for that sort of information these days.
Abu Khadija
I think part of that is that we started as a company in the LLM era. And one of the things about doing something at this massive scale is that it helps to have like, really cheap. You know, I think of them as, I think we talked about this before. I call them like, you know, college undergrads or interns. Yeah, you have an infinitely scalable set of these interns and you can just tell them to go and look at all the code and they're going to get it wrong a lot. They do all the time. They'll come back and say, oh, this is malicious. And then you look at it and it isn't, but they're still pulling, you know, a lot of signal out of the noise. And then we just add humans in the loop. And the combination of AI plus humans is actually really powerful. And this in this space.
Patrick Gray
So you got like a, like an agentic sweatshop, basically.
Abu Khadija
Pretty much, yeah.
Patrick Gray
So, you know, some nice cheap labor. So look, you just mentioned it, and this is, you know, obviously a big part of the conversation today is this idea of reachability analysis, which, look, again, you know, I've been in this discipline for a very long time and I can understand why this is like actually pretty game changing from a volume management perspective. But let's just start by talking about defining what is reachability analysis for us. Please define it for the audience.
Abu Khadija
Yeah, so the main problem with vuln scanners today is they give you too much noise, not enough signal. And what that's really caused by is when they identify a CVE and a dependency. What they're telling you is that, yes, you have a component somewhere in your application that a CVE has been issued on. That doesn't mean that your application is actually vulnerable to, to that vulnerability, that you're actually. That it's exploitable or that it's truly what we call reachable, meaning there's a way for an attacker from the outside to hit your application and then find a path through the code to actually trigger that vulnerable function in the application. So the component has a CVE issued against it, but your application is actually not vulnerable. So that's the problem. That's the problem with vuln scanners today. They don't do a very good job distinguishing that.
Patrick Gray
Well, I mean, that's right. There could be some library that somebody's included, but you're not using the vulnerable function and there's no way to trigger it. There's no way to get some input to it. Right. So you are going to get so many alerts telling you to fix bugs that just don't affect you. Right. But so here's the thing. Okay, reachability analysis sounds great. If there's a way to figure out like, well, you know, are these bugs reachable in our software packages? But that is a very hard problem. There's a reason nobody solved it.
Abu Khadija
Yeah, no, I 100% agree. It's actually an area of computer science that's in active development. And really what it comes down to is there's this fundamental kind of rule about analyzing source code that it's called the halting problem. There's no way to, when you're looking at a piece of source code to really determine what it's going to do at runtime without actually running the code. And so what you have to do when you're doing static analysis is you have to make some assumptions, you have to use some heuristics. And that's where it's really challenging to pick the right heuristics to get good results from this analysis. You often find when folks do this, it's hard to actually roll out the reachability analysis from the legacy vendors, the other vendors that have tried this, because the reachability goes down these kind of false paths in the code and spends all the CPU cycles going down this path that doesn't really matter and then ends up kind of, you end up running out of time and then you have to kind of give up. This will never terminate, the analysis will never terminate. And then they try to solve this by kind of putting in these hacks like, oh, don't actually scan that folder because it's a degenerate case and the scanner just gets confused and goes wild in that direction. So they'll kind of ignore things here. So it's really like not a very rigorous process to get these things rolled out. And then you also have to go like repo by repo within your company and kind of add this extra CI step and do all this tweaking and customizing to get this analysis to run. And it's very expensive. This analysis takes hundreds of gigs of RAM in some cases just because of the effort involved to actually do this analysis. It's a very hard problem and there's a reason why people either don't do it or they say they do it. But when you actually dig a little bit deeper, you learn there's a bunch of asterisks on that. They're not scanning the transitive dependencies. It doesn't actually work on a monorepo in a real world code base. It doesn't work very well in dynamic languages because they're making all these assumptions that they've baked in that work for Java, but not for JavaScript or Python or Ruby or any of the other dynamic languages that are really popular with developers. So there are a lot of caveats with the kind of existing solutions.
Patrick Gray
Yeah. So what's your approach here? And I believe that this involved there was some very small startup that was doing some promising work here and you've sort of acquired them. Right. Like you've done some sort of stock swap thing and brought them into the, into the fold. So why don't you tell us about like that process and like what this person or these people were actually doing that made it so compelling for you?
Abu Khadija
Yeah, so like I said, when we started looking at CVEs and we were hearing from our customers like they wanted us to be the kind of all in one solution for supply chain security, tell us if our dependencies are safe, help us identify problems. We started developing reachability analysis ourselves internally and we used an approach called module based reachability, which is where you kind of take a look at a high level of like, are you including entire files or entire classes in Java applications and that's how you can kind of determine whether or not like a particular class is going to be loaded into the application. And that works pretty well for static languages. It gets you about a 30% noise reduction and it's relatively easy to build. And so that's what we see a lot of like the other vendors kind of doing as they're kind of, you know.
Patrick Gray
Yeah, so it says this, this may be reachable.
Abu Khadija
Yeah, exactly. And so when we were like, okay, we want to do better than this, obviously this is not like this is good to kind of put our, you know, just to kind of put draw a line in the sand and say that we can kind of do reachability. But it didn't feel like a truly socket level quality product that we wanted to put our name on. And so we started looking at, okay, how can we do this better? And it turns out, like I said, this is a really difficult problem. And it was going to take our team probably six, maybe 12 months to kind of get up to speed on the latest research, to maybe bring on board a couple of PhDs that had the relevant experience. And so we looked around and found the Kawana team, which if you don't know them, they are just a brilliant, brilliant set of engineers. They're out of Denmark. They're actually in a, from a town called Aarhus in Denmark. And the company was started by these four folks. One is a professor who's been studying, doing research on reachability, specifically in JavaScript for over 20 years and three of his PhD students. And then it's kind of a small team, about eight folks. And at the time we did this acquisition, by the way Socket ourselves, we were only 30 people. So I never would have thought a 30 person startup would be acquiring an 8 person startup like that just wasn't on the, on the table for me, you know.
Patrick Gray
Yeah. Is it a, is it, is it an acqui hire? Is it a merger? Is it an acquisition? Well, kind of, yes, I guess to all of those sort of. It's unusual to see this, right? Like I, you know, it's not something you see very often, which is a small startup merging with another startup or sort of acquiring another startup.
Abu Khadija
Yeah, I mean, the thing in this case was it was such a match made in heaven. I mean, the Kiwana culture, they're super technical, almost everyone there was an engineer and they had built the best reachability analysis solution that we'd seen. It was better because it was designed from the start for dynamic languages like JavaScript and Python and Ruby and then it was kind of transferred from there to the static languages. So it works really well on the hard to analyze languages, which is unique. They also didn't really build anything else, so they just did reachability and we did basically everything else, but not reachability. And so it was this perfect match. It's like puzzle pieces, right?
Patrick Gray
Worth. Worth more than the sum of the parts. Right. When you put those two things together. So tell me, these Danish mega nerds who you managed to discover, right, who had built this amazing reachability thing, and that's not meant as a pejorative, by the way, to these Danish mega nerds who might be listening. It's very impressive what you've done. Why don't you tell us what their approach is to solving this problem? Because as you said, you've outlined a couple of kind of dead ends or things that get you like 30% of the way. There's, you know, what's very different about their approach.
Abu Khadija
So the first thing is that they analyze the full dependency tree. So that's something you actually don't see in all. It's not a given actually that you're going to get that from some of the other solutions. And that's really important for a few reasons. The first is that when you're dealing with vulnerabilities in modern applications, you have these deep dependency trees. And so a lot of your vulnerabilities are going to be 2, 3, 5, 10 layers deep. And in your dependency tree, I think I mentioned this to you before, Pat, but like A Hello World JavaScript application today has 1000 dependencies. I'm not kidding. You go to get a REACT app going, you want to show hello world, like 1000 JavaScript dependencies to show hello World in your browser.
Patrick Gray
I mean, this is the same. We are living in the era of the 10 megabyte buttons on websites, right?
Abu Khadija
Yes.
Patrick Gray
It's just. There's too much. There's too much everything.
Abu Khadija
Yeah. And the vibe coding is not helping because people are writing more code. So the key thing is you got to analyze the full dependency tree. And one thing that's really important to us just as a company and with both the Kiwan team and Socket, is that we never tell you that a dependency or that a CVE is reachable if it's potentially not reachable. And we never tell you it's not reachable if it's potentially reachable, because both of those lead to really bad outcomes. You don't want to not fix something because the tool told you that it's fine. And then you end up getting hacked. And you don't want to do the opposite. You don't want to over report things as reachable because that's time you're spending on really fake busy work that isn't going to actually meaningfully improve security. And that's the whole reason why this whole reachability analysis thing exists. So you need to analyze the full dependency tree. But that brings kind of. That's sort of the main challenge is how do you do that efficiently? Because that's a lot of code. You know, static analysis tools, often they're slow to run. Developers don't like running them. They can't run them on every time they hit save in their editor because they're too slow to run. Now imagine static analysis, but not just on the 5% of the application that's your code, but now it's on the 100% of code that's your entire dependency tree. So you're talking like 20x as much code in some cases. And so this has to be fast. And so that's a lot of the reason why the other folks have kind of punted on doing that. And they'll just make these guesses or these assumptions about whether or not a CVE that's deep in the tree is actually reachable or not. So they were very focused on accuracy. And that's another reason that we really connected with the team. Just from a cultural perspective.
Patrick Gray
How do you solve that problem? How do you make it sort of performant and how do you deploy this thing? Because you're saying about some of the other things before, it's like a CI CD pipeline integration, it's kind of like nobody wants to do it. So. So how does it, you know, a very risky business sort of question like, why don't you tell us how this thing works.
Abu Khadija
Yeah, yeah, for sure, for sure. So there's a few things though I need to explain because we actually have a couple of options for reachability and it's actually what makes the solution really powerful for folks. So what we've been describing so far is what we call our Tier one reachability. And this is where you do a full analysis of the entire application and dependency tree. And the Koana teams managed to make that perform pretty well, mostly by using really smart heuristics. So when you're looking at code, you have to make assumptions about what you're going to analyze and what you're not going to analyze and where you're going to kind of cut off the analysis because you can't run the code. Fundamentally what we're doing here is we're doing static analysis. So they've just done a really good job of over the years because of Honest's research, like really honing this down so that it works on real world code bases on large monorepos. And it's something that is literally just due to his decades of research in this space. But in addition to that we have another option we call our tier 2 reachability or pre computed reachability. And this is something that's unique to Socket that we've developed with the Kiwan team since they came on board three months ago. And this is a really great option for folks that want to prioritize ease of getting this rolled out above all else. And it's really powerful because it just needs your manifest files to do the analysis. There's no source code access required. And now this is kind of a crazy idea. Think about that for a minute. We're going to tell you if a CVE is reachable from your application, but we're not going to look at the code of your application. Right. So how do you do that? Well, you make an assumption. The assumption is you make that you say we're going to assume for your direct dependencies that you're using all of the functions that are exported from that direct dependency in your application. Now that's an assumption. But if you make that one assumption, we can pre analyze or pre compute the reachability graph for that dependency and the entire dependency chain that goes all the way down before we've even looked at your application. So when you come to Socket and you say these are my dependencies, we can say, great, you have a CVE that's 10 layers deep and there's no possible way to reach it because we've Already looked at that entire dependency graph and there's no way to use that direct dependency in any way. You can call all the functions, you can call it with all the different arguments, you cannot reach the vulnerability. And so you get almost as good a performance, you get 80% reduction instead of 90 plus percent reduction. But you don't need to analyze the application source code. It's incredible.
Patrick Gray
No, no, no, no, I get it because, yeah, like it's a tree, right. And you just analyze it all the way down and you're going to, you're going to know like that thing there's buried way too deep. You can't get to that.
Abu Khadija
Yep, exactly. And to my knowledge, no one's ever done this before, so I think we discovered this. We actually just filed a patent on it. Because, I mean, not that I, I'm a huge fan of software patents. I think they're mostly bad ideas, but this is a novel idea and we want to protect it. The point is, it's really like the first that I've seen of this type of, type of analysis. And it's great because it means that you just connect us to your GitHub or you just hit our API with those manifest files and you're going to get back results. We've literally just onboarded a customer that's in the Fortune 50 that was using. I don't want to say their name just out of respect for other vendors, but actually I'll say you mentioned their name earlier in the show and they couldn't get reachability rolled out and they've been a customer for five years and then they tried socket with the pre computed reachability. They got results immediately and they were shocked. And so they're switching off now and switching to socket fully. So it's insane. This is actually going to be a really big deal for teams that love the idea of reachability. They've been hearing about it for years, but they've just never been able to get it rolled out.
Patrick Gray
Yeah. So I imagine like, just because you've got a bit of flexibility in the way you can roll this thing out, switching over is going to be pretty easy. But if you're dealing with a company like any Fortune 50, they're going to have a lot of repos, a lot of code, a lot of applications. So this is going to result in a lot of output. So how are you handling that? Right. Like, are you providing outputs that are friendly for vulnerability management tools like Nucleus is a sort of more. I guess they're sort of a maturing startup now who do vulnerability management stuff. But like, there's other companies that do that sort of data, sciencey bit of managing vulnerabilities at scale. Like, what are you doing to integrate with that stuff? Because I mean, it's great to find this stuff, but if you're just crapping it out into a console that you have to log in and it's a Socket console, like that's of limited utility.
Abu Khadija
Yeah, all of our biggest customers basically never log into the console. They're API only users. And that's something we wanted to support. We have supported from the very beginning just because we know everyone has like 80 security tools, I think is the last number I heard. And so it has to go into one system. We see actually Anthropic. One of our customers did a Talk@BSides SF recently and they talked about how they're using Socket's API to integrate Socket, but the company doesn't really even have to know that Socket's being used. So they built this dependency tool that internally tells you whether or not a package is allowed or not within the company. And behind the scenes it's just hitting Socket's API and saying, what's the score for this package? What's the supply chain security score? And if it's below 80 out of 100, then they just don't even allow it in the entire company. It's just banned. And that's their, that's their approach. But the developers don't need to know that. They just check the internal tool and it tells them whether or not they're allowed to use it. So everything in socket uses APIs. And so in this case, yeah, we just, you know, you can import that stuff into whatever you want. Nucleus, Nucleus is a, is a great company and a great, like one, you know, great option for ingesting this data.
Patrick Gray
Yeah, like mega triage, I guess that's what they do. So look, the question becomes, right, you've got this stuff out there in the wild, analyzing corporate applications and trying to find what's reachable, what's not. I mean, what have you learned through that process? Because I'd imagine that you're going to see something interesting over and over and over and over again. Whether it's certain classes of bugs are more likely just generally to be reachable, or bugs in certain types of libraries or packages are more likely to be reachable. What are the insights you can share with us about what you've learned by going through this process with a bunch of enterprise apps?
Abu Khadija
I mean, the first Realization is just the amount of dependencies that all these companies have. It blows my mind. As somebody, as an open source maintainer, former open source maintainer myself, I scrutinized the crap out of my dependencies before I brought them into my projects. Because anyone who's installing my code is going to pull in my dependencies. And so I felt like it was my responsibility to make sure that those dependencies were safe.
Patrick Gray
Yeah, most people don't think like that.
Abu Khadija
For us, I know this was the realization that led to starting the company, actually was when I kept seeing these supply chain attacks, you know, I realized no one's looking at the code because if you just opened the file and you just looked at what the dependency did, you know, these things were not disguised very well. They still are not disguised very well. Like you can just see in the code, it steals all your environment variables and sends them off to an IP address. So that's when I realized no one's looking at this stuff. Because if you did, if you spent even one second and opened up the index JS file of the package, just see what is it doing? You'd see right away something looks off about the code. By the way, this is why the LLMs are so effective at identifying those types of anomalies is because one of the reasons is that there's something about when you just eyeball the code, sometimes not all the stuff we find is like this, but in terms of volume, the vast majority is this like low effort crap that's being pumped out by folks. And so that stuff is you just look at the file, which the LLM can do, and it just sort of sees it. It's like something's off here and then it bumps it up to the human for review. But yeah, so that's the first thing is just like there's a lot of dependencies we've seen, you know, like a lot of companies, brands that, you know, right. I mean, they have 100,000 plus dependencies, you know, in, you know, across the organization. And in some cases they have like 50 versions of the same package, you know, across all the different applications. And some tools make this worse, like Dependabot. When folks are using that tool, it tries to bump you to the latest version of packages. It's a really popular tool with developers, but it doesn't take into account what else is already being used in the company. And so you get this kind of diffusion of versions. So you're using just every version of the package that you could possibly be using.
Patrick Gray
Yeah, like 600 forks of the same thing, basically. That kind of deal.
Abu Khadija
Yep, Yep. Pretty much. Yep. And so we see a lot of that. I think the other thing that's interesting about just stuff we've noticed is there's quite a lot of what are called phantom dependencies. It's kind of an interesting thing. So folks don't necessarily always declare the dependencies that they're using, they just import them. And if the dependency happens to be somewhere in the dependency tree, like something else installed it, maybe it was installed by one of their dependencies, so it's not directly dependent upon by them, but they just go and import it.
Patrick Gray
Right.
Abu Khadija
And the tools let you do that and they just say, oh, the dependency is there on the file system, I'll just grab it. But you're not declaring it anywhere as one of the things.
Patrick Gray
So it's not in a manifest or it's just how it got here. Who knows? Who can say?
Abu Khadija
Exactly. There's no version. You don't know what you're. So a lot of tools will take in the manifests and then just miss those dependencies. So that's why they're called phantom dependencies. So that's been interesting, the amount of that that is present. Which is another reason why, you know, doing. Doing reachability can be really helpful.
Patrick Gray
Yeah, because that thing's not going to get you owned. It's just kind of there. Right. I mean, this is sort of like, you know, dormant malware kind of thing. Like, you don't really have to worry about it. EDR doesn't often alert on it, you know, because it's just sitting around, it's on the file system. It's not really a priority. I mean, it's the same thing for this stuff. Right. Like a bug that no one can exploit when you've got, like. I mean, people just don't understand the extent to which vulnerability management teams are backed up. Right. Like they need help to prioritize. So anything that can help them do that. So look, with this reachability analysis, do you have any sort of statistics on when you compare, like an SCA Tools scan that says, well, there's like, you know, say it finds 1000 CVEs. What percentage of them can you rule out with this reachability analysis?
Abu Khadija
Yeah. So with the Tier 1 reachability analysis that came in from the Kiwan acquisition, we get 90% on average. Now, this is different in each programming language, so it's a kind of an average number. You'll have to sort of see for your application what you get but averaged across all of the different languages that we support, it's 90%, which is a really big number. And that just shows like all that work that's been done by these teams over. I mean, I just feel so bad for them and for the developers that are forced to fix this stuff that.
Patrick Gray
They'Ve been patching, stuff that just doesn't give them any risk at all. Yeah, yeah.
Abu Khadija
And I mean, and I get it if you want to stay on the latest package versions due to like, you know, some other reason, like if you have a good reason, like, oh, I want to get the new features that this, that this package has, like, that's a great reason to upgrade. Right. Maybe it fixes a bug. Like, okay, good, good, that's a good reason to upgrade. But if you're, if you're just doing it on a legacy application that's in maintenance mode. Right. You're. And it's so painful to do these upgrades a lot of times because you're actually breaking stuff.
Patrick Gray
You have to re engineer it, work around it. I, look, I've been there. It's hard.
Abu Khadija
Yeah. And sometimes it's fully the security team's job. Like they have no help from the developers.
Patrick Gray
When I say I've been there, by the way, it's like it's people who are doing that stuff for me telling me about this. I am not a developer. But yes, go on. Sorry.
Abu Khadija
But at least you have the developers helping you. That's not always the case. Sometimes it's like their job is go and patch the application and then they're like, I don't even know what this application does. It doesn't have any tests. I can patch the dependency, but then the app doesn't start anymore. Now I'm going to go and just hack at the code until I get it to run again. But I don't know if it's like working correctly anymore. There's no tests, like, how am I supposed to do this? It's a really painful thing to ask people to do. So the 90% is really powerful. Now I want to just be clear though that the other reachability, the tier 2, which is the pre computed approach, that's really easy to roll out, we see on that one about a 60 to 80% reduction. So it's a little less, but it's way, way easier to roll out and it's nice to have both options.
Patrick Gray
So I guess like, look, the reachability thing is very important. It's very cool. But I guess you're seeing this, you know now that you're doing, you're sort of trying to enter that bigger SCA market, which is huge. Right. Instead of just being a bit player who does like something very specific, you're trying to enter that market. It seems like the reachability thing is like your hook. Right. Like that's your, what do they call it, your usp. Your unique selling proposition seems to be around the reachability stuff. Like, is that a. Is that a fair assessment?
Abu Khadija
It's actually the second usp, I guess, the first being the deep supply chain analysis, looking for the malicious packages in.
Patrick Gray
Addition to the saviors. Yeah, that makes sense.
Abu Khadija
Yeah. And I think in a way it's kind of more broadly applicable to more teams, because what we found is people care a lot about the malicious package detection. Absolutely. Especially if they've been affected by an attack or had a close call or something like that, or if they're in a space that just is known for this kind of stuff. Like the cryptocurrency industry is obsessed with socket.
Patrick Gray
Yeah, I was going to say. I can imagine, I can imagine. You got a lot of customers there.
Abu Khadija
Yeah, yeah. Pretty much anyone who's. Anyone there. So, yeah, it's been a pretty, pretty good hook, but it's not part of like SOC 2 standards, you know, it's not part of, you know, like, you don't have to be doing this type of scanning. And so that's where I think the reachability comes in, because everybody has to do deal with vuln management.
Patrick Gray
Yeah. So it was really funny when you mentioned, like at the start of this interview, you were talking about like, well, people aren't really looking for the malicious, you know, malicious packages. They're so focused on CVEs and whatnot. And I'm just sitting there thinking, well, that's the compliance standards. Tell them that they have to. Right. Like, at what point do we start seeing this sort of analysis for, you know, scanning for malicious packages? Like, when is that coming to compliance standards, the big ones like SOC2?
Abu Khadija
I think it's overdue, to be honest, and I was kind of hoping that.
Patrick Gray
Well, me too. And I figured you might have some insight here.
Abu Khadija
Yeah, you know, they did think about adding it to the secure by design initiatives, but they wanted to start with stuff that was super uncontroversial for the first version of that. And then after, you know, Trump took over, I don't know what the future of secure by design is. So I don't know if a V2 is going to come out with the supply chain stuff. In there. Yeah, I think the SBOM stuff is promising, folks. I mean, that is actually getting pretty standardized, especially if you want to sell software to the US Federal government. So I have some hope that at some point, folks are going to look around at all these sbombs they've been collecting and go, what the hell can we do with this stuff? And then they're going to realize, like, oh, knowing the security status of these things, not just the vulnerabilities, but, you know, who the hell is behind this package? Right. Who is the maintainer? Has it been backdoored?
Patrick Gray
And then you can sell API calls to the US Government, Right, Basically.
Abu Khadija
I mean, yeah, and to everyone else who's been collecting SBOMs as well, and like, you know, just sticking them in their compliance tools and not doing anything useful with them.
Patrick Gray
Yeah, but I mean, that's always been the plan with SBoM, right? Is to collect the information first, get it coming in, and then, you know, naturally you'll be able to find useful things to do with it. And checking for malicious packages is going to be one of those things.
Abu Khadija
And not just that, but even, like, the broader set of risks, like, we see a lot of, like, a surprising use case that we've seen that I did not expect was folks looking through their SBoMS and through, you know, the dependencies that we pick up looking for deprecated packages, not because those are a risk today, but because if there was a CVE added, you know, announced or, you know, discovered in one of those. Those packages, there'd be no path forward. Right. There's no maintenance, no further maintenance happening on that package. So now their company is stuck either forking it themselves and making the fix or migrating off of that package to a different one, which can be significant work. So they just want to get ahead of that before they're scrambling. And so we've actually seen things like that, which I would never have guessed that security teams would be interested in deprecated packages, because that, to me, feels much more like an engineering quality kind of a conversation, you know.
Patrick Gray
No, I get it, I get it. I understand 100 why you would want to. You would want to know, you know, because I think, you know, if you're a security team, the reason you would want that info is because a lot of these are not going to be a tremendously big deal. But, you know, you might shake out one or two where you're like, that's deprecated. That's a problem. You know, that's. Is that sort of how it's working?
Abu Khadija
Yeah, they're surprised sometimes at what's deprecated. Or there's even cases of things that are kind of, like, unofficially deprecated. They haven't had a version released in five years, and so they might not have flagged it as deprecated officially, but it's effectively deprecated. And so things like that. We can also.
Patrick Gray
Is the maintainer alive? Who can say.
Abu Khadija
Yeah, and that's actually all too real. People are getting older and this stuff. There are cases that you're joking about, but there's actually real cases like that.
Patrick Gray
Yeah. Where someone is no longer with us. Right. Makes it a little bit difficult to maintain code when you are no longer alive. All right, look, we're going to. We're going to wrap it up there for awesome dj. Thank you so much for joining us to talk through all of that really interesting stuff. I think this reachability analysis stuff is. Yeah, very interesting. And it's. It's something that everybody kind of needs. Right. So I wish you all the best with it. I hope everybody builds something like this, because for too long we've been patching, you know, we've been patching bugs and stuff that we just haven't had to have to do. Yeah. Great to chat to you, my friend. And we'll. We'll do it again soon. Cheers.
Abu Khadija
Thank you, Pat. It's an honor to be here.
Podcast Summary: Risky Business - "Risky Biz Soap Box: How to Measure Vulnerability Reachability"
Release Date: August 14, 2025
Host: Patrick Gray
Guest: Abu Khadija, Founder and CEO of Socket
Duration: Approximately 35 minutes
In this special Soapbox edition of the Risky Business podcast, host Patrick Gray engages in a deep dive conversation with Abu Khadija, the founder and CEO of Socket. Unlike regular episodes, Soapbox editions spotlight sponsors, allowing them to elaborate on their solutions and perspectives within the information security landscape.
"These soapbox editions of the show are wholly sponsored... someone from one of the sponsors gets to come along, talk to us about what they're doing."
— Patrick Gray [00:00]
Abu Khadija outlines Socket's journey since its inception in April 2023. Initially focused on identifying malicious dependencies in software supply chains, Socket has expanded its offerings in response to customer feedback requesting more comprehensive vulnerability assessments beyond just CVEs.
"We were very focused on solving this problem of how do we help companies safely use open source software... now Socket has also grown."
— Abu Khadija [01:32]
The conversation highlights significant shortcomings in current Software Composition Analysis (SCA) tools. Traditional SCA solutions often inundate users with CVE alerts without providing context on whether these vulnerabilities are exploitable within the specific application, leading to alert fatigue.
"They are inundating folks with too many alerts and people feel that these tools are failing them in a key way."
— Abu Khadija [05:03]
Abu emphasizes that existing tools do not effectively differentiate between vulnerabilities that pose real risks and those that do not, primarily because they rely solely on the National Vulnerability Database (NVD) for CVE information.
"They're relying on literally the CVE system to tell you if a package is malicious, which is just not what that is."
— Abu Khadija [05:52]
Reachability analysis emerges as a pivotal solution to the noise problem in vulnerability management. Defined by Abu, reachability analysis assesses whether a vulnerability in a dependency can actually be exploited within the context of an application's architecture.
"Reachability analysis is the idea that there might be a bug in a package, but, is it reachable and how do you go about testing that?"
— Patrick Gray [04:03]
"The main problem with vuln scanners today is they give you too much noise, not enough signal."
— Abu Khadija [07:39]
This method determines if there's a viable path for an attacker to exploit a vulnerability, thereby enabling security teams to prioritize remediation efforts more effectively.
Abu delves into the complexities of implementing reachability analysis, citing the Halting Problem—a fundamental challenge in computer science that complicates determining code behavior without execution. Static analysis tools must make educated guesses and use heuristics, often leading to inefficiencies such as endless analysis paths or excessive resource consumption.
"There's no way to, when you're looking at a piece of source code to really determine what it's going to do at runtime without actually running the code."
— Abu Khadija [08:27]
These technical barriers have made it difficult for traditional vendors to offer reliable reachability analysis, resulting in incomplete or flawed assessments that fail to provide actionable insights.
To overcome these challenges, Socket strategically acquired the Kiwana team from Denmark, a group renowned for their expertise in reachability analysis, especially for dynamic languages like JavaScript, Python, and Ruby. This acquisition allowed Socket to integrate advanced reachability techniques into their platform.
"We found the Kiwana team... they are just a brilliant set of engineers."
— Abu Khadija [11:03]
Socket's approach combines AI-driven analysis with human oversight, ensuring high accuracy in detecting malicious dependencies and assessing vulnerability reachability.
Socket offers two tiers of reachability analysis:
Tier One: Full Dependency Analysis
This comprehensive approach examines the entire dependency tree of an application, ensuring that all potential paths to vulnerabilities are assessed. Despite the complexity, Socket achieves about a 90% reduction in irrelevant alerts through optimized heuristics.
"With the Tier 1 reachability analysis... it's 90%, which is a really big number."
— Abu Khadija [28:22]
Tier Two: Pre-Computed Reachability
A novel solution that does not require access to the application's source code. Instead, it relies solely on manifest files to pre-compute reachability, enabling rapid deployment and an 60-80% reduction in noise.
"It's unique... we can just hit Socket's API and say, what's the score for this package?"
— Abu Khadija [20:12]
This tiered approach allows organizations to choose the level of depth and resource commitment that best fits their needs.
Socket's API-centric design facilitates seamless integration with existing vulnerability management tools, ensuring that security teams can incorporate reachability data into their workflows without disruption. An example provided was a Fortune 50 company that rapidly transitioned to Socket's solution after finding immediate success with the pre-computed reachability tier.
"They built this dependency tool that internally tells you whether or not a package is allowed or not within the company."
— Abu Khadija [23:28]
Through extensive analysis of enterprise applications, Socket has uncovered several critical insights:
Exponential Increase in Dependencies: Modern applications, especially in JavaScript ecosystems like React, often include thousands of dependencies, making manual scrutiny impractical.
"A Hello World JavaScript application today has 1000 dependencies."
— Abu Khadija [15:37]
Phantom Dependencies: These occur when applications import dependencies that are not explicitly declared in manifest files, leading to untracked and unmanaged packages within the dependency tree.
"Folks don't necessarily always declare the dependencies that they're using; they just import them."
— Abu Khadija [27:06]
Version Diffusion Due to Tools Like Dependabot: Automated tools aimed at keeping packages up-to-date can inadvertently introduce numerous versions of the same package across different applications, complicating dependency management.
"Dependabot... it tries to bump you to the latest version of packages... making all these assumptions that they've baked in."
— Abu Khadija [26:25]
These findings underscore the necessity for sophisticated tools like reachability analysis to manage and mitigate the inherent risks in complex dependency ecosystems.
Abu anticipates that reachability analysis will soon become integral to compliance standards such as SOC 2 and Secure by Design initiatives. The increasing adoption of Software Bill of Materials (SBOMs) provides a foundational dataset that can be leveraged to enhance security postures by identifying not just known vulnerabilities but also potential supply chain threats.
"Knowing the security status of these things, not just the vulnerabilities, but, you know, who the hell is behind this package?... it's overdue."
— Abu Khadija [32:11]
Patrick Gray wraps up the discussion by emphasizing the critical role of reachability analysis in modern vulnerability management. Abu Khadija reiterates Socket's commitment to providing accurate, scalable solutions that empower security teams to focus on genuine threats rather than sifting through endless noise.
"It's the first time that you've ever done this type of analysis... it's really like the first that I've seen."
— Abu Khadija [20:22]
This episode of Risky Business provides an insightful exploration into the complexities of vulnerability management within software supply chains. Abu Khadija's expertise sheds light on the innovative solutions Socket is pioneering, particularly in the realm of reachability analysis, which promises to transform how organizations assess and mitigate risks associated with their dependencies.
End of Summary