[00:00:00] Jacob Haimes: This episode was recorded on August 11th, 2026. Ignore Igor's laughing [00:00:09] Jacob Haimes: Welcome to Muckrakers, where we dig through the latest happenings around so-called AI. In each episode, we highlight recent events, contextualize the most important ones, and try to separate muck from meaning. As always, I'm your host, Jacob Haimes, and joining me is my co-host, Igor Krawczuk [00:00:26] Igor: Thanks, Jacob. This week we're talking about the latest plans of using dinosaurs to combat A- AI safety issues, I wish we were talking about this, but it was a very funny, uh, satirical take on recent from OpenAI, uh, that we will link in the show notes as, like, this week's, like, funny little thing for you to check out, 'cause I think we all need that right now. But what we're actually gonna be talking about is the AIs are hacking everyone [00:00:57] Jacob Haimes: I just wanna say that the f- the, the, the figure two is, like, really compelling for the case as to why we should build Jurassic Park, um, given, given, you know, current circumstances. So, so definitely check out the, uh, you know, plot there 'cause it, it's, it's really worth, uh, worth addressing [00:01:14] Igor: if it's answers all of the questions you m- uh, you might have. But let's, [00:01:19] Jacob Haimes: yeah [00:01:20] Igor: talk about the AI is hacking everyone. So do you want to intro it? [00:01:25] The OpenAI/Hugging Face Containment Escape --- [00:01:25] Jacob Haimes: Um, yeah. I, I mean, let's see. In what? Late July, I guess. Uh, OpenAI had one of its models doing an cybersecurity evaluation on, on a Hugging Face thing. Um, it got out of the containment, uh, structure, uh, and was in Hugging Face's, um, like environment, uh, for like a week or something. Uh, Hugging Face identified it, uh, then eventually, uh, OpenAI also identified it, uh, I think independently, but then they corroborated. [00:02:04] Jacob Haimes: Um, and then they announced it. Uh, and then, you know, people were like, "Uh-oh." And, uh, you know, then after that we had the, um... Anthropic eventually said, "Oh, that, you know, we did it, we did that too." Uh, and then now we've also had Meta and some other people say, "Oh, we also did that." And Igor, you mentioned earlier that there's like a leaderboard or something now, which is, of course there is. [00:02:33] Jacob Haimes: Um [00:02:34] Igor: I, I saw it on Twitter. I couldn't find, uh, it, uh, uh, but like, um, basically if you're, um, if you're a responsible and safety-focused frontier lab and your, uh, AI agent hasn't escaped yet and hacked s- uh, something out, uh, out in, in the wild, then, then you're behind the curve now and you will not make it. [00:02:53] Igor: That's the current, the, the current, uh, like top le- level, like SOTA score that people care about [00:03:01] Jacob Haimes: Great. Um, and then there was also, uh, recently, uh, like not a Frontier t- uh, model being deployed by, you know, the labs during a cybersecurity evaluation, uh, thing. But, uh, a user in Australia was using OpenClaw and, uh, it, like, essentially used, uh, an exploit to, uh, remove a different, uh, like, person for a gym, uh, appointment or something like that, uh, and place him there instead. [00:03:38] Jacob Haimes: Um, and so that has happened. And then also UKAC corroborated some of the stuff from, like, OpenAI, uh, and, like, breaking out of environments and things like that. So yeah, a bunch of, bunch of, uh, new models, uh, breaking out of their containments [00:03:56] Igor: only corroborated that, that it broke out like it, they also had, I think, their model break out, if I remember this correctly [00:04:05] Jacob Haimes: UK AC doesn't have their own model [00:04:08] Igor: L- but like they have their own eval and they're like, like, like they, like they were doing, an independent eval and it independently escaped and they contained it b- I think, I think better, than OpenAI. I'm gonna look at the report that they had [00:04:22] Jacob Haimes: Okay. Yeah. So, but yeah, I-- sorry. When I say corroborated, I mean, like, this is their own... It, it's not like, "Oh, we saw the, the logs and it did that." It was like we were doing a separate thing. There was a different instance where something very similar happened. Um, yeah, so that's the super high level, uh, of the situation. [00:04:43] Jacob Haimes: And as always, I'm, I'm sure you're curious to hear our takes. Um, so maybe Igor, you can start off and then we can go from there. [00:04:55] Igor: Jacob's letting me run, uh, into the knife in front of him so, uh, so, so he can look nuanced and take. I [00:05:02] Jacob Haimes: Sure [00:05:02] Igor: like, first, first of all, I think it's good to say, like, w-what the response is across, uh, other media as well, it has ranged from basically people putting out, uh, memes and going full hysteria. Uh, I think, uh, some, uh, of Amira Crowds put out like a TikToker speaking baby talk about AI bad, AI kill everyone, uh, if, if, if, if we, if we don't, uh, uh, stop it. Um, Zvi was going through the, um, and timing as it came out, and I think had like a rational take, but like a very, like, AGI believer take of, uh, of like how, how bad things will get. Um, The Guardian had a good article where basically it told people to like be skeptical because is in a, in a weird way great PR for them. And I think what I'm missing so far is like- What the fuck w-was everyone working on security and infrastructure doing at these companies? Because I work at a startup, uh, still, like I, I'm gonna go on a sabbatical and then like, uh, change my role out, and startups usually struggle with like security and like I'm not gonna claim all the credit here. [00:06:40] Igor: I'm gonna claim a little bit of credit, but our security is better than this. Like, like we don't have like f- over-credentialed everywhere running on like, uh, nodes that are in active contact with, uh, like red teaming be-being done effectively. We don't run like three layers of, of cl- uh, uh, cloud security, uh, cloud software and then like reuse like, uh, credentials bet-between nodes. [00:07:11] Igor: Like there's basic like security playbook thing that, that was like violated here [00:07:21] Jacob Haimes: Right [00:07:21] Igor: nobody seems to be commenting on it like, like on, [00:07:24] Jacob Haimes: it's because, [00:07:24] Igor: side of it [00:07:25] Jacob Haimes: that's because the people who are, um, commenting on it are not security professionals. They are the media and ML, uh, machine learning people, or, or just like AI people, you know? I-- It, like, it, like, it would be, uh, offensive to the machine learning, uh, people if I called them machine learning people because it's, they're like AI people, you know? [00:07:51] Jacob Haimes: Um, but when you actually talk to security professionals, um, like... And, and, uh, at my current role with, with Heron, I've been talking to a lot more and, and getting a lot more context on sort of the perspective here. And, and my take is, um, yeah, they didn't have good security, but no one has good security. Uh, basically, like the, the whole internet is held together by safety pins, and the reason that it has sustained itself to this point is because no one cares enough to actually pull at those, like, junction points. [00:08:32] Jacob Haimes: And so when you get these systems that will do that, like this is sort of inevitable. Like, m-most people don't have good security. Um, like a, like way m- like way more than you think sh-should be reasonable don't have good security. And obviously, like, does it make sense that a gym in Australia, like a, a local gym in Australia doesn't have good security? [00:09:01] Jacob Haimes: Yeah, it makes a ton of sense. Um [00:09:04] Igor: we need to d- differentiate with local gym in Australia and, a- and like what happened at like these billions of like companies, right? The, the gym in Australia was like, there was an unsecured endpoint on probably like some home co- home brew app like that's like one thing. But like you can't just shrug and be just like, "Ah, yeah, sec- security is hard," when you have like billions of dollars and you're [00:09:24] Jacob Haimes: Oh, I'm not saying security is hard. [00:09:26] Igor: of dollars. [00:09:26] Jacob Haimes: I'm not, I'm not even saying security is hard though. I'm saying it's not done. I, I'm saying the, the... Yes, those are two different things, but I do think they're actually very related. And the reason that they're very related is that good, actually good security posture does not increase your return on investment. [00:09:49] Igor: Yeah. [00:09:50] Jacob Haimes: Th- the difference between... [00:09:52] Igor: Sorry. [00:09:53] Jacob Haimes: go ahead. [00:09:54] Cost Centers vs. Profit Centers (Praise for AISI) --- [00:09:54] Igor: The easy way to conceptualize this is, is, well, is like other terms, cost center and profit center. In companies, y- you always want to be in the profit center. put more money in this, you make m- more money. So usually this is, for example, sales. Hiring a sales engineer will increase your sales, which makes you more money, or making like a nice new feature is gonna make you more money. But security just, uh, stops like bad stuff from happening when it happens. And I actually want to also like give credit to AISI, like the UK guys. I'm looking at the timeline. Their thing is reasonable. their thing is like the reasonable amount of like failure that you can expect in, in this world. [00:10:36] Igor: Like, okay, software is very complicated, and having like a airtight security posture takes effort. They acknowledge also that they basically deferred doing that, and they know they didn't do, uh, do like the best practice thing, which is also like props to them. That's a good thing to put into like incident, uh, uh, report. And their thing was still only like thirty-four hours from like, uh, they noticed like, uh, the, the egress and like i- it's running, and then they basically have full shutdown and actually clear it out. Like it is [00:11:15] Jacob Haimes: Yeah, like that's, that seems it... [00:11:17] Igor: here [00:11:18] Jacob Haimes: And that is like a good example of better when, uh, you weren't necessarily expecting an outcome or, or a case like that. Uh, now maybe they should have been because, uh, of when it took place, but, you know, even so, I'd like... It's, it's better, better, better job. Um, and [00:11:37] Igor: uh, just before we start into a little story, like to, to say like why I'm so uh, intense about this and like, I think Jacob hasn't-- d- doesn't disagree with me, with me so much, but he's just like trying to convey like how much this is normalized. But like a good part of it is with security is to prepare for unexpected things. Like if your security posture i- is only working if it-- the stuff that you expected happens, it's a bad security posture. And security people will always have to do this, and then they will always get underfunded and under-resourced, and like people will tell them to, they can't do what they want because of the concerns Jacob is, is talking about. But for me, it's important to say, say, and I think you're also not disagreeing with me, but like we have a large normalization of deviance software where we expect the equivalent of a bridge down every day checking out why the concrete from this one company keeps building crashing bridges, or like why, uh, why like our industry practices keep messing up bridges. [00:12:52] Igor: That's like kind of like i- if this was, wasn't software but physical engineering, people would be in jail with the practices that are normalized because of business pressures [00:13:02] Jacob Haimes: I actually think the bridge analogy is very good and something that I've sort of come to recently, um, in how I think about software in comparison to real engineering, um, which I will, I will make that claim and I will stand by it. Software engineering is not real engineering. Um, and part of the reason is because you, in software at least historically, have not had, uh, constraints, uh, in the same way that you do in the real world. [00:13:29] Jacob Haimes: In the real world, you have physics and of course you have, you know, very small physics and very large physics and things are different in those regimes. But in the regime that we operate in, we are firmly within this realm where physics is sort of, uh, unapologetically there and we have to abide by those constraints [00:13:48] Igor: Small exception, there are regimes where there is actual software engineering. Usually those people use Spark or, uh, or other like contract things like the, um, aerospace, automotive, anything controller anything to do with like robotics and like, um, industrially used and like high, um, reliability engineering, like microcontrollers actually operates under like regulatory and physical constraints also like medical devices. B- uh, because like you only have maybe like a kilobyte of RAM to work with, and you need to make it work reliably under many circumstances, including like, like, like fail-safes [00:14:42] Jacob Haimes: Right, but that's, that, that's not what w- w- sorry, that's not what we're talking about when we say like, you know [00:14:47] Igor: Yeah, but, but like people will nitpick th-this if, if, if we talk too blanket. Like, like there are regimes of software engineering that do exist, but like what most people call software engineering is not. And like software development and like the software and the tech industry for the most part does not do engineering in the way that any other discipline do- uh, does it. Which also, before I hand the word back to you, is part of why it was so attractive in the last 20, 30 years. It was a very open w- open way that has the upsides and downsides of being very open and like not non-regulated. Um, but yeah, back to you [00:15:22] Jacob Haimes: Yeah. So You have these physical constraints that are, are not going to sort of give ever when you're talking about real engineering building a bridge. In software, in many cases, those just don't exist. So you can have something that, like, works usually until it utterly fails under one specific circumstance. [00:15:51] Jacob Haimes: Uh, and you didn't really go through the robust method of saying, "Oh, well, you know, given these assumptions, this can't fail in this way," because you don't need to because, you know, it, it probably won't happen. Um, and I feel like it's almost like when we have access to a system which can exploit these, um, sort of poor design choices, so to speak, uh, which w- aren't the fault of the people who are creating it necessarily. [00:16:26] Jacob Haimes: Um, but these poor design choices can be exploited easily, and then, you know, you essentially need to adapt and, and actually start adhering to more rigorous constraints. And I feel like that is something that will happen eventually. I think we're currently in the time period where, um We have all these legacy systems that don't meet those constraints and, like, w- because of that are very susceptible to these sys... [00:17:00] Jacob Haimes: And, like, I don't think this is new. I don't, I don't think that-- I'm sure that the person that was using the Open Claw, uh, rig was not using the same model as the one that was in, uh, the, like, uh, evaluations where it broke out. And, like, it didn't require that because in most cases it, it doesn't, which to me is, is more of an indicator that, like, we should be prepared more, uh, and, and be thinking about... [00:17:35] Jacob Haimes: Like, I don't know. It, it, it feels like [00:17:37] "You Do Not Get to Fuck Up This Much" --- [00:17:37] Igor: I've, I, I, I would, um, like to do like two things if it's okay. One, like to explain again, because we did this in the Mythos already, but kind of like when we talk about security and hacking, like what it looks like in reality, because for most people it's very abstract, and I think it's, it like, it's good to show how basic it is, it often is. And kind of to distinguish between like the, hacking and the like script kiddy type thing that is like, like, like the bread and butter of actual criminals. And then also separately, um, I think What you're saying is all correct, but maybe this is like good to, uh, to do before I go into the little, uh, expose is like There are companies out there now and there's like startups out there now that have good security posture, that it's-- they're just-- it's a small team. They just have one guy who thinks really hard about how the security works. Everyone else just does what, uh, what this guy says, if it's too hard, they think about how to re-engineer it, then they don't leave people's data. shit is secure. People try to hack them, and they fail And this is a thing that is absolutely possible. There's also software that just never has any RCE and any other, uh, large bug that, that like gets, uh, exploited at scale. That also exists. So this is both possible. And if you're like billion-dollar companies are explicitly trying to derive like the moral authority to, stretch if not flaunt laws to break like the, the implicit social contract of the internet and other places to sell the fact that you're gonna make everyone out of a job, but somehow you're gonna make it out well, and you act the responsible ones, and therefore you should also get like special privileges in regards to compute and that other companies shouldn't have, and you m-make all these cases. You do not get to fuck up this much [00:19:59] Jacob Haimes: Yeah. And I, I guess when I say this, I, I guess part of this is like I feel some of this has already been happening, uh, with regards to different subdomains. So like, um, I would say that the cases where, you know, for example, OpenAI, um, like could have caught the, the person who was contemplating, um, be- like being a shooter before it happened in Canada. [00:20:26] Jacob Haimes: Like a wh- like that, that is also, that is a greater, uh, fuck up than, uh, th- these cybersecurity things. Uh, but like, uh, I guess it, it was more isolated or, or abstract for people in some way, and so like it, it doesn't... Uh, if that kind of thing happens and we're okay with that Then, I, I don't know, it kinda seems like to me, you know, it-- we're not holding them accountable for that in an appropriate way. [00:20:58] Jacob Haimes: B- uh, I guess we should just be better at, at having s- you know, security posture because clearly they don't have our best interests in, uh, sort of what they're doing. And I don't know. I, I just... It just feels to me like Once you've gotten to this point, if you're willing to s-say, "Oh, well, yeah, you know, we should have done XYZ better, but it's okay, we'll let them keep going," like, that's sort of missing the, the point here. [00:21:35] Jacob Haimes: Like, th-they're-- I don't know. I need to formulate that idea more, more concretely, but that is essentially the vibe of, of what I'm, I'm getting at here [00:21:47] Igor: There's now a linked to chatbots Wikipedia page and OpenAI was thinking about alerting the police, but decided not to Um, I think that's important details. And I agree with you. Um, I think it's interesting, but, um, at least in the tech bubble and also in the AI bubble, I haven't seen our angle discussed more. [00:22:17] Doomers, Utter Alignment Failure & How Hacks Actually Happen --- [00:22:17] Igor: I've seen a bit on like message boards by like, um, people who are maybe not so much in the bubble But like this part of like, um, okay, we, we just need to get our shit t-together, I haven't seen that much. But also like the v- they have lost all credibility to be responsible guardians or like, uh, wardens or whatever we want to call them [00:22:47] Jacob Haimes: I mean, this is just like a slightly different framing of some of the like more strongly like doomer versions, right? Uh, or, or am I misinterpreting this? I, like [00:23:02] Jacob Haimes: Do you see what I'm saying? [00:23:04] Igor: Like, [00:23:04] Jacob Haimes: so let-- [00:23:05] Igor: what [00:23:05] Jacob Haimes: So, like, let's say Zvi Mowshowitz or, or, or so is, is saying like, um, "Okay, so this indicates that, um, there is..." And the words he used are like utter alignment failure. But like, uh, sure. Maybe it was, uh, at least in my opinion, like it, it's not something that you can really aim for, uh, in the way that it is being suggested. [00:23:32] Jacob Haimes: But, um, let's like... Essentially what he's saying is the things that have been put together, the system that has been put together to try to keep this on track has completely failed and, um, that this is an, like, uh, an example of what's to come. And, uh, like, I don't know, I, I feel like I'm saying something pretty similar when you get down to the core [00:23:57] Igor: But, but like what are the res-- what, what is the suggestion? What is like the intervention? What is like the th- the thing now? Because like I feel like in your thing there's kind of, uh, uh, at least implicitly, uh, that there should be like a call to like take away the control and the res- responsibility of managing the stuff from the current frontier labs [00:24:22] Jacob Haimes: Sure. So but then that, that leads to like Bernie Sanders' call, uh, for like, what did it... I don't remember what exactly it was, but it was essentially like saying you guys have failed, right? Like, uh, that is just another way to frame what I'm saying with maybe like, uh... I, I don't wanna back the request because I don't remember exactly what it was. [00:24:50] Jacob Haimes: Um, but you... [00:24:52] Igor: to pause development [00:24:54] Jacob Haimes: Right. So like, uh, I, I, I would say it's [00:24:57] Jacob Haimes: Because of how, how that would be implemented and, and feasibility there, like I, I feel like a different approach where, uh, it's not just like, "Oh, big tech giant, please pause the thing," and is actually like more, um, involved would be better. Um [00:25:18] Igor: I mean, [00:25:18] Jacob Haimes: But yeah [00:25:19] Igor: I, I, I don't like this for like... What do you mean? Like, like, if there was political will Then the US government could go in there Tell everyone, tell everyone hands off keyboard. If you keep going, you will go on the terrorist list And that will mean anyone who keeps doing business with them, any compute provider, anyone is now the risk of basically being caught up in that or to basically press a button and cut them off. that is a thing that, uh, that, like, could, could do. Like, the feasibility only hinges on, on the political. D-do we agree with that? Or do you think there's, like, [00:25:59] Jacob Haimes: Yeah, yeah. I, I'm saying, but the call is not, at least in from how I heard about the c- the, the request is just saying, uh, to the developers, "Please do this." It's not saying to the people who, who could, you know, actually institute it, um, "We should do this." And I guess maybe that is just a precursor to that. [00:26:23] Jacob Haimes: Um, and like a w- a way to build towards, um, the political will to do, uh, what you were just saying. Um, but I think that [00:26:37] Jacob Haimes: It is unlikely any- anything will happen, um, as a result of that, uh, I, I guess request. Um And I d- I don't know how to actually execute on, like, what should be done as a result of these, uh, these things that's happened. 'Cause, like [00:27:01] Jacob Haimes: As, as much as I think that one should be, you know, skeptical and that this is a boon for the frontier model developers in, uh, like a PR sense, uh, it, it also doesn't I guess, like the, the underlying issue, which is like this happened, nothing's gonna happen because of it, um, is, is, it should also be seen as like a huge red flag [00:27:37] Igor: Like, is it, is it, is it a point of basically like, okay, we, we could be doing the clear intervention, but, uh, it's politically infeasible and I don't, I don't see it happening. [00:27:48] Jacob Haimes: Uh, I guess I [00:27:49] Igor: you don't know what, what else to d- uh, to ask for or to recommend? [00:27:54] Jacob Haimes: No, I th- I think I, I'm just more My instinct is to, is to like really disagree with, I guess like doomer, uh, perspective. And, and in this case, I feel like... A- and I guess this is generally true as well, but like my failure mode and, uh, the failure mode that is presenting itself right now that they're, uh, up in arms about are similar, which is that like People don't understand the consequences of these systems. [00:28:25] Jacob Haimes: They just sort of, uh, release them, uh, and then, you know, bad things happen because people aren't, uh being more purposeful with, with how systems are deployed. Um, that doesn't mean that it's inevitable that this happens, which I, I think is closer to the doomer perspective. But, uh, it is, like based on what, what we are seeing happen, it seems that that is more likely, uh, because like it-- that's just how it's sort of unfolding [00:29:02] Igor: Yeah, but for me, I think the reason why I'm s-- I'm, I'm like pissed at this and I'm a bit more emotional than you is like there's this passivity in the whole discourse. Like this is all inevitable And that is why it's un- unfolding like this, where I don't know, but it feels like everyone is vibe coding their harnesses and vibe coding the, like, the, the infrastructure setup. And [00:29:32] Jacob Haimes: Yeah, I-- So like we could, we could do better at security, like the-- it is possible to do better at security, but it's not something that has been, like, has been considered needed up to this point. A- and so that's why we're, we're sort of running into-- I don't think addressing, addressing the security does not address the underlying problem. [00:30:03] Jacob Haimes: Um [00:30:04] Igor: like, like, like the thing that is annoying to me is that like, okay, like either these people were, were full of shit over the last two years when they, when they were talking about like, "This is so dangerous. The, the AGI is coming. We need to be very paranoid. We need to be very, very careful to, uh, to, to, uh, to not fuck things up." Or they were incompetent. [00:30:28] Jacob Haimes: Right, yeah [00:30:28] Igor: they were lying or they, they were in-in-in-incompetent or they, uh, like, the, like there's... Uh, because otherwise, like if, if you really think the, the, the AI and the AGI is coming and it's, uh, it's very close, like they keep saying publicly, [00:30:41] Jacob Haimes: then you would have been putting more effort into making sure that the systems are contained, which I, I think is also, is also fair. Um, a- and I guess just sort of demonstrates the dissonance between what people will say that they are behind and how they operate in their day-to-day Like people, people will say that they're, you know, like, "Oh, I think, you know, we're gonna be it AGI in, you know, two years or, or a year or six months or whatever. [00:31:13] Jacob Haimes: But then like their behavior does not meaningfully change as a result of that. And, and I feel like that happens relatively frequently, um, with, uh, people and I, I guess I'm not sure. May- maybe it's that like [00:31:30] Jacob Haimes: They don't actually believe that or they, they want to believe that, but I, I don't know. But like that's a meaningful trend, you know? [00:31:38] Igor: So like part of why I'm raging so much is because I think people should be raging more. But [00:31:43] Jacob Haimes: Sure [00:31:44] Igor: like if we all kind of think, "Eh, what, what do we expect? I guess we'll do a thing," then like, uh, there's no chance at this getting better. But like I, I, I ge- I get your point, uh, as well, but like... And we talked about this I think in the Mythos episode, but like, uh, maybe this is now the time for my like little, um, 'cause like the way hacks happen is basically always random probing. like there's a list of known exploits, there's a list of tech- techniques how to like obscure that you're doing exploit probing If you're, if you're already kind of like advanced, you combine the second list with a third list and try to hide that you're trying to hack somebody. But otherwise, it's, it's literally just like probes until you find something. And then the kind of like novelty comes from doing so in a way that, uh, evades monitors or that, uh, chains exploits. Um, and you use that in, in creative ways. But there's like templates on how to do this, similar to like competitive programming. Like if you get, one thing that I was, uh, involved in both me, I think SSRF is like a reflection technique that lets you like into, the internet by calling another URL, and then you can make calls with the service that is on the, on the host. So like instead of, of, uh, making like a, a post request yourself or like a get request yourself against a URL, make a different get request that has like a parameter that gives the URL that you want our service to call, and then you can basically make a proxy for that. And then using this to explore other things and to find other, uh, uh, places and to read into places that you should be able to read the creative bit basically. But it's also like once you've seen it once, it's like part of your toolbox. And The ingenuity of, uh, the agents was ba- basically, again, this like chaining together these things. So it's like a technical achievement, but it's not like they broke an encryption thing, uh, as often happens. Like this, this was not the case of like a novel z- uh, like zero day in, in the sense of, uh, of like a massive gaping hole. [00:34:21] Igor: It was like the of lots of small exploits that were being done, and those now all need to be fixed basically. It used to be that you could, you could maybe get away with like one or two of, uh, of your, um, uh, apps having like one or two exploits if it was small enough, but now that, that is gone, which is maybe one of the points that you're making. [00:34:46] Igor: But like, because this now and will be scaled up and like the posture will be to do, doing a lot more hands-off stuff, like used to be having like bot traffic was not a normal thing in a network. Now it's gonna be very normal that there's just like constant ongoing chatter the network, which means stuff like this can hide a lot easier [00:35:11] Jacob Haimes: Yeah, but so what, where is... Sorry, what was the relation to what we were talking about before and, like, why you're [00:35:21] Igor: So like, uh, for [00:35:22] Jacob Haimes: Upset [00:35:23] Igor: uh, uh, like you're like a kind of like with the doomers and accepting that this is just the way it's gonna be [00:35:29] Jacob Haimes: Oh no, I'm not for accepting at all. Um, I'm just saying [00:35:34] Jacob Haimes: It seems, it seems to me like [00:35:37] Jacob Haimes: The, the framing of like, this is the evidence that we have been waiting for, I think resonates with me, i- is what I'm saying. In, in a way that typically the doomer perspective does not. Um, and I, and I think that that is worth, um, like highlighting and, and, and bringing forward more. Uh, because I think to, to me at least, that means that potentially that, um [00:36:11] Jacob Haimes: It is more easy to, to demonstrate that, that, that this is, uh, or the, the direction that we're headed is not necessarily a good one. Um, but I- that does not mean that I'm saying that I'm okay with the direction or that I think this is what should happen, uh, or that I'm not upset. I guess, like, part of the maybe lack of anger, um, is because [00:36:41] Jacob Haimes: Part of my anger is, is taken or, or, or, um, maybe not anger, but you know, along those lines, part of my, uh [00:36:50] Jacob Haimes: Yeah. Focus is, is put towards the first cases of instances where systems weren't, you know, d- doing what the person wanted, which is like, uh, closer to the, you know, people who are being impacted by, um, negatively by like, uh, chatbot use in terms of a, a mental health context. So like, um, uh, or other things along those lines, like why didn't we care as much, um, when that was happening? [00:37:19] Jacob Haimes: Um, so part of it is anger about that, and part of it-- or yeah, w- focus on that, and part of it is focus and, and anger at the fact that like people weren't as upset when that happened. Um, and so I'm still very much upset about, about those things and, and want to make changes, and the same is true about this But like, I don't know, like my warning shot has already happened kind of, if that makes sense [00:37:53] "Forcing the End" and What Can Anyone Actually Do? --- [00:37:53] Igor: But was this your warning shot or what was your warning shot? [00:37:56] Jacob Haimes: No, like the, the, i- the instances of, of, uh, people being negatively impacted by chatbot use is, is, is what I'm... Like that was [00:38:10] Jacob Haimes: And so I, I'm [00:38:11] Jacob Haimes: Yeah, l- I like less outraged by this o- like specific instance. Um [00:38:17] Jacob Haimes: If that... Do, does that make sense? [00:38:20] Igor: Ja, ähm I, I got a bit reminded b- um, of a term forcing the end [00:38:28] Jacob Haimes: And what does that mean? [00:38:30] Igor: We discussed it, I think, at some point, um I am linking to like a source that discusses this a bit. I-i-in, in scholarship of, Like Communities with strong convictions that like something would happen, would have like a prophecy or like, like a, like a eschatological or apocalyptic tendency. Um, so, uh, the usual example is like milli-millennialist ideals that like, uh, or the prophet of choice will come back soon. Um, there's often a tendency that like the, the deeper pe-people get into it and the, the more they stake on that belief system, the more they will at some point like try to like make it come true. So, so like if like they might start off warning you that like the prophecy says that like, uh, the streets will, will rain in blood, uh, before the, the pro-prophet returns, then a subset of them will make that come true if it takes too long [00:39:34] Jacob Haimes: Yeah. Th- this is like the, the subset of hyper-Christian, uh, people who like, um... I think there's some- someone in Texas that's trying to, to breed like the, the right kind of cow that's talked about in, um, like, uh, the Bible and, and i- is the reason behind why like ultra-Christian groups in the US are pro-Israel. [00:40:02] Jacob Haimes: Um, you know, that sort of thing, right? Is that what you're referring to? [00:40:06] Igor: Yeah. And like it, it takes man-many diff-different shapes. Um, I think that is also part of what is m-my anger, but like These are the people like, uh, uh, this community talks, talks about and you, you say you now have like a connection, uh, with it or you get it a bit more, but l- uh, but y- maybe you, your warning should already happen, but like y- you can what they're talking about basically, which is like we, we all knew this was gonna happen. And what [00:40:36] Jacob Haimes: No, not quite, but sorry, please continue. I can correct it afterwards [00:40:42] Igor: Correct me now because I think it, it's important for my, for my point that, uh, I get this right [00:40:47] Jacob Haimes: Um So I'm not a doomer now, haven't been. Um [00:40:55] Jacob Haimes: More just saying like [00:40:57] Jacob Haimes: And I, maybe this isn't that different, but like I see, I see how they got there. Uh, i- is maybe that, and like if, if there were certain assumptions a- and, uh, beliefs that I had a- about, um These systems, uh, and, and AI in general, like, uh, I, I could very easily see, um, you know, arriving at the same conclusion [00:41:24] Igor: So what I mean, meant was, was like, you know, like you can like empathize more now than in other, uh, other cases basic- basic- uh, uh, ba- uh, basically [00:41:33] Jacob Haimes: Uh, y- yeah, I can empathize more now, um, in this specific instance. There are al- there's still, like, a, a whole bunch of things that I don't think that's, uh, the case with, but I f- I do feel like the, um, this is an example of a thing that we should be worried about, uh, is valid [00:41:56] Igor: And for me, the problem with that is like the only reason this is an issue is because we, these people are so worried about it that they train the model up to the peak of their ability and when they do evals on it, then they fuck it up. Li-li-like as of, uh, as of right now, like the utility of, uh, of these, uh, systems to do cybersecurity has only happened because the frontier, uh, models, but which are the, the only ones with the capabilities of tr- uh, of, of training them s- have trained them to be good at it and then they fuck up the evals. Where like you can argue if you want to and I don't think, uh, it's true, but this like if we don't do it somebody else will. I don't think this is true, even assume you do it then you should at least like damn sure like do a better job at it and, and [00:42:53] Jacob Haimes: Yeah [00:42:54] Igor: for me. And like what is also missing is, is kind of like what will you do about it? now [00:43:00] Jacob Haimes: Yeah [00:43:01] Igor: thing, uh, that has gotten worse, like what will you do about it? I'm not seeing any reaction. Just like business as usual. We, we continue the same trajectory. There's no adjustment. There's only like more money into more evals which even though like the risk happened during the evals right now. Like I know people doing gain of function research in order to assess the risk. Like [00:43:25] Jacob Haimes: Yeah, I, I think that this is [00:43:27] Jacob Haimes: Correct. I, I guess the, the people that I'm thinking of, like, th- there are definitely people who are, who are doing this, uh, and have the, like, uh, take advantage of the perspective here and, like, end up, uh, being the, "Oh, well, we'll build it first," kind of perspective. There are also people who are, are very much trying to, you know, stop or, or, or, uh, alter the course. It's just they don't have the money backing them. and [00:43:58] Jacob Haimes: Yeah, so I'm, I'm not talking about everyone, uh, within the space. More, more like the, um, theoretical ideology. Um, and then... Or, or like perspective. Um, but I, I very much agree that [00:44:15] Jacob Haimes: It doesn't seem like there's any course correction that is being, uh, discussed in a meaningful way, and that seems like a problem in itself. Um So how do, well, so then how, well, what are we gonna do? Uh, how, how are we gonna, like, do... 'Cause I mean, we can say this and we can put out the podcast and that's great and all, but like, you know, uh, if the, of the people listening to this, uh, that's not that many. [00:44:45] Jacob Haimes: Even in, with the most, uh, generous, um, numbers that we have, which isn't, you know, horrible, but [00:44:54] Igor: Is it, is it, is it a question for me or, uh, do you have your own thing that you want to [00:44:59] Jacob Haimes: No, this is a question. This is an open... I don't, I don't know. I don't, I don't know, like [00:45:03] Igor: I, I mean, I know, I know somebody who works in an AI safety org focused on security. Like, uh, maybe it's like a, a, a, a, a good, uh, place to like, uh, plug that right now [00:45:14] Jacob Haimes: Yeah, but, but like, uh, even as the person that Igor's referring to, what do I, what do I do? Uh Like [00:45:23] Jacob Haimes: Do I just, uh, 'cause like the thing that we're, that the organization I work for is already doing is trying to get people who actually have security experience into the, the role so that they can actually do this kind of work. And that like, yeah, that makes sense and that seems good. But that only goes so far. [00:45:38] Jacob Haimes: Um [00:45:40] Igor: Well, like, part of the reason why, like, I, I do, I do the podcast for me is kind of like to, like, bring this perspective of, like, uh, this passivity that I talked about earlier, but also this kind of like self-fulfilling prophecy is a dynamic that I think is worth breaking. That, like, this idea that it's gonna happen anyway, therefore we need to prepare for it and then we, uh, we need to understand how it could happen, so we need to do it. [00:46:07] Igor: Like, if you see yourself doing this, like, people should, like, stop and they should, like, argue with people that they know th- gently, like, maybe they should also, like, like, don't be like me, may- uh, I'm not that efficient at convincing people in this direction, sadly. I'm working on that. but I think that's, like, a good thing. I think that, like, now that I will have more time, I will be doing more as, like, writing posts but also engaging with other people on, like, the social media and, like, trying to have, like, one-on-one personal conversations on, uh, on this if you're in the bubble and, know, like, if you're in, uh, at work and the water cooler. [00:46:49] Igor: There's, like, grassroots things that you can do that I, I do believe ha- have an impact even if it's like a, like a trickle. And then if, uh, if you are looking for a PhD topic, if you know somebody looking for a, a PhD to- a PhD topic or a master thesis topic or, like, you're gonna be joining, like, the MATS or the other pr- uh, programs and, like, maybe go into this directions. [00:47:14] Igor: Like, this, this might be the Hans, are we the baddies, uh, like, mo-moment for people who know the Mitchell and Webb, uh, gag of, like, if you go into this field, AI safety, maybe bring that into the reading groups that, uh, that you're in and, like, sensitize people to, "Hmm, maybe we shouldn't be doing this," so that, like, the support base grows. [00:47:34] Igor's Pivot to Open-Weight Models and Steelmanning the Case Against Them --- [00:47:34] Igor: And then you kn- if you are in the room for bigger decisions, maybe, like, funnel some money towards Jacob's org, but also, like, funnel some money independently towards, um- A, the security department and B, just like actually questioning like what you're doing if you're a, a convicted person. And then, um, I know that for me, there was a personal update where like, um, I was kinda like Trying to be not fully YOLO with using AI agents, but I was a bit like, uh, more trustworthy that like, um, and OpenAI are at least like putting the best effort in to not fuck up. That has now gone. I am moving all of my agentic coding into like proper sandboxes and VMs and, uh, separately running hosts that are far away from my personal stuff. Like that's, uh, just an update I think people should do that like you should probably consider everything that you put into the LLM risk of being leaked or compromised in some way. Um, and you should probably also, if you are like a normal engineering person, like, um, push harder for self-hosted models and people should like push hard to defend open weights models now because they- they're gonna be [00:49:05] Jacob Haimes: Actually, that's a, that's a... Sorry, they're gonna be used for what? I, I w- I wanna... I was, I was getting ready to close it, but I think we should talk about this for a second [00:49:15] Igor: they go- they, they, they're gonna, they're gonna use this incident as an example for why it's important to clamp down on them. But like Hacking Face also like pointed out that like Kimi K3 was helping them like diagnose like, uh, the OpenAI hack. And if you want to make sure that like whatever system you are running is properly secured like you know the way it fucks up, like having your model that you have personally benchmarked, that you have personally like locked down in the, in the right way and your own harness is gonna be probably critical [00:49:53] Jacob Haimes: Yeah, I think that's, that's fair. So, so coming from the Um, like let's just walk through the, the argument against the open-weight model paradigm that is related to this, uh, instance to make sure we're, we're both on the same page and talking about the same thing. Um, the idea being, well, open-weight models are, you know, a couple months behind frontier models for, um, various reasons and, um, so that means in a couple months, uh, or potentially already and people just haven't, uh, sort of figured out the best ways to, uh, sort of set up the scaffolding, which i- is what I would probably argue. [00:50:37] Jacob Haimes: Um, these systems can be used to, you know, really break out of, of sandboxes and, um, start to cause more large scale, I, I guess like cyber chaos, uh, is maybe a way to, to like phrase it. Like, you know, if every organization or n- not every organization, even if, you know, 20% of organizations that are, you know, real functioning organizations at this point in time have security instances that they now have to shore up, um, or their scheduling becomes invalid, for example. [00:51:22] Jacob Haimes: Like, like that is going to cause, um, significant friction in society as, as we know it. Uh, and, and so then that would be bad if everyone had access to these systems, um, and so therefore we should prevent people from using these systems and accessing these systems, um, without any sort of, uh, guardrails on them. [00:51:49] Jacob Haimes: I- is that-- Did, does that roughly sound correct to you, Igor? Did I miss something? [00:51:54] Igor: Ja [00:51:55] Igor: You ask what we can do, like call your representative or write them and say that like having access to open weight models is critical for you and your business and your security. That's like a legit thing [00:52:06] Jacob Haimes: That is, that is actually, that is actually reasonable. Um, but, a-and also, I mean, look up, look at your current legislation 'cause it, there's, there's a very reasonable possibility that there are, um, either, you know, bills or, or, or laws or, or it's whatever, at whatever level you're looking at being drafted that are like, uh, have, um, statements about open-weight models. [00:52:34] Jacob Haimes: Uh, and if that's the case, then it's actually even, uh, better for you to contact, uh, your representatives because then you can point to a specific thing that is in the works and say, "Hey, I like or I don't like this thing about this thing that already exists and should be in y-your radar." Um, that's very valuable for actually sort of conveying things to, to Congress, uh, or people or other representatives. [00:52:59] Jacob Haimes: Um- [00:53:00] Jacob Haimes: But like, I, I [00:53:01] Jacob Haimes: Are you worried about this? I guess the question that I, I'm more interested in here is, like, are you worried about this at all? Or are you more just like, " Yeah, this is not good, but it's also one of the better defenses, and it would be much worse if, uh A single company were to, you know, have the ability to, uh, I guess, completely lock out, um, others for like entrenchment, uh, arguments [00:53:38] Igor: Uh, you're asking if I'm worried about the availability of open weight models making it easier for, like, the bad people at all, if that worries me at all, or if I'm just like, um, uh, accepting the, the upside, uh, uh, the downside with upside, like [00:53:54] Jacob Haimes: Yeah. I, I guess I'm j- I'm just wondering a little bit more about, like, your, how you think about the trade-offs here, um [00:54:02] Igor: So my take is that, that like the vast majority of, um, like, uh, criminals already have access to frontier models to do all of, uh, the bad stuff via jailbreaks, um, and stolen credit cards and token reselling. Um, and that long as you have like, like if OpenAI risks shut down and jail time for executives if they don't get their shit together as in there's like a strong incentive to it, then, then I might discuss this. Like KYC laws for money laundering shut down businesses. It's actually like, like, like, like a big thing right now of like it, it being like an overdue burden for small companies. That also has an effect, like money laundering is legitimately harder now. It's still not perfect and like there's still like ways around it, which is like a whole other debate of if it's, if it's worth it then. But KYC laws work Having OpenAI like this semi-trustworthy, only allowed, uh, to have the, uh, the good stuff entity with no proper oversight and no regulation, and then not having open weight models does not meaningfully increase like, uh, security at all. Like if you, if you want to make, um, OpenAI like a, a, an infrastructure, like a utility, then regulate it as that. there's a reason why, why banks don't, don't trade at the, um, earnings at, at, uh, and to share multiple, uh, or fraction of, uh, the tech companies do because they're, they're bad businesses in, in, in some ways. uh, but like you can't have it both ways. You can't have like the move fast and break things, uh, things of a, of a, of a tech company and then claim that you're doing like critical infrastructure services for, for security. And so going back to your question, because North Koreans use ChatGPT and Gemini and all the other stuff and have zero problems, uh, getting access there because you can't meaningfully differentiate them, uh, from like, uh, normal customers and jailbreaks work, um- I don't think there's a meaningful increase in having open weight models it's not like running these open weight models is, like, very simple. [00:56:27] Igor: Like, you need to have, like, large amounts of j- of, of GPU compute, and OS4 is b- it's much easier to track, like, large amounts of uh, running, and it's much easier to basically say like, "Hey, OpenRouter, you need to like, uh, KYC people." Like reselling is fine, but you need to KYC your, your vendors and, and your customers. [00:56:48] Igor: Done. Like, uh, want to run your own models, you can Like happens if you, if, if you outlaw open weight models is in three years the ca- the cartels will have data centers and their own ML team [00:57:07] Jacob Haimes: Yeah [00:57:07] Igor: it, it's not that hard to train a useful targeted model [00:57:14] Jacob Haimes: Sure. Yeah, yeah, yeah. No, that, uh, that makes sense [00:57:17] Igor: of answer that l- uh, that like, um, in a world where we actually massively shut down access to a, to language models we destroy the booming AI economy in the name of safety, [00:57:31] Jacob Haimes: Or it destroys itself. [00:57:33] Igor: Sorry? [00:57:34] Jacob Haimes: I said, or it destroys itself [00:57:36] Igor: No, I, I mean, like, like, uh, in this world, I would maybe accept the argument of like more safety f- via not having open weight models. [00:57:43] Jacob Haimes: Oh, okay. I see what you're saying [00:57:45] Igor: but like in the world where we are in right now, where we don't do v- those basic safeguards, like anything that, that goes against open weight is, is basically, uh, just making an argument against, uh, competition for the incumbent that would like to be the only person selling the North Koreans', uh, tokens [00:58:01] Outro & Sign-off --- [00:58:01] Jacob Haimes: Gotcha. Yeah, I mean, I, I think that makes sense. Um Cool. Well, uh, you know, that's the-- I'm sure that there's, that's, there's a whole can of worms there to unpack [00:58:13] Igor: But we do it next episode [00:58:14] Jacob Haimes: At some other time, yeah. Maybe next episode, maybe not, maybe further down the line. Uh, but regardless, that is all the muck that we can stand for this week. [00:58:23] Jacob Haimes: If you liked our takes, uh, please, you know, let us know, share this around and, uh, do some thinking [00:58:31] Igor: like our takes, pl-- uh, let us know as well. My email is easy to find. Email me why I'm wrong [00:58:36] Jacob Haimes: Yeah. Awesome. Well, we'll see you next time [00:58:40] Igor: Have a good one