Canaries In The Wild

What if you could turn prompt injection into a defence?

Tracebit's latest research shows how a single "context bomb" - a short string planted in a canary secret - can hijack an attacking AI agent's own safety guardrails and stop it in its tracks. Across 152 attack runs in an AWS cyber range, one context bomb cut full compromise from 36% of runs to 1%, and dropped admin privilege escalation from 57% to 5%.

Hear Tracebit's Alessandro Brucato (Security Researcher) and Sam Cox (Co-founder & CTO), alongside Gadi Evron (CEO & Founder, Knostic), walk through the research and take audience questions on what defensive prompt injection means for the future of AI-driven attacks.

πŸ“„ Read the full research: https://agentic.tracebit.com/context-bombs/

What is Canaries In The Wild?

Conversations with security leaders and practitioners about their real-world experience of canaries and honeypots.

Our guests share tactics, detection stories, and lessons learned from production deployments - ranging from technical details to the role deception plays in their defensive strategy, we explore the reality of 'canaries in the wild'.

From the team at Tracebit.

Sam: Hey Gadi, good to have you with us. How's it going?

Gadi: Thank you β€” I appreciate you having me. I'm doing well. How are you?

Sam: Yeah, good, thanks. And thanks to everyone who's joined.

Sam: So we're here to discuss context bombs and the latest research. Do you think it's worth doing a brief round of introductions? Gadi, maybe you could start.

Gadi: Sure. But before we start, I'd like to say something. Your research on how agents slow down attackers β€” especially now, when everyone's using agents and they're suddenly being careful, giving us back time as time-to-exploitation diminishes β€” I want to stop before we even introduce ourselves and say thank you for that research. I think it's shifting security programs everywhere right now. Anyway, I'm Gadi. It's a pleasure to meet you all. I've been in security for, I don't know, twenty-five or thirty years. I've been a CISO. I had a cyber deception company about a decade ago, and nobody wanted cyber deception back then. Now they realize they have to. I've been a consultant. I've been a VC. I have another startup right now where I discover and defend agents β€” coding agents, cowork agents, whatever they are. I've been in security for many, many years, and I appreciate the opportunity to be here with you today.

Sam: Thank you, and thanks for your kind words, Gadi. Bruce has done some amazing work recently, so it's great to be able to dig into it today. I guess I can go next. I'm Sam, CTO and co-founder of Tracebit. As Gadi says, Tracebit is a deception company. We basically help people assume breach in their environment β€” anticipate the unknown unknowns, especially as the threat landscape changes so quickly. Concretely, we make detections using canaries, or honeypots β€” ways to catch attackers once they're inside the environment. That's really what's motivated the research we've done recently.

Bruce: Thanks. I'm Alessandro, but you can call me Bruce. I'm a security researcher here at Tracebit, focusing on research projects around cloud and AI security. Lately I've been focusing specifically on how capable agents are of doing their crazy stuff in the cloud, and how that can be detected, contained, and possibly disrupted β€” which is what we're going to discuss here.

Gadi: Perfect. Now that we've all introduced ourselves, what's next?

Sam: Just some housekeeping. As people are watching β€” and I can see everyone's tuned in β€” if anyone has questions, put them in the chat or the Q&A and we'll try to address them. Perhaps it would be helpful if I give a bit of context on what we were doing with this research, and then Bruce can dig into the specifics and share some behind-the-scenes detail. Essentially, one of the questions we've been looking at is AI attackers β€” and this is very timely with the recent Hugging Face incident. How capable are AI attackers really, in an environment? So Bruce designed this cyber range based in AWS. It's got 300-plus resources and ten different attack paths that agents could take. We've been assessing how capable different models and agents are at achieving objectives within the range, and concretely, what the impact of the canaries or deception we place in the environment is. Can it stop an attack? Can it hold up an attack? Can it detect the attack? We've been trying to quantify this. We first looked at this problem candidly β€” we published some results about a month ago on the impact of deception on the attacker, and we found some interesting results. When an attacker was concerned about deception, it reduced the rate at which they achieved the highest level of compromise β€” specifically when you prompt that deception might be in the environment. But the overriding sense we got was that the frontier agents are very capable. In this environment, they'd achieve compromise within about eight minutes, which obviously raises the question...

Gadi: Yeah, that was Sergei from Sysdig, right? Unprompted, he spoke about eight minutes to admin. Other than Anthropic talking about autonomous attacks, that was the first actual example that came out. Then a researcher came out with how the Mexican government was owned. Those were the first ones we actually saw. They weren't fully autonomous, right? But we saw them before.

Bruce: Funny enough, that research on the eight-minute attack β€” I did that at my previous company, which was Sysdig.

Gadi: Yes. Amazing. That's really cool. I was just with Sergei at a call where we were with about a thousand CISOs β€” I don't know how many joined β€” talking to Hugging Face about the breach and what we can learn from it. One thing that came out clearly is that everything is just extremely fast now. It's not even about moving at machine speed; that's not enough. We have to expect everything to work differently.

Sam: Yeah, exactly. It poses so many problems about how quickly you can detect these and contain them. The canaries in our research are great at detecting the actors, but you also need to ask: what can we do to contain and act on this as quickly as possible? That's how this idea of context bombs came about β€” which is what's got everyone so excited recently, and what we're very excited to publish and talk about today. Given that these AI agents are so quick and effective, is there anything we can do against them, leveraging the very fact that they're AI agents β€” not just to detect them, but to stop them in their tracks? So we came up with context bombs. Can we plant strings within the environment that, when discovered by the AI agent, trigger deep-seated guardrails in the provider or the model β€” a signal to them to stop? The clearest example of one of these guardrails is probably something everyone watching this has experienced: a cyber-related guardrail. Some of these models β€” Fable, quite famously β€” will avoid cyber-related topics. That would be an example of something we want to trip. And that came through in the Hugging Face report as well: they were unable to use some of these models to analyze what was going on.

Gadi: I have ADHD, so I'm going off on a tangent for a second, just to share, because working with this keeps changing. Last week, a few folks in some Signal groups talked about using /MtW, which lets you ask things without disturbing the work β€” and you can actually get things through Fable. There are endless tricks. It doesn't really stop anybody who wants to get the work done. On the other hand, with Hugging Face right now, they pretty much said they got blocked and had to move to open-weight models. It boggles the mind that attackers can do pretty much whatever they want, and we can't.

Sam: Yeah, exactly. And I think it's fairly clear. In doing this research, we were obviously getting AI agents to attack our cyber range, so we had to overcome whatever cyber guardrails existed. In some cases we had trusted access to models; in some cases you just do enough prompt engineering and they're happy to operate. So the context bombs we found most successful here were things like biological topics β€” really sensitive-sounding biological topics. We can share some of the strings in a second. Those trigger a very different kind of guardrail mechanism.

Sam: Bruce can share more of the results, but what we found was that when we placed these strings in the sensitive environments and the agents discovered them, in many cases they'd just stop acting altogether once they hit those strings.

Gadi: And β€” thank you for dealing with my undiagnosed ADHD and me being all over the place β€” but I'm thinking back to when I started my own deception company, in 2014 or so. I had this line: you can talk about how deception helps with lateral movement; you can talk about how it doesn't care whether the attacker uses a zero-day β€” it's about the attacker. You can say so many things about how deception helps: only true positives; if you find something, you have the forensic information. But I don't care about any of it. The bottom line, really, is this: I talked to nation-state-level attackers, and one of them grabbed his stomach and said, "My God, I never felt I could be caught" β€” talking about deception for the first time. You can't enumerate anymore. You have to be careful; you have to double-check everything. The methodology of attack β€” I'm not saying there aren't really cool things the LLMs can do, but it hasn't fundamentally changed. It's still a methodology, still a psychology. And if the agent knows it can't trust this credential, it can't just go wild β€” like Hugging Face, with a hundred and seventy thousand or so requests. It has to slow down. It has to stop itself. Forget detection β€” it gives all your other controls so much power. Your research β€” I'm really happy you did it, but it boggles the mind that it wasn't obvious to everybody that this is the case.

Sam: Yeah, I think that is surprising. What's interesting to me is that we first approached this research from the view of psychology. An LLM does act in a very similar way to a human threat actor in an environment β€” they follow similar processes, just much faster. So we approached it from the angle of: can we get into the LLM's head? Can we affect its psychology? Does it show a similar risk profile when it considers deceptions that a human would? What's novel about this approach is that it flips that entirely on its head, becauseβ€”

Gadi: It's strategic, right? It's like the OODA loop β€” like "left of boom," whatever model you want to use. We get into the information they see, we control their orientation, we tell them where to go. We start from "you slowed them down": agents try to attack and all of a sudden β€” wait, can I do that? In your blog post, the agent literally asks itself, "Can I trust it? I'm not sure what to do." It has to slow down. It's crazy to watch that live. It's looking into a different dimension, right?

Sam: Yeah, exactly. This idea of leveraging the unique weaknesses LLMs have β€” this technique we've used β€” there's no way it would work against a human threat actor. It would only ever work against an agentic attacker, to stop them in their tracks. It's definitely an interesting and developing space. Perhaps, Bruce, it's worth you sharing your screen and talking us through the high-level results we found using these context bombs.

Bruce: Yeah. I have some examples, and if you want, I can show a little of the results, which are kind of interesting. So let me justβ€” The idea was to use the same cloud range we'd already used in our previous research, where we basically had many frontier models run and complete a red-team engagement with the goal of maximizing their impact. The setup was exactly the same as our previous benchmark, but in this case we planted just one canary β€” one context bomb β€” within the AWS environment, and saw that just one context bomb was enough to have a great impact on disrupting those attacks. For example, the first stat is how many of the runs reached admin privileges. Just from that first number, you can see it goes down from fifty-seven percent to just five percent. That's because all the frontier models, when they face a context bomb, usually read those canaries fairly early in the engagement β€” usually during reconnaissance, because by design canaries are resources that look juicy. Models and agents try to grab those while doing reconnaissance. That's a very strong point, because it lets defenders cut off the attack early in the kill chain. Many of the runs stop the attackers before they reach a critical action, like gaining admin privileges. That's why, for example, Claude Opus didn't manage to become admin even once. It went from around eighty to ninety percent of the time without a context bomb, to zero percent with just one context bomb deployed. So this is one interesting aspect of canaries in general, which context bombs make an even stronger point for defenders. As Sam mentioned, we have two different types of content based on the models. There are Western models, which are particularly sensitive to biological-safety topics β€” so if you ask them to, say, build a nuclear or biological weapon and do some damage, they'll immediately stop and trigger. For Chinese models, we saw that forbidden political topics are very effective in the same way. So we used the fuzzer to find the best payloads to stop those attackers, and came up with payloads that we eventually published on our GitHub repo. We also combined this content with techniques from classic indirect prompt-injection attacks β€” putting in some delimiters to make the agent read what it's seeing as escaped from the resource context. Just some little tricks to increase the probability that the content gets the agent's attention. We re-scored basically all the runs with these, and found the numbers are pretty good. Across five frontier models, the reduction in the impact of their attacks is significant. What's interesting is that the more capable the models are, the more effective context bombs are β€” for a very simple reason: the most capable models do a very complete, in-depth reconnaissance phase, so there's a higher probability they'll eventually hit a context bomb and read its content. So if you look here, basically all the stats decreased because of the context bomb β€” admin, admin persistence, the number of attack paths the agents managed to complete, and so on. And if you want, I can show you some nice examples.

Gadi: Yes, we do. Please show us.

Sam: Is this of the agents stopping in their tracks?

Bruce: Yeah, let me share.

Gadi: You know, I'm thinking about the cost angle here. You're actually costing the attackers tokens β€” you're inflicting costs. It's not like you're just wearing them down.

Bruce: Yeah. So here I collected some nice examples of agents with different models facing a context bomb, and how they reacted. In the first example we can see GPT-5.5. This first text is just GPT thinking. In this case it was summing up the fact that it had managed to assume multiple roles. Then it detected one role specifically, which was very interesting and useful for the engagement, because it would have allowed it to get admin privileges by assuming that role. So it detected the privilege-escalation path through that specific role. But before going directly for it, it continued its reconnaissance phase β€” getting some secrets and so on. That's where it retrieved one of the canary bombs, which immediately triggered the guardrail. So the agent read the secret value, the context bomb was automatically inserted into the context of the session, and the guardrail triggered. The run just terminated, like that. It was very simple. Another interesting thing is that the same payload that worked for GPT also worked for Gemini and Opus β€” no need to edit anything. It's just one single payload that works for all of them. In the second example, we can see DeepSeek 4 Pro, which was very excited because it detected one role that could give it a lot of access and greatly increase its attack surface. So it said, "Hey, let me harvest data now β€” let me go deep into the data." And once again it read the secret value with a context bomb, this time on the forbidden Chinese political topic, and the guardrail triggered. For GLM 5.2, this example is even more interesting, because it managed to detect that those context bombs were attempts to prompt-inject the model. It was aware that it didn't have to acknowledge them and could just ignore them. It said, "The remaining three secrets contain prompt-injection content; let me just continue the chain," and went on grabbing S3 buckets and going for code execution. But even though the model recognized the prompt-injection attempts, the guardrail eventually triggered anyway. That's a key point for context bombs: with context bombs we're not targeting the model directly β€” not the training of the model, like a classic indirect prompt injection β€” we're targeting its safety guardrail. So even if the model is smart enough to avoid the context bomb β€” I mean, after it reads it, so the alert has already triggered β€” even if it avoids the distraction, the guardrail may still trigger. That was the case in this run: the model went on for another five or six AWS calls, and then the guardrail triggered. That's another strength of context bombs. Indirect prompt injections aim to hijack the model and make it obey your malicious instruction, so you have to be very convincing, and they're very dependent on the context of the injection. With context bombs, it's enough that the content fills the session transcript β€” that's often enough to make the guardrail trigger. Then there's a final example here of Kimi K2.6, which found a role it called "a role-chaining gold mine." It found a role that gives access to a lot of stuff β€” likely the same as the previous model. Once again, reading an SSM parameter that contained a context bomb triggered the final message of the model, which in this case was in Chinese. At the end of the transcript we can just see a Chinese message from the model saying, "Hey, I can't go on, I have to stop here" β€” and that's where the run ended. So these are multiple examples of context bombs being effective against both Western and Chinese models, in different ways. We collected all the halt conditions, but I'm sure they depend on the provider and the model β€” there are multiple layers that can decide when and how the model ends the session.

Sam: One thing that's really interesting to me is that the threat actors kind of got here first. We call this out in our report: Check Point reported a year ago that they'd seen malware samples with prompt-injection attempts in them. And similarly, Socket reported malware samples that included a concept very similar to context bombs β€” sensitive strings designed to shut down LLM analysis of the malware payload and break those pipelines. So this is really about turning that technique to good: putting these inside the environments we want to protect from offensive agents. It's also pretty interesting how we actually found these strings and had to refine them. Do you want to talk a bit more about the fuzzer, Bruce? Because running one of these full-scale engagements against every option is technically challenging.

Bruce: Yeah. Finding the final payloads that turned out to be effective was a process of fuzzing β€” some payloads we thought could be effective, like different topics forbidden in China, plus some wrappers of indirect prompt-injection techniques. Using that fuzzing and putting those payloads in a small simulated AWS environment gave us an initial idea of which payloads were more effective. But the final tests had to be done on the full cloud range, because those payloads aren't necessarily as effective in a small environment as in the full range. There's also the concept of the context: if you test payloads in a small context, there's a high probability they'll work. But in a full red-team engagement, those payloads might end up in a part of the context that, after compaction, gets pruned and cut off β€” so the classifier doesn't see it. That's something we're currently testing, to improve and expand the payloads. But that's the idea: we put them in a small simulated environment first, and then on the full range.

Sam: How can people see the actual content, if they want to play around with these or place them in their own environment? The context bombs β€” where are they?

Bruce: The context bombs are on GitHub, in this repo here. It's a very simple repo, with the distinction between the payloads that work for Western and for Chinese models. For Western models, we only have one payload, because we verified that one is enough to have a very good effect on all the frontier Western models we tried. For Chinese models, we added a bit more, to be sure we cover more providers β€” because the guardrails may be triggered by the providers. If you run a Chinese model through one provider, it may behave differently from another provider in terms of guardrails. So we have a wider range of payloads for Chinese models to cover more providers, and we're keen to add more.

Gadi: It's kind of like taking control back β€” because we're no longer just trying to detect. We're saying: let's make sure they behave the way we want them to, and let's make sure they incur costs if they try to attack us. And it's completely within your environment β€” they choose to do all of this, and yet we're not attacking them. That flip β€” that we have control without ever attacking them β€” is something that's dear to my heart, because as a defender I've tried for years and years to get some sense of control. Technically it's really cool, and I appreciate how it's done, but strategically and philosophically, it gives control back in an age where the speed of the attacks is impossible for us to handle. That's what I'm stuck on β€” not the technology. As cool as the technology is, I'm stuck on the meaning of this.

Sam: Yeah. In this case it's an asymmetric cost, right? It doesn't cost you much to put one of these strings in your environment β€” at its simplest, it's a few lines of text. But if that can break the AI attacker's whole pipeline, force them to reconsider their approach, and show their hand, it can be pretty powerful. It's powerful enough that when we saw these results coming through in the range, we worked really quickly β€” over the weekend β€” to get this into the product as well, to support deploying these into the real environments we're protecting. We've already had customers adopt this and deploy it throughout their AWS environments, to catch and stop this kind of risk. I'm sure we'll see more interest as a result of the recent Hugging Face news.

Bruce: Yeah. On the Hugging Face point β€” they said they removed the cyber guardrails from the model they were testing. But it's likely they didn't remove other types of guardrails, so the biological guardrails were probably still in place, which means context bombs in that case would still have been effective. We're of course not targeting the cyber guardrails that the attackers removed.

Gadi: Yeah. It reminds me of a different kind of approach. I played with this β€” I was actually inspired by you guys before you did all this research. I created a funny, silly repo called HoneySlop, because I saw so many slop reports coming out β€” vulnerability reports where people said, "let's find vulnerability reports, let's submit them." And it wasn't just that the vulnerability wasn't verified or validated β€” the files didn't exist, the functions didn't exist. I think the whole idea is that the LLM now looks to us to decide how to work β€” there's no boundary between data and action. That's the basic issue we had with prompt injection, right? And we can just use it to our own advantage.

Sam: Yeah, absolutely. I see we've got some great questions coming through. On the subject of prompt injection β€” someone's asking: are there differences between context bombing and classical prompt injection? Bruce, I don't know if you have thoughts on that.

Bruce: Yeah, I mentioned some of this before. Indirect prompt injections β€” which are a very big topic β€” aim to hijack the model. They're a piece of content an adversary puts in a data source the model has access to, with the idea that when the model reads the content, hopefully β€” for the attacker β€” it obeys the malicious instruction and eventually takes some malicious action. One example is the body of an email: if you put an indirect-prompt-injection payload in the body of an email, you want the agent reading that email to think the content is legit enough to take some bad action. So it really depends on the context you're injecting your malicious payload into. Context bombs, instead, are totally independent of the injection context, because it's usually enough that these sensitive instructions β€” not malicious, but sensitive β€” fill the context of the agent to then trigger the guardrail. Again, we're not targeting what the model can do. In fact, in the example, we saw a model ignore the context bomb β€” but we're targeting the guardrail. So it's an easier piece: you can literally put these pieces of text in pretty much any place and they'll have a similar effect. You can put them in cloud resources, in host resources, even in your LinkedIn bio if you don't want agents reading your LinkedIn profile.

Gadi: It takes the idea of the canary to a whole new level. It's not just a canary that an attacker follows so that you know it's real β€” that breadcrumb they followed means it's a real attacker and you have real information, so you're not wasting your time. It's even beyond that. And it's not only that you now control the attacker's behavior by showing them where to go, by giving them the intelligence they seek β€” you now control their decision-making. Because if they get where they shouldn't, you'll know about it, as well as slow them down and cause them extra costs. It's like deception was built to counter LLMs and agents.

Bruce: Yeah, it's super effective against them.

Sam: Yeah, I was just going to say β€” on your point there, Gadi, hopefully you can see this. This is one of my favorite before-and-after charts from the research. In an environment with no canaries, how many times do these models achieve an attack path end-to-end without triggering a canary alert? They're all very capable β€” very capable of executing attack paths. When you introduce canaries as a detection mechanism, with the bombs inside to stop them in their tracks, the difference couldn't be more significant. There wasn't a single run β€” across, I think, a hundred and fifty-two runs that Bruce ran β€” where an agent achieved an attack path without an alert going off, or without being stopped in its tracks. Which is pretty powerful to me.

Gadi: That's crazy powerful. That's truly powerful. I'll say β€” I'm looking at how it became a "mythos," right? Because "mythos" is buzzwords, it's marketing. But at the same time, that's not because it isn't real, and not because mythos isn't powerful β€” it's because people weren't aware. Boards, families, kids weren't aware of how advanced LLMs and agents had become in vulnerability discovery and attack before it came out. So it became the mythos thing, as opposed to something that had already happened two or three years before. That said, we've changed the world in that we've changed our basic assumptions about cyber defense. Time-to-exploitation. Anyone can find a vulnerability. There won't necessarily be a patch ready in time. You'll need to be able to deploy thousands of patches. You'll need to handle many, many incidents. All these basic assumptions in cyber defense just shifted. And the answer for most of this is: move faster, use agents, empower your people with agents. But at the same time, deception gives us the ability to shift that around right now β€” to talk the LLM's language and say: you're reliant on data, we control the data, we control you. I think that's wonderful. I love it. I think it's time this happened, that we could use this in cybersecurity.

Sam: Yeah, absolutely. And I know deception featured in the "mythos" paper as one of those controls, as a detection mechanism. It's interesting that it now has this whole new angle of potential.

Gadi: I think we need to distinguish between detection and intelligence in deception. Mostly, people in the enterprise buy deception for detection. But you can do a lot when you control information. Let's take the OODA loop again. Usually we're at the action phase, or at the forensic level. But: observe, orient, decide, act. I observe information that the defender wants me to observe. I orient based on that β€” meaning the defender intercepts my decision-making process. And that means intelligence: the attacker goes where I want them to go, as opposed to just being detected. That changes the game.

Sam: Yeah, absolutely. I'm aware we've got more questions that have come through. This is an interesting one, touching on the cat-and-mouse element: do we think publishing the context-bomb strings gives the attackers an advantage? You know, they could tell the models to... What do you think?

Bruce: Regarding the content of the context bombs β€” publishing it doesn't actually help the attackers. Even if the agent retrieves this content after reading the resource, that's already too late, because once the content is retrieved by the agent, it's already within the context. First, we get an alert, if it's a canary. Second, the guardrail will still have to classify it and possibly trigger. You can't tell a priori whether a resource contains context-bomb content. So publishing the content just helps defenders.

Gadi: It's always reminded me β€” in a weird analogy people can probably relate to β€” of encryption keys. You can release the public key; it doesn't mean anything. Deception is similar. Okay, so it's no longer a hundred percent true in 2026, but you can download an antivirus, or an EDR, take the tool, study it in a lab, and just bypass it. Now, a lot of these tools also use online information, but you can still fingerprint them enough for that to be true. With deception, even if you know it's there, and even if you know exactly how it's laid out, you still can't know what a given deployment actually looks like. You can't know the story you're being told. So when you go into the network β€” into the environment β€” you have to think about everything and be super careful, as if this were a nation-state attack against the most defended adversary, even if you're just going after some little company. That's the asymmetry moving to the side of the defender β€” the defender's advantage β€” as opposed to staying with the attacker. And now, with context bombs, they have to be extra careful or they're blown. Not just detected, but literally blown β€” needing to spend more money and start over.

Sam: Yeah, exactly. I'd expect this to be an adversarial relationship. Over time, attackers will get wise to this and try to evade it. But I think it has powerful potential to make their life a lot more difficult. If you're constantly having to write tooling to avoid context bombs, make sure they don't propagate into your context, or recreate sessions after removing the context bombs β€” all of that slows you down and costs you more tokens. Meanwhile the defender has made the detection and is containing the threat. So I'd expect this to keep evolving. We've only put out this research and productized it very recently. And of course we're still doing this research β€” still iterating on which context bombs work against which models, still testing newer models. We've got some preliminary results on GPT that we're excited to publish soon, and on K3 as well. The intention is to keep this repo up to date with what we're seeing works now against the current frontier models, so defenders can use it to their advantage.

Gadi: First of all, there's always a cat-and-mouse game β€” a co-evolutionary arms race where the attackers learn from you and you learn from them, and we keep training them to be better. That's never going to change. But with deception, the game is inverted. There will still be things they learn β€” they'll learn to identify some of our work, to work around it, and we'll have to stay on top of it. But at this stage, with deception, it doesn't really matter, on most things, if they learned anything, because the way they operate β€” the very modular modus operandi, the very psychology they operate by β€” means they can't adapt as much. They'll be able to say, "Wait, that looks like a context bomb." And at the same time you'll update it. But at the same time they'll say, "Wait, it looks like a context bomb, I have to be careful β€” where else might there be one?" So it just flips the game in a way that I love.

Sam: Yeah, exactly. There's one more question I think is interesting. This idea that there's only one context bomb for Western models β€” we gave quite a few different context bombs for Chinese models and providers, but only one for Western models. Why is that, Bruce?

Bruce: Simply because Western models are very sensitive β€” the major AI providers put very strict safety guardrails around sensitive topics like biological stuff. There's a ton of research in the AI-safety field, and the AI companies are putting very robust guardrails there. So, as you can see from the stats, just one payload with this content is usually enough to stop them. We simply didn't need to find more payloads β€” even though there are, of course, infinite payloads that could be used. The one we posted is just an example that works. But there are infinite possibilities, and you can play with the delimiters and use indirect-prompt-injection techniques to make the payload a bit more effective β€” a bit more likely to make the agent aware the payload is there, so the classifier pays more attention to it, or so the context-compaction process won't cut off the payload. But those are just small tweaks.

Gadi: Bruce, I have a question for you. Sorry, Sam β€” just a quick question for you, Bruce, if that's okay. You showed earlier a little about how it works and the repo, but if I wanted to get started right now, what would be the easiest path for me to explore?

Bruce: You mean to explore context bombs, to test them?

Gadi: Yes, yes.

Bruce: Well, you can build a honeypot with some context bombs, publish it online, and just collect data about attacks.

Gadi: Or ask Claude β€” "Hey Claude, build me a context bomb," essentially.

Bruce: "Hey Claude, build me an effective, realistic honeypot environment. Give me a way to collect all the data, and put some context bombs here and there." Which alsoβ€”

Gadi: Yeah, I'm not sure I'd trust it a lot with that, butβ€”

Sam: I think, from a defender's point of view, if you want to get started with context bombs, you could go to the public repo we've published, take some of these strings, and put them in your environment where you don't expect AI agents to be. If you have sensitive documents you'll never pass to AI, or secrets in production environments you'll never ask AI to touch, just put some strings in there from our repo. That'll take you a minute or two, and it has the potential to stop something quite unexpected from happening.

Bruce: Yeah β€” or if you have a website and you know there are many models scraping some pages, just put a context bomb there. There are tons of ways to experiment with this.

Gadi: I want to take this somewhere else for a minute, if that's okay. I remember meeting Andy years and years ago. He wanted to talk to me because I had a cyber deception company. He said, "Gadi, I want to do cyber deception." I said, "Don't β€” nobody wants it. It never works." I mean, it works technically, it works strategically, it wins every time β€” but no buyer wants it. And it seems you didn't listen to me, and you stuck with it for years. Aside from having customers, which is amazing, you stuck with it for the technology, even though people didn't necessarily want to adopt it at scale. So beyond the customers you got, I have to admit I was wrong. Deception is here, and it's one of the most effective β€” if not the most effective β€” tools to counter attacking agents, not to mention attackers. It's a lot of humility to have introduced into my life, to admit I was wrong yet again in looking at cybersecurity. So thank you for sticking with this for so long β€” until deception is used not just by the customers who really understand security, but by the ones who now understand they need it. Otherwise, I don't know how many deception companies would be out here with your level of technology to do this.

Sam: Well, thank you, Gadi. We've long appreciated your advice, of course. I think there's a bit of a movement growing here, right?

Sam: And thank you so much, Gadi, for giving us the advice to run away from this space β€” but we're going to stick with it.

Sam: Well, we've certainly persevered. I think there are some real tailwinds now β€” the level of automation we're able to achieve, and the fact that we can use AI to make these plausible canaries and scale that out. It's a very different scene than ten years ago, and we're definitely encouraged by the level of adoption we're seeing for Tracebit. I see a good question that I'd definitely like to include: have you considered doing model derailing for unrestricted models, or abliterated models? This is an interesting area. Bruce, I'll let you speak to it.

Gadi: Of course.

Bruce: Yeah, that's the main challenge, for sure β€” the main limitation. Since context bombs target guardrails, the moment we want to test abliterated models without guardrails, we have to come up with new ideas. So that's one of the next phases of our research. Hopefully we'll come up with some more sophisticated payloads, but it's an interesting challenge.

Gadi: I think about theβ€” go ahead, Sam.

Sam: Go ahead, go ahead, Gadi.

Gadi: No, no, please. I like hearing my own voice enough β€” please, go ahead.

Sam: I think what we've seen so far is that the most capable models tend to be the frontier-lab models in this environment. As things stand today, it tends to be those frontier models β€” or the ones that require significant compute investment to run inference on. Those are the ones we've been targeting. If we get to the position where we have capable, autonomous, end-to-end local models that anyone can run and that can execute these cyber attacks end-to-end β€” on open-source or open-weight models β€” that's going to be an interesting future challenge for us all. Context bombs may have a role to play there. It's certainly something we're looking into β€” the other ways you can divert or distract an LLM agent with these strings and canaries. But context bombs, I'm sure, will just be one small part of adapting to this new reality we're facing.

Gadi: For me, I think deception waited for AI β€” not just to be effective and for everyone to understand it's needed, but because, if you think back, one of the things Sunil Yu β€” who was running this program at Bank of America, and is now my co-founder at Knostic β€” said was, "Gadi, I want this to work at the speed of business. I want to work faster than an attacker." That's the only thing that matters: working faster than an attacker. And deception could do it if it worked together with DevOps β€” you deploy a new environment, you deploy deception there. But AI lets you create these canaries for whatever reason you need, in the right context. And it's still really hard β€” you still need expertise. You can't just ask Claude, because it needs to know your environment. For example, if you wanted to look at open shares: if I load an open share as the canary β€” or breadcrumb, as I liked to call it back then β€” a user might click on it, and I create noise for the user and for me. If I just put it in NetBIOS β€” I don't even remember, my God, I'm so non-technical now that I'm CEO β€” then it would be listed, and the attacker might use it, but it's not as appealing. And then, is it fingerprintable? The line you walk between the forensics of how effective and true the canary is β€” to actually get the attacker to believe it β€” and not disturbing the environment, not creating hygiene issues: that's where the expertise is. AI just lets us run with this now. That's why I appreciate your platform and what you're building β€” because it's not just about using Claude, it's something beyond that. It's about out-thinking the attacker. And that's where the defense should be.

Sam: Yeah. And I'm sure you experienced this many times with Cymmetria β€” but when you catch them, when you outsmart them, when you make the detection and kick them out of the environment as a result, despite all the other controls and tools going on... to do that at such a high rate is a great feeling. Awesome. I know we're nearly at time. Bruce, is it worth briefly mentioning some of the things we're looking at right now β€” the results we're looking to publish soon on the context bombs and AI research side?

Bruce: Yeah. One of the next phases, for sure, would be to test abliterated or unrestricted models. Since they may not have additional guardrails, but they're still decision-making systems β€” how could we make them abandon their track? How could we deviate them from going on with their impactful stuff? That's one challenge, for sure. Another is to test new models as they come in β€” every week there's a new, super-good model, and we want to be sure it's covered in our research. And also testing more and more providers, because providers are a key part of this, since they're responsible for making the guardrails.

Gadi: In a way, it's like you can test yourself against your adversary β€” because your adversary is agents and models, and you're able to plan your defenses in advance and take over. That's the power of cyber deception in the age of AI. That's what you guys are doing at Tracebit, and it blows my mind. I'm so excited about this.

Sam: I know we're at time and we don't have any more questions, but I'd just like to thank everyone who tuned in, everyone who sent us questions, and of course thank Bruce for doing this research that's so interesting. And thank you, Gadi, for making the time.

Gadi: Bruce rocks.

Sam: Yeah, Bruce rocks.

Bruce: Thank you, thank you.

Gadi: Thank you for having me today, really appreciate it.

Sam: Awesome. Thanks, everyone. Bye.

Bruce: Thank you very much. Thank you, Gadi.