Pop Goes the Stack

Mechanistic interpretability sounds academic until you try to debug a model with printf and realize there’s nothing to print. In this episode of Pop Goes the Stack, Lori MacVittie is joined by F5 Chief Product Officer, Kunal Anand, to talk about why “mechinterp” is getting serious attention: if we can’t understand how models arrive at decisions, we can’t predict failure modes or build effective controls.

Kunal walks through his deep dive, sparked by a conversation about how much weight individual tokens can carry, especially as context windows grow and models don’t always “use” every part of their capability to produce a plausible response. That rabbit hole led to his blog post, “Your Token is a Wonderland,” where he trained a transformer on his own iMessage history to build a model on a dataset he understood intimately. The point wasn’t novelty; it was debug-ability. With a smaller model, he could inspect attention patterns, layer behavior, and token predictions in a way that’s effectively impossible on trillion-parameter frontier systems.

They discuss what this kind of work reveals: how context changes meaning, why certain tokens get selected, and why model behavior can feel opaque even when outputs look confident. The conversation also ties mechinterp back to practical outcomes, from improving guardrails and refusal behavior to finding ways to reduce hallucinations and avoid high-stakes errors without retraining entire models.

The takeaway is pragmatic: we’re early, and the field is still nascent, but it matters. Understanding internal “circuits” isn’t just intellectual curiosity; it’s a path toward better debugging, safer behavior, and more reliable AI systems. Until then, variability is part of the deal, and “the model said so” still isn’t an explanation.

Read Kunal's blog, Your Token is a Wonderland for his mechanistic interpretability deep dive: https://kunalanand.com/2026-03-19-your-token-is-a-wonderland/

Creators and Guests

Host
Lori MacVittie
Distinguished Engineer and Chief Evangelist at F5, Lori has more than 25 years of industry experience spanning application development, IT architecture, and network and systems' operation. She co-authored the CADD profile for ANSI NCITS 320-1998 and is a prolific author with books spanning security, cloud, and enterprise architecture.
Guest
Kunal Anand
As Chief Product Officer at F5, Kunal leads the efforts to deliver transformative solutions in application security and delivery, overseeing product vision, technology strategy, and execution. His passion for cybersecurity, data, and engineering has shaped his career, from co-founding Prevoty, an application security startup acquired by Imperva, to serving as Chief Technology Officer and Chief Information Security Officer at Imperva. These experiences, along with leadership roles at organizations like NASA’s Jet Propulsion Lab and BBC Worldwide, have prepared him to tackle the evolving challenges of modern technology.
Producer
Tabitha R.R. Powell
Technical Thought Leadership Evangelist producing content that makes complex ideas clear and engaging.

What is Pop Goes the Stack?

Explore the evolving world of application delivery and security. Each episode will dive into technologies shaping the future of operations, analyze emerging trends, and discuss the impacts of innovations on the tech stack.

Lori MacVittie (00:01.794)
Welcome back to Pop Goes the Stack, where the black box is still a box, just with more confidence and fewer explanations. I'm Lori MacVittie, and today we're talking about mechanistic interpretability. Yeah, it's about understanding models to see which gears actually turn. You can laugh, Kunal, who's our

Kunal Anand (00:24.177)
This is great.

Lori MacVittie
main attraction today, right? It's a, you know,

Lori MacVittie (00:29.666)
you're picking up the hood and you're looking underneath and going, why did it say what it did? How, why, how did it get from A to Z? Because it skipped a lot of steps in between. It's trying to understand what's going on inside because there's a lot of attention being paid to how did it make the decision it did? Because if we can't understand

Kunal Anand
Totally.

Lori MacVittie
how it thinks, we're never going to be able to predict, oh, it might screw up this way. And if you can't predict how it might screw up, you can't put controls in place. At least that's the thinking. So, and that's a very simplistic explanation of--what did you call it?--mechinterp? Is that

Kunal Anand (01:10.501)
You got it. That's great.

Lori MacVittie
Alright,

Kunal Anand
You're one of

Lori MacVittie
now I'm a member of the club. That's it.

Kunal Anand
You're one of us. It's great. It's awesome.

Lori MacVittie (01:18.178)
Well, you actually wrote an entire blog post about your time spent like doing this, actually going through and, right, building models and understanding how they actually work, the circuits, what's getting fired, what gets attention when, those kinds of things. So maybe you could kind of explain like what it is you did to do that, because printf doesn't work anymore, I found. Like that's not a valid like debug. I can't, can't do it.

Kunal Anand (01:51.067)
But you know what does work? Taking a PDF of the thing and sending it to a chatbot when you run out of tokens.

Lori MacVittie (01:53.743)
Ha ha ha. You just, well, yeah. Ha ha ha.

Lori MacVittie (02:03.692)
You know. That's the best pro tip ever though. Yes, if you run out of context in a conversation, everyone, save the conversation as a PDF and then share it with a new chat and then you will get your context back. That came

Kunal Anand (02:19.505)
You're welcome.

Lori MacVittie
from Kunal, yes.

Kunal Anand
You are welcome, you are welcome, everybody.

Lori MacVittie
I am grateful. Grateful for that. Grateful for that. It's not really mechinterp, but it's still

Lori MacVittie (02:31.144)
it's still valuable information

Kunal Anand (02:32.753)
Talk to me about your PDF tokens.

Lori MacVittie
coming right to us.

Kunal Anand (02:32.753)
Talk to me about your PDF tokens, everybody.

Lori MacVittie (02:39.298)
We've, we've lost it. That's

Kunal Anand
I, you know what'd be pretty amazing is like we check in, like you and I check in like a few months from now and this like podcast like propagates that. Now that like pro tip

Lori MacVittie (02:39.298)
We've, we've lost it.

Kunal Anand
propagates around and like next thing you know, Anthropic and OpenAI are posting "We don't know why, but PDF tokens have hockey stick to like like...

Lori MacVittie
Like everybody's doing it, right? Yeah, yes.

Kunal Anand (03:03.159)
It's gone parabolic. People are taking their source code and shoving them into PDFs.

Lori MacVittie
Ha ha, into PDFs. Wait. Well now that's a serious question because I haven't looked into it. Is, are PDF tokens? Or is it just that they're counting them? These came from PDF. Do they charge separately for them? Because I would believe that they could.

Kunal Anand (03:23.823)
I think they do.

Lori MacVittie (03:25.144)
Yeah.

Kunal Anand
They definitely do for

Lori MacVittie
Okay.

Kunal Anand
for images, so I assume they would treat it differently. Only question is like does it really matter if you need to preserve context?

Lori MacVittie (03:35.359)
Yeah. No. Ha ha ha. Not really, but it is good to know, right? I mean I think a lot of people don't even understand, you know, to go even back up from mechinterp where we're talking about the innards of how it thinks, but even up to the system

Kunal Anand (03:53.414)
Totally.

Lori MacVittie
level that. Right? When the GPT or Claude or whoever you're talking to generates the code but it doesn't actually execute the code.

Lori MacVittie (04:02.978)
That is

Kunal Anand (04:03.056)
Yeah.

Lori MacVittie
a separate system. And I think understanding that you realize, oh, it sends it to a different system to execute and then tokens go back in. So of course you're getting charged more for things that are happening because they're under the covers, if you will, and you're just not seeing it. So I think there's a lot of that going on that people aren't even aware of in terms of tokenomics and how things are happening.

Kunal Anand (04:28.037)
Yeah, I'll set the context for kind of this like mechinterp deep dive and also kind of what I did and share a little bit about what I learned in the process.

Lori MacVittie
Okay.

Kunal Anand
So, months ago, I was back in Los Angeles and was having dinner with my friends and we're all in the industry and we're talking about AI. And I was sharing with them this interesting thing that I learned, which is with some of these models now because their context windows are so large, there's been this interesting thing and it has to do with the sheer amount of like a combination of like pre-training and post-training that labs are doing with these models, that they're getting more and more capable, where you don't necessarily need to use like all parts of the language model underneath the hood to generate a valid response.

And so a good example of this is sometimes what you provide in the context window plus a little bit of the underlying model is enough to produce a high quality response. Like you can close--think about it in terms of like a programming closure--like you can close over the input that's provided pretty sufficiently. And that just has to do with the fact that we've done a really remarkable job with pre-training and post-training. So I brought that up to my friends and we then were talking about debugging these things.

And like it would be amazing if there was a way to kind of observe and look at the way that these sort of neural networks were firing, these synapses, these digital synapses were actually executing. And that then led to a pretty interesting conversation where we were just talking a lot about how, kind of went to a more philosophical conversation about the universe, which is how we're all sort of made of space dust, right?

Whether you want to believe that or not. Like all like the atoms that make humans are existing in the universe. It's just the combination kind of came together in this way. And then some of us are all music nerds, and someone was like, Well, that means that like every token is really, really rich then, because if every token can change the direction of what

Kunal Anand (06:53.321)
model can generate whether it's as part of inference or any other task, then that means there's so much richness to every point in the context window or every point in, sorry, that's kind of being glommed as part of the transformer. And tongue in cheek, we were riffing and someone was like, oh yeah, it's kind of like that John Mayer song, your body is a wonderland, except it's like your token is a wonderland.

Lori MacVittie (07:19.702)
Ha ha ha.

Kunal Anand
And I was like, and I was like, yeah.

Kunal Anand (07:22.775)
Sweet. And so like

Lori MacVittie (07:23.052)
Ha ha ha.

Kunal Anand
so like I literally went home from dinner and I couldn't stop thinking about it. And then I like devoured all the research, including like a stack of papers here, like this is the Anthropic, like

Lori MacVittie (07:42.12)
Wow.

Kunal Anand
the mathematical framework on circuit design and whatnot. And like I devoured all of this stuff. And sitting in a hotel room, you know, coming back from dinner.

Kunal Anand (07:52.858)
It's of course the most amazingly romantic thing that you can do. And my wife was there and was like, what are you doing? And I'm like, I have to like dig into this because it just like triggered this like part of my brain. And I was like, it would be amazing to really properly go through building one of these models, one of these networks, with the intent of debugging it.

And I've built models before, but I've never built a model on top of a data set that I had intimate knowledge of. Because if I had intimate knowledge of the data set, then I would be able to debug it and just know like this looks right or this doesn't look right based on me. And so I had this idea of why don't I take every single iMessage that I have ever sent and received, every conversation, and train a neural network from it and build a transformer on it.

And I did. And so this whole blog post that I posted a while ago, Your Token is a Wonderland, is basically me walking through the full steps of taking my iMessage history, every iMessage ever, and really putting it all together and then building a network on top of it. And it, I even go into like the esoterics of like

Kunal Anand (09:18.799)
me training on top of like Intel's kit versus NVIDIA's kit, and like the little things that I learned. And I'm sharing all the scri-, I shared, not sharing, I shared all the scripts,

Lori MacVittie (09:26.712)
All the scripts, yep.

Kunal Anand
all the things that like anybody can repro this thing on their own anytime they want. And at the very end of it, I was able to really go and build a network where I can now understand like how I think. It's a weird reflexive thing that kind of ended up happening. And also

Kunal Anand (09:48.25)
the byproduct of it is I've got a model that can talk like me now and I can like prompt it in a certain way. So like if I want a model to like sound professional and technical, I can do that. And if I want a model to shit post, I can do that too now.

Lori MacVittie (10:02.284)
Ha ha ha ha.

Lori MacVittie (10:06.799)
And that's, I mean, that's I think what we're looking for ultimately. I mean, one of the reasons it's so easy to spot AI, right, slop is because it's not reflecting a particular style or tone. All, anyone who writes a lot or just communicates a lot has a tone and voice and style

Kunal Anand
Yeah.

Lori MacVittie
that is uniquely them. And AI is very generic, so you can tell when someone is just using it. Right. And it's not me. So I don't actually want it to write my email or respond to things or, you know, generate, you know, my content 'cause it's gonna lose the essence of me in doing so. But what you're doing is you were doing it for different reasons, trying to figure out right how it works in a way, right? Like what makes it, I guess when I think about it, I think about like what makes it go down one path or not?

I mean way back in, you know, undergrad, we were building neural networks and they're basically like, you know, doubly triply linked lists that connect in webs. And you know the basics, right? Oh yeah, and weights and it decides, but how does the token relate to that? How does the, how do they get the token, right, to be this? And how can I and then, of course, I go to how can I manipulate that

Kunal Anand (11:30.043)
Yeah.

Lori MacVittie
to go, right,

Kunal Anand
Totally.

Lori MacVittie
where I want it to?

Lori MacVittie (11:34.165)
So that's why this is kind of fascinating is that this is an attempt to kind of reverse engineer that without digging into the actual code and going through it. So...

Kunal Anand (11:45.468)
What was really cool was so the model, the network itself, was about five million parameters on top of the messages,

Lori MacVittie (11:53.423)
That's crazy.

Kunal Anand
which also makes sense because it's not like, you know, I'm, I primarily use iMessage from my phone. And yes, I've gotta connect to a computer or whatnot, but I don't typically send text messages from my computer, so my messages are pretty short and but five million parameters is still a pretty decent size. It's not like

Kunal Anand (12:15.451)
probably a power user, but it's what it is. And to your point, when you have a five million parameter network, circuit visualizations become really interesting, like building patching grids where you can do things like activation patching, where you can like visualize like superposition of items in the network. Like a token may mean something different in certain contexts.

Lori MacVittie (12:42.198)
Mm, mm.

Kunal Anand (12:42.797)
And also like it was able to generate all sorts of knowledge or make all sorts of predictions based on things that are in my language and my vocabulary.

Lori MacVittie
Mm-hmm.

Kunal Anand
So for example, my, and you'll see it in the blog post, but my daughter's name is Kara. And my wife and I have a nickname for her. We call her Karebear with a K.

Lori MacVittie (13:03.619)
Mm-hmm.

Kunal Anand
And like it's cool.

Lori MacVittie
It's cute. It's cute.

Kunal Anand
And so like we, so we call her Karebear. And like when, in one of

Kunal Anand (13:12.197)
the visualizations where I'm looking at how it attends to certain positions, I'll like feed it some context and then looking at all the different layers--like four different layers, so five million parameter model, but like four different layers--it's really amazing to look at, oh, like I would have predicted this token as an example. And like to see to see that in there was just truly special.

And why it's predicting certain things in there. It's super cool. So I loved it. It was a super helpful thing for me. I learned so much about the process. I mean I think we tend to in our industry now, and it's just probably because the way things are moving and it's not a slight on people or anything, it's just the industry's moving super quickly and we don't have a lot of time to learn like we used to.

Lori MacVittie (14:07.235)
Mm-hmm. Yeah.

Kunal Anand (14:08.877)
And so like unless you decide like you are going to go and pursue like a PhD in this world or like you really are gonna give up a lot of time to figure this out. For me, like I really gave up weekends of my life, multiple weekends of my life, to like really dig in to do the homework, to do the research, to do the coding, to like get to the outcome and to get to the output; that work itself was significant and I don't think a lot of people are gonna go and do it. But

Lori MacVittie (14:38.158)
No.

Kunal Anand (14:38.789)
but that said, it was super worth it. Like it was, you know, two weekends of crazy amounts of reading and crazy amounts of work. I am not an AI researcher. Like I'm not that. But I'm someone who's curious, and I think other people are curious. And so if you are, and if you ever wanted to build like or want to train a model, first if you just ever wanted to train a model go check out the blog post.

Like I think what I wrote up on there was like a very pragmatic way for people

Lori MacVittie
Yep.

Kunal Anand
to like think about just writing or like making a model. Like if you've ever wanted to do that, or if you just want to like see a very human, readable way for how to like train one of these things, I think you'll get value out of the post. But if you also have any interest in mechanistic interpretability, which is the way that these tokens come together, why certain tokens are predicted, why certain tokens are picked from the residual stream, why things are attended to, how to visualize that, how to patch certain things, like if I actually change this token to something else, does the outcome change in some way?

If you're curious about any of that, go and check this out. And the last point that you brought up, I think is probably the most important, which is we generally don't know how these things work. Like we generally don't know why certain tokens are being predicted when they are. And this was helpful for me to kind of see it, but it also made me take a step back to realize, look, this is a five million parameter model, and like it's already gnarly to like get my head around this. How on earth are we doing this for like trillion parameter models?

And nothing but respect to researchers in this space. So Anthropic and Google and OpenAI, they have loads of researchers who do nothing but spend their time on this. And the reason for it is, if you can debug and understand how these networks work, then we can shape these networks to produce better outcomes. Like if we can understand like why we go down a particular code path, yeah, it can help us. And that that has all sorts of utility. Like for coding example, like if we should never pick this branch or we should never go down this token or predict this token, that's a big deal because it could it could save people time and it could improve the quality of our products.

Kunal Anand (17:01.794)
So again, not a lot of utility yet in this sort of world of mechanistic interpretability,

Lori MacVittie (17:08.153)
Mm-hmm.

Kunal Anand
but it is super interesting and it's always evolving and there's a lot of, a lot of crazy research going on there.

Lori MacVittie (17:17.131)
Well I think that's good because I think we have to know at some point. I mean, and it's built into every engineer. I know why did the code do what it did. And you walk through it and you can see, oh, because it cho-, you know, it skipped this case because the variable was-, it's very understandable, right? We know it, it's logical. And right now we're kind of sitting on, yeah, but it has this token, and it why did it go down that path? It should have gone, I would have gone down this other path.

Why? We have to answer the question, why? And once we understand why, and maybe it's certain word triggers, right? The combination,

Kunal Anand
Yeah.

Lori MacVittie
the tokens that are forcing it, then we can learn to either, you know, not use those tokens in certain contexts or use the right ones. And so I think this is very important research. And eventually just, I mean yes, I want to know, but I think it's good to know so that we can start building, you know, better playbooks, if you will. How do I use this when I'm in this context

Kunal Anand
Totally.

Lori MacVittie
best, you know, and it will give us better, like you said, better output. So you know, I guess the takeaway, right, for, you know, listeners is we don't know what we don't know. And we don't know what we think we know. And we certainly don't know what's going on inside models. Not really. But people are working on it and, you know, you have to recognize that until we figure that out, there's gonna be variability. You're not gonna be able to predict exactly every response.

So you're gonna get differences here and we have to accept that and that's okay

Kunal Anand
Yeah.

Lori MacVittie
because sometimes it gives you some pretty at least amusing to me responses.

Kunal Anand (19:03.355)
Well, you know, the thing that I get excited about when I think about this is this is an area of research that is nascent.

Lori MacVittie
Yeah.

Kunal Anand
And I always love finding these pockets in our industry because it's not like it is, it's not like there's years and years and years of like, you know,

Lori MacVittie
Yeah.

Kunal Anand
history and experience here. Everyone is figuring it out in real time. And this is so interesting to me because obviously we spend a lot of time thinking about and building AI guardrails, right, as an organization. If you actually can understand the directional vectors of these underlying models, like for example, refusal, then it means that safety teams can patch models to block harmful generations, or we can adjust biases without retraining a model.

Lori MacVittie (19:58.458)
Yeah.

Kunal Anand
And at the same time, like imagine being able to,

Kunal Anand (20:02.523)
you know, with techniques like sparse autoencoders--which I did not pursue, that's something that I'm interested in in doing--that could give us more advanced debugging capabilities. Like to pinpoint exactly which circuit is causing like a structural error or a hallucination to what you were describing. It's fun sometimes when AI hallucinates, but in some cases it's okay and like the stakes are low, but in other cases the stakes are really, really high.

Like I think about, you know, people with self-driving cars as an example. I think about people who rely on AI in like mission critical systems. You can't afford to make a mistake. And even in our industry where people are like looking at implementing AI, AI models for improving their products, whether it's like improving the security features of their product or whatever it may be, you can't block a good transaction. You can't stop something good.

Lori MacVittie (20:56.151)
Mm. Yeah.

Kunal Anand (20:56.821)
And so like it's super critical that we figure this out. And I think what we should do is just pay attention to this and watch it evolve. I'm sure as more and more people kind of get into it and as more and more people kind of evaluate the space, it's gonna become really interesting. Like I said, this was a super spontaneous thing. It appealed to me and I was kind of like, Okay, I'm kind of just fell into the rabbit hole.

And I would encourage a lot of people to just not even, if you just wanna skim it and just to know like what's going on here, maybe it'll draw you in. But I think we should stay on top of this, especially because we do care deeply about guardrails, because we do care deeply about the security of these models. I think at some point we'll have some pretty big breakthroughs here. It's just too early right now.

Lori MacVittie (21:45.035)
It is, but I'm glad to see it going. I mean that

Kunal Anand (21:47.931)
Yeah.

Lori MacVittie
that need to know is not just to satisfy, right, our curiosity, it's also to be able to provide better solutions in many different areas. You know, especially in generation when it just no matter how many times you tell it to generate one way, it keeps doing it the other way. And there's no rhyme or reason to it. And you just wanna know why. Like why?

Kunal Anand (22:12.955)
Yeah.

Lori MacVittie
What makes that...

Lori MacVittie (22:14.529)
And a way to fix it. Like, is there a word you can use? Things like that. So hopefully this

Kunal Anand (22:19.014)
Yeah.

Lori MacVittie
field will continue to evolve. And like you said, research will come out of it and it will allow all sorts of advances in other areas just because we know a little more about how they work and therefore how to work with them. So unfortunately for this week, that's all we have time for. So

Kunal Anand (22:36.229)
We ran out of tokens.

Lori MacVittie
that is a wrap

Kunal Anand
We ran out of tokens.

Lori MacVittie (22:39.043)
We are out. We are over. We are out of context. We'll have to PDF this and put it in another conversation, ha ha ha.

Kunal Anand (22:44.635)
Take the transcript, make a PDF out of it,

Lori MacVittie (22:47.703)
Put it put it in another

Kunal Anand
put it into your favorite AI model, and then you too can hallucinate the way you think this conversation could go.

Lori MacVittie (22:56.119)
That's right. That's right. What we might have said next. You make it up, you tell us. So but before you do that, go subscribe because the model said so is not a root cause analysis and we're gonna keep asking for receipts.