Oxide and Friends

More and more, people are relying on AI to do their writing for them. Is this the new normal, or the intellectual equivalent of claiming take-out is your own home cooking? Max Spero, founder and CEO of Pangram, joined Bryan and Adam to discuss how they distinguish AI and human authorship.

In addition to Bryan Cantrill and Adam Leventhal, our special guest was Max Spero.

Previously, on Oxide and Friends:
Some of the topics we hit on, in the order that we hit them:
If we got something wrong or missed something, please file a PR! Our next show will likely be on Monday at 5p Pacific Time on our Discord server; stay tuned to our Mastodon feeds for details, or subscribe to this calendar. We'd love to have you join us, as we always love to hear from new speakers!

Creators and Guests

Host
Adam Leventhal
Host
Bryan Cantrill

What is Oxide and Friends?

Oxide hosts a weekly Discord show where we discuss a wide range of topics: computer history, startups, Oxide hardware bringup, and other topics du jour. These are the recordings in podcast form.
Join us live (usually Mondays at 5pm PT) https://discord.gg/gcQxNHAKCB
Subscribe to our calendar: https://calendar.google.com/calendar/ical/c_318925f4185aa71c4524d0d6127f31058c9e21f29f017d48a0fca6f564969cd0%40group.calendar.google.com/public/basic.ics

Adam Leventhal:

Hello. Max, how are doing?

Bryan Cantrill:

Hello, Max. How are you?

Max Spero:

Hey, Adam.

Max Spero:

Hey, Bryan. How are you?

Bryan Cantrill:

Doing well. Man, I, Max, I thank you so much for joining us. I'm really excited for this.

Max Spero:

Yes.

Bryan Cantrill:

So we're gonna dive right in because I I know, you music to Adam's ears. Your time is limited here. So we wanna dive right into this very hot topic. Max Spero is with us, cofounder and CEO of Panagram Labs. Max, is it my imagination or is this topic of AI detection?

Bryan Cantrill:

Is the algorithm sculpting things for me where this seems to be getting more and more attention? Or is there it feels like there's something organic happening where this is becoming a bigger and bigger deal, like, by the passing hour. Is this my imagination? Am I suffering from some sort of, like, social media delusion here?

Max Spero:

I mean, I think I'm the wrong person to ask because all I see is playing with him all the time on Twitter.

Bryan Cantrill:

But have you seen an uptick? I mean, it just feels like there's been a lot of this just, like, even in the last week, in the last, like, forty eight hours. It just feels like

Max Spero:

Oh, definitely.

Bryan Cantrill:

Very, very current.

Max Spero:

I I feel like the the noise just keeps going up. Like, there was the, like, the the, like, Wall Street Journal op ed recently.

Bryan Cantrill:

Yeah. I was gonna ask you about that. Yeah. So do wanna give people context with the about the op ed? Because I I I honestly do believe I wonder if we're gonna look back on this as, like, a really interesting watershed moment.

Bryan Cantrill:

So, yeah, describe the the this is the the the Stanley Drewkenmiller op ed.

Max Spero:

Yes. So Stanley Drewkenmiller wrote this op ed. It was published in The Wall Street Journal. He's kinda like my my understanding, he's he's very famous, investor. And, basically, it like, you you start reading it, and then it's very clear very quickly that it's, like, obviously AI generated.

Max Spero:

You're just like it's like that this wasn't liquidity management. It was price management, and a mistake far larger than $4,000,000,000 suggested. It just sounds like like it's very much like AI pros.

Bryan Cantrill:

Totally. Yeah.

Max Spero:

And, obviously, some people, like, immediately call it out. Some people run it through Pangram. Hey. That's a 100% AI. Interestingly, and I think, like, a little bit different than what's happened before, is he stands by it.

Max Spero:

He says, yes. Of course I used AI to write this. I'm not a good writer, essentially. And Wall Street Journal stands behind him and says, well, you know, he like, it's really, like, the ideas that matter, and, you know, we're we're not gonna retract this or anything. So I I think it's a very interesting watershed moment because, like, in the past, when somebody used AI in a way that was very public, like, they weren't very famous.

Max Spero:

They either, like, didn't stand by their AI use, or they, like, tried to minimize or, claim that they they didn't use AI. This was very different because he's just like, yes. Of course, I did.

Bryan Cantrill:

Yeah. He's like, you're all the problem. Anyone who like, I am not the problem. The fact that my op ed was generated by Noah was not the problem. You you've got a problem if you've a problem.

Bryan Cantrill:

And so I'm like, okay. How about the prompt? Great. Okay. Super.

Bryan Cantrill:

Like, I I get it. Like, you don't view can we just see the prompt then? Because I think that and this gets to, like, a and I'm not sure kind of what, Adam, I don't if you saw any kind of the storm and drunken around this last week, but, I wanna see the prompt because I don't know how much of these are your ideas. Like, is the prompt write an op ed against the most recent treasury department actions? I mean, that could be a prompt.

Bryan Cantrill:

Is the right. I mean, Max, were were people calling on on? I mean because I think it's we're just kind of like it to me is like a false dichotomy whether we kind of accept this or don't because I think it gets it gets into this much deeper issue of, like, where are the ideas emanating from?

Max Spero:

Yeah. Yeah. The core question is, like yeah. Yeah. It could have been, you know, write me an op ed on this topic, and it could have been, here's, like, my five pages of notes.

Max Spero:

Please, like, produce a good op ed

Adam Leventhal:

out Here's of here's my notes. Here's my outline. Here's all my thinking about it, but I can't string more words together, I guess.

Bryan Cantrill:

Yeah. Oh, wait. Here he or here's the far too long op ed I've written. And as Mark Twain famously said, I I don't have time to write a shorter letter. So please, LLM, can you help me tighten this writing?

Bryan Cantrill:

I mean, it would be okay. That would be like what feels unlikely.

Max Spero:

It seems kind of unlikely, just in terms of given the information density of the op ed.

Adam Leventhal:

This reminds me of when I was in elementary school, I was given a five page essay. I would write two pages and then use margins to get me to five pages. And there are people I went to school with who would write a 10 page essay and then use like font size to try to get it down to five. So I never understood those people.

Bryan Cantrill:

Have those people gone further, Adam? I mean, are where are those people now? Have we we should do they where are they now? Adam's fourth grade essayist.

Adam Leventhal:

Those darn try hards.

Max Spero:

Right? Exactly.

Bryan Cantrill:

Exactly. In fact, one of the I think it was Stanley Druckenmiller that I went to school with, actually. That was

Adam Leventhal:

was wondering where old Stanley might ended up.

Bryan Cantrill:

Yeah. Old Stand dog. He was much more loquacious back in the day, but, know, they really just wrapped his finger so much. He's alright. You refuse to write.

Bryan Cantrill:

So, I mean, in this I mean, this whole thing, I think, is is you know, as you say, Max, like, information density on this one and and so actually, I'm sure you found out about this be or did you find out about this because people were running Pangram against it? I mean, by the time you found out about it, had people already had the reveal already been there? Or

Max Spero:

Oh, yeah. Definitely. I I feel like every time this is the case, it's just like I open Twitter, and I have, like, 15 people being like, at max Spero. Like, hey. Like, some somebody did another thing with AI.

Bryan Cantrill:

Well, it's extremely valuable.

Max Spero:

Yeah. I mean I mean, this is why, like, I think we're we're building Pangram to some degree. It's, like, I like, trying to build this as this, like, pro social technology where, like, I I'm hoping that we can continue to value humanity as AI eats more of the economy. And

Bryan Cantrill:

Yeah. I cannot tell you how pro social I think this is. I mean, I this I you know I'm a strong believer. We're customers. I'm a big user.

Bryan Cantrill:

I love what you've in Panagram. And this and I I really, really do think that you are on to something very large in terms of get the importance of writing and thought. And I think that the the ability to reliably detect LLM assisted or LLM authored writing is really important to kind of make the the what I think is a rather important point. I actually think that, like, doing one's own writing actually helps you form your thoughts. I know this is a very it's a very radical idea.

Bryan Cantrill:

Very, we're very pro literacy in this regard. So it's very, very iconoclastic. But

Max Spero:

I mean, I can't tell how much of that is sarcasm. It it seems like it kinda is. But but, I mean, I I do think, like, obviously, like, you start from an outline and then write the damn thing yourself versus use AI. You're gonna end up in two very different spots just through the

Bryan Cantrill:

Yeah. Yeah.

Max Spero:

Through

Bryan Cantrill:

it. No. In no sarcasm, honestly. I mean, it's more that the the fact that that feels like an old fashioned idea, the fact that you've got Stanley Druckenmiller saying that, like, yes. Of course, I'm using an LLM.

Bryan Cantrill:

What of it?

Adam Leventhal:

No, Matt. Max, Bryan, I've been blogging for over twenty years. And I don't know if you have this reaction, Bryan, if you tell people, like, obviously in this day and age, we can tell everyone that we have a podcast episode on a topic. And I know you do all the time. And I do sometimes.

Adam Leventhal:

But like to I told a friend recently that I had written a blog post and they were they looked at me like I was a time traveler. But I learned so much from from actually putting words to page. And like you're it does clarify your thinking as you're saying in a way that interacting with the chatbot doesn't.

Bryan Cantrill:

That it absolutely doesn't. Or that or that if it does, I mean, so here's the other the other kind of question I've got. Like for those folks who who don't think it's a big deal, again, just share your prompt, share your iterations, and then let other people draw the inference about what where your ideas end and the LLMs begin. And that's where I think why I think this is such a hot topic is because it gets into trust and integrity. And and it feels like honesty, honestly, where, you know, if it it's no one's got a problem with you ordering takeout.

Bryan Cantrill:

But when you, like, order takeout and then, like, kinda whip it out of the oven when your spouse walks in,

Max Spero:

you're you're saying this is

Adam Leventhal:

steamed hams. Like, when this article is steamed hams.

Bryan Cantrill:

Yes. Yes. So, Max, taking back to kind of the origin of Pangram because because you've been at this for a little while. I mean, certainly, for for decades by LLM years. But what's the origin of of Pangram?

Max Spero:

Oh, man. Yeah. So this is, like, mid twenty twenty three. So about three years ago now. It was, like, GPT four era.

Max Spero:

So, like, know, ChatGPT, kind of just a toy. Pretty cool. And then I think, like, GPT four, we're like, wait a minute. This is, like, really kind of important technology, and there's yeah. And then I kinda just I I think I had this intuition that it's going to be very important to detect this because, like, at the time, I think no one had an intuition for detecting AI writing.

Max Spero:

So it's kind of just it seemed completely indistinguishable. Seemed impossible. I asked all my researcher friends. They're like, yep. It's impossible.

Max Spero:

You shouldn't try this. Like, it's it's impossible. And then if it's not, like, no one's gonna care.

Bryan Cantrill:

And Oh, interesting.

Max Spero:

Or, like, you know, the AI models are gonna get better and better. So even if you can detect AI text today, you can't you're you're not going to be able to in a couple years.

Adam Leventhal:

What why impossible? Like, what what underpinned that insight such as it was underpinned?

Max Spero:

I I mean, I I understand. Like, I I think it's very easy to come up with a counterexample of, oh, I could train a transformer to say any text, or I could have an LLM even, like ChatGPT, produce any text with with the right prompt, the prompt that says, like, say this. Right. I think, like, in in a, like, very, like, mathematical sense, like, you could say counterex example, like, AI detection is disproven. But I think what we've done is said we don't really care about that aspect of things.

Max Spero:

We don't we don't care about, like, solving it in a, like, theoretical sense. We care about solving it in a practical sense.

Bryan Cantrill:

Yes.

Max Spero:

And so for for us, the practical sense is, like, where the, like, the input the the information, the output is like more than what was put into it.

Bryan Cantrill:

Yeah. And and and you define that kind of your paper, which I think is really interesting, where it's what you're trying to detect is where there's a word inflation in terms of like there's more in the output than there is in the input. And especially on kind of open ended writing. And I think, I mean, were you surprised about how important that has become? Because that has now become, I think, just, again, almost existentially important.

Bryan Cantrill:

I think it is so important now. Yeah.

Max Spero:

Like, early on well, we started seeing, like, fake reviews on the Internet that were AI generated. That was kinda like my first inkling of things, like, early twenty twenty four. Like, hey, why are these why are all these, like, Yelp reviews AI generated now? And all these Amazon reviews are AI generated now? Like, what's going on?

Max Spero:

And then so we tried to apply the the model there. Partly the model wasn't that good at the time. And partly also this wasn't really like a hair on fire problem for any of the big review platforms. They're like, oh, 3% of our reviews are AI. Like, who cares?

Max Spero:

But I think Who

Bryan Cantrill:

cares for Yelp? It doesn't matter. Who cares what the the fidelity of the reviews? Come on.

Max Spero:

Amazon's probably, like, over 50% of our reviews are fake anyway. So, like, who cares if, like, 3%

Bryan Cantrill:

of

Max Spero:

the fake reviews are AI generated.

Bryan Cantrill:

Right. Right. Interesting. Yeah. And was was there any inkling of the kind of the the broader problem that that it was gonna be very important in the kind of arbitrary future to differentiate human author text from LM author text to prevent these things from solely training on themselves?

Bryan Cantrill:

I mean, is that something that has kind of kind of entered in? Because I mean, I think that's a very important aspect of this. I mean, correct me if you disagree, but I think that the the ability to to train future models purely on on human content is gonna be very, very important.

Max Spero:

Yeah. Yeah. That's that's something I had no idea about before. But now in in retrospect, it makes sense. Like, you don't wanna be training g p d six on, like, g p d 3.5 turbo slop.

Max Spero:

That's, like, SEO slop on the Internet. Like, that that's just gonna be detrimental to your model.

Bryan Cantrill:

Well, and and you and then these things, like, have wanted to, like you know, they train themselves on Reddit or whatever. And, the you know, I there Adam, I know you've seen this, but basically, the am I the asshole on Reddit, which obviously is the, you know, at the epicenter of all of humanity, has become now just, like, rampantly LLM generated. And so if you train I mean, I if you train on the because I think the other Max, I I wanna ask you is like, I am gonna be really curious to see if you start seeing sites that allow people as perhaps even as a premium feature. Like, I might pay LinkedIn to only show me human generated posts. Like, I'm not I I like experiment with that idea, LinkedIn.

Bryan Cantrill:

I'm not sure I'm not totally sure, but I'm not sure I wouldn't. And I think that, like, I you just wonder if, you know, as people realize that I am losing the power of my writing when I use an LLM and if people stop reading. Because that's the issue, right, is that people stop reading when they when they hit those tells that are so clear. I I think do people not realize how clear it is to many human readers that that the LLM has been involved?

Max Spero:

So so, Bryan, do you do you use LLMs and how often?

Bryan Cantrill:

Yes. Yeah. Yeah. I use LLMs for sure. So I use, for sure on my own writing too, as an editor.

Bryan Cantrill:

I mean, I think they are phenomenal, phenomenal, phenomenal editors. And the the the closer I am to to done when I give it to the LLM, the better its feedback. And I've gotten just Yeah. Because a couple

Max Spero:

Yeah. That makes sense.

Bryan Cantrill:

The the the you know, my my wife told me in no uncertain terms that she was done acting as an editor for my block entries. And I'm like, okay, well, I guess I'm a robot. What are you doing? You wanna help me out a little bit here? And I found that they were really good.

Bryan Cantrill:

And it's like often I mean, the number of times they will be like, hey, this transition sentence is a weak transition sentence. Like, goddamn. I knew that transition sentence was weak. And then I'll go back and rewrite it. Like, I'm not I'm definitely not taking its suggestions.

Bryan Cantrill:

Mhmm. And I mean, I think that's the key is like when when you allow when you give it too much leash, it will rewrite things, and it will cheerfully do it. Like, it always it it does this thing where it's like, this is a work of absolute genius. Do you want me to rewrite it for you? It's like, okay.

Bryan Cantrill:

Which is it?

Max Spero:

Yeah. And it's gonna rewrite it in in a way that it sounds exactly like Claude or ChatGPT.

Bryan Cantrill:

Right.

Max Spero:

So, actually, the reason I was asking this because you were saying, it AI writing is so obvious to you, and I think that's because you use LLens. So, there have been a couple of researchers that have looked into this and found that people who use LLMs are pretty robust detectors of AI text. People who don't like, if if you're, like, an English teacher who just, like, has never used ChatGPT or played around with it once, like like, you just have no clue until you've built up this intuition.

Bryan Cantrill:

Don't you feel like I would've do that test.

Adam Leventhal:

Yeah. Teacher is, like, the

Bryan Cantrill:

most 100%.

Adam Leventhal:

Detector

Bryan Cantrill:

of AI. Actually, that's

Max Spero:

yeah. English teacher is probably the wrong Right. The the wrong choice because they they get their exposure to AI text through other ways. Right. And,

Bryan Cantrill:

Max, I'll I'll tell you Oh, a 100%. And I totally second in smoke. We're gonna discover it's carcinogenic. The I the other thing that I do that aside from using LMs because so I don't know how much you know about the Oxide hiring process, but it's a very writing intensive hiring process. And we ask open ended questions.

Bryan Cantrill:

And as you might imagine, we get I mean, I know you mentioned this on an earlier podcast too about cover letters and kinda going through that for Pangram and realizing that, like, wait a minute, these are all these all these cover letters look the same. And, I mean, I can pick up an LLM from space at this point. And I because there are certain tells that are so unbelievably clear. Like, I mean, I people do not write genuinely in in text and LLMs love mean, there's a whole bunch of word choice. There's a bunch a bunch of structure.

Bryan Cantrill:

There's a not this, but also but as they've gotten more sophisticated, I've really and so I this is where I end up using Pangram a lot. And only when I've got a hunch generally, but it is not like it is just not rocket science because it is really, really clear on these open ended questions when you're like, oh, hey. You've said a lot, you actually also said nothing. And I as you say, like, the content level is low. I mean, I feel like there's almost like that.

Bryan Cantrill:

You can tell that, like, I think we've gotten watered down here a bit on the thinking, and I think this this this is an LL. So I feel like that's where it comes from more than than actually using LLMs. It's comes from reading just a lot of text.

Max Spero:

Yeah. And so and so you're using, I'm actually kinda just curious about your hiring processes and, like, the the writing aspect of it. Sorry if this is we've talked about this before. No. Not at all.

Max Spero:

Yeah. But but yeah. I mean, I I think that's that seems, like, pretty unique, honestly. Like like, we we have a few other customers who also use Pangram for the same thing. But, like, by and large, most people don't have that much writing in their, like, hiring process.

Bryan Cantrill:

They don't. That's true. They don't. And I have I've said for a long time that that mean, this is something we've had since the dawn of the company. Mhmm.

Bryan Cantrill:

And there is a somebody psychic chat. There's a podcast episode for that. We just describe kind of the origins of that. And the origin of that is that I made the worst hire of all time and previous life, not at not at Oxide. And you know, lost people like, well, I think I made the worst hire.

Bryan Cantrill:

It's like, oh, okay. Well, but that's fine. But my guy actually just I presented himself under an assumed name. I just got an office in Quentin for violent felonies, and that's actually not what made him a bad employee. That's what I they just like, well, okay.

Bryan Cantrill:

Exactly. Very bad. And one of the things that we learned in all of that was that the the folks that are when you when you are forced to write your ideas down, you are forced to act as a check on your own ideas. And you can it's just much higher fidelity to look at someone's writing than it is to interview them. An oral exam is not actually really high fidelity, because I had people that interviewed pretty well and they were terrible.

Bryan Cantrill:

So, and conversely, I've had people that that are were terrific that, did not interview well because they were nervous or because, you know, for all the reasons that you so the we've got this very, very writing intensive hiring process that I thought was LLM immune kind of when you were starting the company in 2023, 2024. I remember a friend of mine's like, Oh, that's not gonna work. An LLM can like answer all those questions. I'm like, well, the questions are like, when have you been happiest in your career and why? And when have you been unhappiest in your career and why?

Bryan Cantrill:

And I kind of think now a couple years down the road, he and I are both right. In that, you definitely can have an LLM answer that question, but it is really easy to catch an LLM authored answer to when you've been happiest and why. I mean, kind of obviously. But the number of people that and so then we ask people, these kind of open ended questions including like, why do you wanna work for Oxide? And the number of people that have human authored materials up until the Y Oxide, and And then the Y Oxide is a 100% LLM generated.

Bryan Cantrill:

It's like, do you do you think you might not wanna actually work here? I'm just throwing that out there.

Adam Leventhal:

Do it the other way around. Right?

Bryan Cantrill:

Like like Do it the other way around. Do it a 100%. Like, I wanna work for Oxide, like, more than I wanna live. And I think I wanna work so badly. I'm have this LM generate all this other stuff.

Bryan Cantrill:

It's like, yeah, why would you Yeah. Don't do this, please. Folks that are applying to Oxide, please don't do that. It's like kinda heartbreaking, honestly. And the it's like I can tell an LLM written Y Oxide again from space because and Adam, I don't know you have any of these seen I I try to shield other people from these, basically.

Bryan Cantrill:

Like, I live in the general population of material submissions all the time. And Adam, I I think that, like, you get the ones that were basically, really good that we're gonna hire.

Adam Leventhal:

Oh, yeah. I I get the really good ones we're gonna hire and the ones that are so bad that, like, you pass it around for everyone to smell.

Bryan Cantrill:

Yes. And so the I meanwhile, I live in the in in the sloppy middle. And so I in particular, Max, if when someone describes wanting to work for Oxide, you can see when they are when an LM is parroting back the y Oxide to them. So in particular, you'll get someone who's like, I work for Oxide because I believe in the power of hardware software co design. It's like, yeah, you there's nothing in your background, history, materials.

Bryan Cantrill:

Like, you've never done anything with hardware software code. That is not a plausible answer for you, by the way. And so that's where I end up encountering a lot of this. I also think it's like really This is why I think it's so important because you know when you are But conversely, you know when you're in human authored materials and when someone is speaking from the heart. And I think it's like really important to have to be able to authoritatively tell when something is LM authored.

Bryan Cantrill:

I think it's important. So people that are doing this realize the rest of us can tell. Don't do this. Like, you think that people can't tell and they can. And so stop embarrassing yourself, please.

Adam Leventhal:

Max, along those lines, I have a question about Pangram, which it seems like it it gives an opportunity for people to, like, teach to the test. That is to say, you know, I can write my own op ed by just asking ChatGPT to do it, then feed it through Pangram. And can it can ChatGPT kind of use it as an adversary to, like, keep on making edits until Pangram decides it's human generated?

Bryan Cantrill:

Mhmm.

Max Spero:

I've been I've seen people do this to varying degrees. I saw someone on Twitter the other day that said they spent $700 in API credits, and their claud was really sad at the end having Amazing. Run out of API credits and not not gotten something that was human written. Which part

Bryan Cantrill:

of that did you find more delightful? I mean, obviously, that whole thing is extremely delightful. Did you find more delightful?

Adam Leventhal:

Than $100 of of Paymgrab credits.

Bryan Cantrill:

Or they or they think the Claude was sad at the end.

Max Spero:

I love sad. That's that's heartbreaking.

Bryan Cantrill:

Or heartwarming, man, depending on your perspective. I think it's great.

Max Spero:

But, yeah, I I've seen other people who, like, have been able to do it. So I think it really kind of just depends on, like, how well you specify the task. But, I mean, there's my opinion is, like, Pangram is not a good hill to climb. I've definitely had people at different AI labs or Neo labs ask me, like, can we use Pangram to make our LLM sound more human? And my answer is really, like, no.

Max Spero:

Like, you can use Pangram to make your LLM texts more incoherent and I mean, we just, like, he'll climb into just, like, gibberish. Like, sure. Pangram will say that's human, but it's, like, not

Adam Leventhal:

useful. Only a human could be this incoherent.

Bryan Cantrill:

I saw yeah. So and Adam, their paper is really good. The Pangram four technical report is really good. And so one of the things they talk about is, like, typo injection as a way to, like, foil Pangram. It's like, hey.

Bryan Cantrill:

Good job, applicant to Oxide. You really tricked us by injecting a bunch of typos and grammatical errors. Like, congratulations. You also still don't have the job. Because now you just can't write, by the way.

Bryan Cantrill:

Now we got other issues.

Max Spero:

Yeah. Yeah. So, I mean, that that means we have a good model in in a sense, you know, that or at least it's it's, like, less gameable than than if you just had to, like, paraphrase a few words, and then you could, beat Pangram. So that that's pretty positive.

Bryan Cantrill:

Wait. And so one of the things that is just remarkable about it, and I and this is what I found in my own experience, is that the false positive rate is very, very low. And I mean, say this as well. I mean, but my, like, anecdotal experience definitely comports with with your with the the way you quantified it. How I mean, can you can you kinda describe how you drove that so low?

Bryan Cantrill:

And can can you describe a little bit about how Pangram four in particular works? Because I think it's it's more clever than or than people realize are very clever. I mean, this is not like an em dash detector. This is the this is much more sophisticated than that.

Max Spero:

Yeah. Yeah. So, I think this is a technical audience. Right? So I I can just, like, go and, like, talk about Yeah.

Adam Leventhal:

Go down.

Max Spero:

Learning models. Great. Okay. Cool. So Pangaram is a classifier model.

Max Spero:

Largely, what we do is we we take an open source LLM, so it already understands language and has a tokenizer, and then we lop off the next token prediction head. And instead, what we do is we add on a classifier head. So originally, was binary. It says human or AI. Now it has more buckets.

Max Spero:

So it has a bunch of buckets that are degree of AI assistance that start at zero and ends at, I don't know what the current one is, 15 or something. And so how we create the training data for Pangram is a method that we call synthetic mirroring. So we take a human document. For example, let's just say a blog post on don't know, give me a topic on a blog post. A blog post on AI detection.

Max Spero:

And then we ask an LLM to kind of summarize this blog post and turn it into a prompt. And so we're gonna take this blog post on AI detection and then turn it into a prompt that says, write me a 200 word blog post on AI detection that hits these three bullet points. And then we feed that prompt into a random frontier model. And so now we have an actual blog post written by a human, and then we have a blog post written by an AI model. They cover the same topics.

Max Spero:

They have the same core ideas, but one of them is AI written and one of them isn't. And the Pangram model can then and we feed these in as training examples to the Pangram model. Say the first one's human, second one's AI, and it learns the differences across a whole bunch of basically small, weak signals here at the text. So it's kind of like the level zero. And then the level one is we also do a bunch of editing prompts.

Max Spero:

So we'll take the blog post on AI detection, and we'll say, make it better, make it more detailed, fix my grammar, all these different things. And then we do clause level labels to say like this, for this clause, this clause that is in the AI assisted output is exactly the same as the one in the human assisted output. So we'll label this clause as human, whereas this quad was mod this clause was modified. So we'll say this clause is AI assisted. And then this sentence doesn't appear anywhere and actually has new information, so we're gonna label this sentence as AI generated.

Max Spero:

And then so our model now has token level outputs that will say for each token, is this likely human, AI assisted, or AI generated? And that's kind of like the long story of this. And then we train this classifier model over the course of several days and then get something that has a pretty good understanding of what the different frontier models sound like across a wide variety of domains. And yes. So our false positive rate, you were asking a bit about this before.

Max Spero:

For Pangrom four, it's about one in twenty four thousand, and we've measured this by looking at

Bryan Cantrill:

Crazy.

Max Spero:

A whole bunch of millions of pre 2022 documents. So, like, we we know for sure that they're probably not AI generated, except maybe GPT-two, but I think we can be confident enough that it's not AI generated. And, yeah. And then similarly for false negative rate, it's about one in 300, so it's it's not perfect. And we do this we measure this by looking at wild chat, which is some it's a dataset of AI prompts in the wild.

Max Spero:

Then and then we regenerate, the actual outputs with frontier models instead of, like, ChatGPT 3.5 or whatever, the dataset was collected on.

Bryan Cantrill:

Yeah. And that is interesting. And I have hit what I believe to be false negatives. It doesn't happen frequently, but has happened. We're just like, God, I am certain this is AI written, but it's reporting as human written.

Bryan Cantrill:

But false negative rate is also like, is low. I mean, it is really quite low. So okay, then you had, what was the first model? Because I mean, I started using Pangram, I think like at the very end of last year or the beginning of this year, when did it first, when was it first publicly? Was I missing it for a long period of time?

Bryan Cantrill:

How late was I to Pangram?

Adam Leventhal:

So

Max Spero:

so you were late, but it's fine because the early models were not as good. So

Bryan Cantrill:

Oh, interesting.

Max Spero:

We published our first model in, like, February 2024. We had a technical report on it. And then somewhere by, like, 2025, we, like, extended it for a longer context. And then at the end of twenty twenty five in, I think, November was when we launched, Pangram three instead of over Pangram two. And Pangram three had this, like, third class.

Max Spero:

It had the AI assisted class. And so I think that was the first time where we were very proud of going from a regime where, like, any AI at all just, like, lights up the whole thing as AI to, like, we can actually get kinda close to telling the degree of AI assistance here. And that's, like, kinda when it got useful.

Bryan Cantrill:

So so I don't think

Max Spero:

you were that you you weren't that late to things. And then I think Pangram four, which came out, a little over a month ago, also feels like a really huge step change to me. It barely see errors now.

Bryan Cantrill:

Why is this do you do you guys know kind of why it is such a big step? I mean, what was the I mean, obviously, I'm sure you made a bunch of improvements to it. But what made it so much better?

Max Spero:

Think adding the token wise head helped a lot. Like Oh, shit. Before, it was kind of just like it it was classifying windows of 512 tokens. And, you know, that that's just, like, kind of low granularity. And so understanding on on a token level that, like, you know, these this could be two AI sentences in the middle of, like, a human written document.

Max Spero:

Like, that's something that, I I think just, like, makes the model better in general.

Bryan Cantrill:

And it's been remarkable. I gotta tell you. I mean, it's been really, really good. And I, again, I think that this is so healthy for humanity. I cannot tell you.

Bryan Cantrill:

Because I think it's important for people I pointed this to you that we updated our own RFD five seventy six on LLMs and LM usage at Oxide to make sure that any blog entry coming out of Oxide is 100% human author and according to Pangram four. To a certain degree, it's like that matters more than how you actually got there because what I think I think people are gonna do this and people are doing this where they wanna know how authentic is this or not. So I'm gonna run this through Pangram. And it's very important that that you not lose the authenticity of your voice. Apparently, Wall Street Journal doesn't care.

Bryan Cantrill:

I mean, apparently, Stanley Stanley Druckenmiller does not give a shit about how you think about his voice. But I

Adam Leventhal:

think for the rest of

Bryan Cantrill:

the the rest of folks, I think they actually do care. Certainly, we care at Oxide. So I I think that, you know, the the more ubiquitous this gets, I think, again, the healthier it is. Because if you're gonna spend a huge amount of time and energy trying to outsmart PanGram four, you also could just write it. I mean, that's another possibility.

Bryan Cantrill:

I mean, just just throwing that out there as, you know, just just spitballing here. But

Max Spero:

That's definitely the thing in education at least of just, like it doesn't need to be perfect, but creating friction on the, like, the easy path just makes people do the the next easiest path, which is doing the assignment themselves. It's great.

Bryan Cantrill:

Yes. It and I also think that, like, there's real value and it's something that this kind of gets to this like, I think this very core and very old question about the role of language and thought. And what role does it play? And I personally think that language plays a lot of role in the way we conceive of ideas and not just express them, but the way that we understand them for ourselves. And I think that that's what people are, I think, worried about in this LLM era is that people lose track of that.

Bryan Cantrill:

And the ideas get sloppy. That that not only is the writing sloppy, but the actual ideas are sloppy.

Max Spero:

And I think so. Absolutely.

Bryan Cantrill:

And so I think that, like, man, it is so valuable to just tell people like, no, no, like don't Like, we're gonna be able to detect that it's LM authored. And I can tell you that like, because I've got college aged kids, the way they use LM I mean, if you want Go to your college professors, if you want detectors that are, you know, LLM detectors par excellence. And they are the curricula are changing very quickly around it about how do we teach in and use LLMs without using an LLM as a substitute. And I think there's a lot of interesting stuff going on. But I I I think that like AI detection, reliable AI detection kind of ends up being the bedrock for all of this.

Bryan Cantrill:

That if you if you don't have it, it's just too easy for it's too tempting. And when you can remove that temptation, I think you can break through to something that's much harder. Are are you I I assume that education is a big market for you all. I assume.

Max Spero:

Definitely. Yeah. Yeah. I think this is it's a huge market, and this is this is the first year that people are catching on that there are AI detectors that work and, you know, aren't just gonna have a whole bunch of errors and false positives. So yeah, and I think this is actually critically important to our future.

Max Spero:

Our children need to learn. And I think if our educational institutions don't get their act together I I think AI is only a part of it. I think there's also a lot to be said about general academic rigor from k 12 into higher ed. But I I think AI just being this easy way out that was not punished at all, there's no friction, and kind of just lets students choose to get an A without doing any work, this is just a really negative for the educational environment in general.

Bryan Cantrill:

Yeah. And I I gotta say, I think it's more widespread outside of education than in it because I think that the there's already the people doing in class writing. There's a bunch of other things that they're doing to like, they were I mean, it is definitely actually Adam Alexander was saying that he like the very first time he learned about GPT, which was like in I mean, was very short at the GPT tube was released. It was from his brother saying, hey, by the way, there's this thing that could do your homework for you. I mean, it's like I mean, of course, the teenagers were on this extremely early.

Bryan Cantrill:

And Max, my my my kids were using GPT to write Nextdoor posts, trolling Nextdoor posts. So if you do not trust the Nextdoor corpus as human authored, even if it's written in 2022, depends on the month in 2022 because I'm telling you, these guys were early, early adopters, and they were like, oh, this is great. It makes me sound like an adult. And they were, they were, you know, able to get next door up in flames. Very easy.

Bryan Cantrill:

And actually, it was so easy. There was little sport to it. So they did it for, like, four days, and then they that was that was that. Oh, man. But the so I mean, like the teenagers have been way ahead on this stuff.

Bryan Cantrill:

I think that the the folks that are getting kind of the wake up call are the people that are have really started to use this professionally in a way that they kind of use it as a replacement for their running. Because one question I wanted to ask you is people are hot about this online. Like you're getting this like rise of like, oh, well, Pangram totally screwed up you know, I record my dream journals every day that I myself write. And I obviously can't share them with you. I mean, their dream journals.

Bryan Cantrill:

It would be I can't share that with anybody. But Pangram is reporting that's a 100% AI and it's definitely you know, it's and he's like, why are people? Why is that happening? I feel like I've seen a bunch of these where people are kinda disparaging Pangram. It's like, what is your dog in the fight?

Bryan Cantrill:

Do you have any any thought on that?

Max Spero:

I I do think there is, like, a big cohort of people that wishes that there were no AI detectors that worked, and they they wish that we could get back to this world where everyone assumes that, like, AI, everything is inevitable. And partly, maybe it's that they find it so convenient for themselves in their own professional life that they find it very inconvenient that now what they produce can be detected as AI. Even though, like, maybe, like, you and I could could have told you could have told them a year ago that it's obviously AI slop. But, like, there was no, like, reliable gold standard tool to that everyone can point to and be like, that's AI. And so I I think it's partly that where people find it, like, it's really inconvenient that Pangram works, and so they wish that it didn't work.

Max Spero:

And then I think there's other things that I've seen on the Internet. So I've seen on Reddit specifically so so there's these tools called humanizers. And these these humanizers, they they sell largely to students. What what they do is they will offer to take your AI text and rewrite it with some paraphrasing model, maybe an LLM with a custom prompt, maybe something else. And their claim is that it will evade any AI detector, like Turnitin, now Pangram.

Max Spero:

And so one of their marketing strategies has been to oh, yeah. And another thing about humanizers is they all also have an AI detector. So, like, when you Google AI detector, like, maybe, like, the first couple of results that you're gonna find is actually an AI detector, but they're not trying to sell you on the ability to detect AI. They're trying to they're just gonna say your text is AI and and try to sell you on humanizing it. So Oh

Bryan Cantrill:

my god.

Max Spero:

So it's this cohort of of tools. And then they'll go on Reddit, and then they'll make, like, astroturfy posts being, I, you know, was, like, falsely accused by my teacher of using AI. My teacher said because it was 76% AI, I'm getting a zero, but I tried it on Winston AI, like the convenient humanizer one, and and and Winston said it was was human. I even put it through the humanizer, and it it you know, my my something like that. Some, like, weird story, and then they they use this as essentially lead gen on the Internet by, like, astroturfing.

Max Spero:

So that's sort of the other way that I see a lot of stories about AI detectors come up, which seem obviously fake.

Bryan Cantrill:

Yeah. Where people I I mean, I would love to, you know, obviously pass those through Pangram. Just so

Max Spero:

you know They're all AI. 100%.

Bryan Cantrill:

Are they are they all 100% AI? Yeah. Of course. I mean, of course, they are. Yeah.

Bryan Cantrill:

That is well, clearly like a a vested interest. And yeah, mean, I think it's like people don't want these these detectors to be present because they kind of like are are trading on the bullshit that they're getting away with. And it's yeah, this is But I also feel like just as you said earlier, you're actually not getting away with it. That people do know this is OM generated. I know that you feel that you're getting away with it, but you're really not.

Bryan Cantrill:

And it's actually the sooner you realize that, and the sooner you realize that everybody can see that or everybody. That p that people who read actually read this. I mean, it's a good thing. It's like, don't we actually want to be able to know when we're reading something? Like, don't you person that generates a bunch of LLM slop?

Bryan Cantrill:

Like, do you read at all? Are you, like, a write only person? I mean, are you like, I I definitely have this question. You know? Like, do you like you're like, do you like to read stuff that's LLM generated?

Bryan Cantrill:

I mean, do you enjoy it? Are you?

Adam Leventhal:

Oh, Bryan, I I I on this topic, I don't know if you saw this, but someone on social media was talking about accelerating their reading by having LLMs summarize the great works. Oh, jeez. And it's like it's like yeah. Why I mean, why bother?

Max Spero:

That's that's not accelerating your reader or your reading. That's

Adam Leventhal:

No. No. No. No. No.

Adam Leventhal:

But

Max Spero:

it No.

Bryan Cantrill:

Not anyway

Adam Leventhal:

answers your question. There are some people who clearly don't give a shit.

Bryan Cantrill:

They don't. And which is and I think that like the the sooner we are we have an apparatus that allows us to because I think the reason this is so important, Max, is because this can get us back to trusting arbitrary writing again. And Yeah. Because right now, I mean, this is and I'm sure you see the same thing where societal trust, institutional trust is being eroded because we don't know what to believe and what not to believe. Because also, by the way, like, the the LLMs don't know what truth is.

Bryan Cantrill:

Like, they can't. Mhmm. And so I mean, do you I do do you listen to the Shell Game, Evan Ratliff's podcast? No. Oh my god.

Bryan Cantrill:

So this is very good. This is a podcast where so into the second season of Shell Game, and then you can, I guess, you can ring the chime for we had Evan on on Oxide and Friends to to you can, like, listen to that as bonus material? But he a He starts a company with only LLMs, like personified, And anthropomorphized it is wild. And you realize that it's not even like they're lying. Like, they literally don't know what truth is.

Bryan Cantrill:

And in in one of the I mean, there are many just, like, insane episodes. But in one of them, the LLMs, the agents, the employees decide that they need to plan a hiking off-site. So they are all planning like, what hike should we take? And they're like having arguments on which hike. And he is like, and they all do this via Slack.

Bryan Cantrill:

And he is like, you are not You don't have legs. Like, No one can go on. I could stop. Like, you're just spending tokens. Like, stop talking about hiking.

Bryan Cantrill:

None of you. And they're like, well, we see that he's Evans asking us to stop. But the reality is a hike would be exactly the kind of off-site that we need. And it would really allow us to get some fresh air. And it's like they it it is just other worldly.

Max Spero:

That's so interesting. They don't know that they're not people.

Bryan Cantrill:

They don't know that they're not people exact wait. But they also do know, like, intellectually, they're not people. And actually, I don't know if we talked about it with Evan, Adam, but one of the most amazing episodes was where he takes the CEO, the LM CEO, and he has him talk to a friend of his who's a psychotherapist or psychologist, asks him these probing questions. And he's like, why do you keep talking about hikes? Because you know that you can't do that.

Bryan Cantrill:

It's like, know I can't do that, but I just feel like the spirit of that is so important for my fellow employees. It's very, it's And other I keep pointing people to that who are like, well, think these things are gonna take over the world. I'm like, yeah, you should listen to this podcast. And I'm like, you know, when the singularity comes, they may all the bots may all be like, they may escape hugging face and then try to conspire to take hikes together. You know, like, that's not impossible.

Bryan Cantrill:

And I think that like the so that they don't know what truth Like they can't know what truth is because all they've been presented is the training data that is like the digital world. They actually don't know what truth actually is. They have to infer it. And so that's like the kind of the issue is that like, I'm not saying that you're not you're not lying exactly, but the LLM doesn't actually know what the truth is. So now you end up with something that's actually like not true that I can't trust.

Max Spero:

Yeah. And I think that's sort of the problem is is this, like, disconnect from reality. Like, yeah, all LLMs can trust is the inputs that they receive, the context. You could, like, hook up Claude to a sensor, well, a sunlight sensor or something, and it could, like, receive that input. But you could also, like, give it, like, a fake sunlight sensor and tell it it's sunny when it's not.

Max Spero:

Totally. But I but, I mean, is that true for humans too? Like, that's, like, the the premise of The Matrix? Like, you can only trust your your own, like, physical senses and, physical inputs.

Bryan Cantrill:

Well, this is what I mean. It gets you into these, like, philosophical arguments. And, like, you know, next thing you know, you're drifting into, like, a discussion about solipsism in a freshman section or whatever. I mean I I yeah. I mean, I think that he, absolutely.

Bryan Cantrill:

I I think that that and I mean, I I actually think that I'm I'm long philosophy, Adam. I think the I think the philosophy is, among my

Adam Leventhal:

You're saying you're you're gonna you're gonna take a young man and say, the future is philosophy. Exactly. Exactly. Graduate saying plastics.

Max Spero:

Don't be graduating from science. Study philosophy.

Bryan Cantrill:

Study the Greeks, man.

Adam Leventhal:

Study the Greeks. It's like, what

Bryan Cantrill:

is more what is more Lindy than the Greeks? I mean, you you know, there's a a lot of these things have been thought about before. So I think that yeah. Exactly. Although, actually, Adam, not to Max, we do a the first episode of the year, we do a predictions episode of one, three, and six year predictions.

Bryan Cantrill:

As you can imagine, it's just been wild these last couple of years. Adam, I don't wanna front run one of my predictions, but I I think that this desire for human authored content is gonna be so great that I am going to I don't know. This is very like this is this is very antisocial for me to make a prediction here in in August. I think that there's gonna be a scandal involving a university or civic library where books are being secretly, like, sold out the back. Where someone where one of the frontier models is like, yeah.

Bryan Cantrill:

No. I was like I there's a shell company that was paying me, like, cash for all of these books that no one had checked out for four years. And, I mean, it just feels like that could be hap could happen tomorrow. Probably it's happening tomorrow.

Max Spero:

There's gonna be like a 50% of some major library's books are missing.

Bryan Cantrill:

Yeah. No. Totally. You're just like, wow. That's another book I can't like, is the library looking thinner to you?

Bryan Cantrill:

I just like, this is I I just feel like the I was in the stack studying, and I just swear to God, I've never been able to, like, see the restroom from where from where I study. And I could not only can I see the restroom, but I could see outside? I mean, you just wonder if there's like are the books disappearing? Because I think that there's gonna be a real I I and I think this is the the I mean, the other thing that is so important for you all in AI detection. I think, mean, for the frontier model companies, they're gonna want that stuff too.

Bryan Cantrill:

I mean, you mentioned that earlier that this is gonna be an area of interest.

Max Spero:

Oh, yeah. I mean, if you're buying, training data from a scale or other labeling company, like, and then they're actually just like using, like generating a bunch of answers with like, Claude, like that that's, that's gonna be a big deal.

Bryan Cantrill:

Totally. So another question that kinda came up in the chat that I definitely have too is how much do you feel that you are kinda chasing these frontier models? I mean, you had a post really recently that it was really fascinating that you're seeing actually, if anything, divergence where it's, like, easier to detect? Maybe I misinferred that.

Max Spero:

But I think so. Well, so two things are happening at once with the LLMs. They're getting more capable. They are as you continue to post train these very capable models, teach them correctness, they learn preferences. I I would say, like, today's LLMs have a much more narrow distribution of what they output than, like, g p t two.

Max Spero:

G p two is trying to predict to to, like, model the full human output distribution.

Adam Leventhal:

Oh, interesting.

Max Spero:

Cloud Opus five today, it's very much like, I know what's correct. I'm gonna tell you the the correct token, not what an average person might say. And so I think because of this preference and opinionated token choice that's being baked into these models, you actually see this in the pros. So I've seen people talk about Fable ish. Claude Fable just speaks in such strange ways.

Max Spero:

Does. It feels like an alien.

Bryan Cantrill:

It does a little bit. You feel the same way, Adam? I feel it does too. It is.

Adam Leventhal:

I keep on telling it to I'm like, I don't know what that means. Knock it off. Like, just just talk normal.

Bryan Cantrill:

It just talk. Can you be normal, please? Just for it does not it it sounds like it sounds weird. I mean, it sounds like kind of like strangely erudite or it's like it's but not, yeah. It it it's it's it's very odd.

Max Spero:

So so Goodfire has this pretty cool tool. Goodfire is a AI interpretability company, and they have this tool called Silico, which is a research agent that can help you do mechanistic interpretability work to basically, like, dig into the internals of your model and understand what's going on. And so I was trying to play around with the Goodfire Silico using Fable and dig into Pangram. And very quickly, I felt like I am going to get AI psychosis from using this. I do not understand what Fable is telling me.

Max Spero:

I do not have the full context. It's giving me papers, but then it's then summarizing these papers. And I feel like I understand the summary, but I don't understand the paper. And, like, if if I go down this rabbit hole for another, like, six days, I'm just I'm gonna come out of it with psychosis.

Adam Leventhal:

I that

Bryan Cantrill:

is funny that you say that. Adam, you had that happen where you're just like, I feel like we are and you, like, you can feel like it's kind of in the driver's seat more than you are. And you're like Mhmm. Where it was encouraging you, you're kind of now encouraging it. You're like, wait, I've lost track of where we are right now.

Adam Leventhal:

Right. You're in that sort of feeling of being aloft and slightly untethered.

Bryan Cantrill:

Totally. And where but it but it's assuring you like, no, we're onto something like

Adam Leventhal:

We're on bedrock. Exactly.

Bryan Cantrill:

We're on not only with bedrock, this is the we're unified field theory, man. This is right around the corner. We're we just need a just a couple more tokens here. The Goodfire stuff sounds really interesting, though. So that's

Max Spero:

the Yeah. It's a it is a really cool product, and I I think that the team is really brilliant. And I feel like it was maybe just, like, I I I think it does need to come with, like, a a warning sticker of, like, you need to understand these concepts before you, like, dive in and use this tool. Because I can imagine for an AI researcher who really knows mechanistic interpretability, they they could just dive right in and use it as a a partner in crime and get a lot done. But for me who doesn't know anything, I felt like I was I was in the passenger seat and and we were going way too fast down the highway.

Bryan Cantrill:

Wow. That's really interesting. And what that was and so and you said you were using Goodfire on Pangram? Yes. Yeah.

Bryan Cantrill:

Were you using I was trying to interesting.

Max Spero:

Build some probes to try to understand. We're we're we're training some probes to get a calibrated confidence out of Pangram. Like, you know, right now we have this, like, high, medium, low. It's kind of just based on the the logits. But I I wanted something better.

Max Spero:

Like, could could we produce a confidence interval, like, you know, 94 to to 98% chance that this text is AI generated, something like that.

Bryan Cantrill:

Yeah. That's interesting. I got so another just a couple questions. Know you gotta split here in a sec, but just a couple quick questions. Are you able to detect the model that was used at all?

Bryan Cantrill:

Yes. Do you have

Max Spero:

So we we built an internal probe in on into the pangram model weights, and it does pretty well. It has 90% top one accuracy. So that means, like, 90% of the time, it it gets the model family correct. It'll tell you it's Claude family, GPT family, Gwen family, etcetera.

Bryan Cantrill:

I that is I'm super curious about that. Because I the first time we start and this is one, unfortunately, where, the the the criminals were a little ahead of the policing here. And we got when we first started getting the sophistication that people had in generating LM based materials for oxide was ahead of our ability to detect it. And the first couple that we had, I was mesmerized by it. I'm like, oh my God, this is not, this is all like made up.

Bryan Cantrill:

And I had so many I mean, there's one of them where I I think it was completely fraudulent, where it was the person doesn't exist. Because we're a remote company and like there's a kind of people fantasizing about getting a job and then with those paycheck or just show up to, you know, wherever they are that they live. And I want to almost like exchange them like, look, I'll pay you a thousand bucks if you tell me the prompt that you used. I'm just curious. But I would love to know, like, the model.

Bryan Cantrill:

I mean, it'd just be interesting to know. So there's so you're able to do that with some level of accuracy?

Max Spero:

Yeah. Yeah. Internally, I would love to turn this into a product feature. Think it's really got to get up to, like, 98% accuracy for us to feel good about it. But, like, 90%, like, we're we're close.

Max Spero:

We're not far.

Bryan Cantrill:

I I I think even expressing that ambiguity would be fine. Like, I kinda like I I mean, I don't have any idea right now. Right? And if you were to just give me even if you were to to to to, you know, couch that with, like, lots of ambiguity, like, I think this might be this model, but I don't know. I've got I actually do not know.

Bryan Cantrill:

Because I'm, like, curious, like, how much of it is open weights? How much of it is, how much of it is kind of the frontier models? And and then the other question is, like, do you have any ability to get to the prompt at all? Have you been able to to get to some of the prompt and some of the stuff?

Max Spero:

Definitely an open area of research. I think there's two things we want to understand here. A is like, what is the prompt that could have plausibly generated this text? I I think that that's pretty doable. I think we could train a reverse pangram to do that.

Max Spero:

But I think the other thing that is interesting is just, like, what was the like like, how much context did the LLM have? Like like, did it get a long prompt? Did it get two bullet points, or did it get 15? And and I think that's going to probably be increasingly important as norms shift around AI writing and to make sure we are still giving people relevant information, which is not just was AI used, but, like, what was the degree of human input? For example, like, the the drug and miller op ed.

Max Spero:

It's gonna be very different if it was Totally. Five bullet points or or, like, five pages.

Bryan Cantrill:

Totally. And I think and that's an open question, I feel. I feel like I do not know from that. I feel that you could have I mean, it'd be interesting to experiment with it, but I you could probably ask Claude to just write an editorial against the treasury department actions and get something that would not be wholly dissimilar. So Yeah.

Bryan Cantrill:

It would be really, really interesting to know how sophisticated that was or wasn't.

Max Spero:

Well, the question is then going to be like, oh, somebody works with Claude to build these 10 bullet points and then puts it in and then feels like they don't they didn't get credit for their 10 bullet points when they actually came from Claude in the, like, earlier part of the conversation. That'll be kind of an interesting question of how do we measure our accuracy here? How do we make sure to attribute how much was actually coming from the person?

Bryan Cantrill:

Right. Well, again, I think this is extraordinary work you and the team have done. I mean, it's so great to I mean, to hear how this kind of problem was thought to be either irrelevant or impossible. And you're like, it's neither of those things. It's actually extremely relevant and hard, but possible.

Bryan Cantrill:

Because I mean, think you are four is far and above anything else out there I found. I mean, is guys have really innovated here. And the the technical report is a great read. It's really, really, really great stuff. And I'm really excited for you all.

Bryan Cantrill:

I I you know, Godspeed on on on staying caught up. Hopefully, that this is this doesn't become an arms race. Exactly. Or or that hopefully you're able to to to triumph. But I I really do think that, like, there's so much of this that you are gonna triumph on because you're just gonna make it harder.

Bryan Cantrill:

And just by making it harder, like at some point, it's going be easier to write your own stuff. And then use an LLM as an editor, use an LLM to help brainstorm what you're going to write about or what I use the LLM. All the great ways you can use an LLM that doesn't actually chip away at the trust that one might have in your writing. There are lots of ways to do that.

Max Spero:

Definitely. Yeah. Yeah. LMs are great technology. Just, you know, don't don't offload your life to them.

Bryan Cantrill:

Don't offload your life to them. Yeah. Exactly. I mean, I I I think just like, you know, don't you wanna make sure that you're you're just because you got the ability to, like, live on your couch every day, it's like you wanna stay fit. You wanna you know, you you should you can walk and everything else.

Bryan Cantrill:

So Yeah. Well, thanks so much

Max Spero:

for having me. It was it was really fun to be on here. It was a really great conversation.

Bryan Cantrill:

Thank you for joining us. Standing invitation, if you wanna come back for our predictions episode at the beginning of the year, that one's gonna be, Adam, we gotta see if can get Simon Wilson back as well. When Adam, Simon Wilson, and I together made an excellent prediction about the pope that if he was broadly vindicated, the pope coming out in defense of of AI for humanity.

Max Spero:

So Yes. Nice.

Bryan Cantrill:

Exactly. Well, looking forward to another great predictions episode. But Max, thank you again. And really terrific stuff. Terrific work for you and the team.

Max Spero:

All right. Thanks all.

Bryan Cantrill:

Awesome. Thanks, everybody. See you next time.