TBPN

Diet TBPN delivers the best of today’s TBPN episode in 30 minutes. TBPN is a live tech talk show hosted by John Coogan and Jordi Hays, streaming weekdays 11–2 PT on X and YouTube, with each episode posted to podcast platforms right after.

Described by The New York Times as “Silicon Valley’s newest obsession,” the show has recently featured Mark Zuckerberg, Sam Altman, Mark Cuban, and Satya Nadella.

TBPN is made possible by:
Ramp - https://ramp.com
Public - https://public.com
Cisco - https://www.cisco.com
Console - https://www.console.com
CrowdStrike - https://www.crowdstrike.com
Figma - https://www.figma.com
MongoDB - https://www.mongodb.com
NYSE - https://www.nyse.com
Railway - https://railway.com
Shopify - https://www.shopify.com/

Follow TBPN: 
https://TBPN.com
https://x.com/tbpn
https://open.spotify.com/show/2L6WMqY3GUPCGBD0dX6p00?si=674252d53acf4231
https://podcasts.apple.com/us/podcast/technology-brothers/id1772360235
https://www.youtube.com/@TBPNLive

What is TBPN?

TBPN is a live tech talk show hosted by John Coogan and Jordi Hays, streaming weekdays from 11–2 PT on X and YouTube, with full episodes posted to Spotify immediately after airing.

Described by The New York Times as “Silicon Valley’s newest obsession,” TBPN has interviewed Mark Zuckerberg, Sam Altman, Mark Cuban, and Satya Nadella. Diet TBPN delivers the best moments from each episode in under 30 minutes.

Speaker 1:

I need some soundboard.

Speaker 2:

Here

Speaker 1:

we go. Yes. Today on TBPN, we're talking about model Mayhem. Everyone's launching new models. Slow summer, but not for the AI race.

Speaker 1:

You got x AI unveiling Grok 4.5, the first model built specifically for coding and AI agents developing collaboration with Cursor. Talked about it a little bit yesterday, but we have some more benchmarks, some more discussion on the timeline about where this model fits in on the Pareto frontier. Also, why it might be outperforming so well on CursorBench. Lots of debates there. Meta announced Muse Spark, a new agentic coding model with Mark Zuckerberg.

Speaker 1:

Returning to X for the first time in basically a decade. Three years ago, he posted one joke post about launching threads, but he has not been an active user. But the AI vortex sucked him in and he's got a post.

Speaker 3:

Oh, I think he's an active user, John.

Speaker 1:

You think so?

Speaker 3:

He's just not an active poster.

Speaker 1:

He's just not an active contributor. You're calling him a lurker.

Speaker 3:

I'm calling him a lurker.

Speaker 1:

You're calling him a lurker?

Speaker 3:

I'm calling him a lurker. I think he's absolutely glued. You think so? I think so.

Speaker 1:

You really think so?

Speaker 3:

I think so.

Speaker 1:

I feel like I don't know. So busy, so much other stuff going on. I I feel like he I feel like most people

Speaker 3:

The busiest people I know Okay. Are not active on x Yeah. But they they are on x a lot.

Speaker 1:

Sometimes. But there are there's a different class of person.

Speaker 3:

You can just quiz

Speaker 1:

them. Screenshots come to them via Slack or or via text message because they have a team that's monitoring the timeline and then it's delivered. This is the important this is the important stuff.

Speaker 3:

They're calling him Mark Lerkerberg.

Speaker 1:

Lerkerberg. But the other big news, OpenAI just released GPT 5.6. Let's go. Let's A new general purpose model with expanded coding and agent capabilities alongside GPT Live, which we talked about yesterday, a new real time interactive voice experience. Reactions are great to 5.6.

Speaker 1:

Bunch of interesting details here. You had people have been identifying that while there is a frontier and there are just a few companies that are actually on the frontier, frontier is spiky and they have different flavors to them and reasons to pull different tools off the shelf. People are drawing analogies between Fable five being some, you know, recluse genius and five point six being a, you know, collaborative coworker that you love chatting with or something like that.

Speaker 3:

Said Yeah. I don't know how else to describe it, but Fable five is like Kendrick on Good Kid, Mad City and five point six Soul is like Chief Keef on finally

Speaker 1:

Now it makes sense to me. Yeah. Thank

Speaker 3:

you. I just wanted to put it into, you know, 2010

Speaker 1:

Really

Speaker 3:

hip hop Yeah. Like terminology.

Speaker 1:

Really, really clear there. Thanks for clearing that up.

Speaker 3:

I mean, the funny thing is that will be very explicit for like a 100 people in the whole world. This one's for

Speaker 1:

you. The most interesting benchmark to me has always been Arc AGI v three. We've interviewed the team over there many times and had a lot of fun understanding what goes into that that benchmark. And 5.6 SOL scored a massive 7.78%, which is tiny considering that the whole point of Arc AGI is that a human should be able to get 100% on it, and basically any human. So it is a true test of AGI in the sense of, you know, can you give this test to just actually anyone, not the crazy math projects, the crazy hard programming projects, the hacking.

Speaker 1:

All of that stuff is very economically valuable, of course. But there's a more interesting question where, you know, when there's less of a spiky frontier and there's just this question of what is something that anybody can do that AI can't? Because we've been searching for those and the Arc AGI team has done a fantastic job building out these puzzles that AI has historically struggled with. Arc AGI, one, the model sort of climbed two, became a little bit more complicated And now, three, we're starting to see glimpses of progress, although 7.76% isn't 99%. We're nowhere near saturation, but it's still a huge jump.

Speaker 1:

Opus 4.8 had 1.5%. So GPT 5.6 Soul is showing more generalization, more spatial reasoning, more puzzle solving abilities. So fun, fun stuff. The blog post is also very, very fun because it includes games. I'm a big fan of the the GPT 5.6 launch games.

Speaker 1:

I got immediately sucked into the to the the sailing mini game, which is very high fidelity, but also delightful to actually play.

Speaker 3:

Should we play it?

Speaker 1:

Yes. We should definitely play it. Yeah. Saltwind. You you you guys play it.

Speaker 1:

I want production team to see what they can get. I think my time was twenty five seconds.

Speaker 3:

And is this hosted on a on a site?

Speaker 1:

I think this is I I mean, this is hosted on the OpenAI blog. But I think the the idea is that you could vibe code this in the latest GPT 5.6 in the app, in ChatGPT, and then deploy it and have someone. Are you trimming the sails appropriately? Because it looks like you're losing speed. You're losing wind.

Speaker 1:

It's not working. I'm going to smoke you. I got twenty five seconds. Wow. Amateur hour over here.

Speaker 1:

Look at this. Yeah. Yeah. Well, the whole the whole game, which you probably missed, is that there's a little bar there where you have to trim the sails to be in the sweet spot of the wind while you're turning. So as you turn see the bar?

Speaker 1:

There's a recommendation for where you put the sails. You gotta keep that line in See? It's moving over. You gotta you gotta press the

Speaker 3:

Oh, I I

Speaker 2:

I see.

Speaker 1:

Yeah. Exactly. Keep trimming those sails while you steer the ship. This stuff is very, fun.

Speaker 3:

One interesting data point from the livestream, which was just an hour ago. They said, already, Soul has been transforming our research program as one example of GPT 5.6 Soul autonomously post trained 5.6 Luna.

Speaker 1:

Yeah. That's fair.

Speaker 3:

A lot of people are having fun with that. Dylan Fields says, a lot of people wanna compare Fable versus 5.6 Soul. This is a mistake. They're apples and oranges. Despite all the research achievements, we are still very, very early in exploring the tech tree for model training.

Speaker 1:

Cool. Sorry. I'm just getting set up again. Oh, yes. I I I do think that didn't Dylan Eberscotter write something about this?

Speaker 1:

What was the the essay he wrote about interactive memes and this idea of, like, generative AI enabling these vibe coded mini games. Like, we've been seeing a bunch of them with like the Copybearer simulator, the Coconut simulator where it's something that's just a joke that's funny for like a few people. But and normally, you would instantiate that in a in a tweet. Maybe if you were getting really crazy, you'd do a Photoshop edit of a meme. But now, you can go and create a full mini game, something that runs in the browser.

Speaker 1:

And soon, something that runs in Unreal Engine and can actually be distributed on the Steam store. We're already seeing that with like the data center simulators and all these funny simulator games that are going on Steam. All this all the all the advances in the coding model certainly speeds up the ability to actually deliver polished software. I'm I'm I'm particularly excited for like

Speaker 3:

Dylan's title was The Future of Entertainment is Interactive.

Speaker 1:

Yes. Yes.

Speaker 3:

But but, yeah, that that's part of what I honestly love about AIs. There's a lot of things you can make now that never would have made sense

Speaker 1:

Yeah.

Speaker 3:

To make because they would have taken you four days and it was good for like a small laugh. Yeah. Now, you can do it in four minutes. Yeah. And and it's just fun.

Speaker 1:

Yeah. I I I think there's gonna be there's if you have some sort of like small custom some sort of custom functionality in your business, it feels like there's Is this the David Senra simulator? Why is this David Senra?

Speaker 3:

Late nights in a Miami abandoned apartment Backrooms. Complex in 2015 just recording podcasts and reading.

Speaker 1:

This is very creepy like a horror backrooms liminal space game.

Speaker 3:

Stanley Tang Oh, yeah. Co founder and CPO over at DoorDash says, I have an insane magic trick that so far none of the models can figure out including mythos. It's a bulletproof trick that I've shown to a 100 plus people including magicians that couldn't figure it out. It's not anywhere on the internet. Only way to know it is through first principles reasoning.

Speaker 1:

Mhmm.

Speaker 3:

Told everyone I'll believe in AGI when it can crack this trick. Well, GPT 5.6 just did. How? I want him to I want him to actually open like Well, now okay, like give us now that Yeah. Now that a model cracked

Speaker 1:

it a lot of tricks are, like, sleight of hand. So is he uploading a video or something? Like

Speaker 3:

Well, yes. So John Palmer says, I have a hilarious joke that so far none of the models think is funny. It's a bulletproof joke that I've told to a 100 plus people, including comedians, and no one laughed. It's not anywhere on the internet. Only way to know it's funny is a first principle sense of humor.

Speaker 3:

Told everyone I'll believe in AGI when it tells me a joke. The joke is funny. Well, 5.6 just did.

Speaker 1:

Huge, huge news. Huge. GP 5.6 is a Porsche. Fable is like Warp Drive. I had a different experience.

Speaker 1:

If Fable is an f one car, 5.6 Solit Ultra is a Tesla Model X Plaid, does it find things that Fable misses during plannings and coding? Yes, most of the time. But for the hardest problems, does Fable routinely find things that 5.6 doesn't? Also, yes, some of the time. Is 5.6 way faster and affordable?

Speaker 1:

Yes. With an unlimited token budget, what am I currently using 95 plus percent of the time? GPT 5.6 from Sikhi Chen. So interesting take that the Pareto frontier is alive and well and everyone's duking it out for their slice of the AI opportunity. Very interesting seeing how the market share is shifting during a time of acceleration.

Speaker 1:

You have multiple companies that are growing revenues, even accelerating revenues, while market share is declining because the overall market's growing so fast that you that if you're only growing at 300% and someone else is growing at 400%, you're losing market share where you have, like, one of the greatest businesses by modern metrics. Very, very interesting dynamics in AI.

Speaker 3:

It's also funny because yesterday with Ben Thompson, you were like, some slow summer. And then in in the in the span of twenty four hours, you get Fox four five Muse 1.1.

Speaker 1:

Yeah. I mean, this isn't as dramatic as the AI talent wars. It's not as Word. Dramatic as

Speaker 3:

Rippling deal.

Speaker 1:

Yeah. Yeah. This is this is new technology. And and there's only so much to there's only so much of a take to be given around these things. Although AI 2040 launched today, the sequel to AI twenty twenty seven, that's something that's more of a thought provoking piece that you can debate and interrogate and talk through.

Speaker 1:

I'm sure we'll go through some of it because they pose a couple interesting ideas of where AI might go and where they want it to go and how they want the industry to develop, sort of advocating for a slowdown generally. But it's an interesting way they puzzle piece all the different geopolitical chips on the table. Of course, people are joking about the lead is widening because the the Anthropic and OpenAI version numbers over time. GPT six is predicted. And it is it it the the the model numbering we were talking about this this morning that the numbers, they sort of don't mean anything anymore.

Speaker 1:

Do the model numbers mean anything in particular? It used to be the model number was the pre train and then the then the version number was the post train, but then that sort of got flipped around. And now it's just like, are you do you feel like you're competing at a four class or a five class? So I wouldn't be surprised if we saw like Muse, Spark, not release Muse Spark two, but Muse Spark six or five and jump straight. I mean, Samsung wound up doing this where they jumped to the year, like, sort of like the car manufacturers where, you know, there's a five series BMW, but then there's also just the 2027 because that's the actual model year that's relevant.

Speaker 3:

Twenty twenty seven five series.

Speaker 1:

Yeah. Which is sort of odd. And and we're sort of like duking it out between those. Do you have

Speaker 2:

Yeah. Mean, I I think post reasoning models, you just have like a different way to scale the models besides just pretraining. Yeah. So it's hard to bake that all into one number that like is evocative of both those like two ways.

Speaker 1:

Yeah. So the number is becoming closer to the year in the second decade of the twenty first century, basically. It's just like, is this on the frontier in 2026? You'll probably see a six by the end of the year in front of the models that are leading in the year 2026, something like that. I'm very interested with Google strategy because the the rumor is that 3.5 Pro will be coming out this next week, I believe.

Speaker 1:

But I it was it's very odd going into the Gemini app right now and seeing that there's 3.5 Flash, but then you have to go back to 3.1 Pro. I think 3.1 Pro is the most advanced model, but they default you to 3.1 Flash Lite. And I would expect them to jump just forward to four, but I think that they're gonna do 3.5 Pro, but it's been a little bit of a slower cycle there. As silly I mean, obviously, all these numbers don't really mean anything. They're marketing terms, but they I I still think they do actually stick in people's mind, and so there should be some strategy around them.

Speaker 1:

Mark Zuckerberg is on a press tour. He's talking to the legacy media for the first time in a long time. Andrew Bosworth, the CTO of Meta, also did an interview with the head of The Atlantic, dug into some of the launches around the glasses, and then also had a whole discussion in that podcast around the goals of the keystroke logging thing. It was it was interesting. I mean, was framed as like, you know, like a tough interview around surveillance in the workplace, and it certainly the headlines were very scary.

Speaker 1:

I don't know where I sit on it because I've I kind of always assume that everything you do at work is logged in the sense that, like, if you're on a work computer and every web page you visit is is going through the network and monitored for traffic and security purposes, and all the code you write, and all the emails you write, and all the documents stored in the shared document. It doesn't seem that crazy to go to keystrokes, because everything is already so monitored. But he was framing it as more of an experiment, something that they weren't sure was gonna pan out, something that they allowed everyone in, everyone at Meta, so there certain sections of the workforce that were by default opted out. So anyone who was working on confidential or sensitive information was opted out of that program by default. He said he himself, Andrew Bosworth, was opted out of that program because he has a bunch of legal holds because they're getting sued all the time.

Speaker 1:

So they can't be recording everything, I guess, that he's doing because then that would be admissible in court. And so all of a sudden the lawyer who's suing him would say, Okay, great. In the email, you said, we don't want to do this. But before you Let's

Speaker 3:

typed see that your writing process.

Speaker 1:

Exactly. Yeah. Let's see what sentence you typed and then deleted. Like, what word did you use before minimal impact? Did you say medium impact or whatever?

Speaker 1:

You know? So so he was opted out. And apparently, I I think all of the Meta employees who were part of that program were able to turn it off indefinitely. You could toggle it on and off. And the idea was that they wanted to collect information on how work plays out over a twelve- to eighteen month period, and they couldn't get that from any sort of data labeler because they needed to have very high skilled workers actually chopping wood on projects for a long, long time to see how projects go from start to finish.

Speaker 1:

So basically, like how do you compact the longest possible rollout, not just a single chain of code, but an actual series of meetings and decisions and trade offs and everything that goes into making a decision in a white collar workplace, like how do you actually reason through all of that, it's hard to distill that from just, Oh, well the code got written this way, so that's the right way to write the code. The the code might have gotten written that way because a lawyer said, hey, oh, we have to do this. And then the marketer said, oh, well, we have an activation with this person, so we need to integrate it this way. And then the business people came in and said, oh, well, like the margins will be better if we write it this way. And so it's not entirely first principles software engineering all the time when you're actually building real products.

Speaker 1:

So interesting to see him sort of step into the a tough interview and sort of lay out his side of the story. But Mark Zuckerberg is in Bloomberg today pledging aggressive pricing with Meta's first pay to use AI, which is a funny framing for just an API for a model, but that's the way Bloomberg put it. In a crowded market for AI tools, Mark Zuckerberg wants to win on price. Meta Platforms unveiled a version of its most advanced artificial intelligence model, Muse Spark 1.1, that includes a new paid tier for developers, marking the first time Meta has charged businesses for access to its models and providing a new revenue stream. It'll be among the most affordable options in the market, Zuckerberg said in an interview ahead of the release.

Speaker 1:

Since this is not an open source model, this is, I think, the first time that we're doing a real serious API. And the pricing is going to be very aggressive and attractive. Makes sense. I mean, they own the data centers. They're very efficient at building data centers.

Speaker 1:

They should be able to serve a model efficiently. The new model standout improvement is is is in its agentic capabilities, the Meta chief executive officer said. Agents are a big theme of AI this year with the label applied to systems that can can complete multistep tasks on behalf of the user. Zuckerberg described Newspark one point one as having, quote, state of the art or very close to it, agentic reasoning and tool use. The model is also greatly improved when it comes to coding, and Meta employees are using it internally to build products and features for various apps.

Speaker 3:

Yeah. My big question is how how quickly do they move all of their internal workloads Yeah. Onto their own models. Yep. So they're buying they're getting access to models through Google Yeah.

Speaker 3:

Anthropic and OpenAI. Yeah. I think that a lot of companies will look to Meta's own actions Mhmm. As a way to basically validate whether or not they should be using this model themselves. Right?

Speaker 3:

Because it was just within the last month that Google had said, like, hey, we don't have capacity. We don't have enough capacity for all of Meta's demand for our models. Yeah. And so, yeah, they can't they they can't get enough AI elsewhere, at least from some providers. Yeah.

Speaker 3:

And so how much of their workloads will they be able to run themselves is the big question.

Speaker 1:

Yeah. Meta was one of the first companies to sort of reportedly be token maxing and have a leaderboard and all of that. If you have your own model and your own data centers, the incentive to token max is much, much higher because you're just paying the electricity on the cards that you're already depreciating. So you should sort of lean a little bit back into that. Not that you want to be fully token maxing, but you do want your employees using the tools that you've built as efficiently and as effectively as possible.

Speaker 1:

It's just way cheaper to explore when you're not paying margin on another closed source model, and you're you're not paying anything else and you're actually improving the model. So it makes a lot of sense for them to roll this out broadly. The interesting take that Ben Thompson had, which we didn't get to yesterday because we ended up spending the whole interview talking about Xbox, but the interesting dynamic is that when you are willing to sell API access, you're willing to sell compute directly and then you're also using your own tool internally, it creates this economic incentive internally that you you have an incentive to always go with the most profitable, the most the most economically efficient outcome. That can be very good for business, very good for the investments that they made. The the trick is that you can wind up in a little bit of a situation where your business team or your enterprise sales team goes and sells all your compute capacity or all your chips, and then internally, your team is frustrated that they're not making enough progress.

Speaker 1:

So there's a little bit of a dance there, but in general, it's a forcing function on the internal use of their tools to say, Hey, wait. Why is someone willing to pay five times as much than what we're willing with the value that we're creating here? We spent $1,000,000,000 on energy consuming our own LLM and someone showed up and said, Wait. We'd pay you $5,000,000,000 for that same compute power to run a different model and do a different task. It's like, why is their model not economically valuable internally?

Speaker 1:

That would be the question. The flip side is that they do have low cost, so they should be able to say, oh, yeah. We actually did we yeah. We we we inferenced Newspark one point one internally, and we improved the ad model. And boom, we made a bunch of money.

Speaker 3:

And and these are the same trade offs and decisions that every lab is having to make is how much how much compute do we allocate towards research, towards internal use, towards the APIs Yep. Subscriptions

Speaker 1:

Yep.

Speaker 3:

To free plans Yeah. Etcetera.

Speaker 1:

Yeah. There was that funny semi analysis deep dive into Anthropix forecast. And in there, I mean, some staggering numbers, really, really optimistic. But the flip side was was Ed Zitron was taking shots at the fact that they had EBTIT. EBTIT.

Speaker 1:

Earnings before Training. Training Interest. No. Training in training inference and and everything. No.

Speaker 1:

Earnings before training interest and taxes. And what was odd about it was that Ed Zittrain was was was saying it's like the new community adjusted EBITDA, and it is always odd when a new non GAAP metric pops up. In this case, I think it makes a lot of sense because training runs do fit a depreciation profile. It's a little bit different. I don't know why you wouldn't just put it in depreciation, though, like just figure out how to account for training runs through a depreciation schedule.

Speaker 1:

And then maybe it's like a non GAAP depreciation metric, but it's still in there instead of trying to get everyone up to speed on a different a different, like, sounding phrase entirely.

Speaker 3:

Yeah. I was looking back at Yeah. Ben Thompson's earnings transcript or a script

Speaker 1:

Oh, yeah.

Speaker 3:

That he wrote for for Mark Zuckerberg. He has a good segment on why AI matters. Mhmm. Ben writes, forgive the long preamble, but this is necessary context for me to properly explain why AI is so important to Meta and why I'm making the right choice to invest so heavily in both talent and infrastructure. And he goes on and on and on.

Speaker 3:

But he says, what I've come to realize as I've embraced our status as an entertainment provider and ad purveyor is that our nature as a digital business non withstanding, we are remarkably well placed to thrive in an AI era. Remember what we learned about humans. They are obsessed with other humans and they wanna connect with them. That obsession and desire only going to increase as we interact more and more with AI. AI is gonna make our properties more essential, not less.

Speaker 3:

Moreover, AI is a productivity tool, but productivity is not the end all be all of the human experience. I've talked over the last year about building super intelligence that helps you get things done, but that's a business story. What we can do uniquely is gives give people the experiences they want from connection to entertainment to shopping when they are off the clock. The fact that we are investing in AI but not selling solutions to businesses is actually one of our business biggest advantages. So, of course, this is just a a sort of fan fiction for an earnings transcript.

Speaker 3:

Meta is in fact selling to businesses. No. No. But who knows over time how big will the API business be

Speaker 1:

Yeah.

Speaker 3:

Relative to how much value they can unlock across their broader Yeah. Business with all of their infrastructure Well,

Speaker 1:

leave us five stars on Apple Podcast and Spotify. Sign up for a newsletter at tbpn.com. And we will see you tomorrow at 11AM. Sharp. Goodbye.