00:00:06:01 - 00:00:30:10 Unknown Hello and welcome everyone to another episode of the merge I newsroom. Our special format that we bring to you live from our San Francisco office here to talk about the latest developments in the AI world around new model releases, new frameworks, new harnesses, whatever comes out in you. You learned here from us today. I'm super excited because I'm having from here from our applied AI team. 00:00:30:12 - 00:00:49:04 Unknown Everyone has a great deal of experience, in the AI research world. He's previously worked at companies like stability AI. He's also built his, PhD research around, yeah, large language models and natural language processing. So there's really no greater person to talk to about this latest release, which we just experienced from Google coming out last week. 00:00:49:07 - 00:01:10:22 Unknown Gemini 3.1 Pro yeah, this release is really exciting. Especially with the Gemini 3.1 Pro and the deep think and the constant progress and much of progress towards AGI. I think, they have released a lot of results on a lot of benchmarks, which are all very exciting. And, yeah, we can dive in. I'm excited to dive deep into them. 00:01:10:24 - 00:01:37:15 Unknown There's obviously a lot of stats that we can talk about here. One of them is, for example, maximum result of 77% that they achieved on these, Ark AGI two benchmark are price verified testing. What is so special about this benchmark? And, how is Google managing to to get those rocket numbers? Oh, very, very good question. So the Ark AGI the Ark against it is one of the most credible institutes, in my opinion, in terms of benchmarking language models. 00:01:37:15 - 00:02:04:14 Unknown They are a nonprofit organization. And the Ark AGI two benchmarks specifically focus on general intelligence reasoning. And they don't give like texts. It's multimodal inputs and multimodal outputs. And the tasks in our Ark AGI to focus on things like compositional reasoning, where you have to apply multiple rules to solve a problem and continuously adapt. The they also have Ark AGI three, which is just a slightly different tasks. 00:02:04:16 - 00:02:29:17 Unknown But that's for another, the conversation models typically have plateaued, up until the research topic, until we've pushed the frontier. The difference between Ark AGI, the Ark AGI verified badge that they released, I think, late last year or early this year. Insurers, the institution themselves hired a bunch of faculty. The of that kind of advise the assessment process and oversee it from academia. 00:02:29:17 - 00:02:56:14 Unknown And then they also collaborated with, a bunch of nonprofit, or the bunch of research organizations, that produce language models like Google, Zion and Noah's research. And the verified badge focuses on ensuring that the test set they use is completely hidden. It's not public. The academics oversee the evaluation themselves as well. So, it's one of the it's it's about credible results. 00:02:56:16 - 00:03:16:07 Unknown And it's not some and all the results are reproducible. They require everybody to submit their code. And the code is open source. And you cannot you know, you can you can reproduce the results. There's different tiers that they announced. Right. One of them is the Gemini 3.1 Pro API, which we can use today. And then there's also deep thinking, which is something that they kind of teaser. 00:03:16:07 - 00:03:39:26 Unknown Right. What's the difference there. Yeah. So I think deep think is the flagship model that they have. It's a it's available through the Gemini app only on the ultra subscription. So 200 plus dollars a month or something. Or it's pros available at the 20 bucks a month subscription. So the deep think model focuses on very difficult problems, more open ended problems that might require hyper good hypothesis generation. 00:03:39:26 - 00:04:00:11 Unknown And they claim that the model is able to investigate and generate multiple hypotheses at once. Could it be this could this work could that work? And so on. And it tests them and verifies them also, probably at the same time. So it can give you a more comprehensive answer. And it's very useful for solving open ended and difficult problems in the context of software engineering. 00:04:00:11 - 00:04:27:05 Unknown Something like fixing an issue or debugging would be particularly useful for deep thinking. Pro model is the premier reasoning model where you give it a task, and it can decompose that plan, do the task very, very similar to sonnet and opus. More so the sonnet and opus, I would say, whereas the thinking is really used for open ended problems like mathematical proofs, advancing scientific research and things like that. 00:04:27:05 - 00:04:47:17 Unknown And they have a very good blog article showing some use cases of detecting model. Let's talk a little bit about that article, because I think one thing, one of the things that stood out to me was the amount of benchmarks and the discourse that they reached on those benchmarks. And I think as like an observer who might have not done their PhD in that field, you can get a little bit confused. 00:04:47:17 - 00:05:05:02 Unknown At least that's how I feel sometimes when I look at all these results. And it just feels like every time you model comes out, all the benchmarks are broken again. So maybe walk us through what do these, you know, top four benchmarks that we see on the report actually stand for. Like how do they what are they built off and why? 00:05:05:04 - 00:05:26:14 Unknown Is it such a big thing this time that Google is smashing all of them? Basically. Yeah. Let's start with our ask AGI to where deep think did almost 85%. In terms of the task, that's one of the most generally impressive results that we've we see in this blog post, 85% was sort of the grand prize for our AI beforehand, for our AGI too. 00:05:26:14 - 00:05:44:29 Unknown And we almost like, are there. And that's surreal. And it's a very big difference between deep Think and Gemini Pro or Gemini Pro. I think I got like 30%, 31% and dips and got like 80 over 80%. And the best model beforehand was opus 4.6, and I think it was at 68%. So that's a big, substantial leap in terms of that task. 00:05:45:02 - 00:06:02:13 Unknown And maybe we can see on the screen like an example of the task where it's like, you know, it requires really composition, reasoning, symbolic reasoning, where you're looking at an image and you're trying to figure out a puzzle. Humans can do it well with adequate practice, but models really struggled with that before because it's kind of an alien language to them. 00:06:02:13 - 00:06:21:02 Unknown Like you're giving them these symbolic inputs and you have to produce a symbolic output by manipulating these symbols that they don't know what exactly they mean. They have to make sense of it as they go along. And whereas language, they've seen it a lot, they've pre-trained on it, they've seen trillions of tokens. They know how to manipulate thing, which they know how to write, but they cannot manipulate symbols that they don't know what it means. 00:06:21:07 - 00:06:46:12 Unknown This is a true hallmark of intelligence testing, where we're benchmarking models on tasks they've not seen before. And how we got there. I think I could come up with a several hypotheses, probably like models were trained on, like symbolic reasoning languages where, or tasks that require symbolic reasoning, where they're given a bunch of symbols they've never been exposed to before, and they have to manipulate them to solve a puzzle that might have been part of their post training. 00:06:46:12 - 00:07:06:27 Unknown But this is all speculation. I don't work at Google, so yeah, of course. Yeah, that's the first benchmark. And then we can talk about humanity's last exam and code versus particularly because also these two benchmarks are very impressive and more related to software engineering humanities. Last exam, I think 40% of the questions are mathematical related. And some of them are computer science and AI related. 00:07:07:00 - 00:07:33:03 Unknown And so all of the answers can be verified automatically. But they're very, very difficult questions on models almost better up to 40%. And it's very difficult to cheat on them because they retrieval doesn't help us solving them. It requires real reasoning and code. Forces is a platform where you can submit your code solutions, programing solutions. It's very similar to leet code where you have a task specification, inputs and outputs, and you have to write a program in a specific language that solves it. 00:07:33:09 - 00:07:55:08 Unknown And it's compared to against, everybody else who programs on the platforms, humans and machines. And it's a ELO ranked system. So it's similar to chess where you're, it's a rank compared to everybody else with programs. So this substantial leap in improvement means that there's models that are really good at hypothesis generation and coding are really becoming really good compared to humans. 00:07:55:08 - 00:08:24:11 Unknown On the code programing. This is generally impressive. Maybe let's let's take a little step back and think about, okay, what does that actually mean for like developers. Can you translate somehow what you know, all these these impressive results on these massive benchmarks will ultimately yield to for, you know, people like you and me who are trying to implement a task, with one of those models, or especially with Gemini, maybe, I think Gemini is becoming really good at solving produce solving workflows. 00:08:24:11 - 00:08:50:17 Unknown So I think they also are really high on terminal bench. So giving it traditional tooling like CLI, it can perform tasks well. And technologic models keep getting better at performing and tasks and tasks well so long as a task as well defining you know what you want to solve. And I think what that means for developers is I think design becomes more important than conceptualizing the task and defining the task. 00:08:50:17 - 00:09:19:07 Unknown Well, in other words, prompting becomes really, really important as opposed to solving the problem itself or doing the task by hand. And the code versus performance tells us that I think traditional like coding and traditional like implementing traditional algorithms and pattern matching problems to algorithms becomes less important because I can do that. Well, and what that means for us is it becomes more important to define problems, as opposed to solving them. 00:09:19:07 - 00:09:45:29 Unknown And so when it comes to programing, I think there's also the, more, let's say practical viewpoint. For example, we saw on, terminal nul bench 2.0, for example, that there are different scores for the Atlantic terminal competency versus things like, you know, tools, no tools, search and code and all these other different, more elaborate testing methods that we have where we see how models perform in practical applications. 00:09:45:29 - 00:10:09:17 Unknown Can you speak a little bit to that? I think sweep bench verified. Yeah. There are quite a bit of, three related tasks that they've benchmarked against. I think the sweep bench verified is probably the most relevant. This basically is real world repositories that are scraped off of GitHub. And then you have you have to resolve real world issues and real world bugs and, you know, do a PR it's a complete end to end. 00:10:09:17 - 00:10:28:11 Unknown And the model can just go and clone the repository, perform tasks on it, edit the code test, and then submit their apartments when they're ready. However, that again the task is well defined. So you have an issue, that describes a bug or something that needs to be done in the repository, and the agent has to go in and implement it and verify that the implementation works. 00:10:28:14 - 00:10:48:18 Unknown So, this is a substantially substantial improvement to where we were like a year and a half ago when the sweep bench was just released and models were struggling to do this task. Yeah. But now we have the infrastructure in terms of a coding, in terms of terminal in terms of cloud code, in terms of like cloning repositories, verification, running commands on grab. 00:10:48:25 - 00:11:08:06 Unknown So random commands in the terminal, running verification, writing unit tests and so on. And also retrieving code documentation. That was a huge thing. Yeah. The all that really enabled this rapid progress, along with, you know, raw intelligence advancement of of the models themselves and then becoming more capable. We also run some benchmarks here at Code Robot internally. 00:11:08:06 - 00:11:28:14 Unknown Right. So how did Gemini 3.1 Pro perform on those benchmarks compared to other models like, you know, Codex or Opus? Yeah. So generally Gemini we've benchmarked we have preliminary results. So we're still running more evaluations as we speak. Since the model just came out generally, Gemini Pro produces less comments than the other models at what we currently have. 00:11:28:14 - 00:11:54:17 Unknown And then I think it's about by 25% less comments. And it really tries to focus on comments that are critical. We observe relatively similar performance in terms of whether it catches the issue that we intend the model to catch. Like the really the core issue behind, the PR or the core issue behind the, you know, the bug that we're trying to find and then it doesn't really write nit picky comments as much. 00:11:54:17 - 00:12:20:07 Unknown It tries to avoid that. So things that, you know, maybe the developer intended like an issue, like a logger using info instead of one or something like that, like small comments that the developer might be aware of or might not, but it's not really important to address. Gemini doesn't seem to highlight that at all. Which might be useful for some developers and institutions, but maybe not for everybody. 00:12:20:09 - 00:12:39:15 Unknown Another thing I would highlight is Gemini Pro is cheaper than the other frontier models. Cost per token. And I think you know how many tokens it produces. That depends on the workflow and the implementation of the reviewer systems. And that's something that I think people should be aware of as well. But generally it's good for bang for buck. 00:12:39:15 - 00:13:06:28 Unknown It produces good value compared to the cost, compared to the other frontier models, just because it's cheaper. And if you get comparable performance, that's useful when interesting findings that we had while, doing our internal evaluations is that when the Gemini comment captures the Ash event and like we are looking for, the core issue, it tends to be more a little longer and assertive in its claims. 00:13:06:28 - 00:13:30:06 Unknown So it uses less hedging language might should not maybe. Yeah. It seems to be like a general trend. I had David, our VP of AI on here, I think two weeks ago when the, GPT Codex 4.3 F, so 5.3 and up was 4.6 dropped. And he was also noticeably mentioning this trend of models becoming more assertive and more sure of themselves, so to say. 00:13:30:06 - 00:13:54:17 Unknown Is that something that you see as well? Yes. It's it's becoming more assertive. And there seems to be a correlation between the comments that highlight the issue that we intend, like if it captures the issue, we want it to it tends to be more assertive than the comments where it doesn't. So there's there's there's some really creative relationship here between assertiveness and whether it captures the intended ground truth. 00:13:54:19 - 00:14:17:10 Unknown Behind the core issue. Interesting. Yeah. There's also another new thing, that Gemini three point one Pro introduced, which is this custom tool endpoints that where you can set certain thinking controls. So for people building agents, there seems to be is that is there a practical difference between like setting, a model, a smarter model to think longer, shorter? 00:14:17:10 - 00:14:49:12 Unknown And, what difference does that make? Yeah. The, the most obvious trade off here is cost. So the last thinking it does the cheaper it is. Pro is generally really good at a lot of things. As you can see from the benchmarks, generally simple tasks that you might use flash. You could use flash for Gemini flash, I mean, and or like it doesn't require a lot of reasoning, like question answering, maybe like, you could use low thinking, and by default you could use the medium thinking. 00:14:49:15 - 00:15:10:26 Unknown And for more complex problems that require more thinking and more analysis, multiple tool cores, potentially you would want to use higher thinking. Talking about the context window size, right? I think Gemini we talked about this has always been model that has a lot of input token, a lot of input token, window length. Where do you see the benefits? 00:15:10:26 - 00:15:34:04 Unknown When should you know really what when should somebody default to using all that context or when when should somebody maybe think about that, clever context engineering techniques in order to get the most and, what do you see the cost benefit here? I would say generally, as a rule of thumb, if retrieval can solve the problem or a subset of the problem should use it, retrieval is a very powerful tool. 00:15:34:06 - 00:15:52:11 Unknown So long as the model remembers and knows how to retrieve. And most, most models now are trained genetically so they can use tokens very, very well whenever they need to look something up. I would say passing really, really long files to a model. So like having a long document or a long message tends to produce this phenomenon of context, right? 00:15:52:14 - 00:16:11:13 Unknown More often than longer conversations, I would try to break things up like really big code files should be broken up into smaller chunks, maybe send it by function, or maybe line by line. As a post or subset of lines you know, that are relevant for what the model is trying to do, as opposed to fitting out the entire file all at once. 00:16:11:16 - 00:16:32:10 Unknown But generally, you know, models do really well. We're talking about millions of tokens here. The context window is 1 million. So, it does really well. But like, you know, feeding it like, a really big Json file of, you know, 100 million lines probably wouldn't be a good idea when you can just grab things. So, I mean, we talked about, just before as well about this previous paper. 00:16:32:10 - 00:16:55:09 Unknown A needle in the haystack. Right? Yes. How do you see that evolving? I think this paper is now two years old. It's it's quite some time. And I that's that's almost, ancient. So, how did you see the the model development evolve on that? Is it now other models better at finding the, the needle in the haystack, or do we still experience similar problems? 00:16:55:10 - 00:17:18:04 Unknown So I wouldn't say the problem of context, right is solved, but I don't think it's as important because we have infrastructure to retrieve things. So instead of feeding the entire 100 million line Json file to a model, a model can just write a grep statement to look for what it needs from that Json file and answer or perform the task to respond well. 00:17:18:06 - 00:17:45:07 Unknown So I think that perhaps it's a less relevant problem. Because we have technical, we have a digital infrastructure that models can use to answer them. But that's, that's my opinion. Yeah. Maybe people would disagree. Also, it's more expensive to feed and very long inputs to a model. So in terms of cost efficiency and maximizing value, I think it's very useful to build infrastructure for models in such a way that you can optimize cost and performance. 00:17:45:09 - 00:18:07:09 Unknown And retrieval is such a very good tool for this. And models are really good at using them. So I mean, yeah, I guess that that fits our narrative context. Engineering is still king. Yes. That's right. Let's talk about the token economics. You already touched a little bit on, you know, Gemini being more cost effective. Can you elaborate a bit more on what should build us be be looking for when it doesn't make sense? 00:18:07:12 - 00:18:29:26 Unknown The Gemini models typically are a bit cheaper than the other frontier models. Google tends to price them a little better, probably because they have good cloud infrastructure. Generally, the Gemini models do as well as the other friend whom they keep up. They're not behind, and I think like it's a cat and mouse game or somebody releases a new model and it's slightly better than the other models at certain things. 00:18:29:26 - 00:18:51:25 Unknown But mostly they're neck to neck. So it's wise for institutions and companies and builders to be able to adopt, and quickly evaluate all of these frontier models quickly on their use cases or on their domains, and see what works best for them. The best is unclear, and it will continue to be unclear for the foreseeable future in terms of models. 00:18:51:28 - 00:19:14:21 Unknown And there are trade offs, like, now, how much do I want to spend? And, which models cheaper do I want to maximize performance or cut costs on this specific subtask? Having the infrastructure to run evaluations quickly, concisely, and extract insights. And I should use this model here. I should use this model there. Maybe I could use a cheaper model here I think would be very useful for institutions. 00:19:14:21 - 00:19:36:18 Unknown I mean, you already touched, on the outlook a little bit here, which is great. What should we be excited about for developers, especially maybe looking at Google and the models that they'll be releasing soon, is there anything that you're waiting on? I would love to try deploying. Yeah. Like, oh it's it seems like like it can explore multiple ideas at the same time, generate many hypotheses. 00:19:36:18 - 00:19:58:01 Unknown It's very useful for scientific research. It's very useful for, you know, open ended problems like mathematical proofs. It won the Mathematical Olympiad, the gold medal in Mathematical Olympiad like a year ago. And more recently, and it's very good. I really want to use it for debugging and see how it does there. I think, if they release an API for that, that would be fantastic. 00:19:58:01 - 00:20:18:25 Unknown You can experiment with that. And I think another thing that I'm looking forward to from Google, perhaps, is like if they do any integrations with their anti-gravity ID, they released an anti-gravity ID last year, which seems to be really good for prototyping UI specifically. And generally it's an engineering ID, and I think that, you know, deep think there could be very cool. 00:20:18:29 - 00:20:33:27 Unknown Yeah. Great change. Wow. Okay. Well thank you. There you hear it folks. So keep an eye out for that. Thank you so much everyone for taking the time and talking to us. We'll have you on here again I'm sure. Yeah. So, yeah, thanks for joining in for the show and talk to you guys soon. Thank you for having me.