00:00:00:02 - 00:00:18:18 Unknown Most of these models, they're, you know, built by doing gradient descent. And gradient descent is a fantastic algorithm that tells you nothing about what's going on inside the weights. Right? So the awesome at coding and at, you know, writing and all the other tasks that you want to do, but you won't be able to say, and here's why it's great at those things. 00:00:18:18 - 00:00:33:10 Unknown We're in a world where we have this incredible source of intelligence. You don't really have a good explanation for how it works. And that's why we call ourselves Marcia. If you can understand what's going on inside these models, you can then say, here's exactly why this works. Here's exactly why it will fail, and here's how you fix it. 00:00:33:12 - 00:00:54:03 Unknown This is a really, really hard column. It's the kind of thing that you need to like. Found a research lab for raise a ton of money for because it isn't the kind of thing which has a super obvious answer. We think that in order to study anything accurately, you need to measure it very, very careful. What does the day to day look like for a researcher like you trying to discover the steam engine of AI? 00:00:54:10 - 00:01:04:14 Unknown Yeah, honestly. 00:01:04:17 - 00:01:23:19 Unknown Hello and welcome everyone to another episode of the Merch, our podcast brought to you here live from our Code Rabbit San Francisco office. I'm very happy to, cause today I'll be sitting here with Josh, one of the founders of The Martian, which is one of the frontier evaluation and bench model labs that I could say here in San Francisco. 00:01:23:19 - 00:01:41:11 Unknown And we're going to talk a lot about code review, you know, the interpretability of many of those benchmarks that you see around different coding tools and code review tools as well, and a bunch of other things. So happy to have you all here with me. Hello. How are you doing today? You and well, how are you? That's doing. 00:01:41:12 - 00:02:03:09 Unknown That's great. I'm doing fantastic. Maybe for those people who haven't heard about Martian before. Could you give us a little bit about, you know, the company, the mission, the vision? Yeah. So Martian is an interpreter lab. Our goal is to understand what's going on inside of language models. And I think the company has a fascinating background as well. 00:02:03:10 - 00:02:23:14 Unknown Let's let's start by the name. Maybe. Why did you call it? Have you heard of the group of Hungarian American scientists who were called the Martians? I did read up a little bit on it, but I assume most people would have not have, so maybe you can enlighten me. Yeah. These were some of the smartest people of the 20th century, right? 00:02:23:16 - 00:02:38:08 Unknown So John von Neumann, he was the guy who invented the computer architecture. He was one of the folks who invented game theory. I mean, a lot of the modern world was invented by this one guy. And it turns out there was a bunch of others like him. So Paul Erdos, he was the most published mathematician of all time. 00:02:38:10 - 00:03:01:19 Unknown The Ziller invented the atomic bomb. The list goes on. For a while, the things really remarkable about these folks is they were all born in a single neighborhood in Budapest. And so people were like, this is strange. What's the explanation for all this intelligence? And they didn't have a good answer. The reason they're called the Martians is because it was joked that, oh, you know, a bunch of Martians came down to earth, and here we are. 00:03:01:22 - 00:03:21:26 Unknown That came right to Budapest. Yeah, but the situation with Llms is quite similar, right? We're in a world where we have this incredible source of intelligence. We don't really have a good explanation for how it works. And that's why we call ourselves Martian. Right? That is actually a beautiful tangent over to, you know, a deeper dive into what you actually do at Martian. 00:03:21:26 - 00:03:51:18 Unknown So where did you you know, how what kind of need was the company born and what's the what's the story there? Yeah. So I've been doing stuff in NLP since 2015, 2016. And same with my my co-founder Aton. And when GPT two and T5 came out, that was for us, just like Holy moment, because if you'd worked on language modeling kind of before those models came out, the fact that they could just generate paragraphs that were coherent was insane. 00:03:51:23 - 00:04:08:27 Unknown Now, today, that looks kind of silly. It's like, oh, you know, these models are kind of compared to what we have today, right? They definitely couldn't write your code for you or do code reviews for you. But the feeling that we had is, okay, this is going to be really big. And so we started doing more and more research into Llms. 00:04:08:27 - 00:04:27:29 Unknown We actually met at lab in the University of Pennsylvania, and the thing that we realized is we don't really have any idea of what's going on inside these models, and this is kind of serving. Then two needs, one is just the human need to know, right? I think that there is just a fundamental question about how intelligence works. 00:04:27:29 - 00:04:49:09 Unknown And before you couldn't really answer this question without cutting open the brains. Very large number of people, needless to say, didn't happen. We don't have a scientific answer yet, but with Ms. you can actually just go and inspect the internals of the model, see how they work. The second thing, however, is in improving the ability to control and to have reliability from these models. 00:04:49:10 - 00:05:07:07 Unknown Right. Eventually we're going to be doing all sorts of things with these models, not just code generation, but they're going to run almost every facet of the modern world. If we really believe in that vision, the reliability you're going to need from these models would be enormous. And saying, oh, it's a black box. We don't know how it works isn't really going to cut it. 00:05:07:09 - 00:05:23:28 Unknown If you can understand what's going on inside these models, you can then say, here's exactly why this works. Here's exactly why it will fail, and here's how you fix it. Right. And what exactly is your approach to doing that? Because that seems like an inherently complex problem. I'm sure you're not the first to attempt to solve it. Yeah. 00:05:24:01 - 00:05:43:11 Unknown This is a really, really hard problem. It's the kind of thing that you need to found a research lab for raise a ton of money for, because it isn't the kind of thing which has a super obvious answer. Right. If I told you. Oh, yes. Here's exactly how we're going to solve the problem. It's kind of like being in the 19th century and saying, here's exactly how relativity is going to play out, right? 00:05:43:14 - 00:06:05:21 Unknown It's just not an answer that you can give because the task is really, really hard. Right, right. But here's the bet that we are taking. We think that in order to study anything accurately, you need to measure it very, very carefully. Like most of how science developed during the whole of the scientific revolution was built on measuring things more carefully, right? 00:06:05:23 - 00:06:25:13 Unknown Whether it was building better telescopes or measuring the masses of chemicals, or in chemical reactions, or measuring how light propagates through space. And Michelson-Morley experiment, every time you have a real shift in how science works, it comes from measuring things better. And so you want you need to measure how these models work, what they're actually doing really, really carefully. 00:06:25:16 - 00:06:48:28 Unknown The second thing you need to do is explain why the behaviors that are occurring are occurring, right. This is where kind of all the theory and the math and the kind of large amounts of compute come in. Right. And in order to get the large amounts of compute, we're also making a bet on commercialization. We think that if you can commercialize interpretability, you can put vastly more resources towards it. 00:06:48:28 - 00:07:11:01 Unknown So instead of treating it like a side project, which is kind of the inclination of labs who might be focused on scaling, we can treat it as the core of our business. Well, yeah. Incredibly interesting. Maybe we'll dive into the commercialization aspect in a little bit later as well. One other thing that I found very interesting as well, looking more deeply into your company is you specialize particularly on coding. 00:07:11:05 - 00:07:29:06 Unknown Why is that? So this comes back to the measurement aspect that I mentioned before, right? One of the most important behaviors of language models is the fact that they can write code. I mean, like, I wouldn't be on this podcast, you wouldn't be building the company you're building, and most of the folks in Silicon Valley wouldn't be using tools in the way they're using them. 00:07:29:06 - 00:07:48:25 Unknown If models couldn't do that really effectively, and this also shows up in the usage numbers and kind of what you hear about in the news. And really, this is the most impactful application of Llms today. The thing is, we don't have great measurements for what models are actually doing when they generate this code. How reliable are they? What kind of mistakes do they make? 00:07:48:25 - 00:08:07:05 Unknown When do they run into bugs? All of these kinds of things are just ill measured, right? And so by focusing on code generation, you can make much better measurements of the things that actually matter to people. At the same time, it's also a area that's particularly nice to study because you can formally say, here is what a program is. 00:08:07:06 - 00:08:24:02 Unknown And so if you're studying natural language, it's kind of hard to say, oh, well, what are the basic units of language. What are the theories we should develop around this? For code the mathematics is much more settled. If it weren't, you wouldn't have computers, right? And so you can also get a formal notion of this is what a program is. 00:08:24:04 - 00:08:45:13 Unknown And you can rely on these formal notions to analyze what models are doing, which makes the problem of analysis a lot easier. And then circling back. I mean, I think so far there's a lot of applications. We'll talk a bit about benchmarks in a bit. We'll talk some of the other research. But where do you see. You know a research lab like yourselves really getting into that commercial path that we were talking about? 00:08:45:14 - 00:09:12:09 Unknown Yeah, I think there's kind of two pieces of this, right? The first is if you really want to look at where most of the commercial value comes in, it is a kind of moonshot bet, right? The things you can do, if you understand models better are potentially incredibly, incredibly powerful. If you think about what made the current batch of models possible, it was the invention of the transformer. 00:09:12:11 - 00:09:29:22 Unknown You had this architecture which had a lot of nice properties. It's very parallelizable. You can push data through it. It seems to have these very nice inductive biases, and therefore you're able to suddenly learn how to do language and coding and all these other things. Imagine what you would could get if you had something that's like the next transformer architecture. 00:09:29:23 - 00:09:48:27 Unknown There's every model today that's mainstream is relying on essentially the same architecture. Right? No one has made substantial innovation. There is a hard problem. If you could, you get models that generalize way more effectively or way safer, which are intrinsically understandable, which you could tweak the weights of in the same way that you tweak code instead of in a very black box method like stochastic gradient descent. 00:09:48:27 - 00:10:11:17 Unknown And this would enable potentially much more powerful and much safer models than the kinds that you get today. So that's the long term bet. Of course, there are also things you can do in the intermediate term to provide value to the community, right. And those kinds of things like building better benchmarks or helping folks figure out what models they should use, why they should use those models. 00:10:11:18 - 00:10:29:09 Unknown Those are the kinds of things that we can help with more immediately. Okay, so we're doing both. Yeah I see. I mean, super interesting topic. I love the point about, you know, limitations of the current model architecture. I think there's a lot of people building. I've been hearing the term world models popping up more and more recently seems to be popping out more. 00:10:29:12 - 00:10:51:18 Unknown A lot of work going into that. Okay, let's let's circle back a little bit and talk about the benchmark topic. You just recently released a very almost viral benchmark on code, a code review tools, which I think we should talk about a little bit here. So can you give me an idea why did you focus on code review first instead of. 00:10:51:18 - 00:11:13:16 Unknown Some people would have argued. So if you go down the code avenue you do code generation. Yeah. You did code review. Why is that? I would argue they're the same problem. So if you want to measure code generation right. Like what is a good measure for this. Well if your code generation tools always give you code that you can just push straight into prod, I would argue that you've solved code generation, right? 00:11:13:21 - 00:11:36:25 Unknown This is or at least to a first order approximation, a really good solution. Right? And the thing that checks whether you can push into prod is code review. So if you can measure code review effectively, you're essentially measuring the verifier. You're saying can we actually say whether code gen is good right. And so in some ways, I think the highest leverage thing to measure is our ability to verify code. 00:11:36:28 - 00:11:58:20 Unknown I see and and well taking on that continuing on that tangent, what was your approach to that. Because I think that's that's also something that stuck out quite differently. How is that benchmark that you created different from within a lot of them out there? Obviously you have a very scientific background. So please elaborate on on. Yeah. How did you come across it? 00:11:58:21 - 00:12:22:11 Unknown Have you heard of Good Hearts Law? I cannot say I have. So it's a kind of famous law in certain circles. But the core idea is that if you have some kind of measure and you optimize against that measure, it becomes a goal, then it ceases to be a good measure. Right. A very clear example of this was with bench verified. 00:12:22:18 - 00:12:44:10 Unknown So I think earlier last week before we launched the benchmark, the OpenAI team said, hey, don't use bench verified anymore. This is not a benchmark you should report results on. And the reason they said this is because the data from that benchmark essentially leaked into all of the models, right. And this is essentially a consequence of good hearts law, right. 00:12:44:13 - 00:13:05:17 Unknown Like the more that you go and try and train models get better at coding, the more that just by accident you might end up measuring. How well are they doing on sweet bench verified as opposed to how good are they at coding? Yeah, right. And so the thing that we wanted to try to do is build a benchmark which flights get back against good hearts law. 00:13:05:19 - 00:13:23:09 Unknown Right. And this is why our benchmark has two parts an online part and an offline part. Right. In the online part we can say it's a tracker, right. How do people use these tools? We can just track that and look at the real world data and understand what matters to people. And if you understand what matters, then you can understand why you should actually be measuring. 00:13:23:09 - 00:13:46:17 Unknown And then we built an offline benchmark to try and measure that. Right. Yeah. That's that's I think the the entire approach is super fascinating. At the end, there were a bunch of different tools that that had various different performances that you had in there. I think the offline and the online debate continues to fascinate with my colleagues here in the, in the office. 00:13:46:17 - 00:14:07:21 Unknown So can you elaborate a little bit more on what exactly was the choice behind that and the reasoning? So when you're building a benchmark, right, which has an offline and online component, what you're really doing is thinking about two types of potential bias, right? The first bias is this kind of good heart bias that I mentioned earlier, which is that when people start optimizing against it. 00:14:07:22 - 00:14:31:17 Unknown You might not be measuring what you think you're measuring, right? This is what the online benchmark is meant to solve. So someone could ask, well then why don't you just use the online benchmark that's measuring the thing you care about? The main issue run into there is one of selection bias, right? So let's say that I always work on MML stuff and you always work on back end stuff in that world. 00:14:31:18 - 00:14:47:21 Unknown It might be that if I choose to use a tool because there's fewer people doing ML stuff and there's less data, maybe, you know, the tool looks like it performs poorly, not because it's actually worse, but because in some ways I'm like a worse user to be working for, right? If what you care about is the benchmark numbers. 00:14:47:22 - 00:15:06:18 Unknown Right. And so in the online data, it's hard to correct for this directly. Right. Because each each tool is going to be run on different PRS. And so what you really want to do is you want to say, well, could we create something which represents the online PR as well, but where every tool is run on the same set of PRS. 00:15:06:20 - 00:15:26:20 Unknown And so that's what the on the excuse me, the offline benchmark is meant to do. Right. And you open sourced all of these, which is great by the way for ability. Now you talked about saturation already of benchmarks. Are you afraid that the offline set will saturate? Eventually it will 100% saturate eventually if you don't change anything about it. 00:15:26:21 - 00:15:55:26 Unknown So another part of what we're doing is trying to continuously update the offline benchmark and create new data on a consistent basis. That way, you can't just have that benchmark saturate. You can constantly make both harder problems and also problems which better reflect the real world, right? So if you're doing this kind of a fresh release of data on a consistent basis, and it can't be the case, that model just trains on all the old data, and suddenly it looks like it's good because the new data will come out and we'll say, actually, it's bad, right? 00:15:55:28 - 00:16:25:08 Unknown And so this is part of of what we're experimenting with here too, right. The way that we like to talk about it internally is people have spent billions of dollars making the best models out there, but they've only spent thousands of dollars building good benchmarks. Right. What if you close that gap? Maybe that tangent is great to to talk about what makes this benchmark different because you I mean, you observed, I think 300,000 pairs, but also you took a different approach. 00:16:25:08 - 00:16:50:20 Unknown So compared maybe give us a spectrum of where the margins benchmark lands in comparison to some of the other benchmarks that we've seen out there. Yeah, I think that for code review in particular, most of the benchmarks have been created by code review companies. And this is kind of natural, right? The folks who are going to be closest to code review, who see the problems with it, who see that you need an important benchmark there. 00:16:50:22 - 00:17:08:20 Unknown It's going to be the folks who are actually building the tools, because you guys are closest to all of these pieces. But there's a kind of political economy problem with that, which is if someone creates a benchmark and they measure themselves on that benchmark, and then they say that there are number one, do you trust it? Right? And even if they do really great work, the answer is probably no. 00:17:08:25 - 00:17:29:11 Unknown Right. And so what part of what we realized is we should try and create something that's more objective, where it's not a a kind of first party benchmark, but a third party benchmark where people can then have more trust. And alongside being a third party benchmark, we want to make sure people can actually trust it, which is why we go into all of these pieces around the bias and how you fix that. 00:17:29:12 - 00:18:01:02 Unknown Cool. Yeah. That's work. Let's get a bit more into the benchmark itself. So I think there's two dominant axes that you measure things on, which is precision and recall. Can you explain a bit more? You know, what do they do? What do you measure. Why did you decide to to display and highlight these factors. Yeah. So when you have a tool making code reviews, right, there are essentially two things that people might care about, right? 00:18:01:03 - 00:18:23:17 Unknown One is noisy. It's like, oh, maybe it's just going to put all sorts of stuff in here. And now I have to deal with it. And it's just it's a pain, right? And the second issue is going to be thoroughness. Right. And it's a question of okay, you know, do I really trust this tool enough that if I go through and address the issues that it finds, my PR is actually able to be merged, right. 00:18:23:18 - 00:18:43:07 Unknown And these kind of pull in opposite directions, because if you want a noiseless tool, you could just not have a tool at all, right. Zero comments. No noise. Yeah. No, no I'm done at the same time. Not a great code review tool. Yeah. On the flip side, if you have a tool that's like super duper thorough, it's just like the most thorough thing you could possibly imagine, right? 00:18:43:13 - 00:19:02:07 Unknown You could do that by commenting on every line, right? You could comment everything. It's not going to be the ideal experience for the user. And so what you need to do is measure both. You need to say, how noisy is this and how thorough is it. Right. And so that's why you measure precision recall. Because precision says this has less noise. 00:19:02:07 - 00:19:24:10 Unknown And the recall metric says this is more thorough right. And what have you seen. Kind of the divide and the and the tooling. You know I'm sure I mean I've looked at it. I see the winners in certain categories. Other people optimize more I think. Did you see some learnings on the different systems? What can you tell about the companies like how they, you know, engineered their harnesses? 00:19:24:10 - 00:19:50:06 Unknown So one thing that's pretty clear from the data is that most tools have leaned into precision over recall. Right. My guess for why this is is because it's easier to measure precision, right? With precision, you can say, well, what percentage of people actually act on the comments that the tool is creating, right? Or alternatively, what percentage of the comments the tool creates are being acted on by people. 00:19:50:06 - 00:20:12:22 Unknown And this is something that you can then get like confidence in the fact that you're finding bugs. Everything you're finding is a bug, right? However, this doesn't mean that there's no value to recall. As we discussed, recall is something that's quite important to people. One thing I think is quite noteworthy is that Claude code actually optimizes for recall over precision. 00:20:12:23 - 00:20:29:13 Unknown I think part of the reason why is that it's harder in some ways to have really high recall, because it means you're finding more and more of the issues that actually exist, right? And you don't know what issues exist in the code. If you knew what issues existed in the code, you wouldn't need a code review tool. You just fix them, right? 00:20:29:19 - 00:20:57:04 Unknown Yeah. So so it's like a much harder metric to improve. Right. And to improve with confidence. Right. And I think that really the two tools which have optimized for recall the most are code rabbit and cloud code. Well yeah, that's good to know. Let's go back a little bit into maybe The Martian and what you do there. And, you know, you touched briefly on the moonshot. 00:20:57:06 - 00:21:18:08 Unknown I think it'd be great to, to learn a little bit more about. I read one of your blog posts, which I recommend to everyone. Dive into the website. There's a lot of really good stuff in there. And one of them you talked about, you know, this example of humans building bridges in the past centuries and then nowadays breathing bridges without understanding how bridges work, I should say. 00:21:18:09 - 00:21:46:02 Unknown And now we're building llms without truly understanding how they actually work. So take us to the moon, please, and tell us about your vision on interpretability and what comes next. Yeah, I think the moon is a great place to go. Let's say that you are a medieval peasant in the 16 or 17 hundreds. There's a bunch of people up in their ivory towers, and they're asking questions like, why does the moon stay up in the sky? 00:21:46:06 - 00:22:06:24 Unknown Right? What causes bodies to revolve around the Earth? And it might seem like this is kind of a useless question. Like, I can just go and do things in the world, and it doesn't really matter to me whether whether the moon works in a particular way or doesn't work in a particular way. And so you go and you build a bridge and you say, okay, great, I can cross a river. 00:22:06:27 - 00:22:21:00 Unknown The thing is, just building bridges. Based on your trial and error, your intuitive understanding is never going to get you to the point where you have an industrial revolution, or you have the steam engine where you're able to be lifted out of poverty, or you can cross not just that river, but you can go across any river in the world. 00:22:21:04 - 00:22:50:22 Unknown You have steamboats, right? The thing that's required to get there is a more fundamental understanding. And this is really what science is good for, right? You can, in theory, do anything simply through trial and error, but to generalize well to things you've never seen before. You need science, right? This is the way in which, by studying an apple falling from a tree, Newton comes to understand how bridges work, how the moon goes and revolves around the Earth, and all these other wonderful facts about nature. 00:22:50:22 - 00:23:08:13 Unknown And it's only from laws like those ascribed by Newton that you can go and build a steam engine. So what's the neutron law for a well, that's what we're trying to discover. And that's why I find the kind of work we're doing so exciting, because if you can discover that, you realize, actually, we're just building bridges with these lives right now. 00:23:08:13 - 00:23:26:11 Unknown It's a bridge from natural language to the stuff your computer can do. But what we really need is a steam engine, and that's going to take fundamental science, right? I think beautiful metaphor. Now take a tell me practically what does that mean? What does a day to day look like for a researcher like you trying to discover the steam engine of AI? 00:23:26:19 - 00:23:48:17 Unknown Yeah, honestly, it's a surprising amount of unsexy work. Okay. Because as I said before, a lot of doing science really well comes down to measurement. And measurement means looking at the data, understanding it deeply, seeing how are people using these tools, what issues are there? Why do these tools behave and behave in the way that they do now? 00:23:48:17 - 00:24:10:17 Unknown Sometimes it is really fun math stuff and if you like come by our office sometime, you'll see lots of equations on the whiteboards and, you know, stuff like that. And that is the the fun part of it too. But it's really this, this interplay of trying to touch the world as deeply as possible and then trying to, from that, derive principles and techniques which we understand how the models work. 00:24:10:18 - 00:24:30:00 Unknown Right. And it's a lot of fun. Yeah. Well, I'm glad you're having it. I think we can't wait to see what's going to come out of it. With that being said, I think we're closing the loop here. Before we finish up, though, I have a couple of rapid fire questions that I want to ask you to ask every, every guest that comes onto the show. 00:24:30:06 - 00:25:02:01 Unknown So if you're ready for it, perfect, perfect. So first, let's start off your favorite programing language, Haskell. Okay. That's a very I've never heard that answer before. How come I think that the mathematics behind programing is very beautiful. And there's very few languages which get at that beauty as close as Haskell does. Right? There's a great theorem called the Curry Howard correspondence. 00:25:02:04 - 00:25:28:04 Unknown Right. Which tells you how basically all of math is actually equivalent to programing and vice versa. Oh, okay. Right. It's a really beautiful theorem. And I encourage anyone who listens to this to look it up because it is just gorgeous and it totally changes how you look at programing. But this is how like when you hear about OpenAI or these companies working on like formal verification and like proving things, using Llms and stuff like that, it only works because of that really fundamental correspondence. 00:25:28:05 - 00:25:48:01 Unknown And I think that Haskell goes and takes a lot of those ideas to a really beautiful conclusion with, with this kind of math called category theory more practically, like day to day, I use Python and TypeScript. Like, you know, I am also an engineer, not just a lover of science, but like, if I could just programing language that I wanted to, it'd be Haskell. 00:25:48:02 - 00:26:01:07 Unknown Cool. Well, I mean, interesting follow up question to that. What's your favorite coding agent? Oh, that's a hard question. This is like asking me to choose between children. 00:26:01:09 - 00:26:23:17 Unknown Well, just for the sake of this. Different tools are really good for different things. So, I think that probably most of my time is spent in cursor, because I think having the integrated experience in IDE is really nice. Although recently I actually started trying. Devin, I've been pretty positive on that too. Like they do a lot of really kind of nice integrated stuff. 00:26:23:19 - 00:26:44:03 Unknown I use open code in part because of the model selection, or I can just choose a lot more models. And also this really nice feature around forking sessions, right? So basically like another git work tree is that. Yeah. Well in particular it's like I'm doing something in one session and I realized, oh, I actually need to do two things here. 00:26:44:03 - 00:27:04:20 Unknown I need to both like debug this and add this feature in. I can go and fork the session and that's quite nice. Cloud code is a classic. I mean, I think everyone uses and eventually goes back to and forth from Cloud Code. I've been playing around with Codex recently and there's, you know, again, all over the all over the place. 00:27:04:23 - 00:27:25:14 Unknown I don't have my mind is not made up at all. Like there's just too many cool things. Like, I honestly, I really like open source tooling. So like Kilo and Klein and Ru and if any of those I think is interesting is have like a whole bunch of different integrations and stuff. So like they also have like, you know, pieces around like every surface you could want if you want to use in JetBrains, like go for it. 00:27:25:15 - 00:27:45:10 Unknown I used to be a PyCharm guy before cursor came along. Right. So there's just there's just too many like like I think they're all doing interesting stuff. Do you prefer ID over terminal? They're both good for different things. Like even when I use an IDE, I still have a terminal window open. Okay. Good multitasking, I see. Yeah. Of course, the mind is a multitasking. 00:27:45:12 - 00:28:10:23 Unknown Great. Last question. And circling actually making a little bit of a turn here. What's your advice as a researcher for for people who might be listening here, who want to get into that field, where should they start? How should I pick up? And yeah, just a general piece of yeah, yeah, expertise. I think that machine learning is very interesting relative to other areas of science because of the ease with which it can be picked up. 00:28:10:24 - 00:28:31:24 Unknown Right. If you look at something like physics, we spent years and years and years developing theory. Right. And so to get to the very furthest edges of knowledge and physics, you have to walk down this long road in order to, to reach that kind of edge. Whereas in machine learning there isn't much theory. And this is kind of the problem that Martin is trying to solve is like, we just don't know how things work that fundamentally. 00:28:31:25 - 00:28:49:24 Unknown So a lot of the work is more empirical in nature. And so honestly, you can go and find a topic that interests you. You can look for some of the papers around that. You can maybe have a conversation with one of these models. Right. And then you can dive into them. You can start reading them, trying to understand them and start trying to implement stuff. 00:28:49:26 - 00:29:10:00 Unknown Right. And that's honestly the best way to get started. And you'll find that remarkably quickly. You'll start understanding more and more of what the papers are saying and the methods underlying it. And you'll be an MBL researcher before you know it tomorrow. Well, I can't give guarantees on speed, but I can't. I can't say way faster than it takes to become a physics PhD. 00:29:10:02 - 00:29:26:29 Unknown Well said. Well, thank you so much for coming on the show, Josh. It's been a wonderful conversation. I really enjoyed it. So I hope I have you on here again sometime soon. And with that, I'm going to close the show for today. Thank you for tuning in everyone that's been on their life, and stay tuned for the next episodes and read the blog. 00:29:26:29 - 00:29:34:13 Unknown Follow the Martians and we're very excited to see what you're going to come out with next. So thank you. Yeah, it was great. All right. Cool.