Deep conversations with the founders, investors, and operators building real-world AI - robotics, automation, industrial systems & AI infrastructure. Past the headlines, into how these technologies are really built, deployed, and scaled. Hosted by Bogdan Cristei, venture partner and former systems engineer.
THE OPTIM UPDATE
The Missing Data Layer for Physical AI | James Kujareevanich, Vision Lab
Guest: James Kujareevanich, Co-Founder & CEO, Vision Lab
Host: Bogdan Cristei
---
Bogdan Cristei (00:00)
This is The OPTIM Update. I'm Bogdan Cristei, and today I'm talking with James Kujareevanich, Co-Founder and CEO of Vision Lab. They're building the missing data layer for physical AI - going into factories, capturing first-person video of real operators, and turning that into structured training data for frontier labs and robotics companies. The origin story is great. One of the founders was an MIT PhD student who strapped a camera to his head to automate his own lab documentation before egocentric data was even a thing. They built a training tool for humans first, then got pulled into the robotics data business. We get into how the pipeline works, what it takes to get factory owners to trust you, and why synthetic data isn't the replacement people think it is. Without further ado, here is my conversation with James.
Bogdan Cristei (00:46)
Mr. James, how are you today man? Are you ready to talk about robotics and data?
James Kujareevanich (00:51)
Very excited. Thank you for having me here, Bogdan.
Bogdan Cristei (00:54)
Of course, man. It's a pleasure to have you on. Here's what I want to cover today. I want to start with how you ended up in this space because the origin story I think is pretty awesome. Then I want to get into what Vision Lab does and how the data pipeline works. After that, we'll talk about what happened since you started selling to frontier labs. I think that's going to be an interesting topic - what surprised you, what you learned. And then we'll do maybe some hot takes, hard questions at the end, where you go from here.
James Kujareevanich (01:17)
Sure. Sounds great. Let's get going.
Bogdan Cristei (01:21)
So before we get into what you're building, I wanted to start with the backstory because I think your path is not a typical founder arc. Walk us back to where you were before Vision Lab existed. You weren't pitching the industrial data layer for robotics. You were building something else entirely, right? So how did the company you have today emerge from that?
James Kujareevanich (01:38)
It was like a series of fortunate events, right? I came to do my MBA at MIT. I was previously at McKinsey Consulting in Bangkok. I met a PhD friend who was an early adopter of vision models. He would put a camera on his head, capture his lab experiments, and then basically let it auto-populate the lab documentation for him. And I was like, that's a very novel way of using vision models. So I convinced him, let's try to commercialize it, not in the labs environment, but in factories, which is where I have a lot of experience from McKinsey.
And so our first use case was using video to capture people working in the process and then put it into an iPad to be a training platform for operators. So it was starting off to train humans how to do their work properly. Every new joiner could come into this iPad app and then ask any questions like, how do I start this machine? How do I troubleshoot it?
And then we were selling it until it was about last Christmas when we met one of the AI researchers who was like, look, whatever you're teaching humans is going to be very valuable for teaching robots as well. Robots need exactly this kind of knowledge to learn on the frontline. And that gave us an idea. So we started doing some data sharing with this lab. And then it turns out to be a massively big business that we decided to just completely pivot. Now our initial customers, which are the factories, turn out to be our data suppliers. And we're paying them, actually, to help us collect this data. So yeah, from customers to suppliers - it's much easier because we're giving them money.
Bogdan Cristei (03:08)
So your co-founder was just automating his own lab work in his graduate program by putting a camera on his head basically.
James Kujareevanich (03:20)
Exactly. It was a PhD. That was part of his thesis for his PhD submission.
Bogdan Cristei (03:28)
And that was before all this egocentric data thing was hot, right?
James Kujareevanich (03:31)
Exactly. This was back in 2024. Nobody thought about using human data to train robots. That concept didn't even exist, right?
Bogdan Cristei (03:41)
So when we talk about training data for AI, language models had the internet - billions of pages of text that already existed. What's the equivalent for robotics? Why doesn't it exist yet? I'm curious to get your thoughts on it.
James Kujareevanich (03:58)
I would say firstly, there's YouTube, there's TikTok, right? There's so much content available, but a lot of it is noise. There's a lot of data available in public. It's just not structured enough to make any meaningful use case for training a robotics foundation model.
You need clear hands. Not many people take an egocentric video of themselves. Most people like to edit video into nice chunks, right? That's not useful for training because you want to see the whole episode of a task. So the data that exists in public is just not curated enough for training robots. We're pretty much starting from scratch in terms of training robots.
It's just, you have to be intentional about collecting this data. Unless you are my PhD co-founder who happened to be using it for lab documentation.
Bogdan Cristei (04:50)
So now I kind of see the problem from your point of view. Walk me through how you actually solve it. If I'm a frontier lab or a robotics company and I come to you tomorrow, what do I get? What does the data product look like?
James Kujareevanich (05:05)
So I have to say this is a wide, wide field, right? Everybody has different requirements. Everybody has a different research approach. There's old school teleoperation where you're still having a human carry the whole rig and then act out and control the robot, like playing a joystick. And then you have tactile, right? Collecting tactile pressure data, all this tactile sensor data. And then we have this realm of egocentric, which is collecting first-person POV. And some people are asking for a side camera as well. Some people are asking for 360 camera. Some people are asking for depth camera. There's just so many research approaches that you have to talk to them like, what's your research approach? What sort of data is good for you? What kind of annotation do you want?
A popular technique is having hand pose extraction - identifying the key points on the hands. Having action labeling - describing step by step what the humans are doing. So you're using this action labeling and hand pose to teach robots how to perform particular tasks. It's like doing a custom project with each of them.
I wouldn't leak too much of our secrets, but there's a primary direction that a lot of labs are heading. We're betting on the consensus - the 80% of people who are following a similar approach - because we know we can scale up very quickly and we don't need to invest in unique proprietary custom hardware just for one specific client. We have to make a bet on what we want to invest in terms of hardware and devices because it also determines the whole data pipeline - how we clean, how we process, how we annotate it.
Bogdan Cristei (06:49)
Yeah, that makes sense. So there's no standard operating procedure here. It's more about understanding what the demand pull is from the labs and from the people willing to pay for this data and trying to serve them as fast as possible.
James Kujareevanich (06:52)
Not really standard. No. Exactly. And these people are impatient. They'll be like, I need the data tomorrow.
Bogdan Cristei (07:12)
I'm sure. So we talked about factories before. Walk me through a factory capture from start to finish. You show up at a facility - what happens? What does the operator do? How does a typical capture run?
James Kujareevanich (07:27)
So in the beginning, you know, like YC says, do things that don't scale, right? We are the ones who actually go to the factory. We set up the camera on the operator's head. We help them transfer it to the laptop, upload it. We need to start off from scratch, doing things so that you know where all the pain points are.
We took that motto to heart. We literally ran the ground. I actually had a factory - my dad runs a factory. That's the first use case, we just did it in my factory. Made sure that we know how to run it end to end.
And at this point, we have a playbook so we can just ship off the camera, right? Run a two to three week sprint project with a client. They have a project manager who helps us install the camera, download the videos, upload to our drive, our Google Cloud. And then we handle the rest. And then once they're done, just ship the camera back to us. It's a very straightforward process for the factories.
I would say the hardest part is the initial trust - getting them to understand that we're not here to steal the trade secrets. We're just here to capture human motions. And then once you get the sign-off from the owner of the factory, the rest is very mechanical.
Bogdan Cristei (08:31)
Yeah. So getting access to factories is the part that sounds easy on a slide deck but is really hard in practice. How do you actually do that? How do you get factory owners to let you in?
James Kujareevanich (08:55)
It's a trust game, right? We have a very strict data consent form. We are training our hand pose extraction algorithm. We're only using this data and at most licensing it to the frontier labs for their own foundation model. We're not open-sourcing it. We're not leaking it out to the public. And we just let them understand the whole intention of this - we're training GPT for robotics, right? And a lot of people are bought into the mission. Like, I'm helping to contribute to the next physical AI, a ChatGPT equivalent for physical AI.
It all comes down to you need someone who can vouch for you, right? Like, I know this guy, I know James, he's a very reputable guy, trustworthy friend. At this point, we're also expanding more of a contractor network as well. We get people who have - we call them mini influencers - they know a lot of factories and they help us acquire more factories.
Bogdan Cristei (09:40)
So industrial influencers basically. Interesting. And I think you mentioned you're now working across multiple countries, multiple continents. How does that pipeline scale? Is this a boots-on-the-ground operation everywhere, or have you figured out something more scalable?
James Kujareevanich (10:12)
We have some secret sauce to scaling. We go through a lot of partners - partners who are running platforms that have a lot of players already in their own platform. And that's the best way to get from one to a hundred very quickly. As I mentioned, we have boots on the ground in India, in Thailand, Southeast Asia, expanding to Vietnam, Indonesia. So it's a combination of self-run operations and partnerships through the influencer network and the platform players.
And we are very strict. We use computer vision to screen out what is good and what is bad, and then we teach our partners what is good and what is bad.
Bogdan Cristei (10:44)
And I assume you have some quality control across all the different geographies. I saw your co-founder at the Computer Vision Conference. That was cool.
So you've been selling into frontier labs and robotics companies. I'd love to hear what's happened. You've completed some pilot engagements with multiple frontier AI labs - without naming names if you prefer not to. What have you learned from the pilots? What surprised you about what labs actually want? Did they come back and ask for something you didn't expect? How did you change the product based on feedback?
James Kujareevanich (11:42)
So I would say, bringing back my previous point, there's really no consensus. There's no specific standard. We kind of have to customize our whole data ingestion pipeline - the data creation part - to fit each specific use case. Instead of having one main data pipeline, there's a parallel track for each client. Some are more strict about a particular feature, some are less strict. And we need to invest the effort deliberately. We have human experts, factory engineers, to also validate the action labeling.
Some people are less strict. They're like, okay, just make sure you roughly got the action every 10 seconds correct. That's good. So we spend less time on the manual cleaning. Some people want atomic action - it's got to be super exact frame by frame. That requires a lot of human-in-the-loop correction. That's something I learned. It is a project-by-project basis. Nobody actually knows what's the best data. Honestly, everybody has a different gold standard of what is the gold standard of data.
And a lot of it is surprisingly still just vibe checking - eyeballing it. I asked them, how do you know this is good? They're like, it's vibes. The researchers are still vibe checking the data.
Bogdan Cristei (12:57)
It's all vibe-based. But I would imagine the demand numbers from some of the labs are enormous, right? Probably orders of magnitude larger than the pilots. So how do you think about that gap between what they want and what you can deliver today? If you had to 10x your capacity, is that a people problem, a technology problem, or an access problem?
James Kujareevanich (13:16)
It's very interesting because fortunately we just raised - we closed a six million dollar round. And we have the capital to scale. We're scaling. We just doubled the team size in the last two weeks alone. We grew from six to twelve people.
Bogdan Cristei (13:26)
Congrats.
James Kujareevanich (13:39)
We need to revamp the whole back-end infrastructure to accommodate. We're targeting millions of hours of data by the end of this year. So we need to massively make it more scalable. And then on the operation side, we hire more folks in Thailand and Philippines and India to expand on-ground data collection sourcing, expanding the annotation team, the expert annotation team. It's crazy. You literally have to unbottleneck every single part of the whole value chain. It's just a lot of work, a lot of sleepless nights.
Bogdan Cristei (14:00)
And how do you think about the verticals? Are there new verticals on the horizon beyond manufacturing - maybe in hospitality or healthcare?
James Kujareevanich (14:19)
We're exploring. We are doing more partnerships with other verticals now. We are not directly collecting, but we are working with partners like hotels and people in car mechanics shops doing car maintenance. We are definitely expanding. Our goal is to capture all the economically valuable tasks that robots can eventually help support where there's a labor force.
Bogdan Cristei (14:42)
Got it. Okay, so if you speak with some folks out there, including investors, they would ask some hard questions. Feel free to disagree with any of these, but I'm thinking they would say something like - synthetic data is getting better fast. Companies are building simulation environments, using generative models to create training data. Some of them are arguing that you don't need real-world footage at all. Why are they necessarily wrong? Is there a world where synthetic and real-world data are complements, or are they substitutes?
James Kujareevanich (15:14)
I would say people have been talking about synthetic data for a long time. It's just not there. The money speaks for itself. If synthetic data could replace real data, I wouldn't be here speaking to you. If the frontier labs are still looking for real data, it means the synthetic data is not physically accurate enough, or it's just not diverse enough to supplement it.
I know use cases where they use synthetic environments to do policy evaluation - just test things out. But for the training part, it's mostly still real data.
Have you heard of chaos theory? Even if one frame is physically inaccurate, the next frame will compound that error. So the first frame might be 1% off physically. Second frame will be 2% off. And then once you hit the fifth minute it'll be totally gibberish. The error rates compound. So even if you have 99.99% fidelity in the synthetic data, once you hit the hundredth frame it's already very off. It's really hard. How many nines do you need for the synthetic data to be good enough until you can have super realistic simulation of world models?
I would say it'll take a while, but there will be room to play between real data and synthetic data. But right now, very bullish on real data and that's why I'm here. We are open. Our goal is to assist the development of physical AGI. If in three years synthetic data becomes actually useful, we're happy to go into that space. We are looking at where the most value is being generated right now and we are focusing on that part.
Bogdan Cristei (16:56)
So what's valuable today - follow that. Very good.
James Kujareevanich (16:58)
Yes. We don't need to talk about people going to Mars yet.
Bogdan Cristei (17:02)
So then in this business, where does the defensibility live? If I'm a well-funded competitor and I decide to build a factory data network tomorrow, what stops me? Is this a winner-take-most market, or will different data providers own different verticals and geographies?
James Kujareevanich (17:20)
If you draw a parallel to the LLMs space, you don't see one player. There's Scale AI, there are multiple data companies, and everybody kind of becomes known for a specific data niche. Some specialize in scientific data, some provide localized language, different local cultures. And obviously, if you're the frontier labs, you don't want to have all your bets on one vendor. What happens if this vendor has a data breach or they screw up? And also the pricing dynamics - they want to have options to choose from and they want to make sure their suppliers don't die so they have options to pick from.
And also I think people underestimate how hard it is to actually get good data. There are so many bad actors out there. We also source out for different data, and even just raw footage - 80% of it is not usable. Either the device is not up to standard or they're just not capturing the right thing. And that's just 20% of the value chain. The cleaning part, having good annotation, having good hand pose extraction - that is the hard part. People really underestimate how hard it is to get really good, clean data.
That's why people jump in thinking it's easy. They try to go to a price war, selling cheap data. But in the end, that's not what the labs want. They definitely want cheap data, but they want good quality data and they're going to pay a premium for good quality data.
Bogdan Cristei (18:54)
You mentioned hand pose extraction. What's your experience been with getting quality data around that?
James Kujareevanich (19:00)
It's partly algorithm and it's also partly how good the video is. There are a lot of limitations. I'm sure eventually the hand pose algorithms will solve a lot of things. But the best way to ensure good hand pose is just ensuring the motions of the hands are clear. That they don't leave the frame too much. It's clear what people are doing in the video. There's a lot of other things which are our secret sauce which we're not going to share.
Bogdan Cristei (19:29)
Interesting. So when we talk about timelines, what do you think the market or the general public gets wrong about robotics timelines? Are we closer or further than the consensus view?
James Kujareevanich (19:44)
I would say some people are very bullish. If you just follow the PR of the big robotic labs - I'm not going to name names - they make it look like it's coming next year. I would say to have a really fully generalizable, 99.9% reliable robot is probably five years out at least. The bar is really high to deploy robots in the household or on a factory floor. Imagine if it tripped and fell onto a baby at home. That might happen one out of a million times, but that's still a risk. To get to that kind of reliability, I think it will take a while.
I don't think we're anywhere close to being able to generalize. It could be able to do 90% of tasks well, but 90% is nowhere close to deployment level. In the short run, I think people who are focusing on vertical-specific use cases, making it really reliable - that probably has early adoption. You see automation in factories with pick and place. You can get that to a hundred percent reliable because it's literally just hard-coding the positioning.
For generalized robotics, it's so hard because at the end of the day it's still a probabilistic model. It can trip at any time. You need that extra level of security and risk management for robotics.
Bogdan Cristei (20:59)
Let's do some quick hot takes. I'd love to get some of your hot takes. Our common friend Junfan loves to do these. I'm going to copy him a bit. Shout out to Saturday Robotics. Okay, first one. What do you believe about the future of robotics that most people still disagree with?
James Kujareevanich (21:28)
I have to disclaim - I'm not an expert in robots. I'm a businessman, right? So take my hot take with a pinch of salt. I feel like there's a lot of parallel between LLMs and the robotics world model space. I feel like it's going to eventually narrow down into two or three big players, like how you see OpenAI, Gemini, and Claude. These three guys. I think the resources to build a general robotics foundation model will be super resource-intensive. You'll have a few trillion-dollar companies that will crack it. A lot of these startups, I love them, there's so many of them. But I think there'll be a lot of M&A.
It doesn't make sense to have ten foundation models for robotics. So I think my hot take is convergence into super-big players. But I hope there will be new entrants, like how OpenAI and Anthropic didn't exist five years ago. Now they're some of the biggest foundation model companies. So there might be a place where some of these robotics startups become the next OpenAI or Anthropic. It doesn't just go to the incumbents - the Google, the Amazon.
Bogdan Cristei (22:07)
Hundred percent. All right, if Vision Lab works the way you think it will, what does the world look like in ten years?
James Kujareevanich (22:32)
So we have a mission of not just being data creation in the short run. We are one of the biggest data vendors in the industrial data space, but we want it to be two-way. Once the robots get trained, they need to deploy, they need physical RL. We want to be that partner. We want to help bring all these robots globally to test out with our partners and eventually be like the Siemens of robotics. Our vision is we become that deployment layer for all these frontier labs.
I see a lot of reasons why they would need a partner. They focus on what they do best - the tech, the model building - and then we help with the execution. It's similar to how they need Scale AI to do the expert RL for LLMs, right? We want to be that for the physical environment.
Bogdan Cristei (23:18)
Beautiful. All right, last question. What's one piece of advice for a robotics founder starting today?
James Kujareevanich (23:26)
Don't compete with me. I'm kidding. I think it's a very exciting space. There are people, one of my friends is running cross-robotics transfer learning - different grippers transferring to another robotics form factor. There are people betting on synthetic data, eval environments, doing policy training. There are so many parts of the pie. You can be the hardware player - there's a big market for just building the best hands, really dexterous hands. And then there are the model players. And then there are the vertically integrated players. It depends on how much budget you have, I guess.
Pick your poison. It's definitely a very hard space. Every single part of the value chain will be very competitive because an attractive market attracts players. We also work in a very attractive environment - that's why there are so many players as well. Just be the best in whatever you decide to compete in, because otherwise your competitors will eat you up.
Bogdan Cristei (24:32)
All right man, this has been great. Where can people learn more about Vision Lab?
James Kujareevanich (24:38)
You can go to our website, thevisionlab.ai. We actually haven't done any announcement yet - we're going to do it. I think by the time this airs, you'll probably see more of us in the media. And then yeah, just find me. My name is James Kujareevanich. I'm sure Bogdan will type it out. Just DM me on LinkedIn. I'm very active on LinkedIn. Just don't be too salesy. I'll respond to your message.
Bogdan Cristei (25:05)
Really great to have you on, man. Thank you so much for your time today.
James Kujareevanich (25:08)
Thank you so much, Bogdan. It's a really enjoyable conversation.