Deep conversations with the founders, investors, and operators building real-world AI - robotics, automation, industrial systems & AI infrastructure. Past the headlines, into how these technologies are really built, deployed, and scaled. Hosted by Bogdan Cristei, venture partner and former systems engineer.
THE OPTIM UPDATE
Useful Now: The Case for Application-Specific Robots
Guest: Arjun Subramaniam, Founder & CEO, Factory Intelligence
Host: Bogdan Cristei
---
Bogdan Cristei (00:00)
This is The OPTIM Update. I'm Bogdan Cristei, and today's guest is Arjun Subramaniam, Founder and CEO of Factory Intelligence, a physical AI company training tactile foundation models for the real world. Arjun has personally toured more than 70 factories deploying robots, and his co-founder Mike Strope spent 15 years in packaging, including a president seat at a global packaging machinery manufacturer.
We walk through his first deployment in electrical prefab, and we sit in the integration trap question that every robotics-as-a-service company eventually has to answer. Without further ado, here's my conversation with Arjun.
Bogdan Cristei (00:44)
Arjun, welcome in, man. How are you today?
Arjun Subramaniam
Hey Bogdan, nice to meet you.
Bogdan Cristei
Are you ready to talk about intelligence in factories or what?
Arjun Subramaniam
Yes, I am very excited, and you know how I am - I can speak for ages on this topic. So let's get started.
Bogdan Cristei
I will cap you at some point. For now, walk us back to where this all started. You came out of Purdue Industry 4.0 research. You then spent some years deploying robots inside actual real factories - I think you said 70 factories or something like that. And then Mike has also spent a number of years, I think 15 years in packaging. So what did all of that teach you guys that pulled you toward Factory Intelligence as a company?
Arjun Subramaniam
Like you said, I've got a weird, broad history. I spent a lot of time doing research on robots at Purdue, but then I've also done about 70 factory visits across America, specifically in the Midwest. And my co-founder Mike has 15 years of experience doing industrial automation - not only touring factories here in America, but across the world.
I think the biggest takeaway from this is: when you're doing robotics in a lab, it's hard, but it's way easier than you'd think. Because the further away your robot gets from where you are, the complexity grows exponentially. When you're in a lab, you're able to be in a very tight loop. If something breaks, you can fix it.
Bogdan Cristei
When you say "lab," do you mean academic lab? Is that what you mean?
Arjun Subramaniam
Yeah, I mean academic lab, research lab - where most people working in this exact field come from. I'm from the same place too. When you're in a lab, you're constrained on resources. You don't have as many robots, as much compute. But when things break, you're able to quickly iterate on it. It doesn't have to look pretty after it's fixed. It just has to work enough to publish a paper.
Publishing papers is nice. It allows you to make breakthrough research. It allows you to demonstrate things as a proof of concept. But there's no requirement of reliability in a paper. There's no SLA - no service level agreement. There are no customers. There aren't people that get mad when it doesn't really work reliably. You just need to be able to reproduce it for it to be viable research.
But in factories, you need to be able to reproduce it day in, day out, every single second. Because if you don't, some of these production lines are $20,000 an hour. Every minute can cost you greatly. Especially when you go to food manufacturing, produce - these lines go by extraordinarily fast. Dairy lines, ice cream - these products are moving extremely fast, and every minute of delay cascades down the line, where you have milk expiring because your robot screwed up.
So that's the one thing you don't really see in labs. There's a huge gap between demonstration and deployment. That's where the years of experience I have actually deploying robots in real factories serves me greatly. You have to make sure it works as a demonstration, but it also has to work every single day for your customer. And you're far away. Once you deploy, you can't really sit there and babysit it. You can't watch it like a hawk. You can't cherry-pick the times it works versus the times it doesn't. It needs to work when you're not watching. It's like a tree that falls in a forest - it just needs to keep chugging away.
Bogdan Cristei
And your COO is actually a packaging machinery veteran, right? So maybe that speaks to the same topic.
Arjun Subramaniam
Yeah, absolutely. When you think about what I did, it's more systems integration - going project to project, figuring out a customer's pain point and putting together the pieces. But what Mike has a lot of experience in is productization. How do you make a single product that does automation in factories but is reproducible across factories in 60 countries?
His previous company didn't work directly with the customer in most cases. A majority of their business came from distributors - they had over 175 distributors. So in a lot of cases, you're not even having direct contact with your end customer. You need to build a product that's so intuitive, works so well, that you can pass it off to someone who can then implement it. And their reputation is on the line - distributors will no longer carry your product if you can't properly support it, if it doesn't work, if it doesn't integrate well. So he has a lot of experience building machines that can be deployed easily with very low marginal cost in factories.
When you put this experience together, you're able to truly deeply understand a customer's problem, find a large enough market, and then build a reproducible system - and then not only deploy directly with our customer, but build it in a way where we can scale it up through various channels.
However, I'm a huge proponent of being fully vertically integrated, doing as much of it as possible by yourself, because you have so much more control over the entire stack. That's why even when it comes to robotics, we try to be as full-stack as possible. We don't even buy our teleoperation frameworks - we actually build them ourselves from scratch. So any time we need to remote in from India to solve an issue, it's our own software end to end. This allows us to benefit from the entire stack of robotics.
Bogdan Cristei
So a couple of things to get you going here. What's your one-liner for the company? I know that changes all the time, but I'd love to hear what it is now, especially how it relates to your strategy, which you call "Useful Now." And as you talk about that, I'd also love to understand: where do you think the dominant robotics narrative is most wrong right now?
Arjun Subramaniam
So our one-liner is: we're building autonomous robots to do skilled labor. These are tasks that take years to learn, but minutes - if not hours - to get tired of, and seconds to get hurt doing. Things like what electricians do, construction workers, maintenance technicians, the people who have the knowledge to run CNC machines, mills - the machines that produce the stuff we use every single day.
This is called tribal knowledge, or enterprise knowledge. It's stuff where people are retiring but aren't training enough replacements. No one wants to do an apprenticeship anymore. No one wants to go work in a factory. Everyone wants a cushy desk job, or they want to be an influencer. But no one really wants to pick up heavy things and place them, or screw things in hunched over a table all day.
So we're really building robots that can do these tasks. We're putting together the best of physical AI models for their power and flexibility. We're combining that with traditional classical robotics for its reliability and safety. We build purpose-built hardware on top of off-the-shelf arms, and the software actually ties it into our customer systems. So it's a fully integrated part of their business, like a worker they've had for 10 years.
What we believe the industry gets wrong: one, you can't buy your way into a large enough data set, or into a deployment. You have to have paying customers, and your hardware - your robots - need to be doing real work. This is kind of like the Tesla and Waymo self-driving dilemma. Waymo was under Google, so they had billions and billions in advertising money to burn toward this. Whereas Tesla sent actual cars doing real work, useful work for people, into the world. People thought self-driving was solved for the longest time, but the long tail just kept getting longer. People said self-driving was solved for the last 10 years, but that hasn't been the case.
And with robotics, specifically manipulation, that long tail is even longer. When you're in a car, the only contact you have is with the road, and you hope to not touch anything else. But in robotics, you're constantly touching things, constantly pushing on things, exerting forces, things are exerting forces back on you. So the long tail is extraordinarily long. There's such a high dimensionality of things that can happen - not only in a home, not only in a park, but specifically in a factory where there's such a high density of high-dexterity, contact-rich tasks.
So we believe there's not enough money that can be raised to get you to the point of bootstrapping a large enough data set of actual manipulation doing real-world tasks. You have to start with the long tail first - build specialist systems doing real-world work - and then you can generalize enough to build a foundation model for industrial work.
Another thing we believe is that vision is insufficient to do these tasks. People have over-indexed on vision alone because it's been easy to scale. I'm not someone who hasn't learned the bitter lesson - I believe scaling is the only way to truly achieve generalized robot models. But vision gets you to the middle of the bell curve, where you can choose a lot of lab tasks, pull a lot of data from the internet, and scale based on that. But when you want to do real-world tasks where your robot is actually interacting with things that don't change visually, that's where vision fails.
Current models right now are relatively stateless and vision-based. There are a few companies working on enhancing memory for models - Rho-alpha has demonstrated some memory capabilities, Generalist as well, Pi has demonstrated their own version of memory for VLAs. But currently most models have very short memory, and whatever they're doing in their environment has to change visually for the model to understand that this part of the subtask is done and it needs to go to the next.
But for a lot of industrial tasks, that just isn't there. These things aren't deformable. If you're picking up a bag of chips, you know how much pressure you're putting on it based on how much the bag is deforming. But when you're picking up a block of metal, how do you know? When you're inserting a block of metal, how do you know? A lot of these tasks, there isn't a visual differentiator between states. So you need other modalities - force feedback, torque, tactile sensing - for the robot to understand these contact-rich tasks. Insertion, assembly, screwing, welding, bending - really contact-rich tasks that people can do with their eyes closed. Vision-only models just can't do these, or they're extremely slow.
If you've seen the current demos from Acre Robotics, the way they're able to move so fast is because tactile features are just so much lower dimension. You have less data, but the data is more high-signal. That's why, when you're trying to grab a pen out of a bin of pens, you think "grab the pen," and you can just quickly move until you feel a single pen and grab it out. But a vision-only model has to process super high-dimensional, tons of features at every time step to see: is the pen in between my grasp? Have I grabbed it yet? When it picks up - is the pen still in my grasp? Oh no, it's fallen down, I need to do that again. And when there are five boxes per minute coming down a production line, this is way too slow.
You need to combine modalities. You can plan at a high level with vision, but a lot of this feedback can be given through tactile modalities by moving really, really fast. That's why we believe vision alone isn't enough.
But you're not going to be able to bootstrap a large enough tactile data set by teleoperating robots or by deploying enough data collection devices. We're very AGI-pilled on everything. We believe robotics is the way, and we believe in various data sources - egocentric data, pure video data, UMI-style data collection devices like what Sunday is doing and what Generalist is doing. We believe in all these modalities. But to truly scale this up and have a tight feedback loop between your research and your deployments, you need a fleet of robots.
Tesla is able to push so many software updates so frequently because they can test their experiments on a large enough sample size. And once they make progress, they can deploy it and get feedback very, very quick on a large enough sample size. So this loop helps them iterate fast and benefit from scale.
Robotics right now is extremely early. It's an extremely diverse field where people are betting really big on various theses. Some people are scaling VLAs and world models. Some are going heavily on egocentric data. Some are going very heavy on reinforcement learning. Some are focused on data collection devices. There are so many improvements across so many different tracks, and they're all kind of merging toward one thesis. It isn't one thing - like in LLMs, where scale is the only way. There are so many different fields, and to truly benefit from the breakthroughs from each of these tracks, you need to be a full-stack robotics company - building your models, building the hardware, building the infrastructure and software, and having actual customers to benefit from the improvements in models.
In a long way, this is the same bet people are making with agents right now. People are betting on different theses, like "I'm going to make a better model." But as one model gets better, the value of your company decreases. People are blasting on Twitter or X every single day: "build a startup where your market gets bigger as the agents get better, or the base models get better." That's truly what we're working on too. As the foundation models get better, as the backbones get better, as the infrastructure gets better, our business only gets bigger. We're able to tackle bigger tasks, longer-horizon tasks, more dexterous tasks, more complex tasks, more dangerous tasks. So we are uniquely positioned as a full-stack robotics company working on Useful Now tasks to benefit from the entire wave of improvements across the industry.
Bogdan Cristei
That's a claim you have - that you're the only company occupying the application-specific full-stack quadrant. That's a super strong claim. My question would be: why hasn't anyone else built here?
Arjun Subramaniam
It's for that reason - there isn't enough infrastructure to be full-stack. When I say "only," I'd say only in our field. There are very few - you can count on both hands the number of vertical companies working on real-world customer tasks. Everybody right now is trying to mirror what's happening in the software space. They're building models, they're building infrastructure, but no one's putting it all together for real customer tasks.
Bogdan Cristei
The model-first players would argue you lose because their generalized capability arrives before you finish your vertical-by-vertical buildout. What's the best version of their argument, and how do you push back on it?
Arjun Subramaniam
The best version of their argument is that generalization helps build the model. The more data you have, the more diverse data, the more different deployments your robot can do - it all feeds this God brain. What people haven't done is actually deploy these models in real-world use cases. Getting a general laborer is only table stakes. People can hire a 19-year-old kid who hasn't gone to college off the street just as easily as they could get their robot. They're not really looking for that. They're looking for reliable labor that understands their business end to end - not only as good as a human, but superhuman, reliable, working 24/7.
So yes, they can get generalist models, that's fine. But that's like getting a temp off the street. What we're selling is a specialist - someone who has spent 10 years in your business and understands it end to end. The large model players just cannot go to all these different industries and fully understand their capacity end to end. There's a diffusion problem. It's the same question of "will Claude be good enough to take out your agentic startup?" But if you can go end to end, truly understand your customer, and build something around their business that harnesses the power of these models, then you're the true winner. The applications layer truly wins, not the model layer, because that's where the value is generated. And that's where we sit. As the generalist models get better, our specialist applications get even more verticalized, even more specialized, and generate even more value.
Bogdan Cristei
The platform providers would argue you're just selling expensive workcells. How do you push back on that argument?
Arjun Subramaniam
That's a problem with the platform providers too - they have to be general enough to solve the problem for all their different customers. What we can do is work with multiple platform providers, multiple hardware providers, and build the solution that delivers the best labor output for our customers. It doesn't really matter what your hardware does if it doesn't do the right thing for your business. So we're able to help our customers avoid lock-in with platform providers. We can build the right hardware with the right solution and allow them to be flexible in the future with our robotics-as-a-service model.
Bogdan Cristei
One thing I wanted to ask you about: most of the well-funded robotics bets - humanoids, general-purpose foundation models - are still, in my opinion, years from production deployment. That's clear to a lot of us. And I'm wondering, if what you're saying is true, if Useful Now works, what happens to the humanoid timeline? How does that bet in humanoids turn out maybe five years from now if Useful Now kind of works?
Arjun Subramaniam
There's a weird analogy I like to give: when you're trying to climb a really tall tree, there's no point climbing other shorter trees nearby. If you want to actually deploy robots, you need to work on deploying robots, because there's a huge set of problems with reliably deploying robots at scale. Humanoid companies are making the long-term bet of "can I make something general enough to make it easy to deploy?" But when they get to the point of actually deploying, that's when they open the Pandora's box of deployment-related problems.
What we're building right now is that go-to-market motion, that deployment motion - truly understanding what it takes to have robots working at scale right now. So when they get to this bet five years down the line, they're going to be leagues behind what we're working on right now.
But what that changes for them too is that it doesn't really shrink the TAM in what they can achieve - it just expands the opportunities they tackle. There are far more general tasks that can be taken on: the exploration of the seas, building data centers in space. There are so many tasks to be handled. The tasks that right now generate really high value will be captured by specialist robots that can produce things in the best possible way. You don't need a thing with two arms, two legs, and a head to produce wall outlets or to build buildings.
Bogdan Cristei
That's great. Let's talk about a specific example. Your beachhead is electrical prefabrication. Most of our listeners have never been inside a prefab shop. So maybe you can paint a picture for us. What's actually happening in there? What's the day in the life of a prefab worker? And then: why prefab and not live construction sites? Walk us through that entire thinking, please.
Arjun Subramaniam
That's a very good question. Before, what people used to do is you'd have these union electricians with 20, 30 years of experience. They bring all the raw materials on site when they're building a house, and they start assembling things. Like, if you want an outlet that goes on your wall - the stuff you plug your phone charger or appliance into - first they take an outlet, which might come from China. Then they take wires: they bend a live wire, a neutral wire, and a ground wire. They hook them onto the outlets, screw them in, tape it, and then take a conduit wire that goes from the outlet to the junction box, cut it to length, and install it into your wall. So imagine doing that a hundred times for a house, or 10,000 times for a building.
This is the reason buildings used to take so long to produce and are so expensive to produce. People have been working on making this cheaper and cheaper, and one big change that's significantly reduced the cost is this concept of prefabrication. It's about how much of the material and the building can you build offsite in a centralized location with the right machinery and labor. Then you bring it on site with kits and quickly assemble these buildings. That's how buildings that used to take two years now get put together like Legos in a few months. That's how you can see a skyscraper being built in a year.
Bogdan Cristei
So a good next follow-up: why building outlets first? Of all the tasks in a prefab shop, what made that the right starting point? And maybe you can talk a little bit about Prefab-Cell-E1 - what's inside that workcell, what's it doing, what's the operator's role? Could you walk us through that?
Arjun Subramaniam
Right now our robots are deployed with our first customer, Houston Electric in Indiana, building outlets for their electricians that go on site. Prefab-Cell-E1 is our first pilot production cell. It has eight robots in it, and it can produce outlets for $3 per hour, which is 10 times cheaper than equivalent labor and can produce three times as many plugs. So it's essentially nine times as productive.
Bogdan Cristei
What's an outlet, what's included? Just walk us through it. Is it just the three wires that you connect to it? Or why does that need to be manual?
Arjun Subramaniam
So the robot bends three wires. Electricians usually use pliers, but we've given the robot a tool - like a jig - to do it. It bends three wires, puts a hook on it, then picks up the outlet and tails the outlet - that's what the electricians call it - which is really just hooking that wire onto the screw that's been pulled out. Then it takes a screwdriver and drives the screw in, so the wire is secured.
You put three of these wires - a live, a ground, and a neutral - and this is what's wired up to your house. That's why you see three holes or three prongs in your plugs and three holes in the outlet: one's a ground, one's a neutral, one's a live. That's how you get power from the grid. So this gets tailed on, screwed, and then an electrician checks it and tapes it.
Right now, what the cell is doing is prefabricating these outlets. For every house, whoever the engineer is, whoever designs the house, chooses which outlet, what specification, what wire to use. The electrical contractor has to order all this stuff six months before the project gets done, and then build maybe 10,000 of these. So the moment they get an order, they press start on our cell. It pulls all the information from their ERP system, so it knows what plugs to use, what wires to use, what specification to build to, and it builds the exact number of plugs they need on schedule. So they have it ready to go. It's basically from order to stuff shipped out in your truck. It's kind of like a building that produces buildings before you go put it on site.
We started with electrical outlets because that's one of the most contact-rich, one of the most boring, repetitive tasks that every single prefab shop has to do. Every single building in America has to have an outlet every eight feet. This is something we need to fact-check before you put it in, but I'm pretty sure it's every eight feet. And in specific other scenarios, you need more outlets. So these things are in every single building, every single data center. The reason we cannot produce buildings fast enough is because we don't have enough labor to produce this. That's why in China you can build miles of highway per day, but in America it takes so long - because of the labor and the permitting and all this regulatory stuff. Maybe that's a bad point. The reason you can build buildings so cheaply in India, China, and other developing countries is because you have an abundance of labor there. You have all the raw material, but in America we have such a shortage of skilled labor or tradespeople - plumbers, electricians, construction people.
If we can truly build robots that can do their job, or at least assist their job, then we're unconstrained in the number of houses and buildings we can build. It just becomes cheaper, faster. If there's a data center coming up in a state, all the electricians are booked out for years in advance. People are flying electricians out from other countries. This is more like labor on demand. These are skills we can never lose now, because we have robots that can do the task.
Bogdan Cristei
Okay, that's interesting. Let's talk a little bit about the technology. So we were chatting before this call that a month after you kicked off your project, you had this wire-bending model that was generalizing to colors you hadn't seen before - it hadn't been trained on them. So I'm curious - what was happening inside the model when that started working?
Arjun Subramaniam
We noticed that because we're adding all these additional modalities in, we're able to get so much more performance from way less data. This is less than five hours of data on site across all our different tasks, and we're able to build that plug. This is because there's a grounding between what the robot sees and what the robot feels. So it's able to do these tasks with way less data, because there are fewer states it can be in when you have this tactile data as well.
That's why we're able to see such fast improvement in our models. We believe that every single new task will be faster and easier to deploy as the robot sees more materials, as it feels more materials, as it handles more materials. All of this transfers and generalizes, and we get it for free. If not - we're getting paid to get all this data, because our robots are deployed doing real-world labor.
Bogdan Cristei
Okay. Every robotics-as-a-service company eventually runs into the same wall. Each customer needs custom fixtures, custom workholding, custom safety setups, custom support in general. The fully-loaded deployment cost ends up wrecking the margin. So I'm curious, why doesn't that happen to you, and how do you think about that?
Arjun Subramaniam
Exactly. And that's one reason why we're not focused on robots to automate manufacturing and logistics. Rather, we are building skilled labor. These are people that are trained and given the tools to do a very specific task across different sites, across different jobs, across different buildings. We're building the electricians, the machine tenders, the maintenance technicians of the world - who truly understand the job they're doing. They don't really care where they're doing it, who they're doing it with.
So when you get into the nuts and bolts of automating a specific task, this might change between customer to customer. But when you look at the people doing the task, that is relatively common between customers. That's where our platform really helps us quickly make the modifications necessary to integrate well with our customer.
That's why having a good software platform is so important - where you're not treating every deployment like a lab demo, prompting the robot to do what it needs to do, or trying to make it look good for a video. It needs to be well integrated into their ERP system too. Your software has to be modular enough that it understands all the common different software, the data sources like data historians in factories, the common ERP systems - Oracle NetSuite, SAP. We need to interface easily with these. So having a modular platform helps us avoid a lot of these issues, but focusing on very specialized tasks helps avoid and mitigate a lot of these concerns. It's the reason you can hire one person who worked in a different factory into your factory and still retain a lot of the skills they had. What they're working on might be different, but what they're doing is very similar.
Bogdan Cristei
Interesting. And then, at what scale do you actually know whether the deployment economics work? When do you think you'll have that answer?
Arjun Subramaniam
That's not really an issue if we grow profitably. But the scale we're looking for right now is replacing - or at least having - one robot to one skill, like one robot to one electrician in America over the next four years. We want to have just as many robots doing these jobs as humans doing these jobs, because then we can get a large enough sample size, which lets us know if we're just as effective, or if we can be better than humans.
Bogdan Cristei
So if you get a question like, "when does your data flywheel help you become cheaper for deployment N+1 versus deployment N," how do you answer that question? What does that data flywheel need to even look like for that to be a possibility?
Arjun Subramaniam
Right now, for our data flywheel, we're looking for 50% success rate on any task we try to achieve. It's not about scale of robots - it's really about scale of data success. Because if we have more than a 50% success rate, then we can improve exponentially. It's really about getting more than 50% across a large enough scale of robots. Then we can collect the higher-quality rollouts versus lower-quality rollouts and build that self-improving data flywheel. Because if we're less than 50%, we're not improving at a fast enough rate to hit critical mass. But if we're over 50% success rate, then it's more like an H-index - if we're working on five tasks, we want five robots; if we have 50 different tasks, we want 50 different robots. That's the way I'd answer it.
Bogdan Cristei
Okay. So I see more and more that the touch modality is becoming more talked about. Until recently, I felt like it was a little bit underweighted. How do you view this concept of touch becoming more important?
Arjun Subramaniam
I think vision was a great proof of concept. It was great to bootstrap from minus-1 to 0, or 0 to 1. It got us a lot of the way because we were able to benefit from all the internet-scale data, all the data on YouTube, all the images, VLMs. That's why we were able to improve so fast with vision. But I think people are realizing that vision is saturated. We've gotten a lot of the distance we can with vision. And for all the reasons I mentioned, vision is just - you're throwing way too much compute, way too much information at tasks you don't need to.
That's why people are now realizing you can make a lot of headway when you incorporate tactile data. One, you're able to do tasks at real-time speed with compliance, which is something that's been missing from a lot of stuff you've seen in robot learning, where everything is at least six to eight times slowed down, and people are using hockey sticks near the robot because it's not safe for a human to be around. But now you're seeing stuff happen at 1x, if not half-time speed around humans, because you're able to use tactile data. And now you're able to benefit from the visual features and high-level planning you get from vision, but also the low-level features and very fast, high-frequency data you get from touch sensing.
I think touch sensing is now - you know, necessity is the mother of invention. People have been struggling with slow models that can't do very fine manipulation. So touch sensing is coming in to fill that gap of fast action with very fine-grained manipulation.
Bogdan Cristei
Most robotics teams either bet purely on end-to-end neural nets or pure classical control. You're combining both. I wonder why you think that's the right design. And if you can walk us through your architecture, that would be super interesting to hear about as well.
Arjun Subramaniam
What we get by combining the best of what we'd call neuro-classical controls is the flexibility of neural networks - which works really well in the middle of the bell curve, where there's some variability but it's inside the distribution of the network - and the classical controls let us deal with the long-tail scenarios, the stuff that's very hard to bring inside the distribution of the model.
One, we're able to help the robot fail more gracefully and recover from different scenarios using classical controls. But we're also giving safety boundaries to the robot - or affordances - where we have no-go zones, or collision checking for the trajectories planned by the robot, to make sure they're actually viable. If not, how can we edit these trajectories to make them viable, given a 3D representation of what it sees, and other classical robot motion planning methods. So we're able to use trajectories generated by neural networks, but edit them and check them, and generate various trajectories - sampling based on classical methods, using traditional motion planning - to make sure these are actually viable.
We can also simulate out the forces these trajectories exert on certain objects in the task space. So we have a kind of compliant control, where the model is not only outputting the trajectories - where it wants to hit, what trajectories it wants to hit - but also the various forces it wants to apply during that. If you see when you're trying to pick something up, you want to close your hand, but you're closing your hand until you feel something. So that's a kind of closed-loop compliant control. We're doing the same thing with our robots, where there's a higher-level position-based planner, but there's also a lower-level compliance-based planner that can execute these trajectories with the right amount of force until it hits a desired state.
This is what we mean by combining the best of neural networks and classical controls. It's for safety, reliability, but also performance.
Bogdan Cristei
So you have a world action model that takes image, proprioception, tactile, action, future states, all of that as inputs, right? Does that ever get to be too complicated? What does that give you that maybe simpler architectures don't?
Arjun Subramaniam
That's actually a really good question, and it has a weird answer. It actually gives us the most simple way to test our thesis. If you've seen the architecture, it's very similar to what Nvidia published with Cosmos Policy, where they're able to change just the latent space of a video generation model - not the architecture itself - to produce robot actions.
We're able to modify or extend this for tactile sensing both on the input side and on the output side, to test our thesis, and to build the infrastructure necessary to actually incorporate these various modalities into robots, without getting stuck in the random quagmires of architectural changes that lead to training instability. So we didn't have to waste a lot of compute and time wrangling the model - we're able to just add the modalities we want to add.
Of course, this is way slower than we need, because it's a larger model with a lot of additional little things in it that make inference slower. But we want to take our learnings from this initial model and move to more optimized methods - something similar to what Dream Zero does with a causal world model and an IDM. Dream Zero does not use an IDM, but we believe that's just way better for stable training. So: more along the lines of a causal world model, not as large as Dream Zero, but again, extended with tactile sensing from all the learning. It's really a decision to be economical and fast that informed using this model.
Bogdan Cristei
This is a fascinating conversation. We could go forever - I think it's going to be super helpful for folks to just hear your thinking on this. But maybe two or three questions to round it up toward the end. First one: what do you believe about the future of robotics that most people in the industry still disagree with?
Arjun Subramaniam
I think it would probably be that you can buy your way to a large enough data set necessary to train models that aren't only vision-based. There are so many different modalities you need for a robot to actually be human - or AGI, you know, artificial physical general intelligence, whatever. You still need to be able to touch things, you need to be able to hear things, you need to be able to smell things. These are all data modalities you just can't get from the internet alone. So you're going to end up spending a lot of money trying to get this data from the real world. You need to find that data flywheel that actually gets you to large enough data sets without having to spend millions and billions in venture money.
Bogdan Cristei
Yeah. All right. So if your company, Factory Intelligence, works - what does the world look like in five years?
Arjun Subramaniam
It looks like data centers on the moon, mining asteroids. It means we're truly uncapped with our production capacity. People draw parallels to China all the time, but that's a very low bar to set, because we can do so much more once we solve the problem of 24 hours in a day that humans have. We're limited to 24 hours per day per human. Once we can uncap that, that's when we have true abundance - and the ability to take on these moonshot projects with very low marginal costs. So that's why I think there'll be data centers on the moon, mining asteroids, once we can solve the base problem of skilled manipulation.
Bogdan Cristei
Yeah, all right. Where can people find more about you and what you're building?
Arjun Subramaniam
You can go to factoryintelligence.com. We're updating the site daily to give you more information on what we're working on. You can email me at arjun@factoryintelligence.com. We're always looking for top researchers, amazing customers to help us build the future of robots, and also investors who understand what it's like to deploy hardware and what the future could be like when we succeed.
Bogdan Cristei
Awesome. Well, this has been great, man. Thank you so much for your time, Arjun.
Arjun Subramaniam
Thank you for having me, Bogdan. It's always a pleasure to chat.