Deep conversations with the founders, investors, and operators building real-world AI - robotics, automation, industrial systems & AI infrastructure. Past the headlines, into how these technologies are really built, deployed, and scaled. Hosted by Bogdan Cristei, venture partner and former systems engineer.
THE OPTIM UPDATE
Variation Kills Automation: Teaching Robots Stiffness | Sze Cheong & Elijah Ong, Devol Robots
Guests: Sze Yuan Cheong, Co-Founder & CEO, and Elijah Ong, Co-Founder & CTO, Devol Robots
Host: Bogdan Cristei
---
Bogdan Cristei (00:00)
This is The OPTIM Update. I'm Bogdan Cristei, and today I'm talking with Sze and Elijah, the co-founders of Devol Robots. Devol builds the AI that lets industrial robots handle contact - plugging a connector, seating a lens, snapping a part onto a rail - the work factories still need to do by hand after they bought the robots. Sze spent a decade building manufacturing businesses in Malaysia before starting Devol. Elijah turned down a Stanford PhD to build a force-controlled robot arm from scratch at a three-person startup in Austin.
Their model is running in production at a US optics manufacturer, and this spring they published results showing it beating two leading open-source models on contact tasks. We get into why a camera can't tell a robot resting on a table from one pressing into it, how their model learns stiffness from the way a person moves, and whether structure can beat scale when the other side has billions of dollars. Without further ado, here's my conversation with Sze and Elijah.
Bogdan Cristei (01:05)
Hey guys, so awesome to have you on. We've been talking about doing one of these for three, four months now. Super excited to see you here.
Sze Yuan Cheong (01:14)
Great to see you, Bogdan.
Bogdan Cristei (01:15)
Super excited to have you on. So before we start the discussion, let's start with the name.
Sze Yuan Cheong (01:21)
It's Devol. It comes from George Devol. He invented the first industrial robot arm, the Unimate. So the entire idea is that he revolutionized the robotics field and built something that is programmable. And we were thinking, okay, the field hasn't changed for the last 60 years, and we wanted to be the other Devol, the one that changes the status quo of the field.
Bogdan Cristei (01:56)
Interesting. Very cool. Amazing. All right, let's start with a few questions to get going. Sze, let's start with you. When we were chatting, I remember that you've been running factories for quite some time, maybe even a decade, before you ever touched a robot. So I wanted to start there. Help me see the problem. Walk me through a plant that already owns robots. What works? What is still being done by hand, and why?
Sze Yuan Cheong (02:23)
Imagine walking into a factory, any kind of factory, from electronics - consumer electronics that are really small - to the big metal fabrication factories. What you'll see is that if you walk in from the end, the packaging line, the palletizing line, that will most likely be automated if the factory is big. The robot can palletize all the things that are prepackaged. But the things in between - plugging in different parts, assembling your phone, handling metal fabrication parts from station to station - all the things that have a ton of variation, humans are doing all of that. Because variation kills automation.
Bogdan Cristei (03:18)
Yeah. I was just on a call with another founder today, and he's been hearing from factories that the number one reason they can't automate is that too many things change all the time. That's the problem with automation.
So if you're trying to automate with robotics, but the robot can't feel, the factory tends to compensate with fixtures, jigs, integrator hours, things of that nature. What does the workaround actually cost, and who pays for it?
Sze Yuan Cheong (03:49)
The funny thing is, like you said, I started and have been running factories for a decade, and that's basically what we do. The entire field of industrial engineering is trying to figure out how to be more efficient and break tasks down into their smallest units. What it's trying to do is this: let's say I pick something up and put it somewhere. That particular task can be done either by a machine or by a human. And the way we think about it is that if we can have a fixture that gives an absolute position every single time, then we can automate it. Even in that case, you need to tune the machine or the robot every single day to make sure there's no drift. Those are the kinds of things we tend to automate, either with a robot or with a machine.
The human comes in when that kind of variation can't be fixed, when there are a lot of SKU changes. And when you're integrating machines, one of the biggest problems is that those jigs and fixtures require a lot of design - hardware integration, and ultimately programming and fine-tuning for the hardware to work. That entire part is a huge industry by itself, just to make sure the factory actually runs and can automate.
So that, I think, is the biggest challenge we have. And what I notice especially is that younger generations like us tend to go into computer science and now AI. No one really wants to do that kind of work anymore.
Bogdan Cristei (06:00)
Interesting. So let's talk about one of your deployments. I know you're working with a customer in optics. If you can talk a little bit about that without giving away too many secrets - I would assume there are thousands of lens variants, and there's probably a jig for every single process. Every pocket has a slightly different fit. From some of the videos I've seen you guys post, I would imagine that's a nightmare for a conventional robot. I'm curious, what does a failed attempt cost? How do you automate something like that?
Sze Yuan Cheong (06:32)
Like we usually say, what is automatable is, for now, already automated. That basically means all the major processes - washing, ultrasonic cleaning, coating, things like that - are already automated. But think about it: let's say there are 20 process steps. Twenty process steps require basically 20 different types of jigs, for every variation of the lens. So how do you transfer those lenses between those jigs? And for certain processes, such as washing, you don't need such a high-precision jig, so you tend to build a manually tuned jig that's basically a loose fit, to try to fit everything in.
So imagine the nightmare of having a jig that is itself non-deterministic, and moving lenses that are really small and easily scratched from that jig into a jig that has only 10 to 20 microns of tolerance. The lens gets stuck, because it's glass. And if it gets stuck and you try to push it in, you will almost definitely scratch the lens, and you get a reject. So the nightmare is handling all of this uncertainty throughout the entire process.
And the funny thing is, even when a human is trying to do this, they need at least six to nine months of training to be able to do it. It's really hard to do. I couldn't do it, actually. So yeah, it is a nightmare. That's why the theory is that you train a model and let the robot do it.
Bogdan Cristei (08:30)
Okay. So now that we're talking about theory, let's move over to Elijah, because I have a feeling he's got some thoughts on this.
So Elijah, I wanted to paint a picture for you. Let's imagine a robot arm resting on a table, and then the same arm pressing down with, let's say, 50 newtons. I think to a camera, those two are the same picture, right? And that really matters. So could you define stiffness and damping for us, just so we understand? Maybe even think about what my arm is doing when I plug in a charger in the dark, or something like that. Could you talk a little bit about that?
Elijah Ong (09:07)
Stiffness and damping, in general, in classical robotics, allow us to control and regulate forces and position at the same time, through stiffness. And damping is a way for you to regulate the velocity, the target velocity, through forces.
Back to your question about pixels: if you look at a robot pressing down on a table with 50 newtons, what it gives you from a pixel perspective is only spatial understanding. The robot is sort of touching the table, but the image doesn't tell you what it's doing. So information like forces, stiffness, and damping becomes really important to tell us humans what the actual interaction is, what information we're getting through this kind of interaction. That's why we have to use forces, stiffness, and damping to really understand these interactions, so that the robot can learn way more efficiently instead of just through pixel space.
Sze Yuan Cheong (10:25)
I'll ground this with an example, actually. Quick example. Just imagine you return to your house and you're trying to grab your key from your pocket. You have no vision at all, but you can do it, every time, and really, really fast. But if you try to imagine the trajectory your fingers are following in the pocket, you can't. It is really, really intricate. That entire feeling is based on stiffness and damping. They help us capture that kind of interaction. It's a really, really good medium to capture that.
Bogdan Cristei (11:07)
That makes a lot of sense. Talking about force control - Elijah, you turned down a Stanford PhD to join a three-person startup in Austin building a force-controlled arm from scratch. Curious, what did you learn there that you would not have learned in the PhD?
Elijah Ong (11:26)
I was really fascinated by teaching robots to grasp things, to manipulate objects. At the time, I think about ten years ago, roboticists and researchers in the industry were still pondering grasp metrics, kinematic analysis, what kinds of constraints and what kinds of metrics we could learn - whether from human heuristics, through deep learning, or through reinforcement learning. All these kinds of methods were just trying to teach robots to grasp things better.
But I kept running into the same situation: a robot holding a cup, without any other information, purely from the joint angles and the pixels - that doesn't tell you the actual interaction. So after I finished grad school, I got an offer from Stanford for a PhD program, and the thesis, the research direction, was still mostly robotic grasping and manipulation, mainly the same direction. I couldn't really figure out what was going on with it. I thought that direction didn't lead anywhere at the time. So I felt like I had to get out there and try to find the answer myself.
I came across this startup in Texas doing force control, particularly impedance control, and it really fascinated me. So I joined them and helped them build an entire robotic arm from scratch. At the time, they had this very interesting technology called a series elastic actuator. It wasn't common, and they were building their own actuators and sensors.
And I found out that impedance - forces, stiffness, damping - is actually the complementary element, the right element, to serve as a physical representation for the robot to really understand what's going on behind the whole scene. Because of that, I called Sze at, like, 3 a.m.: hey, I found this, it's so interesting, we can finally solve robotic manipulation. To me, impedance is the most critical piece of the puzzle, one that can serve as a physics prior for the robot to learn.
But then there's a fundamental missing part. When we both started the company, we tried to train the model ourselves on forces and stiffness and damping, the impedance parameters and all this, and the model didn't really learn well. So there's a fundamental missing piece. This is another research question I've been thinking about for quite a long time: robot data has been flattened all the time.
One very simple example is poses. When a pose turns from zero degrees to 180 degrees, then suddenly, in terms of Euler angles, it becomes zero degrees again. This kind of phenomenon in robotics is what I call gimbal lock. When we treat this kind of data in a flattened manner, so that we can use it to train a neural net, there's a huge change in the numbers, but in the physical interaction there isn't. There's an information gap there, and it tells us that robot data actually lives in curved space. We should really treat robot data as a curve. It shouldn't live in just a single-dimensional vector. Every piece of robot data should be projected onto its right curve and right manifold, so the data can be learned much more efficiently.
It's also reflected in something we see quite commonly in VLAs nowadays. A lot of VLAs need a ton of pretraining data and a lot of fine-tuning data to teach the robot even one simple pick-and-place task. The majority of them treat the data as single-dimensional, whether it's robot state or robot action. They flatten it, and then they train on it.
So, with all that being said, what we believe is that we need the right physics prior for the robot to understand physical interaction. And underneath, we need to project every piece of robot data onto the right curve, so the model can learn way more efficiently than any other model in the world, and the robot can really understand the actual interaction behind every manipulation.
Bogdan Cristei (16:33)
So maybe let's talk through one particular task, phase by phase. Let's take an Ethernet plug. There's the alignment, the clip compressing, there's a snap, a seating. What's the model doing with stiffness at each moment? And what happens to a position-control policy at, let's say, the snap moment?
Elijah Ong (16:53)
When you pick up the cable from a table and try to move it closer to the port, imagine that motion as moving in free space. The robot is moving in free space. It's compliant, it's soft. But when it gets closer to the port, that's when you need to really maintain the stiffness. That's when you need stiff control guiding the robot toward the port. At the same time, you need to regulate the forces and the damping well, because you don't want it to be too fast. You want it to be a little more reactive when it touches the edges of the port. And when it's sliding in, you need to be a little more precise. You need to maintain your position along that axis, along that trajectory.
So this is where stiff control becomes really important. By learning through impedance, we have this kind of compliance schedule in our model. This is what we call a visual impedance map: at this moment, at this point in time, what kind of stiffness and damping you need to exert in each direction. It's also directional. Stiffness and damping are definitely directional. So you exert that so your motion can be guided really well, to finish the insertion successfully.
Bogdan Cristei (18:14)
And when we were chatting earlier, you described a minimum of a thousand hours to get a task to work well. Talk a little bit about how that compares - your approach versus the big VLA approach.
Elijah Ong (18:29)
With the thousand hours of pretraining, we want to show readers and the audience that a thousand hours is the minimum we can achieve to enable a model to really understand physical interaction. But it doesn't mean we stop there. We're trying to show that a thousand hours gives a robot real physical understanding. We can teach robots to do some physical tasks, and beyond that, contact tasks. In our first paper, IWM 1.0, we show tasks with slowly increasing contact complexity. And we show that with these thousand hours, our model can learn how to do simple assembly, plug insertion, Ethernet cable insertion - those kinds of contact-rich manipulation tasks.
Whereas the typical VLA needs a lot of video data, even robot data, to really understand, based on what the robot sees at the moment, what the best trajectory is, what the best motion is, so that it ends up being successful in the end. But actually, most of the time a typical VLA is open loop. Based on what the robot sees, it just continuously rolls out the action. It isn't really physics-grounded. It doesn't really have closed-loop control, in a sense.
Sze Yuan Cheong (20:04)
That leads to a really fun benchmark we did. With that thousand hours, we're trying to show that with only a thousand hours of training data, the robot learns the physics. How do we show that? We compared it to other state-of-the-art models. The funny thing is, they're already in the million-hour range. And we gave them some post-training data. For ours, we gave 20 trajectories for each task, and for the comparators, we gave them a hundred. The end result is that our success rate is over 90%, and the two comparators we compared against are at around 30% and 40%.
Bogdan Cristei (21:00)
Gotcha. So let's say a manufacturer calls you tomorrow with a new insertion task. What do they buy? What happens on their floor? How long until it's running and it's in production?
Sze Yuan Cheong (21:10)
Let's say they call us with a particular insertion task. First, let's say they have a robot, and the robot is in our ecosystem. They'll use our model and show the robot what to do - basically, a demonstration from their operators. The model identifies what the objective is and what the object is, and from there, there are two paths.
The first path is that the robot plays with the object, remembers the objective, and learns about the interaction. It learns for about an hour, basically trying different ways to do that task. It fails a bunch of times, and then it succeeds. And most importantly, it collects that data for post-training, to come up with a policy that actually works, because we're using the actual interaction data to ground the policy in the physics of that particular task.
The other path is for when they can't do that - they don't have a robot, and they want to train a policy and see it succeed before they deploy. Then we can use other training methods, such as UMI grippers. With about 20 or 30 minutes of collected data, we can post-train the model and come up with a successful policy.
Bogdan Cristei (23:04)
Gotcha. All right, then let's talk a little bit about the bitter lesson. So over to Elijah. The bitter lesson says that scale beats hand-built structure every time. The VLA labs have billions of dollars. If they add force data and keep scaling, do you think they arrive at the same place?
Elijah Ong (23:24)
Well, I'm not saying scaling is not good. It's just that we need to scale in the right paradigm. I think current VLAs and world models are scaling in the wrong direction. What I mean by that is that we shouldn't just look at it as an AI problem. We should really look at it as a robotics control problem. What does a robot actually need? What kind of information, what kind of training data does a robot actually need to really understand the physical world? That's the belief we've held from the beginning, when we started this company.
A physical prior is the first thing. Without any physical prior, a robot won't be able to really understand physical interaction. And second, we shouldn't treat robot data as just numbers. We should really let the data live in curved space, so the robot will really understand what's going on inside the data, inside the interaction, and everything. Then it can learn any interaction behind it way more efficiently. With this kind of paradigm, the scaling law actually works. The scaling law will actually be way more efficient.
Sze Yuan Cheong (24:39)
So I would summarize it as two things we believe to be true. First is the foundation, the model architecture itself. We don't think the current paradigm is the ultimate paradigm. We have not found a way to actually embody all the data and train an efficient model. So, the model architecture.
Second is a lesson from classical robotics. In classical robotics, we always try to abstract. We find the correct abstraction, and then we work with the abstraction. This particular lesson has been forgotten. What we're trying to do is push the boundary on both the model architecture and the abstraction. The stiffness, the impedance and everything would be the abstraction. And then comes a model architecture that can embody this abstraction framework, to train and actually execute the model in the real world.
Bogdan Cristei (25:45)
Gotcha. And correct me if I'm wrong, but if I think of the installed base of robot arms, most of them are position controlled. Your approach requires commanded stiffness. I wonder how much of the world's existing robots you can run on, and what you lose if an arm can't do impedance control.
Elijah Ong (26:06)
Our model can definitely accommodate all kinds of embodiments, regardless of whether it's a force-controlled robot, a position-controlled robot, or even a velocity-controlled robot. In terms of impedance, there are two kinds. The first kind is explicit impedance, which you can collect directly from a force-controlled robot itself, or even from a position-controlled or industrial robot with a force-torque sensor at its end effector. The second kind is what we call derived impedance.
Imagine a trajectory, or multiple trajectories. There's a part of the trajectory where the motion is compliant, and a part where the motion is stiff. Back to the example of Ethernet insertion: when you pick up the cable and move it in front of the port, the compliance schedule is soft. But when you try to insert it, the compliance schedule is stiff. In our model, we're able to extract these kinds of impedance parameters, and they serve as a supervising signal in the model for it to really understand the interaction behind it.
So even today, even if we're using industrial robots or cobots, our model can be deployed, and the robot can use the learned signal, the derived stiffness, to really understand what kind of acceleration and deceleration it needs, in such a way that it ends up in a successful motion.
Bogdan Cristei (27:38)
Gotcha. And where does the model break today? You have deformable objects, you have long-horizon assembly, you have cable routing. Where would it break?
Elijah Ong (27:48)
I would say a very, very extreme example would be threading a needle.
Bogdan Cristei (27:54)
Wow. A thread through a needle. Okay.
Elijah Ong (27:56)
Yeah. Because the thread is very soft, most of the time you rely on position control and the image itself, so this kind of impedance data becomes less critical at that point. But if the sensor is good enough for us to collect and learn the forces that are critical to the motion, then yes, our model can do it. Other than that, for most tasks, we're able to learn all of these interactions.
Bogdan Cristei (28:45)
In closing, I do want to ask maybe two or three questions. These are designed to be more - I don't want to say hot takes, but I'd just love to hear your thoughts. Feel free to answer, both of you, or each of you individually can answer whichever question. So the first one: what is the most technically wrong thing robot learning teams are still doing in 2026?
Sze Yuan Cheong (29:11)
Really hot takes. Do you want to go for it?
Elijah Ong (29:13)
We shouldn't scale. We really shouldn't scale from pixels anymore. We shouldn't just scale from videos anymore. We should really look underneath those videos, that egocentric data. How do we derive the right physical interaction? How do we derive the correct physical representation for the robot to really understand? We shouldn't just do video in, video out, and then predict the action. I truly think that's not the right way to go in 2026.
Sze Yuan Cheong (29:48)
He said one part. I'd say the other part is the model architecture, which I touched on just now. We're actually releasing a paper at the end of this month. It will be submitted to ICLR, and the paper is about our architecture, called Devol-ONE. It's one mixture of transformers that blends in basically everything. Currently there's a divide between VLAs and world action models, and then there's the JEPA route. What we're doing in this particular paper is blending everything together.
Imagine a world action model. There are two distinct routes within the architecture. One part figures out the latent space, and the other part is the action decoder. The fundamental question we asked ourselves is, why couldn't this be sequential? Why couldn't you combine both together? And why couldn't you actually use JEPA? So what we did is take the output of the latent and feed it into the action decoder. We use that to basically have a latent-conditioned action model, and that action model is, in a way, a world model itself.
With this particular architecture, which basically blends VLA, JEPA, and the idea of a world action model - without putting in our framework of geometry and impedance - we tested it on various benchmarks, and it basically ranked number one on every single benchmark. So it gave us a lot of deep thoughts, because what everyone is assuming to be right at the moment might be wrong. We might be spending the money on data and scaling on the wrong side. There might be a lot of fundamental breakthroughs that can still happen, and if we find them, we might be able to find a significantly more efficient way to scale.
Bogdan Cristei (32:26)
I see. So that leads me to my next question, and I have a feeling I know what you're going to say. Five years from now, what will be obvious that sounds non-obvious today?
Sze Yuan Cheong (32:34)
Wow, your questions. Okay, I'll answer first. I think there's a huge divide between classical robotics - what's been learned from decades of development in classical robotics - and where modern AI is right now. And I think one of the very obvious things, which is slowly being rediscovered right now, is that people from the AI world are slowly rediscovering knowledge that was already gained in classical robotics. Which is a bit weird if you think about it. You want to continue?
Elijah Ong (33:17)
Yeah. I've realized that researchers, physical AI people, have started to realize that low-level control actually matters quite a lot. There's some research trying to do a sort of slow and fast brain. The VLM serves as a slow brain, to reason about what's going on in general situations, and the fast brain tries to have closed-loop control: based on the robot's feedback and the image, what kind of action it needs to take. That kind of research is going on. People are starting to appreciate that maybe we really need to think about how to make robots more precise again, instead of just letting the action roll out and letting it be.
That's why VLAs and world action models are still unable to do precise tasks, force-relevant tasks, right now: they ignore this kind of low-level, fundamental control in robotics.
Sze Yuan Cheong (34:23)
And basically, what we're currently solving - what the entire field is currently trying to solve - is merely the planning end. The execution end, on the robotics side, is not solved. And dare I say, no one is actually trying to solve it, or seeing it as an actual problem. Which it is, actually.
Elijah Ong (34:47)
A very typical example would be the action head itself. Everyone is using diffusion, flow matching, and even autoregressive. It's purely open loop. It doesn't have a closed loop at all, and it disregards robot control. They just say, okay, I ask the robot to go there, and it just goes.
Bogdan Cristei (35:10)
Very good. Okay, maybe the last question. Do you have one piece of advice for a robotics founder starting today?
Sze Yuan Cheong (35:16)
It comes from our experience, our own bitter lesson. When we started, we had a very ignorant idea of basically building everything in house. We had this idea, this representation that we wanted to build. We work in the AI space, but as a classical roboticist, you tend to romanticize things and think, okay, if I were to build the best model, I need to build the hardware. Build the entire hardware stack, build our own actuators, maybe build our own motors, definitely build our own drivers and our own custom firmware, so we know how to control everything and optimize everything.
As a startup, you can't really do that. We learned that the hard way. It's good that we got a lot of lessons out of it, but we almost died.
So don't try to do everything yourself. Try to get into the ecosystem, work with your peers, and work with other companies, because robotics is such a huge problem to solve. Find the thing you're good at and solve that. Don't try to solve everything.
Bogdan Cristei (36:45)
Wisdom from pain and suffering. Very good. So with that, where can people learn more about Devol, and who do you want to hear from? Who do you want to reach out to you?
Sze Yuan Cheong (37:00)
Our website is devolrobots.ai. That's D-E-V-O-L, robots with an S, dot A-I. We'd like to work with researchers in this area, roboticists who are interested in control, and obviously customers who have problems they want solved.
And I think the last thing is: anyone who's interested in robotics, we'd actually like to hear from you, even if you have no experience in robotics at all. We've always believed in building a team with variety, bringing in people with very different backgrounds and experience, and mainly trying to find outliers, so that as a team we can think out of the box and build a solution that's fundamentally impactful in this industry.
Bogdan Cristei (38:06)
Well said. Diversity. All right, awesome. Really good to have you guys both on. Thank you so much for your time, and excited to keep the conversation going.
Sze Yuan Cheong & Elijah Ong (38:15)
Thank you, Bogdan. Thank you so much. Always good to talk to you.