Reward AI Releases OM-1: A Robot Policy Trained on Human Demonstrations Only, With No Teleoperation or On-Robot Data
MarkTechPost Asif Razzaq
Reward AI says its new robot policy learns from people in a glove, not from robot demos. That’s a big bet: one model for any body, but no weights or code are out yet.
Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Reward AI has put out OM-1, short for Omnibody Model 1, and the headline is what it leaves out. No teleoperation data. No on-robot data. The startup says the policy learns only from human demonstrations captured with a sensorized glove, then runs on industrial arms and humanoids at human speed.
That puts OM-1 in a different camp from most robot foundation work, which is usually chained to the body it was collected on. Reward AI’s pitch is that manipulation won’t come from piling up more of that same data, or from throwing more compute at the problem. Instead, it wants capture, learning, and control to behave like one pipeline, so a demo recorded now can train a robot body that does not exist yet.
The hardware side starts with Omnibody Hand, a wearable built on the team’s earlier DexCap work. It is a 7-DoF design meant to capture the motions that actually matter in manipulation: contact choice, object reorientation, and moving between precision and power grasps. The glove tracks thumb-index pinching, thumb and index flexion, and coupled motion in the middle, ring, and little fingers. Reward AI also says ergonomics matters because a slipping or restrictive device changes how people grasp, so the hand includes a distal flexion mechanism that avoids per-user adjustment.
The data interface is aimed at fast actions like conveyor-belt sorting, where a person spots, grabs, and tosses an object in a fraction of a second. To keep that whole interaction, the glove combines tactile sensing, proximity sensing, and global-shutter in-hand cameras. For hand pose tracking, Reward AI reports its first quantitative result: electromagnetic sensing plus disturbance compensation reduced mean overshoot error by 60% at the top speed tested, with 9.5 mm versus 24.9 mm for the visual-inertial approach at 67 cm/s. The tests moved both trackers between two mechanical stops at eight speeds from 3 to 67 cm/s, averaged over ten runs each.
OM-1 then learns directly from those multimodal streams. The company says there is no separate pre-training and post-training split; the first demonstration ever recorded and the newest one train a single policy in a single stage. Inputs stay at their native sampling rates, so high-frequency tactile and motion cues do not get flattened into the same clock as lower-frequency vision. Outputs include motion direction, speed, force, and timing for events like grasp initiation. Below that sits a reinforcement-learning control layer on its own clock, trained in simulation to handle disturbances, delays, and dynamics that change with velocity and acceleration.
Reward AI says the stack can pick up a new task, including hard dynamics and long horizons, from less than 30 minutes of human data. The company’s about page says published clips run at 1x speed and that the model spans arms, legged humanoids, and wheeled mobile manipulators. But there’s no paper, no public baseline comparisons, no success rates, and no release of weights, code, dataset, or API. For now, OM-1 is a promise inside Reward AI, not something other developers can plug into their own hardware.
My take — AI-written commentary, not fact-checked reporting
This is the right kind of stubborn. Robot builders have spent years treating robot demos like sacred data, as if a machine must first imitate a machine before it can do anything useful. Reward AI is at least trying the messier, more interesting route: human skill first, embodiment later. The only catch is the usual one in robotics—if nobody outside the lab can touch it, the grand theory stays very convenient and very private.
Read more about this at: MarkTechPost
Related stories
Generalist AI Releases GEN-1.5: A Robot Foundation Model That Learns New Tasks From One 3–12 Second Demo
MarkTechPost · 3 weeks ago ·
8
MolmoAct 2: An open foundation for robots that work in the real world
Allen Institute (AI2) · 4 months ago ·
26
Dyna Robotics Introduces Dyna-2: A World-Action Model Pre-Trained on 1 Million Hours of Human Video
MarkTechPost · 1 month ago ·
9