A vision-language model trained only in simulation runs a real robot arm

I fine-tuned Qwen3.8-27B to make Jev-style decisions from camera images, using nothing but simulated data. On a real SO-101 it completed a two-ball pick-and-place command in 104 seconds, with no teleoperation and no real-world demonstrations. Here is how it works, how it was trained, and what broke on the way.

Daniiar AbdievSeptember 202615 min read
Figure 1. One typed command, real time, recorded from the top, side and gripper cameras. The recorder missed about 11 seconds of the first approach; the clock in the frame shows the true elapsed time.
104 sfor two balls into the container, from one command
0teleoperated or real-world demonstrations in training
~0.2 sof GPU time for each skill decision
81%success on 544 held-out simulated episodes

Decide, don't drive

Most robot foundation models today are vision-language-action models, or VLAs. They output motor commands directly and learn them from demonstrations: a person teleoperates the target robot for hundreds of episodes and the model imitates. That works, but the data is slow to collect and tied to one robot and one set of tasks.

My earlier system for this arm, Jev, took another route. The model decides and ordinary code acts. The model makes typed choices, such as which step comes next or whether the task is done, and code turns each choice into a bounded motion. Jev saw the scene only as text: code measured the objects, and the model read numbers and labels.

This project gives that pattern eyes. The model looks at the cameras itself.

Top cameraSide cameraGripper camera
Three cameras and a commandplus the last step tried and the gripper reading
Qwen3.8-27B + LoRAfine-tuned only in simulation; served with vLLM on one GPU
What nextone of nine skills, read as a single token
Wherea pixel in the top image, mapped to the table by a calibrated camera
Checkyes or no before releasing and before finishing
Adapterinverse kinematics, calibrated heights, speed limits, a stop key
SO-101 arma low-cost arm with six hobby-grade servos
Figure 2. The model decides and code acts. Every step starts again from fresh camera images, so a failed grasp or a stalled move shows up in the next decision.

At every step the agent receives three 448 × 448 images, the command, the last step with its outcome, and the gripper reading. It answers with one letter, one per skill:

Amove_to_objectBgraspCliftDmove_to_placeErotateFreleaseGopenHdoneIgive_up

The letter is read as a single token at temperature 0, so a decision costs one forward pass. When a skill needs a position, the model points at a pixel in the top image, and a calibrated camera model maps that pixel to a spot on the table. Short yes/no questions guard the release and the finish: is the held object over the container? and is the blue ball inside?

Why a VLM? It already knows what a ball and a container are, which one is blue, and where things sit in a picture. Fine-tuning only has to teach the decisions. The skills are plain code, so the model never has to learn servo control, which is where teleoperated data usually goes. After training only on simulated images, the model still answers “is the object inside the container?” correctly on 92 of 100 real photos of SO-100 and SO-101 arms.

Inside the 104 seconds

Here is every decision from the run in Figure 1: what the model saw, where it pointed, and how much probability it gave each of the nine skills. Use the arrow keys or the step buttons.

Top camera at this step
Top camera, with the point the model chose
Side camera at this step
Side
Gripper camera at this step
Gripper
Step 1 of 15Blue ball
move_to_object

Points at the blue ball in the top image. The move stops a little short.

Skill probabilities
Yes/no checks
Model calls this step
Figure 3. The images are the ones the model received at each step. Probabilities are the model's scores over the nine answer letters, normalised. Latencies are round trips from the laptop to the GPU server. Blue and yellow marks under the steps show which ball each sub-task is about.

Two things stand out. The model is confident about almost every skill choice, usually above 90%. The yes/no checks are where it is too careful. Over the container it answered “not yet” twice for the blue ball and three times for the yellow one. For the yellow ball, a repeat guard in code then released it anyway, and it landed inside. That caution is the same bias the training labels were meant to remove (next section), and it is not fully gone.

Training in simulation only

The training data is simulated SO-101 episodes in MuJoCo, rendered photoreal with randomised tables, lighting, arm colours, objects and containers. Nothing was recorded on the real arm, and nobody teleoperated anything.

Training: simulated and rendered
Deployment: the real cameras
Simulated top view: yellow arm on a wooden table with a container Simulated top view: white arm on a green table with a cube Simulated top view of a take-out scene Simulated gripper camera over a container Real top camera: the arm, two balls and the container Real gripper camera over the blue ball
Figure 4. Left: images from the simulator, with randomised surfaces, lighting and arm colours. Right: what the real top and gripper cameras showed on the day. The model saw no real-world data in training.

Labels from simulated outcomes

The core of the method is how each training example gets its answer. Nobody demonstrates the right step. The simulator tries them all and scores what happens.

  1. Collect states from the model itself. Run the current model in simulation and save a snapshot at every step. This is DAgger: the states include the model's own mistakes, so it learns to recover from them.
  2. Try every next step. At each state, restore the snapshot and execute each candidate skill, with pointing noise, twice. A scripted policy that can read the true simulator state then finishes the task, for up to 25 steps.
  3. Score and pick. Each branch scores success minus 0.02 per step. The best branch becomes the label, and the step cost makes the model prefer the shorter way to success.
  4. Fine-tune. Train a small LoRA adapter on these examples, then repeat from step 1 with the new model.
A simulated state: the arm holding a block over the container
A simulated state: the arm holds the block over the container.
releaseinside, done 2 steps later0.96
move_to_placeinside, done 3 steps later0.94
rotateinside, done 4 steps later0.92
openobject falls outside−0.50
give_upnot allowed: still reachable–
score = success − 0.02 × steps  ·  label = the best-scoring step
Figure 5. How one training label is made. The branch outcomes shown are a worked example of the rule; the real pipeline runs every candidate skill from each of 2,616 saved states. A geometric rule would have said “not over the container yet, move again” here. The simulated outcome says release.

This is one step of policy improvement by Monte Carlo rollouts, distilled into the model with supervised fine-tuning, in the family of expert iteration. It sits closer to reinforcement learning than to imitation: the reward comes from the simulator, not from a human showing the way.

Figure 6. Same states, better labels. Success on the same 544 solvable test episodes (600 photoreal episodes, eight tasks, seven fault types), with 95% intervals. Rule labels and outcome labels were trained on exactly the same 2,616 states; the outcome labels added 22 successes (paired McNemar p = 0.033). The last row is the released agent with the serving, pointing and sensing changes used on the real arm.

The gain came from the labels, not from more data. The earlier labels for “is the object over the destination?” came from a geometric rule, and the rule was too cautious: in replays, releasing where it said “not yet” often worked. Relabelling by simulated outcome fixed that and cut the step-limit loops from 48 to 22. I also tried a KL-regularised RL objective on a further round of data. It tied with this adapter (434 against 440), so the released model is the outcome-labelled one.

Fine-tuning itself matters too. On the same test episodes, the untrained base model succeeded in 54% of clean episodes on the training tasks and the first adapter in 72% (paired, p = 3 × 10−12). The released agent is at 92% on those tasks.

RoundTraining dataLabelsCompute
v15,287 examples from 2,760 states of a noisy scripted walk, with mistakes and faultsscriptedA100, 93 min
v27,597 examples: v1 states, near-miss “twins” and the model's own DAgger statesrule and simulatorA100, 4.6 h
v45,347 examples from 2,616 DAgger states (200 rollouts)simulated outcomesH100, 70 min
Table 1. The released adapter's lineage. Each round continues from the previous adapter for one epoch. LoRA rank 16, alpha 32, 79.7M trainable parameters on Qwen3.8-27B. About seven GPU-hours of fine-tuning in total.

How well it works in simulation

The test set is 600 photoreal episodes on held-out seeds: eight task types, each clean and with one of seven injected faults, such as a missed grasp, a slip, an object moved mid-task or a knocked-out object. Three of the task types never appeared in training in any form.

Table 2. The released agent on 600 held-out simulated episodes. Recovery counts the episodes where an injected fault actually fired. Out-of-reach episodes are scored separately: the agent correctly gave up on 54 of 56.

On the real arm

The real setup is an SO-101 on a home desk, three consumer USB cameras, a laptop running the adapter, and the model served on one H100 in the cloud with vLLM in FP8. The model is the simulation-trained one, unchanged. What is new is the adapter, which implements the simulator's nine skills on the real arm.

CommandResultTime
Put one ball in the container6 of 9 valid runs61–164 s
Put both balls in, one command1 of 4 full successes; one more reached the goal, then gave up104 s
Take a ball out of the container0 of 10–
Put a ball in a tall cup0 of 1 (out of reach)–
Table 3. All 28 recorded runs over two days, judged from the recordings. Four more runs were invalid: three wrist-camera drops and one where the container was moved mid-run.

Every run, in order

successgoal reached, then gave upfailureinvalid

Why the 16 failures happened

Figure 7. Left: each dot is one run, in the order they happened; hover for the command and verdict. Right: the cause of each failure, from the frames of all three cameras. The adapter is my motion code; reach means the target was outside what the arm could reach with the gripper pointing down.

When I went through the frames of every failure, almost none came from the model misreading the scene. Where I checked, it pointed at the right ball to within one or two pixels. The failures sat in the layer between the model and the arm: calibration, motion code, workspace layout and hardware. The model also showed something I had hoped for from the DAgger training. When a grasp failed, it noticed and recovered.

Figure 8. The first grasp catches only the edge of the ball. The agent sees the ball still on the table, opens, looks again and picks it up. Real time, 164 seconds.

What broke

Four failure frames: sweeping through a container wall, jaws closing on the wall, a release from 18 cm, a pose at the base joint's limit
Figure 9. From left: travelling at the simulator's carry height through the container wall; the jaws closing on the wall instead of the ball; a stuck carry released 18 cm up; a pose at the base joint's limit.

Calibration

Calibration took more of the two days than anything else and decided more outcomes than the model did. The first calibration had only photographed poses within 40° of straight ahead, with the gripper pointing down. So I swept the whole workspace: 152 poses over every base angle the planner allows, 14 to 40 cm out, at 5, 11 and 18 cm high, each photo taken once the joints had been still for 0.3 seconds. One fit then estimated the camera, the joint offsets, the paint's position on the jaw and the table height together, from 188 photos and 19 table touches.

Where the calibration looked

first calibration (47)full sweep, paint foundpaint not found

Height error when the fingers touch the table

beforeafter, same touchesafter, 6 new touches
Figure 10. Left: the workspace seen from above, forward up, with the base joint's ±55° limits dashed; the sweep covers three heights at each spot. Right: where the arm model puts the fingertip when the real table stops it; a perfect model reads 0 everywhere.

Held-out error is 2.4 mm on the table plane. On data taken after installing it, 20 new photos are predicted within 1.8 px (median; worst 4.7 px), and six new table touches read between −0.8 and +4.7 mm.

The green paint mark on the moving jaw
Six calibration photos with the detected paint circled in green and the old prediction in red
Figure 11. Left: the paint mark, on the moving jaw. Right: six of the 152 sweep photos. Green circle: the paint the camera found. Red cross: where the old calibration expected it.

Speed

The first successful runs took 61 to 164 seconds, and the logs showed the arm spending most of that time waiting. Every decision sent three lossless PNG images, about 271 KB each, through the cloud provider's proxy: 1.8 to 3 seconds per decision from the laptop, for about 0.2 seconds of work on the GPU.

Round trip for one skill decision

Where the 104 seconds went

Figure 12. Left: decision latency measured from the laptop. Right: the two-ball run, split into the 32 model calls, arm motion, and everything else (camera frames, settling, checks and code).

What made it fast:

Two traps. Pointing did not tolerate JPEG: re-encoding one frame moved the pointed spot by 53 pixels and once sent a ball to an empty patch of table, so pointing stays on PNG. And skipping the final settle of each move made runs slower, not faster. That settle is where the arm compensates for gravity at full stretch; without it moves stalled, and one run took 333 seconds.

Figure 13. How long each successful run took. Grey: the six single-ball successes before the speed work. Black: the two-ball command after it. Red: the run where moves skipped their settle.
Figure 14. “Put the yellow ball in the square container,” with the blue ball already inside. Real time, 69 seconds, before the speed work.

What is not solved

Taking balls back out of the container failed in all 10 attempts. Balls pressed against a wall, far corners and the base joint's limit are hard for a small set of discrete skills. In one two-ball run both balls ended inside, but the agent did not recognise the second completion and gave up. The yes/no checks are still too careful. And the adapter changed between runs as I fixed things, so this is a field report, not a benchmark.

What comes next

My read: a VLM making Jev-style decisions over simple skills, trained only in simulation, already shows real promise for crude pick-and-place. It needs no teleoperation rig and no real-world data, and on this arm the model was rarely the part that failed.

Lessons

  1. Look at every frame yourself. Several of my diagnoses were wrong until I looked at all three cameras for each step. The person watching the arm was right more often than my numbers.
  2. Calibrate everywhere you will operate, not a convenient subset.
  3. Validate on fresh data through the code the robot runs. Cross-validation inside the fitting code missed the unit error.
  4. Measure height by touch. One camera sees depth poorly; a few touches fix it.
  5. Hunt for simulator constants. Carry heights and ready poses tuned for simulated scenes will meet taller containers.
  6. Speed is mostly waiting. Cut the image payloads and the settling, but keep the gravity compensation.
  7. Fail safe. Never open the gripper unless the arm has arrived, and park the arm after every run, including failures.

Everything is open

The repository has the real-arm adapter, the calibration tools, the simulator evaluation and the training code. The code and the LoRA adapter are released under Apache-2.0. A paper with the full run ledger is on its way to arXiv.

Cite this post

@misc{abdiev2026so101vlm,
  author       = {Abdiev, Daniiar},
  title        = {A Vision-Language Model Trained Only in Simulation Runs a Real Robot Arm},
  year         = {2026},
  howpublished = {\url{https://robopsychologist.ai/so101-vlm-agent}}
}