RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
Embodied Cognition and Localization Results (Large-Scale Models):
| Benchmark | RynnBrain 1.1 122B-A10B | Qwen3.5 122B-A10B | Gemini 3 Pro | GPT 5.4 | HY-Embodied 0.5 |
|---|---|---|---|---|---|
| VSI-Bench | 75.0 | 66.6* | 48.8* | 49.2* | 68.3 |
| MMSI | 52.0 | 9.2* | 49.2 | 37.5* | 39.2 |
| ERQA | 54.3 | 62.0 | 70.5 | 47.5* | 62.3 |
| RefSpatial-Bench | 79.1 | 69.3 | 65.5 | 26.7* | 57.2 |
* denotes results from authors' reproduction.
Embodied Cognition and Localization Results (9B-Scale Models):
| Benchmark | RynnBrain 1.1 9B | RynnBrain 8B | Molmo2-ER 5B | Qwen3.5 9B |
|---|---|---|---|---|
| VSI-Bench | 74.9 | 70.9 | 74.5 | 61.0 |
| MMSI | 47.0 | 39.6 | 43.8 | 13.4 |
| MindCube | 86.9 | 56.6 | 57.0 | 33.1 |
| RefSpatial-Bench | 67.2 | 59.2 | 52.5 | 37.3 |
Embodied Cognition and Localization Results (2B-Scale Models):
| Benchmark | RynnBrain 1.1 2B | RynnBrain 2B | Cosmos 3-Edge 2B | Qwen3.5 2B |
|---|---|---|---|---|
| VSI-Bench | 72.9 | 70.5 | 59.2 | 44.1 |
| MMSI | 40.5 | 34.1 | 32.3 | 28.3 |
| MindCube | 61.7 | 50.1 | - | 37.5 |
| RefSpatial-Bench | 58.5 | 52.7 | 48.4 | 30.0 |
Scaling Analysis:
- A comparison of RynnBrain 1.1 models (2B, 9B, 122B-A10B) against their corresponding Qwen3.5 baselines revealed three distinct scaling regimes for embodied capabilities.
- General embodied cognition: Both RynnBrain 1.1 and Qwen3.5 improve monotonically with scale, with the performance gap narrowing.
- Reasoning-intensive cognition: RynnBrain 1.1 improves steadily (+38.6%), while Qwen3.5 exhibits negative scaling (-39.2%). The performance gap widens from 18.2 to 50.8 points.
- Embodied localization: RynnBrain 1.1 improves significantly (+24.4%), while Qwen3.5 stagnates (+1.3%). The performance gap widens from 24.3 to 45.4 points.
3D Grounding and Contact Point Prediction:
- The 2B and 9B models were trained with explicit 3D supervision. On the WildDet3D-Val benchmark, the 9B model achieved an AP of 15.3 and mAP of 10.1, outperforming the 2B model's 13.9 AP and 9.1 mAP.
- On the FoundationPose-Val benchmark, the 9B model achieved an AP of 41.9 and mAP of 31.9, compared to the 2B model's 38.6 AP and 28.7 mAP.
- For contact point prediction on a held-out test set, the 9B model achieved a success rate of 84.2%, while the 2B model achieved 78.3%.
Real-Robot VLA Evaluation:
- RynnBrain-VLA policies were deployed on Astribot-S1 and Tianji-Wuji robots.
- Policies initialized from pretrained RynnBrain 1.1 consistently outperformed VLA policies built from the base Qwen models.
- Joint multi-task and multi-embodiment training improved the average process score and final success rate over separately fine-tuned per-task policies.
The authors present RynnBrain 1.1, a family of embodied foundation models that demonstrate a clear upward trend in capabilities with increasing scale. The 122B-A10B model achieves state-of-the-art results on several embodied cognition and localization benchmarks. The introduction of contact point prediction and native 3D grounding improves alignment with robotic manipulation. Finally, the RynnBrain-VLA variant shows strong real-world performance and benefits from joint multi-task, multi-embodiment training, highlighting the value of embodied pretraining for downstream robotic policies.
Don't read this site daily. Get it in your inbox.
The daily brief and Sunday deep dive — distilled, scored, and opinionated. For builders only.