Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Dataset Scale and Composition:
- Total Duration: ~2,000 hours of egocentric manipulation video.
- Contributors: 500+ contributors.
- Devices: 400+ consumer smartphone models.
- Content: 400+ scenes and 8,000+ tasks.
- Annotations: Videos are annotated with text descriptions, MANO-based hand poses, camera trajectories, and temporally localized atomic actions.
Comparison with Other Egocentric Datasets:
| Dataset | Hours | Tasks | Hand Pose | Language | Camera Traj. | Capture Devices | Toolchain |
|---|---|---|---|---|---|---|---|
| Ego4D | 3,670 | N/A | ✗ | ✓ | ✗ | 7 cameras | ✗ |
| EPIC-KITCHENS | 100 | N/A | ✗ | ✓ | ✗ | 2 models | ✗ |
| EgoDex | 829 | 194 | ✓ native | ✓ | ✓ | 1 model | ✗ |
| EgoLive | 1,680 | 346 | ✓ | ✓ | ✓ | 1 custom model | ✗ |
| OpenEgo | 1,107 | 290 | ✓ unified | ✓ | Partial | Mixed / N/A | ✗ |
| EgoScale | 20,854 | N/A | ✓ | ✓ | ✗ | Mixed / N/A | ✗ |
| Open-AoE | 2,000 | 8,000+ | ✓ MANO | ✓ | ✓ | 400+ models | ✓ |
Visual Diversity Analysis (vs. OpenEgo, EgoDex, EgoXtreme):
Based on 3,000 balanced random trials sampling 2,000 CLIP embeddings from each dataset. Open-AoE achieved the highest mean score on all metrics.
| Metric | Open-AoE | Description |
|---|---|---|
| Effective Rank | 97.43 | Effective number of feature directions with substantial variance. |
| Participation Ratio | 40.47 | Inverse concentration of the covariance spectrum. |
| Meaningful Coverage | 19.89 | Number of visual clusters with at least 5 samples. |
| Normalized Cluster Entropy | 0.712 | Evenness of sample distribution across a shared codebook. |
| Effective Clusters | 16.29 | Entropy expressed as an equivalent number of uniform clusters. |
| kNN Domain Mixing | 0.0775 | Average proportion of neighbors from other datasets (k=20). |
Semantic and Temporal Annotation Analysis:
- Vocabulary Size: Open-AoE has 32,407 distinct natural-language action descriptions, compared to 26,864 for OpenEgo and 111 for EgoDex. It also includes 175 action verbs, 8,030 object strings, and 135 scene labels.
- Temporal Coverage: Annotations cover 99.99% of the evaluated timeline, compared to 50.1% for OpenEgo.
- Annotation Density: Open-AoE has a density of 13.97 segments per minute, with a mean segment duration of 9.64 seconds.
Annotation Consistency Analysis:
Using Idefics2 to score visual-text alignment on a 1-5 scale, Open-AoE demonstrated significantly higher consistency.
| Dataset | Sequence-Macro Score (out of 5) |
|---|---|
| Open-AoE | 4.583 |
| OpenEgo | 3.029 |
| EgoDex | 2.916 |
| EgoXtreme | 2.015 |
Open-AoE aims to provide practical, open infrastructure for embodied intelligence research by lowering the barriers to both data contribution and reuse. By releasing a large-scale, diverse dataset alongside a full capture-process-reconstruct-train toolchain, the project turns low-cost smartphone video into a valuable resource for training embodied models, studying human-to-robot transfer, and advancing world modeling.
Don't read this site daily. Get it in your inbox.
The daily brief and Sunday deep dive — distilled, scored, and opinionated. For builders only.