SmolVLA is a compact, efficient vision-language-action model that achieves competitive performance at reduced computational costs and can be deployed on consumer-grade hardware.
Model source
Source excerpt
SmolVLA is a compact, efficient vision-language-action model that achieves competitive performance at reduced computational costs and can be deployed on consumer-grade hardware.
Sources
1 sourceVerified Aug 6
Model artifacts
1 artifactSource excerpts
3 excerpts--- base_model: lerobot/smolvla_base datasets: vpraise00/raise-medicine-insert-pill-sft5-statefix-vlmclean-v004only-100ep-14d-trainmix library_name: lerobot license: apache-2.0 model_name: smolvla pipeline_tag: robotics tags: - smolvla - lerobot - robotics --- # Model Card for smolvla <!-- Provide a quick summary of what the model is/does. --> [SmolVLA](https://huggingface.co/papers/2506.01844) is a compact, efficient vision-language-action model that achieves competitive performance at reduced computational costs and can be deployed on consumer-grade hardware. <p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/640e21ef3c82bd463ee5a76d/aooU0a3DMtYmy_1IWMaIM.png" alt="smolvla architecture" width="85%"/> </p> <!-- A short demo is worth more than any description! Record a GIF/video of the policy running on your robot, upload it to this repo, and embed it here: <p align="center"> <img src="https://huggingface.co/<hf_user>/<policy_repo_id>/resolve/main/demo.gif" width="60%"/> </p> --> This policy has been trained and pushed to the Hub using [LeRobot](https://github.com/huggingface/lerobot). Learn how to train and run it in the [LeRobot smolvla guide](https://huggingface.co/docs/lerobot/main/en/smolvla), or browse the [full documentation](https://huggingface.co/docs/lerobot/index). --- ## Model Details - **License:** apache-2.0 - **Fine-tuned from:** [lerobot/smolvla_base](https://huggingface.co/lerobot/smolvla_base) - **Robot type:** `so101_follower` - **Cameras:** `top`, `left_wrist`, `right_wrist` ## Inputs & Outputs The policy consumes these observation features and produces these action features. **Inputs** | Feature | Type | Shape | | --- | --- | --- | | `observation.state` | STATE | `(14,)` | | `observation.images.camera1` | VISUAL | `(3, 480, 640)` | | `observation.images.camera2` | VISUAL | `(3, 480, 640)` | | `observation.images.camera3` | VISUAL | `(3, 480, 640)` | **Outputs** | Feature | Type | Shape | | --- | --- | --- | | `action` | ACTION | `(14,)` | ## Training Dataset - **Repository:** [vpraise00/raise-medicine-insert-pill-sft5-statefix-vlmclean-v004only-100ep-14d-trainmix](https://huggingface.co/datasets/vpraise00/raise-medicine-insert-pill-sft5-statefix-vlmclean-v004only-100ep-14d-trainmix) - **Episodes:** 100 - **Frames:** 67328 - **Frame rate:** 30 FPS - **Task(s):** "Pick up the pill and...
Source context: 6 downloads · 0 likes · Pipeline robotics · Library lerobot · Repo vpraise00/smolvla-raise-medicine-insert-pill-sft5-statefix-vlmclean-v004only-100ep-14d
Source context: 45 downloads · 0 likes · Pipeline robotics · Library lerobot · Repo vpraise00/smolvla-raise-medicine-insert-pill-sft5-statefix-vlmclean-v004only-100ep-14d