SmolVLA

SmolVLA는 HuggingFace에서 개발한 경량화된 Vision-Language-Action (VLA) 모델입니다. 자연어 명령을 이해하고 로봇의 동작을 제어할 수 있는 강력한 기반 모델입니다.

SmolVLA의 특징
  • 경량화: 450M 파라미터로 일반 GPU에서도 학습 가능
  • 다중 입력: 카메라 뷰, 로봇 상태, 자연어 명령 동시 처리
  • 빠른 파인튜닝: 적은 데이터로도 새로운 작업 학습 가능

개요

SmolVLA는 로봇공학을 위해 특별히 설계된 기반 모델로, 다음 세 가지 입력을 통합하여 처리합니다:

  1. 멀티뷰 카메라 입력: 여러 각도의 시각 정보
  2. 로봇 상태 정보: 현재 센서모터 상태
  3. 자연어 명령: 수행할 작업에 대한 텍스트 지시

모델 아키텍처

SmolVLA는 다음과 같은 입력을 통합하여 처리합니다:

  • 비전 입력: 멀티뷰 카메라로부터의 이미지
  • 로봇 상태: 현재 관절 각도 및 센서 정보
  • 언어 명령: 자연어로 된 작업 지시
  • 출력: 로봇의 다음 동작 시퀀스

사용 방법

1. 설치

# SmolVLA 의존성 설치
uv pip install -e ".[smolvla]"

2. 사전 학습 모델 불러오기

from lerobot.policies.smolvla import SmolVLAPolicy

# 기본 모델 불러오기
policy = SmolVLAPolicy.from_pretrained("lerobot/smolvla_base")

3. 데이터셋 준비

데이터 수집 권장사항
  • 최소 50개 에피소드 수집
  • 다양한 물체 위치와 상황 포함
  • 명확한 자연어 명령 레이블링
# 자연어 명령과 함께 데이터 수집
python -m lerobot.record \
    --robot.type=so101 \
    --dataset.repo_id=${HF_USER}/my_vla_dataset \
    --dataset.num_episodes=50 \
    --dataset.single_task="Pick up the red block and place it in the basket"

4. 파인튜닝

# SmolVLA 파인튜닝
python -m lerobot.scripts.train \
    --dataset.repo_id=${HF_USER}/my_vla_dataset \
    --policy.type=smolvla \
    --policy.pretrained_model="lerobot/smolvla_base" \
    --output_dir=outputs/smolvla_finetuned \
    --training.num_steps=20000 \
    --training.batch_size=16
# 고급 설정으로 학습
python -m lerobot.scripts.train \
    --dataset.repo_id=${HF_USER}/my_vla_dataset \
    --policy.type=smolvla \
    --policy.pretrained_model="lerobot/smolvla_base" \
    --policy.freeze_vision_encoder=false \
    --policy.learning_rate=5e-5 \
    --policy.weight_decay=0.01 \
    --training.num_steps=30000 \
    --training.batch_size=32 \
    --training.gradient_accumulation_steps=4 \
    --wandb.enable=true

5. 추론 및 실행

# 학습된 모델로 로봇 제어
python -m lerobot.record \
    --robot.type=so101_follower \
    --robot.port=/dev/ttyUSB0 \
    --policy.path=${HF_USER}/smolvla_finetuned \
    --policy.instruction="Pick up the blue cube"

성능 최적화

학습 팁

학습 시간 단축
  • A100 GPU: ~4시간 (20,000 steps)
  • RTX 3090: ~8시간
  • 배치 크기: GPU 메모리에 따라 조정

데이터 증강

# 다양한 자연어 표현 사용
instructions = [
    "Pick up the red block",
    "Grab the red cube", 
    "Get the red object and put it in the basket",
    "Move the red block to the container"
]

ACT와의 비교

특징SmolVLAACT
자연어 명령✅ 지원❌ 미지원
파라미터 수450M~100M
학습 시간더 김빠름
일반화 능력높음보통
메모리 사용량높음낮음

실제 활용 예시

복잡한 명령 처리

# 조건부 작업
"If there is a red block, pick it up, otherwise pick the blue one"

# 순차적 작업
"First pick up the cup, then place it on the shelf"

# 상대적 위치
"Move the object to the left of the basket"
SmolVLA 활용 시나리오
  • 다양한 물체 조작: 자연어로 유연한 작업 지시
  • 멀티태스크 로봇: 하나의 모델로 여러 작업 수행
  • 인간-로봇 협업: 직관적인 음성/텍스트 명령

문제 해결

일반적인 이슈

메모리 부족

해결책:

  • 배치 크기 감소
  • Gradient accumulation 사용
  • Mixed precision training 활성화
낮은 명령 이해도

해결책:

  • 더 다양한 자연어 표현으로 데이터 증강
  • Instruction template 일관성 유지
  • 파인튜닝 스텝 증가

추가 리소스


SmolVLA는 지속적으로 개선되고 있습니다. 최신 업데이트는 HuggingFace LeRobot 저장소를 확인하세요.