SPACE 002 / LEARNED POLICY
Go1
locomotion.
A joystick-conditioned quadruped policy trained with smoke-gated PPO, evaluated locally, and opened in SimRig’s browser preview.
episode—
step—
reward—
playback—
TaskFollow planar velocity and yaw commands over flat terrain.
Training8,192 parallel environments with local and remote smoke gates.
ArtifactDownloaded policy, metrics, configuration, logs, and checkpoints.