Video to motion

Act it out. Get it on your rig.

Record the motion on your phone and get an R15 animation from it. Faster than describing something specific, and far faster than keying it by hand — particularly for timing you can feel but can't write down.

mp4 or mov · up to 10 seconds · up to 100 MB · one person in frame

This browser could not start WebGL, so the 3D preview is unavailable. The animation itself is unaffected — download the .rbxm or apply it from the Studio plugin.
Dance·30 fps·72 frames
0.00s · f0drag to orbit · scroll to zoom
L
R

Before you upload

What makes a clip work

These checks run on your side before the upload starts and again on the server before any GPU time is spent, so a bad clip costs you nothing but the time to re-shoot it.

Shoot it like this

  • One person, fully in frame, head to feet
  • Steady camera — a tripod or a propped-up phone
  • Even lighting, clear separation from the background
  • The whole action inside ten seconds

These get rejected

  • More than one person visible
  • Limbs cut off by the frame edge
  • Heavy motion blur from a dark room
  • The subject small or far from the camera

Rejections name the actual reason — “two people detected in frame”, not “processing failed”. You should never have to guess what to change.

The path

Where your video goes

01

Checked in your browser

Duration, size and container format are validated before a byte is uploaded. If the clip is twelve seconds long you find out immediately, not after a two-minute upload.

02

Straight to storage

The file uploads directly to object storage over a presigned URL. It never passes through our application servers, which is why a 100 MB clip doesn't tie anything up.

03

Framing checked server-side

Resolution, duration and the number of visible people are probed before the GPU job starts. Rejections here cost you nothing.

04

Deleted on schedule

The inference job reads the file once. Source videos are then deleted on a fixed retention schedule, and you can opt to have them removed immediately after processing.

Timing you can't write down

A hesitation before a swing, a weight shift before a step. These are the things prompts are bad at and a five-second phone clip is perfect for.

Your own performance

Character comes from how a specific person moves. Recording yourself is the shortest path to an NPC that doesn't move like every other NPC.

Nothing kept by default

Uploads exist to be read once by the inference job. Retention is stated plainly in the privacy policy and configurable in settings.

Questions

Video specifics

Do I need a special camera or a mocap suit?+
No. A phone on a tripod in a reasonably lit room is what this is designed for. No markers, no suit, no depth sensor.
Can two people be in the shot?+
No. Multiple people in frame is the most common rejection reason. Clear the background of anyone who might walk through.
What if part of me is out of frame?+
Limbs cut off by the frame edge produce guesswork where the joint should be, so those clips are rejected rather than silently guessed at. Step back and get head to feet in shot.
How long can the clip be?+
Ten seconds. Longer performances work better generated as separate clips and stitched, which also keeps each re-run cheap.
Is my video used to train anything?+
No. Uploads are read once by the inference job for your generation and then deleted on the retention schedule.