Model weights (Apache 2.0): RekaAI/Reka-Inverse-Dynamics-Model on Hugging Face
Interactive World Models are different. They are visually more realistic, and can model more complex phenomena such as folding materials or squeezing towels. By embedding digital replicas into these dynamic environments, robots can execute commands through direct interaction.
Here is the challenge though. True interactivity requires a world model to accurately respond to low-level commands (e.g., moving left or right). To solve this, we introduce the Reka Inverse Dynamics Model (RIDM), which extracts low-level commands directly from real-world video data to act as ‘captions’ describing physical changes in consecutive frames.
RIDM is a compact, efficient model that processes short clips and extracts low-level motor commands like moving forward, sideways, turning or tilting. Even though the model is entirely trained on video games, by decoupling visuals from motion, its predictions transfer to real-world videos.
Predicted key probability
- Wforward0.85
- Amove left0.34
- Sback0.01
- Dmove right0.76
- Shiftwalk0.01
turn (yaw)16.9° left
tilt (pitch)1.4° up
First window of the clip
The data
We recorded gaming footages at 1280 × 720, 48 frames per second, with 11 ground-truth key commands, the mouse movement and the exact camera angles for every frame. The annotations are provided by the engine.
The model
Our model disentangles visuals from motion by projecting inputs into optical flow space. Since our goal is to use RIDM to automatically annotate real-world footages, not 3d gameplays, the method needs to transfer from games to real-world. We also compare that approach with a direct one that is working in the pixel space. To differentiate both, we use names: flow-based and pixel-based model.
Training both models on random gaming footages will not work either. Successful approach requires data filtering and resampling. For instance, some actions such as backward or sideways strafe movements are much more rare than simple forward movements. Suitable resampling helps in learning these less common movements.
Both models are compact. The flow-based model has 2.8m (2,785,499) parameters; with only 1.8m (1,795,337) trainable parameters and about 990k frozen that constitutes RAFT-small. The pixel-based model is fully trainable and has around 9.8m (9,836,063) parameters.
Click the view · W A S D · Shift · arrows
Pixels · what the camera sees
Optical flow · what the flow model seesdirection
This interactive figure needs WebGL, which this browser has turned off.
The label · what the game records
WASDShift
Turn0°/sTilt0°/s
Walking forward
Evaluation
We tested the generalisation of both approaches on videos that neither method has seen during training. As expected, the flow-based model works better, with the biggest gains obtained when transferring from gaming footages to real-world video clips.
In evaluation, we used three datasets:
- Real-world videos with camera angle labels,
Real-world videos with camera angle labels,
- Real-world GoPro footage,
Real-world GoPro footage,
- Synthetic, 3d gaming footages.
Synthetic, 3d gaming footages.
Note that real-world datasets hold only forward and turn actions.
Each dataset asks two questions; totalling six different configurations. Did the camera go forward or turn? If it turned, which way?
- The flow-based model works better on five out of the six configurations,
The flow-based model works better on five out of the six configurations,
- On real-world video the gap is the largest. The flow-based model gets 86 % while the pixel-based model gets only 31 %,
On real-world video the gap is the largest. The flow-based model gets 86 % while the pixel-based model gets only 31 %,
- The pixel model is only better to predict GoPro turn direction, and only by a single clip.
The pixel model is only better to predict GoPro turn direction, and only by a single clip.
Table 1 — Two models benchmarked on six different evaluation configurations. Since the ground-truth label derived from sensors is inaccurate in real datasets — Real video and GoPro — we manually inspected predictions to calculate accuracies. For gaming footages, we use available ground-truth annotations.
Evaluation data
Task
Flow model
Pixel model
Real video, 78 clips
Forward, turn left or right
85.9 %
30.8 %
Real video, 47 turn clips
Turn left or right
91.5 %
51.1 %
GoPro, 72 clips
Forward, turn left or right
75.0 %
58.3 %
GoPro, 25 turn clips
Turn left or right
72.0 %
76.0 %
Gaming footage, 5,216 flow pairs and 5,164 pixel pairs
Forward, turn left or right
84.5 %
77.5 %
Gaming footage, 3,716 flow turn pairs and 3,674 pixel turn pairs
Turn direction
84.9 %
70.9 %
Our RIDM, of course, occasionally fails in some situations:
- it reports false turns on static frames and on dark frames,
it reports false turns on static frames and on dark frames,
- when it has to disentangle side step and opposite yaw turns. For example when the person is stepping right and turns camera left at the same time,
when it has to disentangle side step and opposite yaw turns. For example when the person is stepping right and turns camera left at the same time,
- it reports walk keys on a camera that stands still and zooms in / out,
it reports walk keys on a camera that stands still and zooms in / out,
- the camera stabilization on action-camera footage hides part of the rotation.
the camera stabilization on action-camera footage hides part of the rotation.
We believe that training on more data, especially on the data containing camera effects like zooming or stabilisation, would improve the model’s performance on the failures above.
Visualisations
We show a few real-world egocentric footages together with model predictions. We report positive prediction together with the model failures.
Predicted key probability
- Wforward0.98
- Amove left0.01
- Sback0.00
- Dmove right0.78
- Shiftwalk0.00
turn (yaw)5.2° left
tilt (pitch)1.4° up
First window of the clip
Predicted key probability
- Wforward0.00
- Amove left0.03
- Sback0.12
- Dmove right0.02
- Shiftwalk0.83
turn (yaw)2.8° right
tilt (pitch)1.5° down
First window of the clip
Predicted key probability
- Wforward0.05
- Amove left0.10
- Sback0.03
- Dmove right0.07
- Shiftwalk0.78
turn (yaw)3.5° left
tilt (pitch)1.7° up
First window of the clip
Predicted key probability
- Wforward0.02
- Amove left0.04
- Sback0.01
- Dmove right0.01
- Shiftwalk0.91
turn (yaw)1.4° right
tilt (pitch)10.7° down
First window of the clip
Predicted key probability
- Wforward0.98
- Amove left0.16
- Sback0.00
- Dmove right0.12
- Shiftwalk0.00
turn (yaw)3.6° left
tilt (pitch)1.3° up
First window of the clip
Predicted key probability
- Wforward0.02
- Amove left0.33
- Sback0.02
- Dmove right0.00
- Shiftwalk0.69
turn (yaw)5.3° right
tilt (pitch)1.3° down
First window of the clip
Predicted key probability
- Wforward0.00
- Amove left0.01
- Sback1.00
- Dmove right0.02
- Shiftwalk0.00
turn (yaw)1.1° left
tilt (pitch)1.1° down
First window of the clip
Predicted key probability
- Wforward0.42
- Amove left0.20
- Sback0.02
- Dmove right0.23
- Shiftwalk0.43
turn (yaw)48.8° left
tilt (pitch)2.3° down
First window of the clip
Predicted key probability
- Wforward0.42
- Amove left0.05
- Sback0.01
- Dmove right0.86
- Shiftwalk0.08
turn (yaw)19.3° right
tilt (pitch)1.5° down
First window of the clip
Predicted key probability
- Wforward0.09
- Amove left0.13
- Sback0.03
- Dmove right0.09
- Shiftwalk0.64
turn (yaw)5.9° left
tilt (pitch)1.5° down
First window of the clip
Predicted key probability
- Wforward0.01
- Amove left0.12
- Sback0.56
- Dmove right0.06
- Shiftwalk0.22
turn (yaw)5.0° left
tilt (pitch)1.5° up
First window of the clip
Predicted key probability
- Wforward0.06
- Amove left0.23
- Sback0.17
- Dmove right0.12
- Shiftwalk0.57
turn (yaw)16.6° left
tilt (pitch)1.7° up
First window of the clip
Extras
Real-world footage is based on Wikimedia Commons, licensed CC BY 3.0.
Weights are stored in safetensors format and can be downloaded under . Each folder also holds config.json and inference.py.
- Flow model: flow/model.safetensors (7,189,300 bytes)
Flow model: flow/model.safetensors (7,189,300 bytes)
- Pixel model: pixel/model.safetensors (39,353,324 bytes)
Pixel model: pixel/model.safetensors (39,353,324 bytes)
License
The weights and inference code in this repository are released under the Apache License 2.0. This applies to Reka’s model weights and code only. Pixel model: self-contained; no third-party model weights are required. Flow model: requires optical flow computed by RAFT-small, which is not included and must be obtained separately from Torchvision. The Torchvision and RAFT code is BSD-3-Clause. The pretrained RAFT weights were trained on datasets that are subject to their own terms, which may restrict commercial use. You are responsible for reviewing and complying with those terms before using the flow model in a commercial setting. We make no representation about their applicability. These models are provided “as is,” without warranty.





