Reka Inverse Dynamics Model for Interactive World Models

Источник: Reka AI

Reka Inverse Dynamics Model for Interactive World Models

Source: Reka AI

RIDM extracts low-level commands like moving forward, sideways, turning or tilting from real-world video, for interactive world models. It is trained entirely on video games.

•Updated: October 2, 2026

Model weights (Apache 2.0): RekaAI/Reka-Inverse-Dynamics-Model on Hugging Face

Interactive World Models are different. They are visually more realistic, and can model more complex phenomena such as folding materials or squeezing towels. By embedding digital replicas into these dynamic environments, robots can execute commands through direct interaction.

Here is the challenge though. True interactivity requires a world model to accurately respond to low-level commands (e.g., moving left or right). To solve this, we introduce the Reka Inverse Dynamics Model (RIDM), which extracts low-level commands directly from real-world video data to act as ‘captions’ describing physical changes in consecutive frames.

RIDM is a compact, efficient model that processes short clips and extracts low-level motor commands like moving forward, sideways, turning or tilting. Even though the model is entirely trained on video games, by decoupling visuals from motion, its predictions transfer to real-world videos.

Predicted key probability

  • Wforward0.85
  • Amove left0.34
  • Sback0.01
  • Dmove right0.76
  • Shiftwalk0.01

turn (yaw)16.9° left

tilt (pitch)1.4° up

First window of the clip

The data

We recorded gaming footages at 1280 × 720, 48 frames per second, with 11 ground-truth key commands, the mouse movement and the exact camera angles for every frame. The annotations are provided by the engine.

The model

Our model disentangles visuals from motion by projecting inputs into optical flow space. Since our goal is to use RIDM to automatically annotate real-world footages, not 3d gameplays, the method needs to transfer from games to real-world. We also compare that approach with a direct one that is working in the pixel space. To differentiate both, we use names: flow-based and pixel-based model.

Training both models on random gaming footages will not work either. Successful approach requires data filtering and resampling. For instance, some actions such as backward or sideways strafe movements are much more rare than simple forward movements. Suitable resampling helps in learning these less common movements.

Both models are compact. The flow-based model has 2.8m (2,785,499) parameters; with only 1.8m (1,795,337) trainable parameters and about 990k frozen that constitutes RAFT-small. The pixel-based model is fully trainable and has around 9.8m (9,836,063) parameters.

Click the view · W A S D · Shift · arrows

Pixels · what the camera sees

Optical flow · what the flow model seesdirection

This interactive figure needs WebGL, which this browser has turned off.

The label · what the game records

WASDShift

Turn0°/sTilt0°/s

Walking forward

Evaluation

We tested the generalisation of both approaches on videos that neither method has seen during training. As expected, the flow-based model works better, with the biggest gains obtained when transferring from gaming footages to real-world video clips.

In evaluation, we used three datasets:

  • Real-world videos with camera angle labels,

Real-world videos with camera angle labels,

  • Real-world GoPro footage,

Real-world GoPro footage,

  • Synthetic, 3d gaming footages.

Synthetic, 3d gaming footages.

Note that real-world datasets hold only forward and turn actions.

Each dataset asks two questions; totalling six different configurations. Did the camera go forward or turn? If it turned, which way?

  • The flow-based model works better on five out of the six configurations,

The flow-based model works better on five out of the six configurations,

  • On real-world video the gap is the largest. The flow-based model gets 86 % while the pixel-based model gets only 31 %,

On real-world video the gap is the largest. The flow-based model gets 86 % while the pixel-based model gets only 31 %,

  • The pixel model is only better to predict GoPro turn direction, and only by a single clip.

The pixel model is only better to predict GoPro turn direction, and only by a single clip.

Table 1 — Two models benchmarked on six different evaluation configurations. Since the ground-truth label derived from sensors is inaccurate in real datasets — Real video and GoPro — we manually inspected predictions to calculate accuracies. For gaming footages, we use available ground-truth annotations.

Evaluation data

Task

Flow model

Pixel model

Real video, 78 clips

Forward, turn left or right

85.9 %

30.8 %

Real video, 47 turn clips

Turn left or right

91.5 %

51.1 %

GoPro, 72 clips

Forward, turn left or right

75.0 %

58.3 %

GoPro, 25 turn clips

Turn left or right

72.0 %

76.0 %

Gaming footage, 5,216 flow pairs and 5,164 pixel pairs

Forward, turn left or right

84.5 %

77.5 %

Gaming footage, 3,716 flow turn pairs and 3,674 pixel turn pairs

Turn direction

84.9 %

70.9 %

Our RIDM, of course, occasionally fails in some situations:

  • it reports false turns on static frames and on dark frames,

it reports false turns on static frames and on dark frames,

  • when it has to disentangle side step and opposite yaw turns. For example when the person is stepping right and turns camera left at the same time,

when it has to disentangle side step and opposite yaw turns. For example when the person is stepping right and turns camera left at the same time,

  • it reports walk keys on a camera that stands still and zooms in / out,

it reports walk keys on a camera that stands still and zooms in / out,

  • the camera stabilization on action-camera footage hides part of the rotation.

the camera stabilization on action-camera footage hides part of the rotation.

We believe that training on more data, especially on the data containing camera effects like zooming or stabilisation, would improve the model’s performance on the failures above.

Visualisations

We show a few real-world egocentric footages together with model predictions. We report positive prediction together with the model failures.

Predicted key probability

  • Wforward0.98
  • Amove left0.01
  • Sback0.00
  • Dmove right0.78
  • Shiftwalk0.00

turn (yaw)5.2° left

tilt (pitch)1.4° up

First window of the clip

Predicted key probability

  • Wforward0.00
  • Amove left0.03
  • Sback0.12
  • Dmove right0.02
  • Shiftwalk0.83

turn (yaw)2.8° right

tilt (pitch)1.5° down

First window of the clip

Predicted key probability

  • Wforward0.05
  • Amove left0.10
  • Sback0.03
  • Dmove right0.07
  • Shiftwalk0.78

turn (yaw)3.5° left

tilt (pitch)1.7° up

First window of the clip

Predicted key probability

  • Wforward0.02
  • Amove left0.04
  • Sback0.01
  • Dmove right0.01
  • Shiftwalk0.91

turn (yaw)1.4° right

tilt (pitch)10.7° down

First window of the clip

Predicted key probability

  • Wforward0.98
  • Amove left0.16
  • Sback0.00
  • Dmove right0.12
  • Shiftwalk0.00

turn (yaw)3.6° left

tilt (pitch)1.3° up

First window of the clip

Predicted key probability

  • Wforward0.02
  • Amove left0.33
  • Sback0.02
  • Dmove right0.00
  • Shiftwalk0.69

turn (yaw)5.3° right

tilt (pitch)1.3° down

First window of the clip

Predicted key probability

  • Wforward0.00
  • Amove left0.01
  • Sback1.00
  • Dmove right0.02
  • Shiftwalk0.00

turn (yaw)1.1° left

tilt (pitch)1.1° down

First window of the clip

Predicted key probability

  • Wforward0.42
  • Amove left0.20
  • Sback0.02
  • Dmove right0.23
  • Shiftwalk0.43

turn (yaw)48.8° left

tilt (pitch)2.3° down

First window of the clip

Predicted key probability

  • Wforward0.42
  • Amove left0.05
  • Sback0.01
  • Dmove right0.86
  • Shiftwalk0.08

turn (yaw)19.3° right

tilt (pitch)1.5° down

First window of the clip

Predicted key probability

  • Wforward0.09
  • Amove left0.13
  • Sback0.03
  • Dmove right0.09
  • Shiftwalk0.64

turn (yaw)5.9° left

tilt (pitch)1.5° down

First window of the clip

Predicted key probability

  • Wforward0.01
  • Amove left0.12
  • Sback0.56
  • Dmove right0.06
  • Shiftwalk0.22

turn (yaw)5.0° left

tilt (pitch)1.5° up

First window of the clip

Predicted key probability

  • Wforward0.06
  • Amove left0.23
  • Sback0.17
  • Dmove right0.12
  • Shiftwalk0.57

turn (yaw)16.6° left

tilt (pitch)1.7° up

First window of the clip

Extras

Real-world footage is based on Wikimedia Commons, licensed CC BY 3.0.

Weights are stored in safetensors format and can be downloaded under . Each folder also holds config.json and inference.py.

  • Flow model: flow/model.safetensors (7,189,300 bytes)

Flow model: flow/model.safetensors (7,189,300 bytes)

  • Pixel model: pixel/model.safetensors (39,353,324 bytes)

Pixel model: pixel/model.safetensors (39,353,324 bytes)

License

The weights and inference code in this repository are released under the Apache License 2.0. This applies to Reka’s model weights and code only. Pixel model: self-contained; no third-party model weights are required. Flow model: requires optical flow computed by RAFT-small, which is not included and must be obtained separately from Torchvision. The Torchvision and RAFT code is BSD-3-Clause. The pretrained RAFT weights were trained on datasets that are subject to their own terms, which may restrict commercial use. You are responsible for reviewing and complying with those terms before using the flow model in a commercial setting. We make no representation about their applicability. These models are provided “as is,” without warranty.

What this article says

Something is unclear? Ask about the article — I will explain in plain words.

Do not want to dig deeper? We will sort it out for you.