Dev48
Language
  • About
  • Services
  • Industries
  • Technologies
  • Articles
  • Contacts
Book a call
    Home/Articles/Rekadaily 10k collecting 10000 hours of egocentric household manipulation data
Dev48

© 2026 · All rights reserved.

RekaDaily-10k: Collecting 10,000+ Hours of Egocentric Household Manipulation Data

Источник: Reka AI

RekaDaily-10k: Collecting 10,000+ Hours of Egocentric Household Manipulation Data

Source: Reka AI

RekaDaily-10k: 10,000+ hours of unscripted first-person household recordings from paid collectors, released under Apache 2.0 in a raw tier and a captioned tier.

September 27, 2026•Updated: September 27, 2026

Update, Sep 16, 2026: Both tiers are now fully live. The raw tier finished at 10,865 hours, and the processed and captioned tier is available for the first time: 10,200 hours across 6,373,064 clips, totalling 74.2 TB. About 3,000 hours of the raw tier also include IMU sensor data (acceleration, rotation, and precise frame timing) captured alongside the video, flagged via has_imu in the metadata.

To understand and simulate the physical world, omni world models and vision-language-action models need more than text. They need high-quality visual data with descriptions of what is happening, recorded directly in the chaotic environments where real behavior occurs.

Search for cooking videos and you get an enormous library of edited, staged, tripod-mounted footage cut to keep the interesting parts. What a machine needs in order to learn a physical task is the opposite: one continuous first-person view of somebody actually doing it, at the speed they actually do it, in the mess they actually live in. Nobody uploads that, because nobody would watch it.

That footage has to be commissioned. Teleoperated data is precise but slow to produce, and it tends to inherit the tidiness of the space it was recorded in. Synthetic scenes scale but smooth over the clutter of real homes. Real first-person recordings sit in between, and the reason there are not more of them is that somebody has to pay people to make them.

Introducing RekaDaily-10k

To meet our own requirements for scale and quality, and to satisfy the needs of the broader industrial community, we built Claru, Reka’s foundational data engine, which sources egocentric video through a global network of paid collectors.

Today we are releasing RekaDaily-10k: unscripted, first-person recordings of everyday household life, captured across homes by paid collectors recording their own routines, with a significant share in 4K. We are contributing this dataset to the research community as part of our open ecosystem initiative, under Apache 2.0. Both raw and processed, captioned tiers are now fully live on Hugging Face: the raw tier at 10,865 hours, and the processed and captioned tier at 10,200 hours.

The dataset ships in two tiers:

  • Raw tier. Unfiltered, raw footage, allowing teams to implement their own clipping, filtering, and annotation workflows. It has 10,865 hours. Roughly 3,000 hours of the raw tier ship with IMU data alongside the video: acceleration, rotation, and precise frame timing, sampled up to 200 times a second and synced to the footage. Clips with this data are flagged via has_imu in the raw tier’s metadata; the remaining ~7,800 hours are video only.

Raw tier. Unfiltered, raw footage, allowing teams to implement their own clipping, filtering, and annotation workflows. It has 10,865 hours. Roughly 3,000 hours of the raw tier ship with IMU data alongside the video: acceleration, rotation, and precise frame timing, sampled up to 200 times a second and synced to the footage. Clips with this data are flagged via has_imu in the raw tier’s metadata; the remaining ~7,800 hours are video only.

  • Processed and captioned tier. Footage that has been through our processing pipeline, cut into shorter clips and captioned, for teams that want language supervision out of the box.

Processed and captioned tier. Footage that has been through our processing pipeline, cut into shorter clips and captioned, for teams that want language supervision out of the box.

The Apache 2.0 licence covers commercial use and redistribution. Each clip ships as video with a text caption.

The Egocentric Landscape

The egocentric ecosystem has grown quickly, and this release is meant to add to it with more videos showing egocentric activities. Ego4D established the modality and remains the reference corpus for daily life. Egocentric-10K and the larger releases that followed it showed how far first-person data can scale, and they cover industrial work in real production environments. EPIC-KITCHENS is still the gold-standard of ego-centric benchmarks of human activities.

What RekaDaily-10k adds: Unscripted household activity recorded in real homes with detailed descriptions. Roughly 1,670 hours are in native 4K, which is a higher resolution than most large egocentric corpora carriers. Together, this makes it an ideal complement for teams building domestic AI where language supervision matters.

Where The Footage Comes From

is a paid collection network of more than 100,000 people recording the physical world. Collectors join a project, pass a qualification assessment, then record, submit, and get paid per accepted hour. The work spans domestic life, commercial environments and skilled trades across several regions. This release draws on the household portion of that network, recorded by a subset of those collectors on phones in head mounts.

Environment Diversity

Because every collector records in their own home, the number of distinct environments rises with the number of people who contribute rather than with the number of hours recorded. Different kitchens, appliance models, cabinet layouts, floor plans, lighting conditions and degrees of clutter. Different outlets and switch plates, different packaging on the shelves, different signage languages, different weather through the windows. That kind of variation is difficult to produce any other way.

Authentic and Unscripted

The footage is unscripted, capturing the dirt and noise of the real world, and thus making the dataset valuable. Collectors record real activity instead of performing a task list, so sessions run long, hands leave frame, tasks get abandoned and resumed later, and people walk in and interrupt. That is roughly what a deployment environment looks like, and it is the part that is hard to arrange deliberately.

Everyday Household Routines

What people recorded is mostly household work and the routine around it. Laundry from the pile to the folded stack. Kitchen cleanup, dishes, unloading groceries, wiping surfaces. Reorganising rooms, closets and drawers. Sweeping, taking out the trash, watering plants, clearing a table, unboxing something new, changing a bulb or a battery. The short fiddly two-handed jobs are here precisely because no one would commission a teleop session for them, and they are still on everybody’s actual list.

Consent and Privacy

Paid Collector Consent

This is footage recorded inside people’s homes, and publishing it openly raises the stakes on how it was obtained. Collectors are paid contractors who opt in, and every session is recorded with the wearer’s knowledge and agreement.

Bystanders

The wearer is not the only person a home camera sees, so bystanders are handled separately. Collectors are instructed to record only with the agreement of other adults present and to keep others out of frame where that is not possible, and footage that shows identifiable non-participants is flagged for review before it can be released.

PII Screening

Personally identifiable information cannot be reduced to a yes or no question. Recording real environments at this volume will inevitably lead to PII passing in front of the lens; what matters is identifying it and deciding what to do.

From Submissions to Captioned Video

10,865 hours of unedited phone footage make up the raw tier. The processed tier runs quality controls, cuts the video into shorter clips, deduplicates shots, and captions them. All done at a large scale.

Quality control runs first, on every submission. It was introduced partway through collection, so it covers most but not all of the dataset we are releasing now; everything going forward passes through it. The checks cover the failure modes that actually recur: hands never visible, illumination too poor to make anything out, frozen or duplicated frames, wrong orientation, solid-colour padding, clips too short to use, footage that does not match the assigned activity. A second pass reads the content and asks whether it is what it claims to be, flagging non-POV, staged or screen-recorded video, and time inflation where somebody stalls to pad a per-hour payout. Every video also gets a perceptual fingerprint, because at this volume near-duplicates stop being hypothetical.

The quality thresholds came from data rather than instinct. We took a large sample of historical submissions that already carried a final human verdict and replayed them through the pipeline to compare machine flags against what reviewers had decided. Where the two disagreed we moved the thresholds toward the reviewers rather than toward numbers that look strict on paper. One principle is written into the code. A false rejection is the worst outcome, so borderline video goes to review instead of the bin, anomalous scores raise a warning rather than a rejection, and no automated check can reject a submission without clearing a confidence floor.

Captioning is where the length of these recordings becomes the whole problem. A caption has to reflect where an activity sits in the arc of a session rather than what happened to be in one sampled frame, which matters when somebody loads the washer, wanders off, and comes back twenty minutes later to unload it.

A caption has to say where an activity sits in a session, not what happened to be in one sampled frame

ONE UNSCRIPTED SESSION · 42 MIN

loads the washer

elsewhere in the house

unloads the washer

folding

one sampled frame

0:00

42:00

FROM THE SAMPLED FRAME

“A person walks through a kitchen.”

True of that instant and useless as supervision — nothing in it says a wash cycle is running.

FROM THE WHOLE SESSION

“Loads the washer, leaves for twenty minutes, returns to unload it and folds the load.”

The activity placed in the arc of the session, with the order and the pauses intact.

For the wider picture of how footage like this gets prepared for world model training, our data platform team wrote up the full pipeline in World Model Data Pipeline.

What You Can Train On It

Two things that can make this corpus useful.

The first is language tied to real activity. A caption on a stock clip describes a scene. A caption over a continuous session describes what a person was doing and in what order, with the pauses, mistakes and corrections still in frame. That is the supervision instruction-conditioned models run on, and the interesting part is not the captioning. It is having ten thousand hours of unscripted first-person footage worth captioning in the first place.

The second is coverage of ordinary domestic environments. If you are training a household robot, a world model, or a video generation model that has to respect how objects behave when handled, the distance between your training data and a real kitchen matters. This is thousands of real kitchens at the hours of day when people are actually in them.

The raw tier is deliberately unopinionated, and it is there because our processing choices should not be forced on you. We clip, filter and caption to serve the projects we run, and any team with its own pipeline will want different boundaries, different thresholds and its own annotation schema. Releasing the footage as collected means you can start upstream of every decision we made.

Dataset Fields

Each clip in the processed and captioned tier ships with the following fields.

Field

Description

clip_id

A unique ID for this specific clip. It combines the ID of the original source video with a suffix showing which segment, scene, and clip number it is within that video.

video_id

A unique ID for the original, full-length source video this clip was cut from. Several clips can share the same video_id if they came from the same recording.

subset

The name of the collection this clip belongs to within the dataset, for example “daily_activities”.

width

The width of the video frame, in pixels.

height

The height of the video frame, in pixels.

num_frames

The total number of frames in the clip.

fps

The clip’s frame rate, in frames per second.

duration_s

The length of the clip, in seconds.

motion_class

How much movement happens in the clip. One of four values: “minimal”, “low”, “medium”, or “high”.

quality_tier

A score from 1 to 20 based on technical video quality (blur, noise, compression artifacts), where 20 is the highest quality and 1 is the lowest.

text

A written description of what’s happening in the clip, generated automatically from the video content.

text_quotes

Any text visible on screen in the clip, such as a sign, label, or caption, extracted automatically using text recognition. It’s an empty string when there’s no meaningful on-screen text.

has_significant_text

Whether the clip has meaningful on-screen text. True means text_quotes contains something; false means it’s empty.

chats

Example question-and-answer exchanges about the clip, generated automatically from the video content, for example someone asking what’s happening and getting a description back.

activities

One or more everyday activities shown in the clip, for example “pet care” or “laundry”. These labels come from the person who recorded the footage.

activity_timeline

A breakdown of the activities across the full source video, each with a start and end time in seconds. These times refer to the original, full-length video, not just this clip. Also labeled by the person who recorded the footage.

country

The two-letter code for the country where the footage was recorded, based on where the person recording it was located.

Get The Data

The dataset is on Hugging Face under Apache 2.0, ungated.

Check out on Hugging Face

Following our June release of RekaCS2-10k, 10,000 hours of egocentric Counter-Strike 2 footage with per-frame action annotations and its accompanying open-source renderer, we are releasing RekaDaily-10k to advance progress in physical AI and foster an open research ecosystem.

Work With Us

Physical AI depends on footage captured in authentic, real-world environments. We operate a global network of over 100,000 collectors and are continuously scaling, making our data engine a powerful, extensible resource for your projects.

If this release aligns with your goals but you require a different setting, specific activities, a particular region, or specialized annotation, we can help. We provide custom data collection using the same proven pipeline. We have built bespoke datasets to spec for frontier research labs before, delivering completed projects with turnaround times measured in weeks.

This release is a snapshot of what our network produces daily. If you have specific requirements for data that doesn’t exist yet, whether for training or benchmarking, talk to us and we will build a programme to collect it.

Citation

← All articles

More in AI & Machine Learning

All →
Lambda to build new data center in Mayes County, Oklahoma, generating half a billion dollars in tax revenue over next decade
Lambda

Lambda to build new data center in Mayes County, Oklahoma, generating half a billion dollars in tax revenue over next decade

Google tests buying from Walmart-owned Flipkart through Gemini and AI Mode in IndiaПресса
Gemini

Google tests buying from Walmart-owned Flipkart through Gemini and AI Mode in India

OpenAI expands review of model behavior after more rogue agent incidents emerge
Пресса
OpenAI

OpenAI expands review of model behavior after more rogue agent incidents emerge

Apple faces $5.7 billion patent infringement verdict over iPhone and Apple Watch hapticsПресса
Apple

Apple faces $5.7 billion patent infringement verdict over iPhone and Apple Watch haptics

Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledgeПресса
OpenAI

Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledge

Proaction boosts sales 60% and saves 75+ hours with Codex
OpenAI

Proaction boosts sales 60% and saves 75+ hours with Codex

More from Reka AI

Beyond Recognition: How Our Models Reason About Video
Reka AI

Beyond Recognition: How Our Models Reason About Video

Real-Time Video Generation
Reka AI

Real-Time Video Generation

Evolution of LLMs: Omni-World Models
Reka AI

Evolution of LLMs: Omni-World Models

Reka EdgeQ, On-Device
Reka AI

Reka EdgeQ, On-Device