Data sourcing

We help teams get the strange, real data that is not already on Hugging Face.

ReplayAI focuses on source-driven datasets: messy real-world documentation, spatial material, video, and interaction traces that are useful for models learning to understand environments and act inside them.

Built-environment records

Architectural plans, landscape layouts, housing documentation, mechanical sheets, infrastructure drawings, and related project material where rights and source permissions allow.

Spatial and world-model data

Real spaces, plans, constraints, layout variants, and visual context for systems that need to reason about rooms, paths, buildings, interfaces, and physical structure.

Video and agent traces

Ego-view recordings, screen or interface interaction data, task traces, and custom collection plans for vision, robotics, and computer-use agents.

Processing and preparation

Filtering, redaction, conversion, organization, annotation planning, and GPU-scale preparation for teams turning raw material into training-ready datasets.

Rights, source permissions, and redaction are reviewed first.

eric@replayai.org