Real humans make
real data. Real data
makes real AI.
How people communicate, work, and move - in the wild.
Explore datasets ↓Mission
Our mission is to bring
real human behavior into AI -
the conversations, decisions,
and movements.
Featured Datasets
Conversations
176B messages · 10.8T tokens · the largest natively multimodal dataset
Texts, Images, Voices, Videos - all entangled
89+ languages · reactions, forwards, threads, replies
Available datasets
Video
2.7B+ clips · 141M+ hours · 70+ PB
Short-form + long-form. Category subsets available.
Delivered in 2 days
Speech
390M voice messages · 3.2M hours of audio
65+ topic categories, diverse accents. Category subsets available. Plus 29.5M+ audio files · 56+ TB header-declared inside archive attachments (on demand).
Delivered in 2 days
Images
8.2B+ images · 99 formats
JPEG, PSD, RAW/DNG, HEIC, WebP & more. Includes 131M+ PNG, 46.7M+ JPEG and 18.8M+ SVG inside archive attachments (extracted on demand).
Delivered in 2 days
Documents
72M+ files · 400+ TB
PDF, DOCX, XLSX, PPTX & more. Category subsets available. Includes 10.8M+ documents · 19+ TB header-declared inside archive attachments (extracted on demand).
Delivered in 2 days
Code
235M+ files · 2B+ lines of code
C/C++ (61%), Python (15%), JavaScript, Lua, Shell - game engine source, SDKs.
Delivered in 2 days
Game Assets
161M+ files · 17+ TB
Textures (49M+), sound effects (6.8M+), 3D models, Unity projects, Minecraft worlds. Plus 3.6M+ 3D meshes · 49+ TB (STL, OBJ, FBX, BLEND, GLTF) and 1M+ CAD files (DWG, DXF, STEP, SLDPRT) inside archive attachments (on demand).
Delivered in 2 days
Books
3.2M+ books · 36+ TB
EPUB, MOBI, CBR/CBZ (comics), FB2 - multilingual long-form content. Includes 304k+ books · 2.5+ TB header-declared inside archive attachments (extracted on demand).
Delivered in 2 days
Corporate Data
Messenger, task tracker, meetings transcriptions, emails — how work actually happens inside companies. Sourced through direct enterprise partnerships with consenting organizations.
Delivered on demand
On-Chain Trading
5.4B+ swaps · 39M+ wallets · 7M+ tokens · 2.3 TB
Solana + EVM DEX trades, liquidity pools, mints, supply changes. Real-time pipeline.
Delivered on demand
Robotics
Video + IMU sensors · 3 tiers from GoPro to full sensor rig.
Delivered on demand
MIDI Music
2.55M+ files · 8,900+ source archives
MID, MIDI - symbolic music, note-level and instrument-level, 2015-2026. Held inside archive attachments; extraction on demand.
Delivered on demand
Subtitles
2.6M+ files · 210+ GB header-declared
SRT, ASS/SSA, VTT - timed dialogue from 435k+ source archives. Held inside archive attachments; extraction on demand.
Delivered on demand
Browse more datasets or design one with us
We offer additional proprietary datasets not listed here. Contact us to request a sample, explore more options, or collaborate on a new dataset.
Access
How to access
our datasets
-
1.
Request samplesWe will set up a quick call to understand your use case and then send you relevant data samples.
-
2.
Purchase accessEnter a data license agreement for the dataset and use-cases your team needs.
-
3.
Receive dataFor off-the-shelf datasets, we will grant your team access within 2 days.
-
✦
Experiment with usWe frequently partner with research teams to design new shapes of data for any use case. Contact us for more information.
Team