Real humans make
real data. Real data
makes real AI.

How people communicate, work, and move - in the wild.

Explore datasets
Mission

Our mission is to bring
real human behavior into AI -
the conversations, decisions,
and movements.

Conversations
176B messages · 10.8T tokens · the largest natively multimodal dataset
Texts, Images, Voices, Videos - all entangled
89+ languages · reactions, forwards, threads, replies
Video
2.7B+ clips · 141M+ hours · 70+ PB
Short-form + long-form. Category subsets available.
Delivered in 2 days
Speech
390M voice messages · 3.2M hours of audio
65+ topic categories, diverse accents. Category subsets available. Plus 29.5M+ audio files · 56+ TB header-declared inside archive attachments (on demand).
Delivered in 2 days
Images
8.2B+ images · 99 formats
JPEG, PSD, RAW/DNG, HEIC, WebP & more. Includes 131M+ PNG, 46.7M+ JPEG and 18.8M+ SVG inside archive attachments (extracted on demand).
Delivered in 2 days
Documents
72M+ files · 400+ TB
PDF, DOCX, XLSX, PPTX & more. Category subsets available. Includes 10.8M+ documents · 19+ TB header-declared inside archive attachments (extracted on demand).
Delivered in 2 days
Code
235M+ files · 2B+ lines of code
C/C++ (61%), Python (15%), JavaScript, Lua, Shell - game engine source, SDKs.
Delivered in 2 days
Game Assets
161M+ files · 17+ TB
Textures (49M+), sound effects (6.8M+), 3D models, Unity projects, Minecraft worlds. Plus 3.6M+ 3D meshes · 49+ TB (STL, OBJ, FBX, BLEND, GLTF) and 1M+ CAD files (DWG, DXF, STEP, SLDPRT) inside archive attachments (on demand).
Delivered in 2 days
Books
3.2M+ books · 36+ TB
EPUB, MOBI, CBR/CBZ (comics), FB2 - multilingual long-form content. Includes 304k+ books · 2.5+ TB header-declared inside archive attachments (extracted on demand).
Delivered in 2 days
Corporate Data
Messenger, task tracker, meetings transcriptions, emails — how work actually happens inside companies. Sourced through direct enterprise partnerships with consenting organizations.
Delivered on demand
On-Chain Trading
5.4B+ swaps · 39M+ wallets · 7M+ tokens · 2.3 TB
Solana + EVM DEX trades, liquidity pools, mints, supply changes. Real-time pipeline.
Delivered on demand
Robotics
Video + IMU sensors · 3 tiers from GoPro to full sensor rig.
Delivered on demand
MIDI Music
2.55M+ files · 8,900+ source archives
MID, MIDI - symbolic music, note-level and instrument-level, 2015-2026. Held inside archive attachments; extraction on demand.
Delivered on demand
Subtitles
2.6M+ files · 210+ GB header-declared
SRT, ASS/SSA, VTT - timed dialogue from 435k+ source archives. Held inside archive attachments; extraction on demand.
Delivered on demand

Browse more datasets or design one with us

We offer additional proprietary datasets not listed here. Contact us to request a sample, explore more options, or collaborate on a new dataset.

Contact us →
Access

How to access
our datasets

  1. 1.
    Request samples
    We will set up a quick call to understand your use case and then send you relevant data samples.
  2. 2.
    Purchase access
    Enter a data license agreement for the dataset and use-cases your team needs.
  3. 3.
    Receive data
    For off-the-shelf datasets, we will grant your team access within 2 days.
  4. Experiment with us
    We frequently partner with research teams to design new shapes of data for any use case. Contact us for more information.
Team

Founding team

Vadims Casecnikovs
Vadims Casecnikovs
LinkedIn ↗
Artem Brustovetskii
Artem Brustovetskii
Lev Chizhov
Lev Chizhov
LinkedIn ↗