Five kinds of data, collected to order.

Tell us the task, the place, and the people. We recruit contributors, capture the data, check every file, and deliver it in the format your pipeline expects.

First-person & task video Field & street video Speech & audio Images Text & writing

First-person & task video

Footage filmed from the worker's point of view while they do real tasks in real homes and workplaces. Built for robotics, embodied AI, and activity recognition.

Common requests

  • Household chores: cooking, cleaning, laundry, dishes
  • Skilled hand work: repairs, crafts, assembly, food prep
  • Workplace tasks in shops, markets, warehouses, and kitchens
  • Multiple people doing the same task, for variation

Typically used by robotics labs, humanoid and home-robot teams, and computer-vision companies.

Clip metadata
Capture
Head-mounted phone, smart glasses, or action camera
Resolution
1080p to 4K, 30 to 60 fps
Labels
Task, step, and object tags on request
Delivery
Per approved hour of footage

Field & street video

Footage of the world as it really looks: dense motorbike traffic, intersections, markets, farms, and building sites. Filmed fixed, handheld, or from a vehicle.

Common requests

  • Mixed motorbike and car traffic at intersections
  • Street scenes, markets, and crowds
  • Agriculture, rice fields, and plantations
  • Construction sites and outdoor workplaces

Typically used by autonomous driving, mapping, smart-city, and agriculture AI teams.

Clip metadata
Capture
Fixed, handheld, drone, or vehicle-mounted
Resolution
1080p to 4K
Labels
Scene, object, and event tags on request
Delivery
Per hour of footage or per location
Clip metadata
Capture
Studio-quiet rooms or on-location
Format
WAV, 16 to 48 kHz
Labels
Transcripts, speaker age, gender, region
Delivery
Per recorded hour or per utterance

Speech & audio

Recordings from native speakers, scripted or natural, in quiet rooms or real environments. Any language, with particular depth in Indonesian and regional languages that have little public data.

Common requests

  • English and accented English, plus Indonesian, Javanese, Balinese, Sundanese, and other under-represented languages
  • Natural two-person conversations on set topics
  • Read speech from scripts, commands, or wake words
  • Accents, age groups, and noisy real-world settings

Typically used by speech-recognition, voice assistant, translation, and local language-model teams.

Clip metadata
Capture
Phone or controlled studio setup
Format
RAW or high-quality JPEG
Labels
Classes, bounding boxes, or attributes
Delivery
Per approved image

Images

Photos shot to a consistent spec, at volume. From retail shelves to handwritten notes to graded collectibles.

Common requests

  • Products, packaging, and retail shelves
  • Food and dishes, including local cuisine
  • Documents, receipts, and handwriting
  • Collectibles and trading cards, including graded examples

Typically used by retail, document AI, food recognition, and grading or authentication teams.

Text & writing

Original writing, translations, and question-and-answer pairs from native speakers, for languages the internet barely covers.

Common requests

  • Native writing and handwriting in English, Indonesian, and regional languages
  • Translations to and from English
  • Question-and-answer and instruction pairs
  • Local knowledge, culture, and everyday topics

Typically used by teams building or fine-tuning multilingual language models.

Clip metadata
Source
Native writers and translators
Format
JSON, CSV, or plain text
Review
Checked by a second native speaker
Delivery
Per item or per word

Need something not listed here?

Most projects start as a custom request. If it can be filmed, recorded, photographed, or written by real people, we can scope it.

Send us your spec