Field notes / Process

How to write a data collection spec a vendor can actually deliver

Masana3 June 20265 min read

Most problems in a data collection project start before anyone presses record. A vague spec produces a vague dataset, and the gap only shows up once footage is on your drive and the budget is spent. A good spec is short, specific, and testable. This note covers what to put in it, what to leave out, and gives you a template you can copy.

Why specs fail

Specs usually fail in one of three ways. The first is describing the model instead of the data. "We need data to train a kitchen manipulation policy" tells a vendor what you want to build, not what to capture. Two vendors can read that sentence and deliver two completely different datasets, and both can argue they met the brief.

The second is leaving acceptance undefined. If the spec does not say what makes a clip good or bad, then every rejection becomes a negotiation. You end up arguing about whether a partially occluded hand counts, or whether a clip that starts two seconds late is usable.

The third is asking for everything at once. A spec that lists every environment, every device and every demographic slice you might ever want is hard to price and hard to staff. It is better to write a tight spec for the first batch and expand it once you have seen real data.

What a deliverable spec covers

A vendor needs answers to a fixed set of questions before they can quote honestly and recruit the right people. Work through them in this order.

1. The task and what counts as one unit

Describe the activity in concrete steps, and define the unit you are paying for: one clip, one approved hour, one image, one transcribed page. Say where a unit starts and ends. For task video, that might be "from the moment the hand first touches the object until it is released and the hand leaves frame."

2. Capture setup

Name the device class and mounting (head-mounted phone, smart glasses, chest mount, handheld, studio microphone), resolution, frame rate or sample rate, orientation, and anything that must stay fixed. If you need a second camera, IMU data, or depth, say so here. If you do not care, say that too, because it widens the pool of contributors and usually lowers cost.

3. Diversity and coverage

List the dimensions that matter to your model: environments, lighting, object types, accents or languages, age ranges of adult contributors, handedness. Give target proportions or minimums where they matter and mark the rest as "best effort." Be explicit about what must not appear, such as brand logos, minors, screens with personal information, or specific locations.

4. Quality and acceptance criteria

This is the most important section. Write rules a reviewer can apply without asking you: the subject is in frame for the whole task, no motion blur that hides the hands, audio without clipping, no background music, transcripts match speech word for word. For each rule, say whether a failure means reject, fix, or flag.

5. Metadata, labels and format

List every metadata field you need per item and its allowed values. Specify file formats, naming conventions, folder structure, and the manifest you expect. For video this is often MP4 with a JSON sidecar per clip and a CSV manifest; for robot learning teams it may be a structured format such as RLDS or LeRobot. Our note on robot learning data formats covers the trade-offs.

6. Consent and compliance

State what contributor consent must cover (training, commercial use, redistribution to your customers if relevant), how bystanders are handled, and what records you need to receive. If you have regulatory constraints in a specific market, put them here rather than in a side email. See what to ask your vendor about consent.

7. Volume, timeline and pilot

Give the pilot size, the full target volume, and when you need each. Say who reviews the pilot on your side and how quickly they will respond, because review delays on the buyer side are a common cause of slipped timelines.

A spec template you can copy

Fill in what you know and leave the rest marked as open questions. An honest "TBD, need advice" is more useful to a vendor than a guess.

DATA COLLECTION SPEC  v0.1
Project name:
Owner / reviewer:            (name, email, review turnaround)

1. PURPOSE (one paragraph)
   What the data is for, and what it is NOT for.

2. TASK DEFINITION
   Activity:                 (concrete steps)
   Unit of delivery:         (clip / approved hour / image / item)
   Unit start:
   Unit end:
   Target unit length:       (range, e.g. 30-120 s)

3. CAPTURE SETUP
   Device class / mount:
   Resolution / frame rate:
   Audio:                    (required / optional / none)
   Extra sensors:            (IMU, depth, second view, none)
   Must stay fixed:
   Flexible / don't care:

4. COVERAGE
   Environments:             (list + target share or minimum)
   Lighting:
   Objects / categories:
   Contributors:             (adults only; languages, regions, handedness)
   Must NOT appear:          (minors, logos, faces of bystanders, PII on screens)

5. ACCEPTANCE CRITERIA
   Rule                              | Failure action
   ----------------------------------|----------------
   Subject in frame entire unit      | reject
   Hands/objects not obscured        | reject
   No audio clipping                 | fix or reject
   Metadata complete and valid       | fix
   (add rules)                       |

6. METADATA (per item)
   Field          | Type   | Allowed values / format
   ---------------|--------|------------------------
   item_id        | string |
   contributor_id | string | pseudonymous
   environment    | enum   |
   device         | enum   |
   duration_s     | float  |
   (add fields)   |        |

7. DELIVERY
   File format:              (e.g. MP4 / WAV / JPEG)
   Metadata format:          (JSON sidecar, CSV manifest, RLDS, LeRobot)
   Naming convention:
   Transfer method:

8. CONSENT AND COMPLIANCE
   Consent scope:            (training, commercial use, sublicensing)
   Release language:         (contributor's own language)
   Bystander handling:       (blurring, exclusion)
   Records required:

9. VOLUME AND TIMELINE
   Pilot size:
   Pilot review window:
   Full volume:
   Delivery cadence:

10. OPEN QUESTIONS
   -

Mistakes worth avoiding

  • Using adjectives as criteria. "High quality," "natural," and "diverse" are not testable. Replace each with a rule or a target.
  • Over-fixing the setup. Every constraint you add shrinks the contributor pool. Fix only what your model actually depends on.
  • Hiding the edge cases. If you already know that left-handed contributors or low light are hard for your model, say so. A vendor can target those cases deliberately.
  • Treating metadata as an afterthought. Missing fields are expensive to backfill. Define them before collection starts.
  • Skipping version control. Specs change after a pilot. Number each version and note what changed so both sides are working from the same document.

Use the pilot to tighten the spec

The first version of a spec is a hypothesis. The pilot is where you test it. Review the pilot batch against your acceptance criteria, note every case where a reviewer hesitated, and turn each hesitation into a clearer rule. Most specs reach a stable version after one round of pilot feedback. Our guide to running a data pilot walks through the checks in detail.

At Masana we usually start from a draft spec like the one above and fill the gaps together before quoting, so the price reflects what will actually be delivered.

Have a rough spec, or just a task description? Send it over and we will help turn it into something a collection team can deliver against.

Share your spec