A data pilot is a small, paid batch collected under the same conditions as the full project. Its job is not to prove that a vendor can produce nice-looking samples. Its job is to surface every problem that would be expensive at scale, while it is still cheap to fix. Here is how to set one up and what to check before you approve more volume.
What a pilot should answer
Before the pilot starts, write down the questions it needs to answer. Typical ones are:
- Does the data match the spec, and is the spec itself right?
- Can the vendor hit the acceptance criteria consistently, not just on a hand-picked sample?
- Does the data load cleanly into your pipeline without manual fixes?
- Is the consent and compliance paperwork complete and usable by your legal team?
- Are the cost per unit and the delivery pace realistic for the full volume?
If you cannot say how the pilot will answer each question, the pilot is too vague. Tie each question to a check you will run and a person who will run it.
Set it up so the results mean something
Use the real conditions. The pilot should use the same devices, the same kinds of contributors, and the same environments planned for the full run. A pilot shot by the vendor's most experienced staff in one ideal location tells you little about week six of production.
Spread it across the coverage plan. If the spec calls for several environments or languages, include some of each. Problems often hide in one slice: one location with bad lighting, one accent the transcribers struggle with.
Agree on acceptance criteria in writing. Reviewers on both sides should apply the same rules. If your criteria are not yet written as testable rules, fix that first; our note on writing a data collection spec includes a template.
Book your own review time. Pilots stall most often because the buyer's team has not set aside time to review. Name the reviewer and the review window before delivery.
The checks to run
1. Technical integrity
Run every file through an automated pass before anyone watches or listens to it. Confirm that files open, durations match the metadata, resolution and frame rate or sample rate match the spec, and nothing is truncated or corrupted. Check that file names follow the convention and that every file in the manifest exists, and every file delivered is in the manifest.
2. Content against the spec
Review a sample by hand, and make it a random sample rather than the first items in the folder. For task video, check that the task is fully visible from start to end, that the hands and objects are not obscured, and that the camera is mounted as specified. For speech, listen for clipping, background noise, and whether the prompts were read or spoken naturally as required. For images and text, check framing, legibility and variety.
Keep a simple log: item, rule broken, severity. Patterns in that log matter more than any single failure.
3. Metadata and labels
Validate metadata against the schema automatically: required fields present, values in allowed ranges, IDs unique, timestamps consistent. Then spot check that the metadata is actually true. An environment field that says "kitchen" on a clip filmed in a garage is worse than a missing field, because it passes validation. If the pilot includes annotations or transcripts, compare a sample against your own careful review and note where disagreements cluster.
4. Pipeline fit
Load the pilot into the pipeline you will actually train with. If you asked for a structured format such as RLDS or LeRobot, confirm it loads with your tooling and that observations, actions and episode boundaries are where you expect them. See robot learning data formats explained for what to look for. A dataset that needs a custom conversion script for every batch will cost you engineering time for the rest of the project.
5. Consent, privacy and exclusions
Ask for the consent records that correspond to the pilot contributors and confirm they cover your intended use. Check the footage for things the spec excludes: minors, brand logos, visible personal information on screens or documents, and unblurred bystander faces in field footage. One missed face in a pilot is a process issue to fix; the same miss at full scale can mean re-reviewing everything.
6. Diversity and coverage
Tabulate the pilot by the coverage dimensions in your spec. Even at small volume you can see whether a slice is missing entirely or whether contributors are concentrated more than planned. Ask the vendor how they will hit the targets at full volume, not just whether they will.
Reading the results
When the checks are done, sort every issue into one of three buckets:
- Vendor execution issues. The spec was clear and the data did not meet it. Ask what will change in their process, not only for replacements.
- Spec issues. The data met the spec, but the spec was wrong or ambiguous. These are the most valuable findings of a pilot. Update the spec and version it.
- Your own pipeline issues. The data is fine, but your tooling needs changes. Better to find these now than in the middle of a training run.
Share the issue log with the vendor in full. A good vendor wants specific, item-level feedback, and it is the fastest way to get the next batch right.
Deciding whether to scale
You are ready to scale when the acceptance rate on the pilot is at a level you can live with, the remaining issues have a clear fix, the spec has been updated, and the data loads into your pipeline without hand edits. If the pilot raised more questions than it answered, run a second small batch on the revised spec rather than jumping to full volume. A second pilot costs far less than a large batch you cannot use.
When you do scale, keep a lighter version of the same checks running on every delivery: automated integrity and metadata validation on everything, and a random manual sample on each batch. Quality tends to drift as volume grows and contributor pools widen, and continuous checks catch the drift early.
This is why Masana works pilot-first on every project: a small batch, reviewed against agreed criteria, before anyone commits to volume. You can read more about the process on our how it works page.
Planning a collection and want to start with a pilot? Tell us what you need and we will propose a pilot scope and acceptance criteria.
Plan a pilot