Field notes / Data quality

Why AI struggles outside North America and Europe, and what data fixes it

Masana5 August 20265 min read

AI models often look impressive in demos and then stumble when they meet the real world outside a handful of wealthy countries. A vision model misreads a kitchen in Jakarta, a speech system loses track of a Manila accent, a robot policy freezes in a cramped market stall. The cause is usually not the model architecture. It is the data the model learned from, and where that data came from.

The evidence: models learn the world their data shows them

This is not a hunch. Researchers have measured it. In 2017, a Google team analyzed two large, widely used public image datasets in "No Classification without Representation" and found they showed an observable Amerocentric and Eurocentric representation bias. Classifiers trained on them performed noticeably differently on images from different locales.

In 2019, researchers at Facebook AI tested publicly available object recognition systems on Dollar Street, a dataset of household items photographed in homes across many countries and income levels. Their paper, "Does Object Recognition Work for Everyone?", found the systems performed relatively poorly on items common in lower-income households. The authors traced the drop mainly to two things: objects looking different within the same category (their example was dish soap) and objects appearing in a different context (toothbrushes outside a bathroom).

Those two findings are the heart of the problem. Models do not just learn what a thing is. They learn what it usually looks like and where it usually appears in their training data. When that data comes mostly from one part of the world, everywhere else becomes an edge case.

Why the skew exists

Nobody set out to build biased datasets. The skew comes from how data is usually gathered:

  • Web scraping follows the web. Content that is easy to scrape, captioned in English and hosted on large platforms is overrepresented.
  • Collection follows the researchers. Many datasets are recorded by universities and companies near where they are based, with participants who are easy to reach.
  • Labels follow the labelers. Category names and annotation guidelines often reflect one cultural frame, so objects or activities that do not fit get mislabeled or left out.
  • Cost follows convenience. Running collection in unfamiliar countries means local recruiting, translation, consent paperwork and logistics. It is easier to skip.

Some large projects have made real efforts to spread out. Ego4D, a major egocentric video dataset, was recorded by 931 camera wearers in 74 locations across 9 countries. That is a meaningful step, but even broad research datasets cannot cover every environment, language and task a commercial product will meet.

Where it shows up in practice

Vision and robotics

Homes, shops, streets and workplaces differ a lot across regions. Kitchens may be outdoors or shared, cooking may happen at floor level, tools and utensils differ, and lighting ranges from harsh equatorial sun to a single bulb. Traffic mixes motorbikes, carts, pedestrians and animals in ways a model trained on wide suburban roads has never seen. For robots learning from human demonstration, the way people actually perform a task, such as washing clothes by hand or using a wok, may simply be missing. See our explainer on egocentric video data for why first-person task footage matters here.

Speech and language

Speech models struggle with accents, dialects and code-switching, where speakers mix two or more languages in one sentence. Code-switching is everyday speech in much of Southeast Asia. Background noise also differs: open-air markets, scooters and rain on tin roofs are real acoustic conditions that clean studio recordings do not prepare a model for.

Text and handwriting

Handwritten forms, receipts, signage and informal messaging in regional languages and scripts are thin in most training corpora. Optical character recognition and document models often fail on exactly the documents that matter for local deployments.

What data actually fixes it

More data from the same sources will not close a geographic gap. What helps is data that deliberately represents the places and people a model will serve. In practice that means:

  • Collected in the target region, by local contributors. Real homes, real streets, real workplaces, not staged sets designed to look local.
  • Varied within the region. Urban and rural, different income levels, different building types, different times of day and weather. One city is not a country.
  • Task-specific. If a robot will fold laundry in Indonesian homes, collect people folding laundry in Indonesian homes, with the clothes, surfaces and habits that come with it.
  • Linguistically honest. Speech and text that include local languages, dialects, accents and natural code-switching, with metadata that records them.
  • Richly tagged. Location type, environment, device, language and contributor attributes captured as metadata, so you can measure performance by slice and see exactly where the model still fails.
  • Properly consented. Regional data is only useful if you can use it. Contributors should sign releases in their own language, and bystanders in field footage should be protected. Our consent guide covers what to ask.

How to start without boiling the ocean

You do not need to collect data from every country at once. A practical approach:

  1. Evaluate by region first. Build a small, geographically tagged test set for your target markets and measure where performance drops. This tells you whether you have a problem and how big it is.
  2. Prioritize by deployment. Start with the regions where you have users or plan to launch, and the tasks where errors are most costly.
  3. Run a pilot. Collect a small, targeted batch, fine-tune or evaluate, and check whether the gap narrows. Our guide to running a data pilot explains the quality checks.
  4. Scale what works. Expand collection only along the dimensions that moved your metrics.

Southeast Asia is a good example of why this matters. It is home to hundreds of millions of people, many languages and a wide range of living environments, yet it is thinly represented in most public training data. Masana is based in Bali, Indonesia, and its deepest contributor network is in Indonesia and the wider region, which is why we see this gap up close.

If your model needs to work in Indonesia or the rest of Southeast Asia, we can collect consented video, speech, images and text in real local settings, starting with a small pilot.

See our Southeast Asia coverage