
Author:
Casper Grathwohl, President
Enterprise AI teams routinely spend months evaluating model architectures, weeks debating infrastructure choices, and an afternoon picking their language data provider. That imbalance has a name: the evaluation asymmetry, and it is one of the most reliable predictors of production failure.
The reason it persists is that language data failures are invisible until they are embarrassing. A model trained on the wrong data does not produce obvious errors. It produces fluent, confident, subtly wrong outputs, in the domain where precision matters most, in the language your user speaks, or in the register your application requires.
By the time that failure surfaces, it has already eroded user trust, and the remediation is a full retraining cycle that traces back to a data decision made months earlier.
Oxford Languages draws on more than 150 years of experience building the language resources that production AI systems run on. Here is our advice on how to evaluate a language data provider for enterprise AI.
Before you start: know what you are actually buying
Language data is not a commodity. The same label, “English language dataset”, can describe a web-scraped text or structured dictionary data with lexical metadata. These are not the same product, and they will not produce the same model.
Before any evaluation, get specific about what your application actually requires:
- Does it need word sense understanding in context, or surface-level vocabulary?
- Does it serve non-English speakers, and if so, which languages at what quality level?
- Is it operating in a specialist domain: legal, medical, financial, or educational?
- Does it require static datasets, or ongoing access to live lexical data through dictionary APIs?
- Is it subject to regulatory requirements that create data provenance obligations?
The answers determine which evaluation criteria carry the most weight. This is because a voice AI platform and an enterprise search tool need language data for structurally different reasons. Treating them identically in evaluation produces the wrong result for both.
How to evaluate a language data provider
Audit how the data was actually built
Most providers describe their data in terms of size and coverage. Neither tells you how it was made, and that is the most consequential question.
Data built on web crawls reflects what the public web tends to overrepresent, including English, informal register, dominant word senses, and content optimized for visibility rather than accuracy. Data built through expert lexicographic processes reflects how language is actually used across contexts, registers, and communities.
These approaches produce meaningfully different models, and when a provider cannot explain that difference in specific terms, it reveals something about the data itself.
When auditing data construction, ask:
- Who created the data? Was it in-house linguists and lexicographers, or automated pipelines?
- How are word senses handled? Is there explicit disambiguation, or does the data treat all uses of a word as equivalent?
- How are definitions ordered? Is it by frequency of current usage, or historical precedence?
- What does quality assurance look like before data is released?
At Oxford Languages, our editorial process is built on continuous corpus analysis and expert lexicographic review. This is the same methodology that has tracked language change for over a century, now structured for AI and NLP training requirements.
Verify the source before you sign anything
Legal risk from unlicensed training data is no longer hypothetical. Litigation against major AI labs shows that even small amounts of unlicensed content in training datasets can create exposure for downstream applications trained on those datasets. The EU AI Act requires documentation of training data sources for high-risk AI systems. France’s CNIL also issued guidance in June 2025 requiring documentation of data sources, collection criteria, and data minimization practices for AI training pipelines.
If a provider cannot produce chain-of-custody documentation at the evaluation stage, they will not produce it when your compliance team asks. That is the time to find out, not after you have trained a model.
Require from any shortlisted provider:
- Documented chain of custody for all data sources
- Clear commercial licensing that explicitly covers AI training use
- GDPR and data privacy compliance confirmation
- Audit-ready provenance documentation
Our dictionary datasets are built from curated and licensed sources with the documentation enterprise legal teams need.
Test multilingual quality, not coverage claims
Headline coverage figures are best treated as a starting point. Quality in specific languages is what determines whether your product works for non-English speakers.
Some widely used multilingual web-crawled datasets contain little to no usable text for the languages they claim to cover, while many others include a significant proportion of low-quality sentences that are easy to spot even without specialist knowledge. These datasets often pass automated checks, yet the underlying issues still persist.
Do not evaluate aggregate language counts. Evaluate specific languages your product serves, in the domains and registers your users actually use. Ask:
- Was the non-English data sourced with native speaker involvement, or machine translation?
- Can the provider show language-specific quality metrics, not just overall coverage?
- Does the data capture regional and dialectal variation, or only dominant forms?
Good to know: Oxford Languages works directly with external, native-speaking consultants for all the languages it offers, including rare and underrepresented languages. This results in high-quality, accurate, and idiomatic language data.
Check domain coverage against your actual use case
General language data underperforms in specialist domains, and it underperforms quietly, in ways that do not appear in testing but do appear when professional users start asking precise questions.
Legal, medical, financial, and technical applications require vocabulary that is correctly classified and appropriate to the professional register. A model trained on general web text will produce fluent responses in these domains that are subtly imprecise in the ways professionals immediately notice. This is the most expensive failure mode in specialist applications, because the model has lost user trust before anyone has identified the data as the cause.
Before shortlisting, confirm:
- Does the data include specialist vocabulary for your target domain?
- How are terms handled when they carry different meanings across contexts?
- Is there explicit register tagging, formal, informal, technical, or colloquial?
Our Language Data for AI is structured specifically for the precision requirements of domain-specific AI applications.
Require a maintenance and versioning commitment
A dataset is not a static asset. Language changes continuously, new terms enter use, existing terms shift meaning, and domain vocabulary evolves. A provider that releases data once and considers it done is not a long-term partner. They are a liability waiting to surface in production.
Static datasets create a specific failure mode: models that perform well on historical benchmarks and degrade as the gap widens between training data and current usage. In fast-moving domains, such as technology, finance, and healthcare, that degradation is not gradual. It is sudden.
Require any shortlisted provider to show:
- A documented update cadence and what triggers a data release
- A versioning system so your team can track what changed between releases
- A process for capturing emerging vocabulary and usage shifts
Our data is maintained as a living system, continuously updated through corpus analysis and editorial review, the same process that has tracked language change for over a century.
Evaluate integration fit and support quality
High-quality data in a format your pipeline cannot use is still a problem. Integration friction is one of the most common reasons enterprise data partnerships fail, even when the underlying data is good.
WellSaid Labs needed phoneme-level pronunciation data for their AI voice platform, not dictionary definitions. They integrated our Pronunciations Data because the data structure matched the specific technical requirement. That fit between requirement and product is what makes an integration work without a prolonged engineering workaround.
When evaluating integration, ask:
- What formats does the data come in? Do they match your pipeline requirements?
- Is there dictionary API access, and how comprehensively is it documented?
- What does onboarding look like for a team integrating the data for the first time?
- Is there dedicated support during and after integration?
The Oxford Dictionaries API gives development teams direct, documented access to structured lexical data. Our customer success team supports integrations across use cases. This tends to come in useful because voice AI, enterprise search, and educational tools use language data in structurally different ways.
Demand reference customers, not just case studies
A case study is a provider’s best version of their own story. A direct conversation with a reference customer in your category is evidence.
Ask any shortlisted provider to connect you with teams who have deployed their data in production at a comparable scale. If they cannot or will not, that is itself information. Look for deployments in your industry or use case category, not just adjacent ones. Focus on long-term production use rather than pilots, and teams willing to discuss what did not work, not just what did.
What to look for in references:
- Deployments in your industry or use case category, not just adjacent ones
- Long-term production use, not pilots
- Teams willing to discuss what did not work, not just what did
Our case studies cover production deployments across voice AI (WellSaid Labs), enterprise dictionary platforms (Kielikone), assistive technology (HumanWare), multilingual consumer apps (MobiSystems), and educational tools (C-Pen). These are long-term partnerships, and the teams behind them can speak to the data in production.

“For Kielikone, partnering with Oxford Languages has been instrumental in the evolution of our high-quality language tools. Oxford University Press’s meticulously curated language data helps us bring a level of trust and authority that our clients value, distinguishing our offerings from those relying solely on automated data. The collaborative spirit and open communication between us have made working together an inspiring experience over the years. We look forward to continuing this mutually rewarding partnership for many years to come.”
Kaarina Hyvönen, VP Operations
The evaluation process in practice
A structured approach to provider evaluation looks like this:
- Define requirements before approaching vendors. The seven criteria above only matter in proportion to your specific use case. Decide which ones are non-negotiable before you see a demo.
- Ask for data samples in your target language and domain. Do not evaluate English data if you are building a multilingual product.
- Test against your hardest cases. Demos are optimized for vendor strengths. Your production failures will happen at the edge cases. Design your evaluation around those.
- Request provenance documentation before shortlisting. If a vendor cannot produce it at the evaluation stage, they will not produce it when your compliance team asks.
- Talk to reference customers in your category. A case study is a starting point. A direct conversation with a comparable team is the actual evidence.
Set your criteria before vendor demos, weight them according to your use case, and hold every provider to the same evidence standard.
The decision that compounds
Language data decisions are not easy to undo. When a model is trained on the wrong data, fixing it requires a full retraining cycle. What seemed like a manageable shortcut during procurement becomes a costly production problem with a data decision at its root.
This asymmetry is not inevitable. Criteria for evaluating language data are well understood, and the evidence needed to assess quality, coverage, and compliance is available. The scrutiny can be applied before a commitment is made rather than forced by a production failure afterward.
Providers worth choosing are the ones that can answer hard questions clearly, before the contract is signed, not after the model is trained.
If you’re assessing language data for a specific use case, take a closer look at our datasets, see how other teams are using them in production, or get in touch with our team to discuss your requirements.

Casper Grathwohl is the President of Oxford Languages at Oxford University Press. Over the past 25 years, Casper has led the transformation of Oxford’s language program from a publisher of print dictionaries into one of the premier language data suppliers in the world, partnering with global tech giants as well as local advocacy groups to build out digital resources in under-represented languages and help usher in a new era of responsible, authoritative AI.
Let’s connect on LinkedIn!


You must be logged in to post a comment.