
Author:
Casper Grathwohl, President
There is a trap built into how most teams evaluate language data. A dataset can score well on standard evaluations, pass deduplication checks, and look clean in a schema, and still produce models that fail in ways that matter. Call it the good enough trap: the gap between data that passes tests and data that survives production.
Oxford Languages draws on more than 150 years of linguistic research and editorial expertise. Here’s what separates good language data from great language data, and why the difference compounds the longer a system runs.
The problem with “good enough”
Most language data in use today was assembled for volume first, quality second. The assumption is that scale compensates for noise. The research increasingly says otherwise.
A 2024 study at the North American Chapter of the Association for Computational Linguistics found that selecting for quality rather than quantity in instruction tuning substantially boosted model performance, challenging the assumption that more training data is always better. Careful data filtering can improve model performance to a degree that is often comparable to increasing compute.
The difference between good and great language data is structural. And it compounds the longer a system runs.
Good language data has a ceiling
Good language data is clean, deduplicated, and formatted for use. It performs well on standard benchmarks and integrates without major pipeline friction. For many applications, it works, right up until it doesn’t.
The problem is where it stops. Most datasets in circulation today are built through web-scale collection pipelines that prioritize:
- Volume over verification
- Coverage over consistency
- Speed over governance
That optimization for collection, not for what language actually requires, introduces problems that don’t show up in testing:
- Mislabelled or contextually incorrect data
- No traceability back to the source
- Bias embedded in the source material itself
These weaknesses are invisible in controlled benchmarks. They surface in production, where edge cases, domain-specific queries, and non-English inputs are the norm, not the exception.
Good data passes tests. Great data survives reality.
Great language data is intentionally built, not assembled

The defining difference lies in intent.
Great language data is not assembled from raw web text and filtered for noise. It is designed through a structured process that combines linguistic expertise, real-world evidence, and editorial governance.
At Oxford Languages, that process is anchored in three things:
- Large, continuously updated corpora capturing real-world usage across registers and contexts
- Expert analysis of linguistic patterns, meaning, and variation, carried out by trained lexicographers and linguists, not automated classifiers
- Editorial review cycles that validate accuracy and relevance before data is published or licensed
This ensures that data reflects how language is actually used, not how it appears in isolated, noisy, or unrepresentative sources. It is the difference between a dataset that describes language and one that accurately models it.
Five properties that define great language data
These aren’t marketing criteria. They are the properties that determine whether language data holds up when a model moves from benchmark to production. Each one marks a structural difference between data assembled at scale and data built with intent.
1. Linguistic integrity over surface fluency
A sentence can be grammatically correct and semantically wrong. Most large-scale datasets optimize for the former.
Fluency is easy to measure. Accurate meaning and correct usage within a given register are not. Great language data captures words in their correct grammatical and semantic context, not just their most frequent context.
This requires trained lexicographers. They distinguish a word’s core meaning from its incidental patterns in a corpus. Without this, models learn what language looks like around a concept, not what it actually means.
The practical cost is real. Lexical data for NLP applications that lack this level of precision will underperform in exactly the domains where accuracy matters most: medical, legal, and technical content.
2. Provenance is non-negotiable
If you cannot answer where your data came from, you cannot audit it, trust it, or defend its use in a regulated context.
This has moved from good practice to a legal requirement:
- The EU AI Act requires transparency documentation for training data in high-risk AI systems
- France’s CNIL issued guidelines in June 2025 requiring proof of data sourcing for AI training
- The OECD’s 2025 regulatory outlook flags IP risks when enterprises train on scraped text
Opaque datasets create compounding risk. If a source violates copyright or contains unlicensed personal data, every model trained on it is affected. The remediation cost is not captured in the original acquisition budget.
Our datasets are built from curated and licensed sources. That means full traceability and the legal certainty needed for long-term commercial and enterprise use.
3. Expert curation over automated shortcuts
Automation can collect language at scale. It cannot replace the judgment required to interpret it accurately.
Consider sense disambiguation, for instance. The word “bank” means something different to a geologist, a software engineer, and a mortgage officer. Automated systems learn co-occurrence patterns. Only trained lexicographers can assign structured sense labels that tell a model which meaning applies and why.
Skipping this step has documented consequences. A 2025 study in NPJ Digital Medicine tested four leading LLMs on psychiatric patient cases. Treatment recommendations changed materially when race was introduced. Models omitted medication recommendations or suggested guardianship based on racial identifiers alone. The root cause in every case was bias entering at the data layer. Automated pipelines cannot catch what they are not designed to look for.
Our lexicographers and language experts analyze corpus data to select examples in their correct grammatical and semantic context. They resolve ambiguity through structured editorial review. This is slower than automated collection, and it produces data that holds up where automated data does not.
4. Representation without distortion
Great data reflects language as it is used, not as it is simplified.
Many datasets claim global coverage and deliver English-dominant outputs with thin multilingual layers on top.
Volume does not solve the quality gap. Datasets built on machine-translated content deliver the surface form of a language, not the substance. They miss:
- Cultural context and community-specific usage
- Regional and dialectal variation
- Formal versus informal register distinctions
Models trained on this type of data produce outputs that are technically correct and contextually wrong. That failure is most costly in customer-facing, educational, and healthcare applications.
Oxford Languages collaborates with native-speaking expert consultants for all the languages it offers. This results in high-quality, accurate, and natural-sounding language data.
5. Continuous evolution, not static snapshots
Language changes constantly, new words emerge, meanings shift, and their usage evolves.
Static datasets become outdated quickly.
Great language data is maintained as a living system:
- Updated through ongoing corpus analysis
- Refined through editorial review
- Expanded to reflect new developments
Our dictionary datasets are built on exactly this kind of continuous research. The same process that has tracked the English language through centuries of change now applies to the data that language technology systems learn from.
How to evaluate a language dataset before you commit
Most data quality problems are not visible in a demo or a sample file. They surface after integration, when the model is running on real queries from real users. These questions are designed to expose weaknesses before that happens:
- Where did this data come from? Can the provider show a documented chain of sources and licences?
- Who curated it? Was there a human linguistic review, or was it filtered automatically? Oxford Dictionaries API provides human-curated lexical data to ensure highest quality.
- How is it maintained? Is it updated on a regular cycle, or is it a static release?
- Does it cover your domain? General web text and specialist language data are not the same thing.
- Does it represent your users? If your product serves non-English speakers, check how the non-English data was built and by whom.
A provider that cannot answer these questions clearly is signalling something about the data itself.
Our case studies show how organizations across voice AI, language learning, and enterprise search have applied these criteria in practice, and what they found when they moved from volume-first to quality-first data.
Data quality is infrastructure, not procurement
Good language data gets a model to production. Great language data keeps it performing there.
The organizations building the most reliable language applications treat data quality as an architectural decision, not a procurement one. What data a system trains on determines what it is capable of, and what it cannot do, for the lifetime of the deployment. That decision deserves the same rigour as any other infrastructure choice.
The good enough trap is easy to fall into. The cost of falling into it shows up later — in retraining cycles, in production failures, in users who quietly stop trusting the product. The teams that avoid it are the ones that asked harder questions about their data before deployment made those questions unavoidable.
If you are evaluating language data for an AI or NLP application, explore our full range of datasets, from lexical data for NLP to pronunciation data and multilingual datasets, or speak to our team about what your specific use case requires.

Casper Grathwohl is the President of Oxford Languages at Oxford University Press. Over the past 25 years, Casper has led the transformation of Oxford’s language program from a publisher of print dictionaries into one of the premier language data suppliers in the world, partnering with global tech giants as well as local advocacy groups to build out digital resources in under-represented languages and help usher in a new era of responsible, authoritative AI.
Let’s connect on LinkedIn!


You must be logged in to post a comment.