
Author:
Nilanjana Banerji, Product Manager
Listening to the market
At Oxford Languages we provide lexical and language datasets for a wide range of technologies and applications. We offer dictionary data in more than 60 languages, and these are made up of a number of different components. Providing dictionary content in digital format, in XML or via API, that supports integrated dictionary experiences in operating systems, eReaders, educational software, accessibility tools, and more.
This has become the preferred medium to find the meaning of a word, or synonyms etc., over printed dictionaries. Our dictionary-derived datasets have grown into a product line because, in listening to the market, we have realized not everyone needs the complex structure or robust metadata that a full dictionary dataset provides.
One example of this is licensing our lexical content to word game developers in a format made for their specific requirements, the Word Games Developer Toolkit.
But first, how a word game works
There are two use cases for word games:
- Word validation
- Word selection
Word validation ensures that a word played by a player in a game is a real word, as in Scrabble. Most games display a short definition to prove the validity of any word that is challenged, and so developers require a large list of words and definitions to be able to validate the words in the game.
Games such as wordsearches or games where users have to guess a five-letter word, already have pre-selected words in the background as answers that the player is looking for, which requires word selection. For that, a developer would need additional metadata such as wordform frequencies to be able to select familiar or unfamiliar words according to difficulty levels. They might want to exclude words like proper nouns, abbreviations, or vulgar and offensive words. Some of this data exists in our dictionary, but there’s also additional information in our dictionary data that word game developers don’t need, like etymologies or example sentences.
What goes into building a word game dataset? Behind the scenes of our process
In conversation with mobile game developers, we came up with the idea of selecting relevant features from our dictionary databases and then adding in additional features to deliver it all in one easy format for developers to use.
We curated and assembled the “right data” from multiple sources: our existing monolingual dictionaries and morphologies, wordform frequencies from language corpora, and newly generated metadata for proper nouns, abbreviations, and sensitive words, to help developers create appropriate, family-friendly games.
We worked with our data and engineering teams to make sure that our product is suitable for developers to use and with editors and native speakers to ensure accuracy and appropriateness of the content, particularly with sensitivity in mind.
The Word Games Developer Toolkit was initially created in five languages: French, German, Italian, Spanish, and Portuguese, following on with British English and American English as requested. With product insights developed through customer feedback and market discovery, we were able to transform this product rapidly into a developer-focused, market-ready solution with standard features and word coverage that would work across a range of word games.
Customer success stories
With the toolkit, we’re able to support customers’ wide range of requirements for their games, including word validation and definitions. We value all such partnerships and were happy to hear UniWords CEO Alexander Stobbe say:
“We could have found a word list in the public domain. However, such a word list would not inspire the same confidence as data provided by Oxford Languages.”
Read our insightful case study with UniWords and learn more about how innovative developers are reimagining the classic word game experience by blending learning, strategy, and play through an intuitive app.
If you’re looking for authoritative language data to power word games and puzzle experiences, explore the Word Games Developer Toolkit and access curated lexical resources designed for developers, or get in touch with our team to find the right solution for your product.

Nilanjana Banerji is a Product Manager at Oxford Languages, specializing in dictionaries and language data. She has over fifteen years’ experience developing lexical datasets and digital language products used in education, assessment, digital literacy, and education technology platforms. Her work bridges lexicography, language research, and product strategy, with a focus on making high‑quality language data accessible and impactful in modern digital and AI‑enabled contexts.
With a background in corpus research and a DPhil from the University of Oxford, Nilanjana brings a strong research‑informed perspective to product development.
Let’s connect on LinkedIn!



You must be logged in to post a comment.