Loading Ahlan…
Loading Ahlan…
Every source Ahlan’s offline translation corpus draws from, with its license.
Ahlan’s offline translator ships a corpus of dictionary entries, idioms, and parallel sentences drawn from the following open language-data sources. Each source is redistributed here under its own upstream license, unchanged in substance and with its license text preserved with every record. Where a license requires attribution (CC-BY, CC-BY-SA, WordNet), this page is the notice.
Corpus records are checked at ingest against a hard allow-list that only accepts licenses which unambiguously permit commercial redistribution and modification. GPL / LGPL / AGPL, non-commercial (NC) licenses, and any source shipping without a written license grant are refused.
Locale metadata (not translation pairs): language autonyms, script names, cardinal plural rules, number-formatting symbols, date-field display names, ISO 4217 currency names, and ISO 3166-1 country / UN M.49 region names. Powers accurate locale-aware UX (language picker autonyms, plural rendering, number/date/currency formatting, country picker) without leaking language-name records into the translation lookup path.
Coverage: 174 locales
Frequency-ranked wordlists derived from OpenSubtitles corpora. Top 5,000 most-common words per language, used for baseline frequency data.
Coverage: ~50 languages
Cross-lingual word translations extracted from Wiktionary.
Coverage: many ↔ many
Aligned WordNets across dozens of languages. Each per-language WordNet retains its own license as recorded on every record.
Coverage: many
Professionally translated parallel data for under-represented languages: GATITOS token/short-phrase lexicon, SmolSent sentence translations, and SmolDoc sentence-aligned document translations. First-ever data for several registry languages (Filipino, Maithili, Waray, Kabardian, Santali, Dogri, Fula, Quechua, and more).
Coverage: 71 registry languages ↔ en
Morphological analyses, lemmas, and usage examples across 100+ languages. Includes the UD-PUD parallel corpus.
Coverage: many
Structured extract of Wiktionary. Currently ingested: per-language pos-{phrase, proverb, intj, adv, conj, num, verb, adj} extracts for English (with cross-language translation tables covering 169 target languages including everyday verbs, adjectives, adverbs, conjunctions, and numerals); per-language pos-{phrase, proverb, intj} extracts for 32 native-language sources (French, Spanish, German, Italian, Portuguese, Japanese, Russian, Chinese, Arabic, Hindi, Korean, Turkish, Vietnamese, Thai, Polish, Dutch, Swedish, Norwegian, Finnish, Danish, Ukrainian, Hebrew, Persian, Indonesian, Malay, Czech, Romanian, Hungarian, Greek, Bulgarian, Slovak, Slovene); full-dictionary English gloss pairs for 63 low-resource languages (Magahi, Komi-Zyrian, Sanskrit, Amharic, Yoruba, Xhosa, Zulu, Hausa, Tigrinya, Punjabi, Nepali, Tibetan, Gulf/Moroccan/Egyptian Arabic, Swahili, Scottish Gaelic, Ilocano, Breton, Turkmen, Avar, and 43 more); and idiom + proverb records extracted from those same full dictionaries (first idioms/proverbs ever for Cebuano, Yoruba, Maltese, Gujarati, Chichewa, Gulf Arabic, and 26 other languages).
Coverage: 169 target languages via English + 32 native-language sources
Records under CC-BY-SA or its earlier versions retain their share-alike terms when redistributed. Downstream users of Ahlan’s corpus files must preserve the same license on those records and provide the same attribution.
If you are the maintainer of a source listed here and would like to update the attribution or request removal, please contact us via the details on the Legal Contact page.
Last updated: August 2026. Corpus contents and their licenses are enforced by scripts/claude-license-guard.mjs on every pull request.