🇨🇳 中文 reference data

Fully supported — read, tap-to-define, mine vocabulary, and practice frequency-banded cloze in Mandarin. Chinese writes no spaces, so the reader splits each sentence into words before you tap one: 我喜欢读书 reads as 我 · 喜欢 · 读书, not five loose characters. Because a character does not show how it sounds, pinyin sits above every word while you read, and each reading retires once you mark that word known. Entries are keyed on Simplified, and Traditional resolves to the same entry, so 這 and 这 both answer zhè.

What's inside

145,875 entriesDictionary
220,906English senses
145,361Pinyin readings
7,967 sentencesCloze bank

Provenance & licensing

Built entirely from open data — nothing here is derived from copyrighted dictionaries. See the methodology for how frequency drives vocabulary learning in Lector.

  • Dictionary and pinyin: Chinese Wiktionary entries extracted through Kaikki.org. Every entry carries its Standard Mandarin pinyin, and Traditional headwords are filed as aliases of the Simplified form they convert to. CC BY-SA 4.0
  • Cloze sentences: Mandarin–English sentence pairs from Tatoeba, segmented with jieba, frequency-ranked with wordfreq and filtered against the on-device dictionary. Traditional-script rows are excluded. CC BY 2.0 FR
  • Script conversion and readings: OpenCC decides the Traditional-to-Simplified key for every entry, and pypinyin ranks the reading a character takes when it has several. Both run at build time, so neither ships in the app. Apache-2.0 (OpenCC), MIT (pypinyin)