lector / funding

Funding the data

Lector is a reader, but the language packs are the product. Volunteers and small research groups built the data in those packs. Almost none of them get paid for it.

The commitment

Lector pays $100 per month, or 10% of profit, whichever is greater.

The floor applies now. Lector is not yet profitable, so the maintainer pays the floor from a personal giving budget. The percentage takes over when 10% of profit is more than $100 per month.

Profit means revenue after payment-processor fees, refunds, hosting and inference costs. This page publishes the payments. It does not publish the accounts.

Lector pays once per year, in a single transfer to each recipient. Bank fees and administration make small monthly transfers wasteful.

Where it goes

The annual total at the floor is $1,200. That splits three ways, not eight. A share of $1,200across every project on the list is too small. The administration then costs more than the money, for both sides.

Every language pack draws its example sentences from Tatoeba. Association Tatoeba is a French non-profit and donations pay for it.

Bank transfer to Association Tatoeba

30%Directed work fund

Some projects Lector depends on have no legal entity and cannot take a donation. eSpeak NG and jieba are two. Paid work on them is possible instead.

Paid work on upstream projects

kaikki.org supplies the dictionaries for most packs. It asks for nothing, so Lector asked what would help.

Offered to the maintainer directly

What Lector paid

Total paid to date: $600. The commitment is set in US dollars. A recipient abroad takes its own currency, so the row shows both figures. Each receipt is the evidence for one row.

DateRecipientAmountMethodEvidence
22 August 2026TatoebaThe example sentences in every language pack. Association Tatoeba is a French non-profit.$600paid as €513.78 EURPayPal9RH44382EK247412GReceipt

Who cannot take money

Lector depends on these projects and cannot pay them. Each entry gives the reason. This list matters as much as the one above, because it shows where the gaps in open-data funding actually are.

Anki

Anki runs as a business, and its FAQ says that it cannot easily accept donations. It asks people to buy AnkiMobile instead. Lector buys licences as a normal expense and does not count them in the 10%.

Anki's donation FAQ

wordfreq

The frequency bands in every pack come from wordfreq. Robyn Speer ended the project, because generated text polluted the web corpora it sampled. No work remains to fund.

Why wordfreq will not be updated

Wikimedia

Wiktionary and Wikipedia dumps feed the dictionaries and the frequency blends. The Wikimedia Foundation is very well funded, so a payment there does the least good of any option here. Lector corrects Wiktionary entries instead, through the directed work fund.

Universities and standards bodies

MeCab and UniDic come from NINJAL. UDPipe comes from Charles University. OPUS comes from the University of Helsinki. The Unicode Consortium supplies ICU, CLDR and Unihan. Their gift processes cost more in administration than these amounts justify today. They become candidates when the percentage beats the floor.

One problem money cannot fix

The frequency data has no upstream any more. wordfreq is finished, and the open web corpora it drew on are degraded. Every pack's frequency banding rests on a source that stopped. Lector must find another source or build one, and no payment changes that.

Tracked as issue #543

Corrections

If Lector uses your data and this page is wrong about it, open an issue. If you know a better way to reach a project named here, open an issue.

FUNDING.md in the repository