Funding the data
The commitment
Lector pays $100 per month, or 10% of profit, whichever is greater.
The floor applies now. Lector is not yet profitable, so the maintainer pays the floor from a personal giving budget. The percentage takes over when 10% of profit is more than $100 per month.
Profit means revenue after payment-processor fees, refunds, hosting and inference costs. This page publishes the payments. It does not publish the accounts.
Lector pays once per year, in a single transfer to each recipient. Bank fees and administration make small monthly transfers wasteful.
Where it goes
The annual total at the floor is $1,200. That splits three ways, not eight. A share of $1,200across every project on the list is too small. The administration then costs more than the money, for both sides.
Every language pack draws its example sentences from Tatoeba. Association Tatoeba is a French non-profit and donations pay for it.
Bank transfer to Association Tatoeba
Some projects Lector depends on have no legal entity and cannot take a donation. eSpeak NG and jieba are two. Paid work on them is possible instead.
Paid work on upstream projects
kaikki.org supplies the dictionaries for most packs. It asks for nothing, so Lector asked what would help.
Offered to the maintainer directly
What Lector paid
Total paid to date: $600. The commitment is set in US dollars. A recipient abroad takes its own currency, so the row shows both figures. Each receipt is the evidence for one row.
Who cannot take money
Lector depends on these projects and cannot pay them. Each entry gives the reason. This list matters as much as the one above, because it shows where the gaps in open-data funding actually are.
Anki
Anki runs as a business, and its FAQ says that it cannot easily accept donations. It asks people to buy AnkiMobile instead. Lector buys licences as a normal expense and does not count them in the 10%.
wordfreq
The frequency bands in every pack come from wordfreq. Robyn Speer ended the project, because generated text polluted the web corpora it sampled. No work remains to fund.
Wikimedia
Wiktionary and Wikipedia dumps feed the dictionaries and the frequency blends. The Wikimedia Foundation is very well funded, so a payment there does the least good of any option here. Lector corrects Wiktionary entries instead, through the directed work fund.
Universities and standards bodies
MeCab and UniDic come from NINJAL. UDPipe comes from Charles University. OPUS comes from the University of Helsinki. The Unicode Consortium supplies ICU, CLDR and Unihan. Their gift processes cost more in administration than these amounts justify today. They become candidates when the percentage beats the floor.
One problem money cannot fix
The frequency data has no upstream any more. wordfreq is finished, and the open web corpora it drew on are degraded. Every pack's frequency banding rests on a source that stopped. Lector must find another source or build one, and no payment changes that.
Corrections
If Lector uses your data and this page is wrong about it, open an issue. If you know a better way to reach a project named here, open an issue.