Pecha Lab

Open, machine-readable texts of the Tibetan Buddhist canon.

般若译藏

Mission

Why this exists

A pecha (པེ་ཆ་) is a traditional Tibetan book — long, loose leaves of scripture, copied and recopied by hand for a thousand years. Pecha Lab is an independent, non-commercial initiative that carries that work into open data: complete, validated, machine-readable editions of the Tibetan Buddhist canon, freely usable by anyone.

The texts belong to all Tibetan Buddhist traditions equally. We serve two audiences at once: communities and scholars who need trustworthy digital access to the canon, and researchers building natural-language-processing tools for Tibetan — a language of millions of speakers that remains severely under-resourced in modern computing. Careful open data advances both preservation and research.

The Dataset

The complete Derge canon, in Unicode

Our first release is a validated Unicode corpus of the complete Derge (sde dge) edition — both the Kangyur (the Buddha-word) and the Tengyur (the treatises) — built from the Buddhist Digital Resource Center’s public catalogue and etext archive, segmented into sentence-like units ready for research use.

4,484works
236,122text segments
288M+Tibetan characters
CompleteDerge Kangyur + Tengyur

Kangyur: 1,114 of 1,114 works — including 100 works recovered from a clean parallel digitization of the same edition after our validation gate flagged legacy-font corruption. Tengyur: 3,370 works. Every published text passes strict Unicode validation.

Method

How it is built

Roadmap

What comes next

Contact

Get in touch

Questions, corrections, or collaboration ideas: hello@pechalab.com