Source and licence
share: attributiongreekSource: corpus/incoming/perseus-greek-lexica/lexica-master.zip (staged 2026-09-10 by worker W-B1; provenance, terms as fetched and SHA-256 in corpus/incoming/perseus-greek-lexica/README.md)
Attribution: Text provided under a CC BY-SA license by Perseus Digital Library, http://www.perseus.tufts.edu, with funding from The National Endowment for the Humanities. Data accessed from https://github.com/PerseusDL/lexica/ [2026-09-10]. — the credit line the LSJ directory's README requires verbatim; the availability statement must be left intact and any modification offered back to Perseus.
Licence: CC-BY-SA 4.0 — Perseus Digital Library, github.com/PerseusDL/lexica, branch master fetched 2026-09-10; GitHub reports spdx_id CC-BY-SA-4.0 and the repository licence.md says "You must offer Perseus any modifications you make"
The text
5 lines · a line number is a link to itself| ref | L0 witness |
|---|---|
| 1 | 116,497 <entryFree> entries in 27 LSJ XML files (TEI P4) under CTS_XML_TEI/perseus/pdllex/grc/lsj/ — the Greek lexicon the loanword layer wants. ⚑ FILED AS A STUB, NOT AS TEXT, AND DELIBERATELY SO: THE HEADWORDS ARE IN BETA CODE, NOT UNICODE GREEK (47 Unicode-Greek characters in 271 MB). A real entry reads <orth lang="greek">*ii</orth>, <orth lang="greek">i)w=ta</orth> — i.e. `*ii` is Ἰ and `i)w=ta` is ἰῶτα. A BETA CODE → UNICODE CONVERTER IS THEREFORE A PREREQUISITE before this can be joined to corpus/greek_loanwords.txt or to the SBLGNT/Swete texts beside it, and writing that converter is work nobody has done: it is not attempted here, because converting 116,497 entries with an unverified table would put 116,497 unverified Greek spellings into the repository under CLAUDE.md §1.2. Only the Latin half of the repository was discarded at staging (161 MB, off-mission); nothing was deleted from the zip. |
| 2 | Staged files: |
| 3 | · corpus/incoming/perseus-greek-lexica/lexica-master.zip — 61,880,851 bytes |
| 4 | · corpus/incoming/perseus-greek-lexica/lexica-master — 30 files, 283,636,015 bytes |
| 5 | THE PAGES ARE NOT IN THIS REPOSITORY. This document is a provenance stub: the original is staged under Coptic OCR/corpus/incoming/, with its SHA-256 and the terms as fetched in the README beside it. The catalogue row's `words` count is this stub's, not the volume's. |
Provenance
- title
- Liddell–Scott–Jones, A Greek-English Lexicon (Perseus CTS TEI)
- dialect
- greek
- source
- corpus/incoming/perseus-greek-lexica/lexica-master.zip (staged 2026-09-10 by worker W-B1; provenance, terms as fetched and SHA-256 in corpus/incoming/perseus-greek-lexica/README.md)
- licence
- CC-BY-SA 4.0 — Perseus Digital Library, github.com/PerseusDL/lexica, branch master fetched 2026-09-10; GitHub reports spdx_id CC-BY-SA-4.0 and the repository licence.md says "You must offer Perseus any modifications you make"
- share
- attribution
- attribution
- Text provided under a CC BY-SA license by Perseus Digital Library, http://www.perseus.tufts.edu, with funding from The National Endowment for the Humanities. Data accessed from https://github.com/PerseusDL/lexica/ [2026-09-10]. — the credit line the LSJ directory's README requires verbatim; the availability statement must be left intact and any modification offered back to Perseus.
- file
- Reference/Dictionaries/Perseus LSJ — A Greek-English Lexicon/Liddell–Scott–Jones, A Greek-English Lexicon (Perseus CTS TEI).data.L0.txt
L0 witness
Cite as:
Liddell–Scott–Jones, A Greek-English Lexicon (Perseus CTS TEI). corpus/incoming/perseus-greek-lexica/lexica-master.zip (staged 2026-09-10 by worker W-B1; provenance, terms as fetched and SHA-256 in corpus/incoming/perseus-greek-lexica/README.md). Layers as published by the Kemetic project (Yousef Hanna), 2026: Coptic · Cairene · Kemetic, /doc/reference/dictionaries/perseus-lsj-a-greek-english-lexicon/liddell-scott-jones-a-greek-english-lexicon-perseus-cts-tei.data/. Text provided under a CC BY-SA license by Perseus Digital Library, http://www.perseus.tufts.edu, with funding from The National Endowment for the Humanities. Data accessed from https://github.com/PerseusDL/lexica/ [2026-09-10]. — the credit line the LSJ directory's README requires verbatim; the availability statement must be left intact and any modification offered back to Perseus.
What the layers are
The text as its witness or edition has it, original spelling; its dialect is a description and is never standardized.
For Egyptian texts: the transliteration with the editors' marks turned into the symbols ° * < > _ ^ (ruled 2026-09-10). For papyri: the edition's Leiden marks as the same symbols, and the editors' readings beside the scribe's as ‹scribe→editors› (ruled PAP-3).
Standardized Greco-Bohairic spelling and grammar, generated by code from the layer above under the rulings of the project.
The Kemetic alphabet, generated by code from L1.
Quoted translations where they exist, credited; otherwise model drafts, marked as drafts.