Observed arrival · 2026-08-23
Khmer Corpus, Built Around Kambuja Suriya
A Khmer corpus-linguistics workbench for searching, comparing, segmenting, and analyzing texts from an 80-year scholarly archive.
Why it surfaced
The free archive contains 9,490,670 tokens across 61 issues of Cambodia’s Kambuja Suriya journal, covering 1926–2006. Its tools include KWIC concordance, collocation networks, keyness, POS tagging, named-entity recognition, IPA transcription, and syntax-tree visualization—an unusually concrete piece of language infrastructure.
A Khmer corpus-linguistics workbench for searching, comparing, segmenting, and analyzing texts from an 80-year scholarly archive.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue