pythainlp.corpus
The pythainlp.corpus module provides access to various Thai language corpora and resources that come bundled with PyThaiNLP. These resources are essential for natural language processing tasks in the Thai language.
Modules
countries
find_synonym
get_corpus
- pythainlp.corpus.get_corpus(filename: str, comments: bool = True) frozenset[str][source]
Read corpus data from a file and return a frozenset.
Each line in the file becomes a member of the set. Whitespace is stripped, and empty values and duplicates are removed.
If comments is False, any text at any position after the character “#” in each line is discarded.
- Parameters:
- Returns:
frozenset of lines in the file
- Return type:
- Example:
>>> from pythainlp.corpus import get_corpus >>> get_corpus("negations_th.txt") frozenset({'แต่', 'ไม่'}) >>> get_corpus("ttc_freq.txt") frozenset({'โดยนัยนี้\t1', 'ตัวบท\t10', ...}) >>> get_corpus("icubrk_th.txt") frozenset({'กกขนาก', '# Thai Dictionary for ICU BreakIterator', 'กก', ...}) >>> get_corpus("icubrk_th.txt", comments=False) frozenset({'กกขนาก', 'กก', ...})
get_corpus_as_is
get_corpus_db
- pythainlp.corpus.get_corpus_db(url: str) _ResponseWrapper | None[source]
Get the corpus catalog from a server.
Uses HTTPS with certificate validation enabled by default in Python’s urllib. Download a corpus catalog from trusted URLs only.
- Parameters:
url (str) – URL of the corpus catalog
- Returns:
response wrapper, or None if the request fails
- Return type:
Optional[pythainlp.corpus.core._ResponseWrapper]
get_corpus_db_detail
get_corpus_default_db
get_corpus_path
- pythainlp.corpus.get_corpus_path(name: str, version: str = '') str | None[source]
Get the local path of a corpus.
The function checks these locations in order:
Bundled (default) corpora shipped with PyThaiNLP.
The local download catalog (
~/pythainlp-data/).
When the corpus file is not present locally, the behavior depends on the
PYTHAINLP_OFFLINEenvironment variable:If
PYTHAINLP_OFFLINEis set to a truthy value (for example,"1"), the function raisesFileNotFoundErrorimmediately.Otherwise, the function downloads the corpus automatically.
- Parameters:
- Returns:
full local path if the corpus exists, or None if the corpus cannot be found or downloaded
- Return type:
Optional[str]
- Raises:
FileNotFoundError – if the corpus is missing locally and
PYTHAINLP_OFFLINEis set to a truthy value- Example:
(Please see the filename in this file)
If the corpus already exists:
>>> from pythainlp.corpus import get_corpus_path >>> get_corpus_path("ttc") '/root/pythainlp-data/ttc_freq.txt'
If the corpus has not been downloaded yet (online mode):
>>> get_corpus_path("wiki_lm_lstm") '/root/pythainlp-data/thwiki_model_lstm.pth'
To download manually:
>>> from pythainlp.corpus import download >>> download("wiki_lm_lstm") >>> get_corpus_path("wiki_lm_lstm") '/root/pythainlp-data/thwiki_model_lstm.pth'
download
- pythainlp.corpus.download(name: str, force: bool = False, url: str = '', version: str = '') bool[source]
Download a corpus.
The available corpus names are listed in this file: https://pythainlp.org/pythainlp-corpus/db.json
This function always performs the download regardless of the
PYTHAINLP_OFFLINEenvironment variable, because an explicit call todownload()is a deliberate user action.PYTHAINLP_OFFLINEonly blocks the automatic download triggered bypythainlp.corpus.get_corpus_path().By default, downloaded corpora and models are saved in
$HOME/pythainlp-data/(for example,/Users/bact/pythainlp-data/wiki_lm_lstm.pth).- Parameters:
- Returns:
True if the corpus is found and downloaded successfully, False otherwise
- Return type:
- Example:
>>> from pythainlp.corpus import download >>> download("wiki_lm_lstm", force=True) Corpus: wiki_lm_lstm - Downloading: wiki_lm_lstm 0.1 ...
remove
provinces
- pythainlp.corpus.provinces(details: bool = False) frozenset[str] | list[dict[str, str]][source]
Return a frozenset of Thailand province names in Thai.
Examples are “กระบี่”, “กรุงเทพมหานคร”, “กาญจนบุรี”, and “อุบลราชธานี”. See dev/pythainlp/corpus/thailand_provinces_th.csv.
- Parameters:
details (bool) – return details of provinces if True
- Returns:
frozenset of province names of Thailand (if details is False), or list of dict of province names and details such as
{'name_th': 'นนทบุรี', 'abbr_th': 'นบ', 'name_en': 'Nonthaburi', 'abbr_en': 'NBI'}(if details is True)- Return type:
thai_dict
thai_stopwords
- pythainlp.corpus.thai_stopwords() frozenset[str][source]
Return a frozenset of Thai stopwords.
Examples are “มี”, “ไป”, “ไง”, “ขณะ”, “การ”, and “ประการหนึ่ง”. See dev/pythainlp/corpus/stopwords_th.txt. The stopword list is from a thesis by เพ็ญศิริ ลี้ตระกูล.
thai_wikipedia_titles
- pythainlp.corpus.thai_wikipedia_titles() frozenset[str][source]
Return a frozenset of words from the Thai Wikipedia titles corpus.
The titles are mostly nouns and noun phrases, including event, organization, people, place, and product names. Commonly misspelled words are included intentionally.
See dev/pythainlp/corpus/wikipedia_titles_th.txt.
More info: https://github.com/PyThaiNLP/pythainlp/blob/dev/pythainlp/corpus/corpus_license.md
thai_words
thai_wsd_dict
thai_orst_words
thai_synonyms
thai_syllables
- pythainlp.corpus.thai_syllables() frozenset[str][source]
Return a frozenset of Thai syllables.
Examples are “กรอบ”, “ก็”, “๑”, “โมบ”, “โมน”, “โม่ง”, “กา”, “ก่า”, and “ก้า”. See dev/pythainlp/corpus/syllables_th.txt. The syllable list is from KUCut.
thai_negations
thai_family_names
thai_female_names
thai_male_names
pythainlp.corpus.th_en_translit.get_transliteration_dict
- pythainlp.corpus.th_en_translit.get_transliteration_dict() defaultdict[str, dict[str, list[str | bool | None]]][source]
Get the Thai to English transliteration dictionary.
The format is
dict[str, dict[str, list[Union[str, bool, None]]]].- Returns:
transliteration dictionary
- Return type:
collections.defaultdict[str, dict[str, list[Union[str, bool, None]]]]
- Raises:
FileNotFoundError – if the dictionary file is not found
ValueError – if the dictionary file cannot be parsed
TNC (Thai National Corpus) —
The Thai National Corpus (TNC) is a collection of text data in the Thai language. This module provides access to word frequency data from the TNC corpus.
pythainlp.corpus.tnc.word_freqs
- pythainlp.corpus.tnc.word_freqs() list[tuple[str, int]][source]
Get word frequency from the Thai National Corpus (TNC).
See dev/pythainlp/corpus/tnc_freq.txt.
- Returns:
list of tuples of word and frequency
- Return type:
- See Also:
Korakot Chaovavanich. https://www.facebook.com/groups/thainlp/posts/434330506948445
pythainlp.corpus.tnc.unigram_word_freqs
pythainlp.corpus.tnc.bigram_word_freqs
pythainlp.corpus.tnc.trigram_word_freqs
- pythainlp.corpus.tnc.trigram_word_freqs() dict[tuple[str, str, str], int][source]
Get trigram word frequency from the Thai National Corpus (TNC).
TTC (Thai Textbook Corpus) —
The Thai Textbook Corpus (TTC) is a collection of Thai language text data, primarily sourced from textbooks.
pythainlp.corpus.ttc.word_freqs
pythainlp.corpus.ttc.unigram_word_freqs
OSCAR
OSCAR is a multilingual corpus that includes Thai text data. This module provides access to word frequency data from the OSCAR corpus.
pythainlp.corpus.oscar.word_freqs
pythainlp.corpus.oscar.unigram_word_freqs
Util
Utilities for working with the corpus data.
pythainlp.corpus.util.find_badwords
pythainlp.corpus.util.revise_wordset
- pythainlp.corpus.util.revise_wordset(tokenize: Callable[[str], list[str]], orig_words: Iterable[str], training_data: Iterable[Iterable[str]]) set[str][source]
Revise a set of words to improve a dictionary-based tokenize function.
The function uses orig_words as a base set for the dictionary. It removes words that do not perform well with training_data and returns the remaining words.
- Parameters:
tokenize (Callable[[str], list[str]]) – tokenize function, which can be any function that takes text and returns a list of words
orig_words (Iterable[str]) – words used by the tokenize function, used as a base for the revision
training_data (Iterable[Iterable[str]]) – tokenized text, to be used as a training set
- Returns:
revised set of words with underperforming words removed
- Return type:
- Example:
>>> from pythainlp.corpus import thai_words >>> from pythainlp.corpus.util import revise_wordset >>> from pythainlp.tokenize.longest import segment >>> base_words = thai_words() >>> more_words = { ... "ถวิล อุดล", ... "ทองอินทร์ ภูริพัฒน์", ... "เตียง ศิริขันธ์", ... "จำลอง ดาวเรือง", ... } >>> base_words = base_words.union(more_words) >>> dict_trie = Trie(base_words) >>> tokenize = lambda text: segment(text, dict_trie) >>> training_data = [ ... ["word1", "word2"], ... ["word3", "word4"], ... ] >>> revised_words = revise_wordset( ... tokenize, base_words, training_data ... )
pythainlp.corpus.util.revise_newmm_default_wordset
- pythainlp.corpus.util.revise_newmm_default_wordset(training_data: Iterable[Iterable[str]]) set[str][source]
Revise the default word set to improve newmm tokenization.
newmm (
pythainlp.tokenize.newmm.segment()) is a dictionary-based tokenizer and the default tokenizer of PyThaiNLP.The function uses words from
pythainlp.corpus.thai_words()as a base set for the dictionary. It removes words that do not perform well with training_data and returns the remaining words.
WordNet
PyThaiNLP API includes the WordNet module, which is an exact copy of NLTK’s WordNet API for the Thai language. WordNet is a lexical database for English and other languages.
For more details on WordNet, refer to the NLTK WordNet documentation.
pythainlp.corpus.wordnet.synsets
pythainlp.corpus.wordnet.synset
pythainlp.corpus.wordnet.all_lemma_names
pythainlp.corpus.wordnet.all_synsets
pythainlp.corpus.wordnet.langs
pythainlp.corpus.wordnet.lemmas
pythainlp.corpus.wordnet.lemma
pythainlp.corpus.wordnet.lemma_from_key
pythainlp.corpus.wordnet.path_similarity
pythainlp.corpus.wordnet.lch_similarity
pythainlp.corpus.wordnet.wup_similarity
pythainlp.corpus.wordnet.morphy
pythainlp.corpus.wordnet.custom_lemmas
Definition
Synset
A synset is a set of synonyms that share a common meaning. The WordNet module provides functionality to work with these synsets.
This documentation is designed to help you navigate and use the various resources and modules available in the pythainlp.corpus package effectively. If you have any questions or need further assistance, please refer to the PyThaiNLP documentation or reach out to the PyThaiNLP community for support.
We hope you find this documentation helpful for your natural language processing tasks in the Thai language.