pythainlp.corpus

The pythainlp.corpus module provides access to various Thai language corpora and resources that come bundled with PyThaiNLP. These resources are essential for natural language processing tasks in the Thai language.

Modules

countries

pythainlp.corpus.countries() → frozenset[str][source]

Return a frozenset of country names in Thai.

Examples are “แคนาดา”, “โรมาเนีย”, “แอลจีเรีย”, and “ลาว”. See dev/pythainlp/corpus/countries_th.txt.

Returns:

frozenset of country names in Thai

Return type:

frozenset[str]

find_synonym

get_corpus

pythainlp.corpus.get_corpus(filename: str, comments: bool = True) → frozenset[str][source]

Read corpus data from a file and return a frozenset.

Each line in the file becomes a member of the set. Whitespace is stripped, and empty values and duplicates are removed.

If comments is False, any text at any position after the character “#” in each line is discarded.

Parameters:
  • filename (str) – filename of the corpus to be read

  • comments (bool) – keep comments

Returns:

frozenset of lines in the file

Return type:

frozenset[str]

Example:
>>> from pythainlp.corpus import get_corpus
>>> get_corpus("negations_th.txt")
frozenset({'แต่', 'ไม่'})
>>> get_corpus("ttc_freq.txt")
frozenset({'โดยนัยนี้\t1', 'ตัวบท\t10', ...})
>>> get_corpus("icubrk_th.txt")
frozenset({'กกขนาก', '# Thai Dictionary for ICU BreakIterator', 'กก', ...})
>>> get_corpus("icubrk_th.txt", comments=False)
frozenset({'กกขนาก', 'กก', ...})

get_corpus_as_is

pythainlp.corpus.get_corpus_as_is(filename: str) → list[str][source]

Read corpus data from a file as it is and return a list.

Each line in the file becomes a member of the list. Member values and their order are not modified. To strip whitespace or remove comments, use get_corpus() instead.

Parameters:

filename (str) – filename of the corpus to be read

Returns:

list of lines in the file

Return type:

list[str]

Example:
>>> from pythainlp.corpus import get_corpus_as_is
>>> get_corpus_as_is("negations_th.txt")
['แต่', 'ไม่']

get_corpus_db

pythainlp.corpus.get_corpus_db(url: str) → _ResponseWrapper | None[source]

Get the corpus catalog from a server.

Uses HTTPS with certificate validation enabled by default in Python’s urllib. Download a corpus catalog from trusted URLs only.

Parameters:

url (str) – URL of the corpus catalog

Returns:

response wrapper, or None if the request fails

Return type:

Optional[pythainlp.corpus.core._ResponseWrapper]

get_corpus_db_detail

pythainlp.corpus.get_corpus_db_detail(name: str, version: str = '') → dict[str, Any][source]

Get details about a corpus from the local catalog.

Parameters:
  • name (str) – corpus name

  • version (str) – corpus version (empty string means any version)

Returns:

details about the corpus, or an empty dict if not found

Return type:

dict[str, Any]

get_corpus_default_db

pythainlp.corpus.get_corpus_default_db(name: str, version: str = '') → str | None[source]

Get the corpus path from default_db.json.

To edit default_db.json, edit pythainlp/corpus/default_db.json.

Parameters:
  • name (str) – corpus name

  • version (str) – corpus version (empty string means latest)

Returns:

path to the corpus, or None if the corpus does not exist on the device

Return type:

Optional[str]

get_corpus_path

pythainlp.corpus.get_corpus_path(name: str, version: str = '') → str | None[source]

Get the local path of a corpus.

The function checks these locations in order:

  1. Bundled (default) corpora shipped with PyThaiNLP.

  2. The local download catalog (~/pythainlp-data/).

When the corpus file is not present locally, the behavior depends on the PYTHAINLP_OFFLINE environment variable:

  • If PYTHAINLP_OFFLINE is set to a truthy value (for example, "1"), the function raises FileNotFoundError immediately.

  • Otherwise, the function downloads the corpus automatically.

Parameters:
  • name (str) – corpus name

  • version (str) – corpus version (empty string means latest)

Returns:

full local path if the corpus exists, or None if the corpus cannot be found or downloaded

Return type:

Optional[str]

Raises:

FileNotFoundError – if the corpus is missing locally and PYTHAINLP_OFFLINE is set to a truthy value

Example:

(Please see the filename in this file)

If the corpus already exists:

>>> from pythainlp.corpus import get_corpus_path
>>> get_corpus_path("ttc")
'/root/pythainlp-data/ttc_freq.txt'

If the corpus has not been downloaded yet (online mode):

>>> get_corpus_path("wiki_lm_lstm")
'/root/pythainlp-data/thwiki_model_lstm.pth'

To download manually:

>>> from pythainlp.corpus import download
>>> download("wiki_lm_lstm")
>>> get_corpus_path("wiki_lm_lstm")
'/root/pythainlp-data/thwiki_model_lstm.pth'

download

pythainlp.corpus.download(name: str, force: bool = False, url: str = '', version: str = '') → bool[source]

Download a corpus.

The available corpus names are listed in this file: https://pythainlp.org/pythainlp-corpus/db.json

This function always performs the download regardless of the PYTHAINLP_OFFLINE environment variable, because an explicit call to download() is a deliberate user action. PYTHAINLP_OFFLINE only blocks the automatic download triggered by pythainlp.corpus.get_corpus_path().

By default, downloaded corpora and models are saved in $HOME/pythainlp-data/ (for example, /Users/bact/pythainlp-data/wiki_lm_lstm.pth).

Parameters:
  • name (str) – corpus name

  • force (bool) – force the download

  • url (str) – URL of the corpus catalog

  • version (str) – corpus version (empty string means latest)

Returns:

True if the corpus is found and downloaded successfully, False otherwise

Return type:

bool

Example:
>>> from pythainlp.corpus import download
>>> download("wiki_lm_lstm", force=True)
Corpus: wiki_lm_lstm
- Downloading: wiki_lm_lstm 0.1
...

remove

pythainlp.corpus.remove(name: str) → bool[source]

Remove a corpus.

Parameters:

name (str) – corpus name

Returns:

True if the corpus is found and removed successfully, False otherwise

Return type:

bool

Example:
>>> from pythainlp.corpus import (
...     remove,
...     get_corpus_path,
... )
>>> remove("ttc")
True
>>> get_corpus_path("ttc")
None

provinces

pythainlp.corpus.provinces(details: bool = False) → frozenset[str] | list[dict[str, str]][source]

Return a frozenset of Thailand province names in Thai.

Examples are “กระบี่”, “กรุงเทพมหานคร”, “กาญจนบุรี”, and “อุบลราชธานี”. See dev/pythainlp/corpus/thailand_provinces_th.csv.

Parameters:

details (bool) – return details of provinces if True

Returns:

frozenset of province names of Thailand (if details is False), or list of dict of province names and details such as {'name_th': 'นนทบุรี', 'abbr_th': 'นบ', 'name_en': 'Nonthaburi', 'abbr_en': 'NBI'} (if details is True)

Return type:

Union[frozenset[str], list[dict[str, str]]]

thai_dict

pythainlp.corpus.thai_dict() → dict[str, list[str]][source]

Return a Thai dictionary with definitions from Wiktionary.

See thai_dict.

Returns:

Thai words with part-of-speech (POS) type and definition

Return type:

dict[str, list[str]]

thai_stopwords

pythainlp.corpus.thai_stopwords() → frozenset[str][source]

Return a frozenset of Thai stopwords.

Examples are “มี”, “ไป”, “ไง”, “ขณะ”, “การ”, and “ประการหนึ่ง”. See dev/pythainlp/corpus/stopwords_th.txt. The stopword list is from a thesis by เพ็ญศิริ ลี้ตระกูล.

See Also:

เพ็ญศิริ ลี้ตระกูล. การเลือกประโยคสำคัญในการสรุปความภาษาไทยโดยใช้แบบจำลองแบบลำดับชั้น. กรุงเทพมหานคร : มหาวิทยาลัยธรรมศาสตร์; 2551.

Returns:

frozenset of Thai stopwords

Return type:

frozenset[str]

thai_wikipedia_titles

pythainlp.corpus.thai_wikipedia_titles() → frozenset[str][source]

Return a frozenset of words from the Thai Wikipedia titles corpus.

The titles are mostly nouns and noun phrases, including event, organization, people, place, and product names. Commonly misspelled words are included intentionally.

See dev/pythainlp/corpus/wikipedia_titles_th.txt.

More info: https://github.com/PyThaiNLP/pythainlp/blob/dev/pythainlp/corpus/corpus_license.md

Returns:

frozenset of Thai words

Return type:

frozenset[str]

thai_words

pythainlp.corpus.thai_words() → frozenset[str][source]

Return a frozenset of Thai words.

Examples are “กติกา”, “กดดัน”, “พิษ”, and “พิษภัย”. See dev/pythainlp/corpus/words_th.txt.

Returns:

frozenset of Thai words

Return type:

frozenset[str]

thai_wsd_dict

pythainlp.corpus.thai_wsd_dict() → dict[str, list[str] | list[list[str]]][source]

Return a Thai word sense disambiguation dictionary.

The definitions are from Wiktionary. See thai_dict.

Returns:

Thai words with their senses

Return type:

dict[str, Union[list[str], list[list[str]]]]

thai_orst_words

pythainlp.corpus.thai_orst_words() → frozenset[str][source]

Return a frozenset of Thai words from the Royal Society of Thailand.

See dev/pythainlp/corpus/orst_words_th.txt.

Returns:

frozenset of Thai words

Return type:

frozenset[str]

thai_synonyms

pythainlp.corpus.thai_synonyms() → dict[str, list[str] | list[list[str]]][source]

Return Thai synonyms.

See thai_synonym.

Returns:

Thai words with part-of-speech (POS) type and synonyms

Return type:

dict[str, Union[list[str], list[list[str]]]]

thai_syllables

pythainlp.corpus.thai_syllables() → frozenset[str][source]

Return a frozenset of Thai syllables.

Examples are “กรอบ”, “ก็”, “๑”, “โมบ”, “โมน”, “โม่ง”, “กา”, “ก่า”, and “ก้า”. See dev/pythainlp/corpus/syllables_th.txt. The syllable list is from KUCut.

Returns:

frozenset of Thai syllables

Return type:

frozenset[str]

thai_negations

pythainlp.corpus.thai_negations() → frozenset[str][source]

Return a frozenset of Thai negation words.

Examples are “ไม่” and “แต่”. See dev/pythainlp/corpus/negations_th.txt.

Returns:

frozenset of Thai negation words

Return type:

frozenset[str]

thai_family_names

pythainlp.corpus.thai_family_names() → frozenset[str][source]

Return a frozenset of Thai family names.

See dev/pythainlp/corpus/family_names_th.txt.

Returns:

frozenset of Thai family names

Return type:

frozenset[str]

thai_female_names

pythainlp.corpus.thai_female_names() → frozenset[str][source]

Return a frozenset of Thai female names.

See dev/pythainlp/corpus/person_names_female_th.txt.

Returns:

frozenset of Thai female names

Return type:

frozenset[str]

thai_male_names

pythainlp.corpus.thai_male_names() → frozenset[str][source]

Return a frozenset of Thai male names.

See dev/pythainlp/corpus/person_names_male_th.txt.

Returns:

frozenset of Thai male names

Return type:

frozenset[str]

pythainlp.corpus.th_en_translit.get_transliteration_dict

pythainlp.corpus.th_en_translit.get_transliteration_dict() → defaultdict[str, dict[str, list[str | bool | None]]][source]

Get the Thai to English transliteration dictionary.

The format is dict[str, dict[str, list[Union[str, bool, None]]]].

Returns:

transliteration dictionary

Return type:

collections.defaultdict[str, dict[str, list[Union[str, bool, None]]]]

Raises:

TNC (Thai National Corpus) —

The Thai National Corpus (TNC) is a collection of text data in the Thai language. This module provides access to word frequency data from the TNC corpus.

pythainlp.corpus.tnc.word_freqs

pythainlp.corpus.tnc.word_freqs() → list[tuple[str, int]][source]

Get word frequency from the Thai National Corpus (TNC).

See dev/pythainlp/corpus/tnc_freq.txt.

Returns:

list of tuples of word and frequency

Return type:

list[tuple[str, int]]

See Also:

pythainlp.corpus.tnc.unigram_word_freqs

pythainlp.corpus.tnc.unigram_word_freqs() → dict[str, int][source]

Get unigram word frequency from the Thai National Corpus (TNC).

Returns:

dict mapping words to their frequencies

Return type:

dict[str, int]

pythainlp.corpus.tnc.bigram_word_freqs

pythainlp.corpus.tnc.bigram_word_freqs() → dict[tuple[str, str], int][source]

Get bigram word frequency from the Thai National Corpus (TNC).

Returns:

dict mapping word pairs to their frequencies

Return type:

dict[tuple[str, str], int]

pythainlp.corpus.tnc.trigram_word_freqs

pythainlp.corpus.tnc.trigram_word_freqs() → dict[tuple[str, str, str], int][source]

Get trigram word frequency from the Thai National Corpus (TNC).

Returns:

dict mapping word triples to their frequencies

Return type:

dict[tuple[str, str, str], int]

TTC (Thai Textbook Corpus) —

The Thai Textbook Corpus (TTC) is a collection of Thai language text data, primarily sourced from textbooks.

pythainlp.corpus.ttc.word_freqs

pythainlp.corpus.ttc.word_freqs() → list[tuple[str, int]][source]

Get word frequency from the Thai Textbook Corpus (TTC).

See dev/pythainlp/corpus/ttc_freq.txt.

Returns:

list of tuples of word and frequency

Return type:

list[tuple[str, int]]

pythainlp.corpus.ttc.unigram_word_freqs

pythainlp.corpus.ttc.unigram_word_freqs() → dict[str, int][source]

Get unigram word frequency from the Thai Textbook Corpus (TTC).

Returns:

dict mapping words to their frequencies

Return type:

dict[str, int]

OSCAR

OSCAR is a multilingual corpus that includes Thai text data. This module provides access to word frequency data from the OSCAR corpus.

pythainlp.corpus.oscar.word_freqs

pythainlp.corpus.oscar.word_freqs() → list[tuple[str, int]][source]

Get word frequency from OSCAR Corpus (words tokenized using ICU).

Returns:

list of tuples of word and frequency

Return type:

list[tuple[str, int]]

pythainlp.corpus.oscar.unigram_word_freqs

pythainlp.corpus.oscar.unigram_word_freqs() → dict[str, int][source]

Get unigram word frequency from OSCAR Corpus (words tokenized using ICU).

Returns:

dict mapping words to their frequencies

Return type:

dict[str, int]

Util

Utilities for working with the corpus data.

pythainlp.corpus.util.find_badwords

pythainlp.corpus.util.find_badwords(tokenize: Callable[[str], list[str]], training_data: Iterable[Iterable[str]]) → set[str][source]

Find words that do not work well with a tokenize function.

Parameters:
  • tokenize (Callable[[str], list[str]]) – tokenize function

  • training_data (Iterable[Iterable[str]]) – tokenized text, to be used as a training set

Returns:

words that do not work well with the tokenize function

Return type:

set[str]

pythainlp.corpus.util.revise_wordset

pythainlp.corpus.util.revise_wordset(tokenize: Callable[[str], list[str]], orig_words: Iterable[str], training_data: Iterable[Iterable[str]]) → set[str][source]

Revise a set of words to improve a dictionary-based tokenize function.

The function uses orig_words as a base set for the dictionary. It removes words that do not perform well with training_data and returns the remaining words.

Parameters:
  • tokenize (Callable[[str], list[str]]) – tokenize function, which can be any function that takes text and returns a list of words

  • orig_words (Iterable[str]) – words used by the tokenize function, used as a base for the revision

  • training_data (Iterable[Iterable[str]]) – tokenized text, to be used as a training set

Returns:

revised set of words with underperforming words removed

Return type:

set[str]

Example:
>>> from pythainlp.corpus import thai_words
>>> from pythainlp.corpus.util import revise_wordset
>>> from pythainlp.tokenize.longest import segment
>>> base_words = thai_words()
>>> more_words = {
...     "ถวิล อุดล",
...     "ทองอินทร์ ภูริพัฒน์",
...     "เตียง ศิริขันธ์",
...     "จำลอง ดาวเรือง",
... }
>>> base_words = base_words.union(more_words)
>>> dict_trie = Trie(base_words)
>>> tokenize = lambda text: segment(text, dict_trie)
>>> training_data = [
...     ["word1", "word2"],
...     ["word3", "word4"],
... ]
>>> revised_words = revise_wordset(
...     tokenize, base_words, training_data
... )

pythainlp.corpus.util.revise_newmm_default_wordset

pythainlp.corpus.util.revise_newmm_default_wordset(training_data: Iterable[Iterable[str]]) → set[str][source]

Revise the default word set to improve newmm tokenization.

newmm (pythainlp.tokenize.newmm.segment()) is a dictionary-based tokenizer and the default tokenizer of PyThaiNLP.

The function uses words from pythainlp.corpus.thai_words() as a base set for the dictionary. It removes words that do not perform well with training_data and returns the remaining words.

Parameters:

training_data (Iterable[Iterable[str]]) – tokenized text, to be used as a training set

Returns:

revised set of words with underperforming words removed

Return type:

set[str]

WordNet

PyThaiNLP API includes the WordNet module, which is an exact copy of NLTK’s WordNet API for the Thai language. WordNet is a lexical database for English and other languages.

For more details on WordNet, refer to the NLTK WordNet documentation.

pythainlp.corpus.wordnet.synsets

pythainlp.corpus.wordnet.synset

pythainlp.corpus.wordnet.all_lemma_names

pythainlp.corpus.wordnet.all_synsets

pythainlp.corpus.wordnet.langs

pythainlp.corpus.wordnet.lemmas

pythainlp.corpus.wordnet.lemma

pythainlp.corpus.wordnet.lemma_from_key

pythainlp.corpus.wordnet.path_similarity

pythainlp.corpus.wordnet.lch_similarity

pythainlp.corpus.wordnet.wup_similarity

pythainlp.corpus.wordnet.morphy

pythainlp.corpus.wordnet.custom_lemmas

Definition

Synset

A synset is a set of synonyms that share a common meaning. The WordNet module provides functionality to work with these synsets.

This documentation is designed to help you navigate and use the various resources and modules available in the pythainlp.corpus package effectively. If you have any questions or need further assistance, please refer to the PyThaiNLP documentation or reach out to the PyThaiNLP community for support.

We hope you find this documentation helpful for your natural language processing tasks in the Thai language.