pythainlp.generate

The pythainlp.generate module provides classes and functions for generating Thai text using n-gram language models.

N-gram generators

class pythainlp.generate.Unigram(name: str = 'tnc')[source]

Text generator using Unigram

Parameters:

name (str) – corpus name * tnc - Thai National Corpus (default) * ttc - Thai Textbook Corpus (TTC) * oscar - OSCAR Corpus

__init__(name: str = 'tnc') → None[source]
counts: dict[str, int]
word: list[str]
n: int
prob: dict[str, float]
gen_sentence(start_seq: str = '', N: int = 3, prob: float = 0.001, output_str: bool = True, duplicate: bool = False) → list[str] | str[source]

Generate a sentence using the unigram model.

Parameters:
  • start_seq (str) – word to begin sentence with

  • N (int) – number of words

  • prob (float) – minimum word probability threshold

  • output_str (bool) – output as string

  • duplicate (bool) – allow duplicate words in sentence

Returns:

list of words or a word string

Return type:

Union[list[str], str]

Example:
>>> from pythainlp.generate import Unigram
>>> gen = Unigram()
>>> gen.gen_sentence("แมว")
'แมวเวลานะนั้น'
class pythainlp.generate.Bigram(name: str = 'tnc')[source]

Text generator using Bigram

Parameters:

name (str) – corpus name * tnc - Thai National Corpus (default)

__init__(name: str = 'tnc') → None[source]
uni: dict[str, int]
bi: dict[tuple[str, str], int]
uni_keys: list[str]
bi_keys: list[tuple[str, str]]
words: list[str]
prob(t1: str, t2: str) → float[source]

Compute bigram probability P(t2 | t1).

Parameters:
  • t1 (str) – first word

  • t2 (str) – second word

Returns:

probability value

Return type:

float

gen_sentence(start_seq: str = '', N: int = 4, prob: float = 0.001, output_str: bool = True, duplicate: bool = False) → list[str] | str[source]

Generate a sentence using the bigram model.

Parameters:
  • start_seq (str) – word to begin sentence with

  • N (int) – number of words

  • prob (float) – minimum word probability threshold

  • output_str (bool) – output as string

  • duplicate (bool) – allow duplicate words in sentence

Returns:

list of words or a word string

Return type:

Union[list[str], str]

Example:
>>> from pythainlp.generate import Bigram
>>> gen = Bigram()
>>> gen.gen_sentence("แมว")
'แมวไม่ได้รับเชื้อมัน'
class pythainlp.generate.Trigram(name: str = 'tnc')[source]

Text generator using Trigram

Parameters:

name (str) – corpus name * tnc - Thai National Corpus (default)

__init__(name: str = 'tnc') → None[source]
uni: dict[str, int]
bi: dict[tuple[str, str], int]
ti: dict[tuple[str, str, str], int]
uni_keys: list[str]
bi_keys: list[tuple[str, str]]
ti_keys: list[tuple[str, str, str]]
words: list[str]
prob(t1: str, t2: str, t3: str) → float[source]

Compute trigram probability P(t3 | t1, t2).

Parameters:
  • t1 (str) – first word

  • t2 (str) – second word

  • t3 (str) – third word

Returns:

probability value

Return type:

float

gen_sentence(start_seq: str | tuple[str, str] = '', N: int = 4, prob: float = 0.001, output_str: bool = True, duplicate: bool = False) → list[str] | str[source]

Generate a sentence using the trigram model.

Parameters:
  • start_seq (Union[str, tuple[str, str]]) – word or bigram to begin sentence with

  • N (int) – number of words

  • prob (float) – minimum word probability threshold

  • output_str (bool) – output as string

  • duplicate (bool) – allow duplicate words in sentence

Returns:

list of words or a word string

Return type:

Union[list[str], str]

Example:
>>> from pythainlp.generate import Trigram
>>> gen = Trigram()
>>> gen.gen_sentence()
'ยังทำตัวเป็นเซิร์ฟเวอร์คือ'

Usage

Choose the generator class or function for the model you want, initialize it with appropriate parameters, and call its generation methods. Generated text can be used for chatbots, content generation, or data augmentation.

Example

::

from pythainlp.generate import Unigram

unigram = Unigram() sentence = unigram.gen_sentence(“สวัสดีครับ”) print(sentence)