pythainlp.lm

The pythainlp.lm package provides language models and language modeling utilities.

Modules

pythainlp.lm.calculate_ngram_counts(list_words: list[str], n_min: int = 2, n_max: int = 4) → dict[tuple[str, ...], int][source]

Calculate n-gram counts for the given word list.

Parameters:
  • list_words (list[str]) – list of words

  • n_min (int) – minimum n-gram size (default is 2)

  • n_max (int) – maximum n-gram size (default is 4)

Returns:

dictionary mapping n-grams to their counts

Return type:

dict[tuple[str, …], int]

pythainlp.lm.remove_repeated_ngrams(string_list: list[str], n: int = 2) → list[str][source]

Remove repeated n-grams from a word list.

Parameters:
  • string_list (list[str]) – list of words

  • n (int) – n-gram size (default is 2)

Returns:

list of words with repeated n-grams removed

Return type:

list[str]

Example:
>>> from pythainlp.lm import remove_repeated_ngrams
>>> remove_repeated_ngrams(
...     ["เอา", "เอา", "แบบ", "ไหน"], n=1
... )
['เอา', 'แบบ', 'ไหน']
class pythainlp.lm.Qwen3[source]

Generate Thai text using the Qwen3-0.6B language model.

A small but capable language model from the Qwen family of Alibaba Cloud, optimized for various NLP tasks including Thai language processing.

__init__() → None[source]

Initialize Qwen3 without a loaded model.

load_model(model_path: str = 'Qwen/Qwen3-0.6B', device: str = 'cuda', torch_dtype: torch.dtype | None = None, low_cpu_mem_usage: bool = True, revision: str | None = None) → None[source]

Load the Qwen3 model.

Parameters:
  • model_path (str) – model path or Hugging Face model ID

  • device (str) – device (cpu, cuda, or other)

  • torch_dtype (Optional[torch.dtype]) – data type of the model, for example torch.float16 or torch.bfloat16

  • low_cpu_mem_usage (bool) – reduce CPU memory usage while loading

  • revision (Optional[str]) – git revision id (branch, tag, or commit hash). Pin to a full commit hash for secure downloads.

Example:
>>> from pythainlp.lm import Qwen3
>>> import torch
>>> model = Qwen3()
>>> model.load_model(
...     device="cpu", torch_dtype=torch.bfloat16
... )
generate(text: str, max_new_tokens: int = 512, temperature: float = 0.7, top_p: float = 0.9, top_k: int = 50, do_sample: bool = True, skip_special_tokens: bool = True) → str[source]

Generate text from a prompt.

Parameters:
  • text (str) – text of the prompt

  • max_new_tokens (int) – maximum number of new tokens

  • temperature (float) – sampling temperature (higher is more random)

  • top_p (float) – cumulative probability for nucleus sampling

  • top_k (int) – number of top tokens to sample from

  • do_sample (bool) – use sampling instead of greedy decoding

  • skip_special_tokens (bool) – skip special tokens in the output

Returns:

generated text

Return type:

str

Example:
>>> from pythainlp.lm import Qwen3
>>> import torch
>>> model = Qwen3()
>>> model.load_model(
...     device="cpu", torch_dtype=torch.bfloat16
... )
>>> result = model.generate("สวัสดี")
>>> print(result)
chat(messages: list[dict[str, Any]], max_new_tokens: int = 512, temperature: float = 0.7, top_p: float = 0.9, top_k: int = 50, do_sample: bool = True, skip_special_tokens: bool = True) → str[source]

Generate text using chat format.

Parameters:
  • messages (list[dict[str, Any]]) – list of messages, each a dictionary with role and content keys

  • max_new_tokens (int) – maximum number of new tokens

  • temperature (float) – sampling temperature (higher is more random)

  • top_p (float) – cumulative probability for nucleus sampling

  • top_k (int) – number of top tokens to sample from

  • do_sample (bool) – use sampling instead of greedy decoding

  • skip_special_tokens (bool) – skip special tokens in the output

Returns:

generated response

Return type:

str

Example:
>>> from pythainlp.lm import Qwen3
>>> import torch
>>> model = Qwen3()
>>> model.load_model(
...     device="cpu", torch_dtype=torch.bfloat16
... )
>>> messages = [
...     {"role": "user", "content": "สวัสดีครับ"}
... ]
>>> response = model.chat(messages)
>>> print(response)

Submodules

  • pythainlp.lm.phayathaibert

  • pythainlp.lm.wangchanberta

  • pythainlp.lm.ulmfit