pythainlp.util
The pythainlp.util module serves as a treasure trove of utility functions designed to aid text conversion, formatting, and various language processing tasks in the context of Thai language.
Modules
- pythainlp.util.analyze_thai_text(text: str) dict[str, int][source]
Count characters in Thai text by descriptive name.
Process the text character by character and map each Thai character to its descriptive name, or to itself (for consonants and digits).
- Parameters:
text (str) – Thai text to be analyzed
- Returns:
dictionary mapping character names to their counts
- Return type:
- Example:
>>> from pythainlp.util import analyze_thai_text >>> analyze_thai_text("คนดี") {'ค': 1, 'น': 1, 'ด': 1, 'สระ อี': 1} >>> analyze_thai_text("เล่น") {'สระ เอ': 1, 'ล': 1, 'ไม้เอก': 1, 'น': 1}
Analyzes a string of Thai text and returns a dictionaries, where each values represents a single classified character from the text.
- pythainlp.util.abbreviation_to_full_text(text: str, top_k: int = 2) list[tuple[str, float | None]][source]
Convert Thai text with abbreviations to full text.
Uses KhamYo to handle abbreviations. See more: KhamYo.
- Parameters:
- Returns:
list of
(full_text, cosine_similarity)tuples- Return type:
- Example:
>>> from pythainlp.util import ( ... abbreviation_to_full_text, ... )
>>> text = "รร.ของเราน่าอยู่"
>>> abbreviation_to_full_text(text) [ ('โรงเรียนของเราน่าอยู่', tensor(0.3734)), ('โรงแรมของเราน่าอยู่', tensor(0.2438)) ]
The abbreviation_to_full_text function is a text processing tool for converting common Thai abbreviations into their full, expanded forms. It’s invaluable for improving text readability and clarity.
- pythainlp.util.arabic_digit_to_thai_digit(text: str) str[source]
Convert Arabic digits to Thai digits.
For example, 1, 3, 10 become ๑, ๓, ๑๐.
- Parameters:
text (str) – text with Arabic digits such as ‘1’, ‘2’, ‘3’
- Returns:
text with Arabic digits converted to Thai digits such as ‘๑’, ‘๒’, ‘๓’
- Return type:
- Raises:
TypeError – if text is not a str
- Example:
>>> from pythainlp.util import arabic_digit_to_thai_digit >>> text = "เป็นจำนวน 123,400.25 บาท" >>> arabic_digit_to_thai_digit(text) 'เป็นจำนวน ๑๒๓,๔๐๐.๒๕ บาท'
The arabic_digit_to_thai_digit function allows you to transform Arabic numerals into their Thai numeral equivalents. This utility is especially useful when working with Thai numbers in text data.
- pythainlp.util.bahttext(number: float) str[source]
Convert a number to Thai text in Baht currency format.
Add the suffix “บาท” (Baht). The precision is fixed at two decimal places (0.00) to fit the “สตางค์” (Satang) unit. This function works similarly to the
BAHTTEXTfunction in Microsoft Excel.- Parameters:
number (float) – number to be converted
- Returns:
text representing the amount of money in Thai currency format
- Return type:
- Raises:
TypeError – if number is not a numeric type
- Example:
>>> from pythainlp.util import bahttext >>> bahttext(1) 'หนึ่งบาทถ้วน' >>> bahttext(21) 'ยี่สิบเอ็ดบาทถ้วน' >>> bahttext(200) 'สองร้อยบาทถ้วน'
The bahttext function specializes in converting numerical values into Thai Baht text, an essential feature for rendering financial data or monetary amounts in a user-friendly Thai format.
- pythainlp.util.check_khuap_klam(word: str) bool | None[source]
Check whether a Thai word is a consonant cluster (Kham Khuap Klam).
- Parameters:
word (str) – Thai word to check
- Returns:
Trueif the word is a true consonant cluster (คำควบกล้ำแท้),Falseif it is a false consonant cluster (คำควบกล้ำไม่แท้), orNoneif it is not a consonant cluster- Return type:
Optional[bool]
- Example:
>>> from pythainlp.util import check_khuap_klam
>>> # True consonant clusters (คำควบกล้ำแท้) >>> print(check_khuap_klam("กราบ")) # True >>> print(check_khuap_klam("ปลา")) # True >>> print(check_khuap_klam("เพราะ")) # True >>> print(check_khuap_klam("ตรง")) # True
>>> # False consonant clusters (คำควบกล้ำไม่แท้) >>> print(check_khuap_klam("จริง")) # False >>> print(check_khuap_klam("ทราย")) # False >>> print(check_khuap_klam("เศร้า")) # False
>>> # Not a consonant cluster >>> print(check_khuap_klam("แม่")) # None >>> print(check_khuap_klam("ตา")) # None
The check_khuap_klam function checks whether a Thai word is a consonant cluster (Kham Khuap Klam, คำควบกล้ำ). It returns
Truefor a true consonant cluster (คำควบกล้ำแท้),Falsefor a false consonant cluster (คำควบกล้ำไม่แท้), orNoneif the word is not a consonant cluster.
- pythainlp.util.censor_profanity(text: str, replacement: str = '*', custom_words: set[str] | None = None, engine: str = 'newmm') str[source]
Replace profanity words in text with a replacement character.
- Parameters:
- Returns:
text with profanity words censored
- Return type:
- Example:
>>> from pythainlp.util import censor_profanity
>>> print(censor_profanity("สวัสดีครับ")) สวัสดีครับ
>>> print( ... censor_profanity("text with profanity word") ... ) text with *** word
>>> # Add custom profanity words >>> print( ... censor_profanity("คำใหม่", custom_words={"คำใหม่"}) ... ) ******
The censor_profanity function replaces profanity words in Thai text with a replacement character (default: “*”). Users can provide custom profanity words in addition to the built-in list for content moderation and filtering.
- pythainlp.util.collate(data: Iterable[str], reverse: bool = False) list[str][source]
Sort words (almost) according to Thai dictionary order.
This implementation ignores tone marks and symbols.
- Parameters:
- Returns:
list of words, sorted (almost) according to Thai dictionary order
- Return type:
- Example:
>>> from pythainlp.util import collate >>> collate(["ไก่", "เกิด", "กาล", "เป็ด", "หมู", "วัว", "วันที่"]) ['กาล', 'เกิด', 'ไก่', 'เป็ด', 'วันที่', 'วัว', 'หมู'] >>> collate( ... ["ไก่", "เกิด", "กาล", "เป็ด", "หมู", "วัว", "วันที่"], reverse=True ... ) ['หมู', 'วัว', 'วันที่', 'เป็ด', 'ไก่', 'เกิด', 'กาล']
The collate function is a versatile tool for sorting Thai text in a locale-specific manner. It ensures that text data is sorted correctly, taking into account the Thai language’s unique characteristics.
- pythainlp.util.contains_profanity(text: str, custom_words: set[str] | None = None, engine: str = 'newmm') bool[source]
Check whether text contains profanity words.
- Parameters:
- Returns:
True if the text contains profanity, False otherwise
- Return type:
- Example:
>>> from pythainlp.util import contains_profanity
>>> print(contains_profanity("สวัสดีครับ")) False
>>> print(contains_profanity("คำหยาบคาย")) True if the word is in the profanity list
>>> # Add custom profanity words >>> print( ... contains_profanity("คำใหม่", custom_words={"คำใหม่"}) ... ) True
The contains_profanity function checks if Thai text contains profanity words. It returns True if profanity is detected and False otherwise. Users can provide custom profanity words for enhanced content moderation.
- pythainlp.util.convert_years(year: str, src: str = 'be', target: str = 'ad') str[source]
Convert a year from one era to another.
- Parameters:
- Returns:
converted year
- Return type:
- Raises:
NotImplementedError – if
srcortargetis not a supported eraValueError – if
yearis not an integer string
- Options for
srcandtarget: be - Buddhist Era
ad - Anno Domini
re - Rattanakosin Era
ah - Anno Hegirae
Warning: This function works properly only after 1941, because Thailand changed its calendar in 1941. Historians need to take the correct calendar into account.
- Example:
>>> from pythainlp.util import convert_years >>> # Convert Buddhist Era (BE) to Anno Domini (AD) >>> convert_years("2566", src="be", target="ad") '2023' >>> # Convert AD to BE >>> convert_years("2023", src="ad", target="be") '2566' >>> # Convert BE to Rattanakosin Era (RE) >>> convert_years("2566", src="be", target="re") '242' >>> # The same era returns the year as a normalized integer string >>> convert_years("2566", src="be", target="be") '2566'
The convert_years function is designed to facilitate the conversion of Western calendar years into Thai Buddhist Era (BE) years. This is significant for presenting dates and years in a Thai context.
- pythainlp.util.count_thai_chars(text: str) dict[str, int][source]
Count Thai characters by type.
The types are consonants, vowels, lead_vowels, follow_vowels, above_vowels, below_vowels, tonemarks, signs, thai_digits, punctuations, and non_thai.
- Parameters:
text (str) – text to be counted
- Returns:
dictionary of character counts by type
- Return type:
- Example:
>>> from pythainlp.util import count_thai_chars >>> count_thai_chars("ทดสอบภาษาไทย") {'vowels': 3, 'lead_vowels': 1, 'follow_vowels': 2, 'above_vowels': 0, 'below_vowels': 0, 'consonants': 9, 'tonemarks': 0, 'signs': 0, 'thai_digits': 0, 'punctuations': 0, 'non_thai': 0}
The count_thai_chars function is a character counting tool specifically tailored for Thai text. It helps in quantifying Thai characters, which can be useful for various text processing tasks.
- pythainlp.util.countthai(text: str, ignore_chars: str = ' \t\n\r\x0b\x0c0123456789!"#$%&\'()*+,-./:;<=>?@[\\]^_`{|}~') float[source]
Calculate the proportion of Thai characters in text.
Deprecated since version 5.3.2: Use
count_thai()instead.- Parameters:
- Returns:
proportion of Thai characters in the text (percentage)
- Return type:
The countthai function is a text processing utility for counting the occurrences of Thai characters in text data. This is useful for understanding the prevalence of Thai language content.
- pythainlp.util.dict_trie(dict_source: str | Iterable[str] | Trie) Trie[source]
Create a dictionary trie from a file or an iterable.
- Parameters:
dict_source (Union[str, Iterable[str], pythainlp.util.Trie]) – path to a dictionary file, an iterable of words, or a
pythainlp.util.Trieobject- Returns:
trie object
- Return type:
- Raises:
TypeError – if
dict_sourceis not a non-empty string, an iterable of words, or a trie
The dict_trie function implements a Trie data structure for efficient dictionary operations. It’s a valuable resource for dictionary management and fast word lookup.
- pythainlp.util.digit_to_text(text: str) str[source]
Spell out digits in Thai.
- Parameters:
text (str) – text with digits such as ‘1’, ‘2’, ‘๓’, ‘๔’
- Returns:
text with digits spelled out in Thai
- Return type:
- Raises:
TypeError – if text is not a str
- Example:
>>> from pythainlp.util import digit_to_text >>> digit_to_text("เบอร์โทร 0812345678") 'เบอร์โทร ศูนย์แปดหนึ่งสองสามสี่ห้าหกเจ็ดแปด' >>> digit_to_text("123") 'หนึ่งสองสาม' >>> digit_to_text("๕๖๗") 'ห้าหกเจ็ด'
The digit_to_text function is a numeral conversion tool that translates Arabic numerals into their Thai textual representations. This is vital for rendering numbers in Thai text naturally.
- pythainlp.util.display_thai_char(ch: str) str[source]
Prefix an underscore (_) to a high-position vowel or a tone mark.
The underscore eases readability.
- Parameters:
ch (str) – character to be displayed
- Returns:
“_” + ch for a high-position vowel or a tone mark, otherwise the character itself
- Return type:
- Example:
>>> from pythainlp.util import display_thai_char >>> display_thai_char("้") '_้'
The display_thai_char function is designed to present Thai characters with diacritics and tonal marks accurately. This is essential for displaying Thai text with correct pronunciation cues.
- pythainlp.util.emoji_to_thai(text: str, delimiters: tuple[str, str] = (':', ':')) str[source]
Convert emojis to their Thai meanings.
- Parameters:
- Returns:
text with emojis converted to their Thai meanings
- Return type:
- Example:
>>> from pythainlp.util import emoji_to_thai >>> emoji_to_thai("จะมานั่งรถเมล์เหมือนผมก็ได้นะครับ ใกล้ชิดประชาชนดี 😀") 'จะมานั่งรถเมล์เหมือนผมก็ได้นะครับ ใกล้ชิดประชาชนดี :หน้ายิ้มยิงฟัน:' >>> emoji_to_thai("หิวข้าวอยากกินอาหารญี่ปุ่น 🍣") 'หิวข้าวอยากกินอาหารญี่ปุ่น :ซูชิ:' >>> emoji_to_thai("🇹🇭 นี่คือธงประเทศไทย") ':ธง_ไทย: นี่คือธงประเทศไทย'
The emoji_to_thai function focuses on converting emojis into their Thai language equivalents. This is a unique feature for enhancing text communication with Thai-language emojis.
- pythainlp.util.eng_to_thai(text: str) str[source]
Convert text typed in the wrong layout to Thai Kedmanee.
The text was typed with the English-US Qwerty keyboard layout.
- Parameters:
text (str) – text typed with the wrong layout (Thai typed using an English keyboard)
- Returns:
Thai text, corrected from the wrong keyboard layout
- Return type:
- Example:
Intentionally type “ธนาคารแห่งประเทศไทย”, but got “Tok8kicsj’xitgmLwmp”:
>>> from pythainlp.util import eng_to_thai >>> eng_to_thai("Tok8kicsj'xitgmLwmp") 'ธนาคารแห่งประเทศไทย'
The eng_to_thai function serves as a text conversion tool for translating English text into its Thai transliterated form. It is beneficial for rendering English words and phrases in a Thai context.
- pythainlp.util.find_keyword(word_list: list[str], min_len: int = 3) dict[str, int][source]
Count word frequencies in a list of words, excluding stopwords.
- Parameters:
- Returns:
dictionary of words and their raw counts
- Return type:
- Example:
>>> from pythainlp.util import find_keyword
>>> words = [ ... "บันทึก", ... "เหตุการณ์", ... "บันทึก", ... "เหตุการณ์", ... " ", ... "มี", ... "การ", ... "บันทึก", ... "เป็น", ... " ", ... "ลายลักษณ์อักษรและ", ... "การ", ... "บันทึก", ... "เสียง", ... "ใน", ... "เหตุการณ์", ... ]
>>> find_keyword(words) {'บันทึก': 4, 'เหตุการณ์': 3}
>>> find_keyword(words, min_len=1) {'บันทึก': 4, 'เหตุการณ์': 3, ' ': 2, 'ลายลักษณ์อักษรและ': 1, 'เสียง': 1}
The find_keyword function is a powerful utility for identifying keywords and key phrases in text data. It is a fundamental component for text analysis and information extraction tasks.
- pythainlp.util.find_profanity(text: str, custom_words: set[str] | None = None, engine: str = 'newmm') list[str][source]
Find all profanity words in text.
- Parameters:
- Returns:
list of profanity words found in the text
- Return type:
- Example:
>>> from pythainlp.util import find_profanity
>>> print(find_profanity("สวัสดีครับ")) []
>>> print( ... find_profanity("text with profanity words") ... ) ['profanity_word1', 'profanity_word2']
>>> # Add custom profanity words >>> print( ... find_profanity("คำใหม่", custom_words={"คำใหม่"}) ... ) ['คำใหม่']
The find_profanity function identifies and returns a list of all profanity words found in Thai text. Users can provide custom profanity words to enhance detection capabilities for content moderation.
- pythainlp.util.ipa_to_rtgs(ipa: str) str[source]
Convert IPA phonemes to Royal Thai General System of Transcription (RTGS).
The conversion follows the rules listed in https://en.wikipedia.org/wiki/Help:IPA/Thai
- Parameters:
ipa (str) – IPA phonemes
- Returns:
converted RTGS text
- Return type:
- Example:
>>> from pythainlp.util import ipa_to_rtgs
>>> print(ipa_to_rtgs("kluaj")) 'kluai'
The ipa_to_rtgs function focuses on converting International Phonetic Alphabet (IPA) transcriptions into Royal Thai General System of Transcription (RTGS) format. This is valuable for phonetic analysis and pronunciation guides.
- pythainlp.util.isthai(text: str, ignore_chars: str = '.') bool[source]
Check whether every character in text is a Thai character.
Deprecated since version 5.3.2: Use
is_thai()instead.- Parameters:
- Returns:
True if every character in the text is Thai, otherwise False
- Return type:
The isthai function is a straightforward language detection utility that determines if text contains Thai language content. This function is essential for language-specific text processing.
- pythainlp.util.isthaichar(ch: str) bool[source]
Check whether a character is a Thai character.
Deprecated since version 5.3.2: Use
is_thai_char()instead.- Parameters:
ch (str) – character to check
- Returns:
True if the character is a Thai character, otherwise False
- Return type:
The isthaichar function is designed to check if a character belongs to the Thai script. It helps in character-level language identification and text processing.
- pythainlp.util.maiyamok(sent: str | list[str]) list[str][source]
Expand Maiyamok.
Deprecated since version 5.0.5: Use
expand_maiyamok()instead.Maiyamok (ๆ) (Unicode U+0E46) is a Thai character indicating word repetition. This function preprocesses Thai text by replacing Maiyamok with the repeated word.
- Parameters:
sent (Union[str, list[str]]) – sentence, as text or list of words
- Returns:
list of words
- Return type:
- Example:
>>> from pythainlp.util import maiyamok
>>> maiyamok("คนๆนก") ['คน', 'คน', 'นก']
The maiyamok function is a text processing tool that assists in identifying and processing Thai character characters with a ‘mai yamok’ tone mark.
- pythainlp.util.nectec_to_ipa(pronunciation: str) str[source]
Convert NECTEC phonemes to International Phonetic Alphabet (IPA).
- Parameters:
pronunciation (str) – NECTEC phonemes
- Returns:
converted IPA phonemes
- Return type:
- Example:
>>> from pythainlp.util import nectec_to_ipa
>>> print(nectec_to_ipa("kl-uua-j^-2")) 'kl uua j ˥˩'
References
Pornpimon Palingoon, Sumonmas Thatphithakkul. Chapter 4 Speech processing and Speech corpus. In: Handbook of Thai Electronic Corpus. 1st ed. p. 122–56.
The nectec_to_ipa function focuses on converting text from the NECTEC phonetic transcription system to the International Phonetic Alphabet (IPA). This conversion is vital for linguistic analysis and phonetic representation.
- pythainlp.util.normalize(text: str) str[source]
Normalize and clean Thai text.
This function applies these rules:
Remove zero-width spaces
Remove duplicate spaces
Remove spaces before tone marks and non-base characters
Reorder tone marks and vowels to standard order/spelling
Remove duplicate vowels and signs
Remove duplicate tone marks
Remove dangling non-base characters at the beginning of text
This function calls
remove_zw(),remove_dup_spaces(),remove_spaces_before_marks(),remove_repeat_vowels(), andremove_dangling(), in that order.To customize the selection or the order of rules, call those functions separately.
Note: for Unicode normalization, see
unicodedata.normalize().- Parameters:
text (str) – text to be normalized
- Returns:
normalized text
- Return type:
- Example:
>>> from pythainlp.util import normalize >>> normalize("เเปลก") # starts with two Sara E 'แปลก' >>> normalize("นานาาา") 'นานา'
The normalize function is a text processing utility that standardizes text by removing diacritics, tonal marks, and other modifications. It is valuable for text normalization and linguistic analysis.
- pythainlp.util.now_reign_year() int[source]
Return the current reign year of the 10th King of Chakri dynasty.
- Returns:
reign year of the 10th King of Chakri dynasty
- Return type:
- Example:
>>> from pythainlp.util import now_reign_year >>> text = "เป็นปีที่ {reign_year} ในรัชกาลปัจจุบัน"\ ... .format(reign_year=now_reign_year()) >>> print(text) เป็นปีที่ 11 ในรัชกาลปัจจุบัน
The now_reign_year function computes the current Thai Buddhist Era (BE) year and provides it in a human-readable format. This function is essential for displaying the current year in a Thai context.
- pythainlp.util.num_to_thaiword(number: int | None) str[source]
Convert an integer to Thai text.
- Parameters:
number (Optional[int]) – integer to be converted
- Returns:
text representing the number in Thai
- Return type:
- Example:
>>> from pythainlp.util import num_to_thaiword >>> num_to_thaiword(1) 'หนึ่ง' >>> num_to_thaiword(11) 'สิบเอ็ด'
The num_to_thaiword function is a numeral conversion tool for translating Arabic numerals into Thai word form. It is crucial for rendering numbers in a natural Thai textual format.
- pythainlp.util.num_to_thaiword_float(number: float) str[source]
Convert a floating-point number to Thai text.
Convert the integer part with
num_to_thaiword(). Read the decimal point as “จุด”. Read each digit after the decimal point individually, without place descriptions.- Parameters:
number (float) – number to be converted
- Returns:
text representing the number in Thai
- Return type:
- Raises:
TypeError – if number is not a numeric type
ValueError – if number is not finite
- Example:
>>> from pythainlp.util import num_to_thaiword_float >>> num_to_thaiword_float(123.45) 'หนึ่งร้อยยี่สิบสามจุดสี่ห้า' >>> num_to_thaiword_float(3.14159) 'สามจุดหนึ่งสี่หนึ่งห้าเก้า'
The num_to_thaiword_float function converts a floating-point number to Thai text. The integer part uses
num_to_thaiword(), the decimal point is read as “จุด”, and each digit after the decimal is read individually.
- pythainlp.util.rank(words: list[str], exclude_stopwords: bool = False) Counter[str] | None[source]
Count word frequencies in a list of Thai words.
Stopwords can be excluded from the count.
- Parameters:
- Returns:
counter of word frequencies, or None if
wordsis empty- Return type:
Optional[collections.Counter[str]]
- Example:
Include stopwords when counting word frequencies:
>>> from pythainlp.util import rank
>>> words = [ ... "บันทึก", ... "เหตุการณ์", ... " ", ... "มี", ... "การ", ... "บันทึก", ... "เป็น", ... " ", ... "ลายลักษณ์อักษร", ... ]
>>> rank(words) Counter({'บันทึก': 2, ' ': 2, 'เหตุการณ์': 1, 'มี': 1, 'การ': 1, 'เป็น': 1, 'ลายลักษณ์อักษร': 1})
Exclude stopwords when counting word frequencies:
>>> rank(words, exclude_stopwords=True) Counter({'บันทึก': 2, ' ': 2, 'เหตุการณ์': 1, 'ลายลักษณ์อักษร': 1})
The rank function is designed for ranking and ordering a list of items. It is a general-purpose utility for ranking items based on various criteria.
- pythainlp.util.reign_year_to_ad(reign_year: int, reign: int) int[source]
Convert a reign year to an AD year.
Return the AD year for a reign year of the 7th to 10th King of Chakri dynasty, Thailand. For instance, the AD year of the 4th reign year of the 10th King is 2019.
- Parameters:
- Returns:
AD year of the given reign and reign year
- Return type:
- Example:
>>> from pythainlp.util import reign_year_to_ad >>> print( ... "The 4th reign year of the King Rama X is in", ... reign_year_to_ad(4, 10), ... ) The 4th reign year of the King Rama X is in 2019 >>> print( ... "The 1st reign year of the King Rama IX is in", ... reign_year_to_ad(1, 9), ... ) The 1st reign year of the King Rama IX is in 1946
The reign_year_to_ad function facilitates the conversion of Thai Buddhist Era (BE) years into Western calendar years. This is useful for displaying historical dates in a globally recognized format.
- pythainlp.util.remove_dangling(text: str) str[source]
Remove Thai non-base characters at the beginning of text and after spaces.
This is a common typo, especially in input fields of a form, as these non-base characters can be visually hidden from the user who may type them in accidentally.
A character to be removed must meet both conditions:
it is a tone mark, above vowel, below vowel, or non-base sign
it is at the beginning of the text or after spaces
- Parameters:
text (str) – text to be normalized
- Returns:
text without dangling Thai characters at the beginning and after spaces
- Return type:
- Example:
>>> from pythainlp.util import remove_dangling >>> remove_dangling("๊ก") 'ก' >>> remove_dangling("คำ ่ที่สอง") 'คำ ที่สอง'
The remove_dangling function is a text processing tool for removing dangling characters or diacritics from text. It is useful for text cleaning and normalization.
- pythainlp.util.remove_dup_spaces(text: str) str[source]
Remove duplicate spaces.
Replace multiple spaces with one space. Replace multiple newline characters and empty lines with one newline character.
- Parameters:
text (str) – text to be normalized
- Returns:
text without duplicated spaces and newlines
- Return type:
- Example:
>>> from pythainlp.util import remove_dup_spaces >>> remove_dup_spaces("ก ข ค") 'ก ข ค'
The remove_dup_spaces function focuses on removing duplicate space characters from text data, making it more consistent and readable.
- pythainlp.util.remove_repeat_vowels(text: str) str[source]
Remove repeating vowels, tone marks, and signs.
This function calls
reorder_vowels()first, so that a double Sara E becomes Sara Ae and is not removed.- Parameters:
text (str) – text to be normalized
- Returns:
text without repeating Thai vowels, tone marks, and signs
- Return type:
- Example:
>>> from pythainlp.util import remove_repeat_vowels >>> remove_repeat_vowels("นานาาา") 'นานา' >>> remove_repeat_vowels("ดีีีี") 'ดี'
The remove_repeat_vowels function is designed to eliminate repeated vowel characters in text, improving text readability and consistency.
- pythainlp.util.remove_tone_ipa(ipa: str) str[source]
Remove Thai tones from IPA phonemes.
- Parameters:
ipa (str) – IPA phonemes
- Returns:
IPA phonemes with tones removed
- Return type:
- Example:
>>> from pythainlp.util import remove_tone_ipa
>>> print(remove_tone_ipa("laː˦˥.sa˨˩.maj˩˩˦")) laː.sa.maj
The remove_tone_ipa function serves as a phonetic conversion tool for removing tone marks from IPA transcriptions. This is crucial for phonetic analysis and linguistic research.
- pythainlp.util.remove_tonemark(text: str) str[source]
Remove all Thai tone marks from text.
Thai script has four tone marks indicating four tones as follows:
Down tone (Thai: ไม้เอก _่ )
Falling tone (Thai: ไม้โท _้ )
High tone (Thai: ไม้ตรี _๊ )
Rising tone (Thai: ไม้จัตวา _๋ )
Using a wrong tone mark is a common mistake in Thai writing. Text without tone marks can be used for approximate string matching.
- Parameters:
text (str) – text to be normalized
- Returns:
text without Thai tone marks
- Return type:
- Example:
>>> from pythainlp.util import remove_tonemark >>> remove_tonemark("สองพันหนึ่งร้อยสี่สิบเจ็ดล้านสี่แสนแปดหมื่นสามพันหกร้อยสี่สิบเจ็ด") 'สองพันหนึงรอยสีสิบเจ็ดลานสีแสนแปดหมืนสามพันหกรอยสีสิบเจ็ด'
The remove_tonemark function is a utility for removing tonal marks and diacritics from text data, making it suitable for various text processing tasks.
- pythainlp.util.remove_zw(text: str) str[source]
Remove zero-width characters.
These invisible characters may cause unexpected results from the user’s point of view. Removing them makes string matching more robust.
Characters to be removed:
Zero-width space (ZWSP)
Zero-width non-joiner (ZWNJ)
- Parameters:
text (str) – text to be normalized
- Returns:
text without zero-width characters
- Return type:
- Example:
>>> from pythainlp.util import remove_zw >>> remove_zw("สวัสดีครับ") 'สวัสดีครับ' >>> remove_zw("ภาษาไทย") 'ภาษาไทย'
The remove_zw function is designed to remove zero-width characters from text data, ensuring that text is free from invisible or unwanted characters.
- pythainlp.util.reorder_vowels(text: str) str[source]
Reorder vowels and tone marks to the standard logical order.
This function reorders or transforms characters in text according to these rules:
Sara E + Sara E -> Sara Ae
Nikhahit + Sara Aa -> Sara Am
tone mark + non-base vowel -> non-base vowel + tone mark
follow vowel + tone mark -> tone mark + follow vowel
- Parameters:
text (str) – text to be normalized
- Returns:
text with vowels and tone marks in the standard logical order
- Return type:
- Example:
>>> from pythainlp.util import reorder_vowels >>> reorder_vowels("เเปลก") # two Sara E become Sara Ae 'แปลก' >>> reorder_vowels("ก้ำ") # reorder tone marks and vowels 'ก้ำ'
The reorder_vowels function is a text processing utility for reordering vowel characters in Thai text. It is essential for phonetic analysis and pronunciation guides.
- pythainlp.util.rhyme(word: str) list[str][source]
Find Thai words that rhyme with a word.
- Parameters:
word (str) – Thai word
- Returns:
list of rhyming Thai words
- Return type:
- Example:
>>> from pythainlp.util import rhyme >>> rhyme("จีบ") ['กลีบ', 'กีบ', 'ครีบ', 'คีบ', 'งีบ', ... ]
The rhyme function is a utility for find rhyme of Thai word.
- pythainlp.util.sound_syllable(syllable: str) str[source]
Classify the sound of a Thai syllable as live or dead.
- Parameters:
syllable (str) – Thai syllable
- Returns:
type of the syllable (“live” or “dead”)
- Return type:
- Example:
>>> from pythainlp.util import sound_syllable >>> sound_syllable("มา") 'live' >>> sound_syllable("เลข") 'dead' >>> sound_syllable("ฤๅ") 'live'
The sound_syllable function specializes in identifying and processing Thai characters that represent sound syllables. This is valuable for phonetic and linguistic analysis.
- pythainlp.util.syllable_length(syllable: str) str[source]
Detect the vowel length of a Thai syllable.
- Parameters:
syllable (str) – Thai syllable
- Returns:
length of the syllable (“long” or “short”)
- Return type:
- Example:
>>> from pythainlp.util import syllable_length >>> syllable_length("มาก") 'long' >>> syllable_length("คะ") 'short'
The syllable_length function is a text analysis tool for calculating the length of syllables in Thai text. It is significant for linguistic analysis and language research.
- pythainlp.util.syllable_open_close_detector(syllable: str) str[source]
Detect whether a Thai syllable is open or closed.
- Parameters:
syllable (str) – Thai syllable
- Returns:
“open” or “close”
- Return type:
- Example:
>>> from pythainlp.util import syllable_open_close_detector >>> syllable_open_close_detector("มาก") 'close' >>> syllable_open_close_detector("คะ") 'open'
The syllable_open_close_detector function is designed to detect syllable open and close statuses in Thai text. This information is vital for phonetic analysis and linguistic research.
- pythainlp.util.text_to_arabic_digit(text: str) str[source]
Convert a digit spelled out in Thai to an Arabic digit.
- Parameters:
text (str) – digit spelled out in Thai
- Returns:
Arabic digit such as ‘1’, ‘2’, ‘3’ if the text is a digit spelled out in Thai (ศูนย์, หนึ่ง, สอง, …, เก้า), otherwise an empty string
- Return type:
- Raises:
TypeError – if text is not a str
- Example:
>>> from pythainlp.util import text_to_arabic_digit
>>> text_to_arabic_digit("ศูนย์") 0 >>> text_to_arabic_digit("หนึ่ง") 1 >>> text_to_arabic_digit("แปด") 8 >>> text_to_arabic_digit("เก้า") 9
>>> # For text that is not digit spelled out in Thai >>> text_to_arabic_digit("สิบ") == "" True >>> text_to_arabic_digit("เก้าร้อย") == "" True
The text_to_arabic_digit function is a numeral conversion tool that translates Thai text numerals into Arabic numeral form. It is useful for numerical data extraction and processing.
- pythainlp.util.text_to_num(text: str) list[str][source]
Convert numerals spelled out in Thai text to numbers.
Return a list of words, with each spelled-out numeral replaced by its numeric value as text.
- Parameters:
text (str) – Thai text with spelled-out numerals
- Returns:
list of words with numerals converted to numbers
- Return type:
- Example:
>>> from pythainlp.util import text_to_num >>> text_to_num("เก้าร้อยแปดสิบจุดเก้าห้าบาทนี่คือจำนวนทั้งหมด") ['980.95', 'บาท', 'นี่', 'คือ', 'จำนวน', 'ทั้งหมด'] >>> text_to_num("สิบล้านสองหมื่นหนึ่งพันแปดร้อยแปดสิบเก้าบาท") ['10021889', 'บาท']
The text_to_num function focuses on extracting numerical values from text data. This is essential for converting textual numbers into numerical form for computation.
- pythainlp.util.text_to_thai_digit(text: str) str[source]
Convert a digit spelled out in Thai to a Thai digit.
- Parameters:
text (str) – digit spelled out in Thai
- Returns:
Thai digit such as ‘๑’, ‘๒’, ‘๓’ if the text is a digit spelled out in Thai (ศูนย์, หนึ่ง, สอง, …, เก้า), otherwise an empty string
- Return type:
- Raises:
TypeError – if text is not a str
- Example:
>>> from pythainlp.util import text_to_thai_digit
>>> text_to_thai_digit("ศูนย์") ๐ >>> text_to_thai_digit("หนึ่ง") ๑ >>> text_to_thai_digit("แปด") ๘ >>> text_to_thai_digit("เก้า") ๙
>>> # For text that is not Thai digit spelled out >>> text_to_thai_digit("สิบ") == "" True >>> text_to_thai_digit("เก้าร้อย") == "" True
The text_to_thai_digit function serves as a numeral conversion tool for translating Arabic numerals into Thai numeral form. This is important for rendering numbers in Thai text naturally.
- pythainlp.util.thai_digit_to_arabic_digit(text: str) str[source]
Convert Thai digits to Arabic digits.
For example, ๑, ๓, ๑๐ become 1, 3, 10.
- Parameters:
text (str) – text with Thai digits such as ‘๑’, ‘๒’, ‘๓’
- Returns:
text with Thai digits converted to Arabic digits such as ‘1’, ‘2’, ‘3’
- Return type:
- Raises:
TypeError – if text is not a str
- Example:
>>> from pythainlp.util import thai_digit_to_arabic_digit >>> text = "เป็นจำนวน ๑๒๓,๔๐๐.๒๕ บาท" >>> thai_digit_to_arabic_digit(text) 'เป็นจำนวน 123,400.25 บาท'
The thai_digit_to_arabic_digit function allows you to transform Thai numeral text into Arabic numeral format. This is valuable for numerical data extraction and computation tasks.
- pythainlp.util.thai_strftime(dt_obj: datetime, fmt: str = '%-d %b %y', thaidigit: bool = False) str[source]
Convert
datetime.datetimeinto Thai date and time format.The formatting directives are similar to
datetime.strftime().This function uses Thai names and the Thai Buddhist Era for these directives:
%a - abbreviated weekday name (such as “จ”, “อ”, “พ”, “พฤ”, “ศ”, “ส”, “อา”)
%A - full weekday name (such as “วันจันทร์”, “วันอังคาร”, “วันเสาร์”, “วันอาทิตย์”)
%b - abbreviated month name (such as “ม.ค.”, “ก.พ.”, “มี.ค.”, “เม.ย.”, “พ.ค.”, “มิ.ย.”, “ธ.ค.”)
%B - full month name (such as “มกราคม”, “กุมภาพันธ์”, “พฤศจิกายน”, “ธันวาคม”)
%y - year without century (such as “56”, “10”)
%Y - year with century (such as “2556”, “2410”)
%c - date and time representation (such as “พ 6 ต.ค. 01:40:00 2519”)
%v - short date representation (such as “ 6-ม.ค.-2562”, “27-ก.พ.-2555”)
This function passes other directives to
datetime.datetime.strftime().- Note:
The Thai Buddhist Era (BE) year is simply converted from AD by adding 543. This is not accurate for years before 1941 AD, due to the change in Thai New Year’s Day.
This function is an interim solution, since the Python standard
localemodule (which relies on the Cstrftime()) does not support the “th” or “th_TH” locale yet. If supported, we can calllocale.setlocale(locale.LC_TIME, "th_TH")and then use the nativedatetime.datetime.strftime().
This function aims to be platform-independent and to support as many extensions as possible. See these links for
strftime()extensions in POSIX, BSD, and GNU libc:Python https://docs.python.org/3/library/datetime.html#strftime-strptime-behavior
JavaScript’s implementation https://github.com/samsonjs/strftime
strftime() quick reference https://strftime.net/
- Parameters:
dt_obj (datetime.datetime) – date and time to be formatted
fmt (str) – string containing date and time directives
thaidigit (bool) – represent numbers in Thai digits if True, otherwise in Arabic digits (default)
- Returns:
date and time text, with month in Thai name and year in Thai Buddhist Era. The year is simply converted from AD by adding 543 (not accurate for years before 1941 AD, due to the change in Thai New Year’s Day).
- Return type:
- Example:
>>> from datetime import datetime >>> from pythainlp.util import thai_strftime
>>> datetime_obj = datetime(year=2019, month=6, day=9, \ ... hour=5, minute=59, second=0, microsecond=0)
>>> print(datetime_obj) 2019-06-09 05:59:00
>>> thai_strftime(datetime_obj, "%A %d %B %Y") 'วันอาทิตย์ 09 มิถุนายน 2562'
>>> thai_strftime(datetime_obj, "%a %-d %b %y") # no padding 'อา 9 มิ.ย. 62'
>>> thai_strftime(datetime_obj, "%a %_d %b %y") # space padding 'อา 9 มิ.ย. 62'
>>> thai_strftime(datetime_obj, "%a %0d %b %y") # zero padding 'อา 09 มิ.ย. 62'
>>> thai_strftime(datetime_obj, "%-H นาฬิกา %-M นาที", thaidigit=True) '๕ นาฬิกา ๕๙ นาที'
>>> thai_strftime(datetime_obj, "%D (%v)") '06/09/62 ( 9-มิ.ย.-2562)'
>>> thai_strftime(datetime_obj, "%c") 'อา 9 มิ.ย. 05:59:00 2562'
>>> thai_strftime(datetime_obj, "%H:%M %p") '05:59 AM'
>>> thai_strftime(datetime_obj, "%H:%M %#p") '05:59 am'
The thai_strftime function is a date formatting tool tailored for Thai culture. It is essential for displaying dates and times in a format that adheres to Thai conventions.
- pythainlp.util.thai_strptime(text: str, fmt: str, year: str = 'be', add_year: int | None = None, tzinfo: ZoneInfo | None = zoneinfo.ZoneInfo(key='Asia/Bangkok')) datetime[source]
Parse Thai date and time text into a
datetime.datetime.- Parameters:
text (str) – text containing date and time
fmt (str) – string containing date and time directives
year (str) – era of the year in the text (ad for Anno Domini, be for Buddhist Era)
add_year (Optional[int]) – year to add to a two-digit year (default is None)
tzinfo (Optional[zoneinfo.ZoneInfo]) – time zone (default is Asia/Bangkok)
- Returns:
parsed date and time
- Return type:
- Supported directives in
fmt: %d - Day (1 - 31)
%B - Thai month (03, 3, มี.ค., or มีนาคม)
%Y - Year (66, 2566, or 2023)
%H - Hour (0 - 23)
%M - Minute (0 - 59)
%S - Second (0 - 59)
%f - Microsecond
- Example:
>>> from pythainlp.util import thai_strptime
>>> thai_strptime("15 ก.ค. 2565 09:00:01", "%d %B %Y %H:%M:%S") datetime.datetime(2022, 7, 15, 9, 0, 1, tzinfo=zoneinfo.ZoneInfo(key='Asia/Bangkok'))
The thai_strptime function focuses on parsing dates and times in a Thai-specific format, making it easier to work with date and time data in a Thai context.
- pythainlp.util.thai_to_eng(text: str) str[source]
Convert text typed in the wrong layout to English-US Qwerty.
The text was typed with the Thai Kedmanee keyboard layout.
- Parameters:
text (str) – text typed with the wrong layout (English typed using a Thai keyboard)
- Returns:
English text, corrected from the wrong keyboard layout
- Return type:
- Example:
Intentionally type “Bank of Thailand”, but got “ฺฟืา นด ธ้ฟรสฟืก”:
>>> from pythainlp.util import thai_to_eng >>> thai_to_eng("ฺฟืา นด ธ้ฟรสฟืก") 'Bank of Thailand'
The thai_to_eng function is a text conversion tool for translating Thai text into its English transliterated form. This is beneficial for rendering Thai words and phrases in an English context.
- pythainlp.util.to_idna(text: str) str[source]
Encode text with IDNA, as used in Internationalized Domain Names.
- Parameters:
text (str) – Thai text
- Returns:
IDNA-encoded text
- Return type:
- Example:
>>> from pythainlp.util import to_idna >>> to_idna("คนละครึ่ง.com") 'xn--42caj4e6bk1f5b1j.com'
The to_idna function is a text conversion tool for translating Thai text into its International Domain Name (IDN) for Thai domain name.
- pythainlp.util.thai_word_tone_detector(word: str | None) list[tuple[str, str]][source]
Detect the tone of each syllable in a Thai word.
This function converts the word to pronunciation with
pythainlp.transliterate.pronunciate().- Parameters:
word (Optional[str]) – Thai word, or None
- Returns:
list of (syllable, tone) tuples, one for each syllable. Tone values are
l(low),m(mid),h(high),r(rising),f(falling), or an empty string if it cannot be detected. Return[]if the word is None or empty.- Return type:
- Example:
>>> from pythainlp.util import thai_word_tone_detector >>> print(thai_word_tone_detector("คนดี")) [('คน', 'm'), ('ดี', 'm')] >>> print(thai_word_tone_detector("มือถือ")) [('มือ', 'm'), ('ถือ', 'r')] >>> print(thai_word_tone_detector(None)) []
The thai_word_tone_detector function specializes in detecting and processing tonal marks in Thai words. It is essential for phonetic analysis and pronunciation guides.
- pythainlp.util.thaiword_to_date(text: str, date: datetime | None = None) datetime | None[source]
Convert Thai relative date to
datetime.datetime.- Parameters:
text (str) – Thai text containing a relative date
date (datetime.datetime) – reference date (default is datetime.datetime.now())
- Returns:
date and time if it can be calculated, otherwise None
- Return type:
Optional[datetime.datetime]
- Example:
>>> from datetime import datetime >>> from pythainlp.util import thaiword_to_date
>>> thaiword_to_date("พรุ่งนี้", datetime(2024, 1, 31)) datetime.datetime(2024, 2, 1, 0, 0) >>> thaiword_to_date("เมื่อวาน", datetime(2024, 1, 31)) datetime.datetime(2024, 1, 30, 0, 0) >>> print(thaiword_to_date("ไม่มีคำนี้")) None
The thaiword_to_date function facilitates the conversion of Thai word representations of dates into standardized date formats. This is important for date data extraction and processing.
- pythainlp.util.thaiword_to_num(word: str) int[source]
Convert a numeral spelled out in Thai to an integer.
- Parameters:
word (str) – numeral spelled out in Thai
- Returns:
integer value of the numeral
- Return type:
- Raises:
TypeError – if word is not a str
ValueError – if word is empty or is not a valid Thai numeral
- Example:
>>> from pythainlp.util import thaiword_to_num >>> thaiword_to_num("ศูนย์") 0 >>> thaiword_to_num("สองล้านสามแสนหกร้อยสิบสอง") 2300612
The thaiword_to_num function is a numeral conversion tool for translating Thai word numerals into numerical form. This is essential for numerical data extraction and computation.
- pythainlp.util.thaiword_to_time(text: str, padding: bool = True) str[source]
Convert Thai time in words to a time string (H:M).
- Parameters:
- Returns:
time string
- Return type:
- Raises:
ValueError – if the text is not a valid Thai time
- Example:
>>> from pythainlp.util import thaiword_to_time >>> thaiword_to_time("บ่ายโมงครึ่ง") '13:30'
The thaiword_to_time function is designed for converting Thai word representations of time into standardized time formats. It is crucial for time data extraction and processing.
- pythainlp.util.time_to_thaiword(time_data: time | datetime | str, fmt: str = '24h', precision: str | None = None) str[source]
Spell out time as Thai words.
- Parameters:
time_data (Union[datetime.time, datetime.datetime, str]) – a
datetime.timeobject, adatetime.datetimeobject, or a string inH:MorH:M:Sformat (24-hour clock)fmt (str) –
output format
24h - 24-hour clock (default)
6h - 6-hour clock
m6h - modified 6-hour clock
precision (Optional[str]) –
precision of the spelled-out time
m - always spell out to the minute
s - always spell out to the second
None - spell out only non-zero parts (default)
- Returns:
time spelled out as Thai words
- Return type:
- Raises:
TypeError – if time_data is not a datetime.time, datetime.datetime, or str
ValueError – if time_data is an empty string or does not match the H:M or H:M:S format
- Example:
>>> from datetime import time >>> from pythainlp.util import time_to_thaiword >>> time_to_thaiword("8:17") 'แปดนาฬิกาสิบเจ็ดนาที' >>> time_to_thaiword("8:17", "6h") 'สองโมงเช้าสิบเจ็ดนาที' >>> time_to_thaiword("8:17", "m6h") 'แปดโมงสิบเจ็ดนาที' >>> time_to_thaiword("18:30", fmt="m6h") 'หกโมงครึ่ง' >>> time_to_thaiword(time(12, 3, 0)) 'สิบสองนาฬิกาสามนาที' >>> time_to_thaiword(time(12, 3, 0), precision="s") 'สิบสองนาฬิกาสามนาทีศูนย์วินาที'
The time_to_thaiword function focuses on converting time values into Thai word representations. This is valuable for rendering time in a natural Thai textual format.
- pythainlp.util.tis620_to_utf8(text: str) str[source]
Convert TIS-620 text to UTF-8 text.
- Parameters:
text (str) – TIS-620 encoded text
- Returns:
UTF-8 encoded text
- Return type:
- Example:
>>> from pythainlp.util import tis620_to_utf8 >>> tis620_to_utf8("¡ÃзÃÇ§ÍØµÊÒË¡ÃÃÁ") 'กระทรวงอุตสาหกรรม'
The tis620_to_utf8 function serves as a character encoding conversion tool for converting TIS-620 encoded text into UTF-8 format. This is significant for character encoding compatibility.
- pythainlp.util.tone_detector(syllable: str) str[source]
Detect the tone of a Thai syllable.
The tone is one of:
l: low
m: mid
r: rising
f: falling
h: high
empty string: cannot be detected
- Parameters:
syllable (str) – Thai syllable
- Returns:
tone of the syllable (l, m, h, r, f), or an empty string if it cannot be detected
- Return type:
- Example:
>>> from pythainlp.util import tone_detector >>> tone_detector("มา") 'm' >>> tone_detector("ไม้") 'h'
The tone_detector function is a text processing tool for detecting tone marks and diacritics in Thai text. It is essential for phonetic analysis and pronunciation guides.
- pythainlp.util.words_to_num(words: list[str]) float[source]
Convert a list of Thai numeral words to a floating-point number.
- Parameters:
- Returns:
float value of the words
- Return type:
- Example:
>>> from pythainlp.util import words_to_num >>> words_to_num(["ห้า", "สิบ", "จุด", "เก้า", "ห้า"]) 50.95
The words_to_num function is a numeral conversion utility that translates Thai word numerals into numerical form. It is important for numerical data extraction and computation.
- pythainlp.util.thai_consonant_to_spelling(c: str) str[source]
Convert a Thai consonant to its spelling.
- pythainlp.util.spell_words.spell_syllable(text: str) list[str][source]
Spell out a syllable in Thai word distribution form.
- Parameters:
text (str) – Thai syllable
- Returns:
list of spelled-out syllable components
- Return type:
- Example:
>>> from pythainlp.util.spell_words import spell_syllable >>> spell_syllable("แมว") ['มอ', 'วอ', 'แอ', 'แมว']
The pythainlp.util.spell_words.spell_syllable function focuses on spelling syllables in Thai text, an important feature for phonetic analysis and linguistic research.
- pythainlp.util.spell_words.spell_word(text: str | None) list[str][source]
Spell out a word in Thai word distribution form.
- Parameters:
text (Optional[str]) – Thai word, or None
- Returns:
list of spelled-out word components, empty list if text is None or empty
- Return type:
- Example:
>>> from pythainlp.util.spell_words import spell_word >>> spell_word("คนดี") ['คอ', 'นอ', 'คน', 'ดอ', 'อี', 'ดี', 'คนดี'] >>> spell_word(None) []
The pythainlp.util.spell_words.spell_word function is designed for spelling individual words in Thai text, facilitating phonetic analysis and pronunciation guides.
- pythainlp.util.to_lunar_date(input_date: date) str[source]
Convert a solar date to a Thai lunar date.
- Parameters:
input_date (datetime.date) – solar date
- Returns:
Thai lunar date text
- Return type:
- Example:
>>> from datetime import date >>> from pythainlp.util import to_lunar_date >>> to_lunar_date(date(2024, 1, 1)) 'แรม 5 ค่ำ เดือน 1' >>> to_lunar_date(date(2024, 12, 31)) 'ขึ้น 2 ค่ำ เดือน 2'
The to_lunar_date function focuses on converts the solar date to Thai Lunar Date.
- pythainlp.util.th_zodiac(year: int, output_type: int = 1) str | int[source]
Convert a Gregorian year to its Thai zodiac year name.
- Parameters:
- Returns:
zodiac name or number of the year
- Return type:
- Example:
>>> from pythainlp.util import th_zodiac >>> # Get Thai zodiac name >>> th_zodiac(2024, output_type=1) 'มะโรง' >>> # Get English zodiac name >>> th_zodiac(2024, output_type=2) 'DRAGON' >>> # Get zodiac number >>> th_zodiac(2024, output_type=3) 5
The th_zodiac function is converts a Gregorian year to its corresponding Thai Zodiac name.
- class pythainlp.util.Trie(words: Iterable[str])[source]
Trie data structure for efficient prefix-based word search.
A trie (prefix tree) is a tree-like data structure that stores a collection of strings. It enables fast retrieval of words with common prefixes, which suits dictionary-based tokenization and autocomplete features.
- Parameters:
words (Iterable[str]) – words to initialize the trie with
- Example:
>>> from pythainlp.util import Trie >>> trie = Trie(["สวัสดี", "สวัส", "ดี", "ครับ"]) >>> "สวัสดี" in trie True >>> trie.prefixes("สวัสดีครับ") ['สวัส', 'สวัสดี'] >>> trie.add("สวัสดีตอนเช้า") >>> len(trie) 5
The Trie class is a data structure for efficient dictionary operations. It’s a valuable resource for managing and searching word lists and dictionaries in a structured and efficient manner.
- __init__(words: Iterable[str]) None[source]
Initialize the trie with words.
- Parameters:
words (Iterable[str]) – words to initialize the trie with
- add(word: str) None[source]
Add a word to the trie.
Remove spaces in front of and following the word.
- Parameters:
word (str) – word to be added
- remove(word: str) None[source]
Remove a word from the trie.
Do nothing if the word is not found.
- Parameters:
word (str) – word to be removed
- pythainlp.util.longest_common_subsequence(str1: str, str2: str) str[source]
Return the longest common subsequence of two strings.
- Parameters:
- Returns:
longest common subsequence
- Return type:
- Example:
>>> from pythainlp.util.lcs import longest_common_subsequence >>> longest_common_subsequence("ABCBDAB", "BDCAB") 'BDAB'
The longest_common_subsequence function is find the longest common subsequence between two strings.
- pythainlp.util.morse.morse_encode(text: str, lang: str = 'th') str[source]
Convert text to Morse code (Thai and English are supported).
- Parameters:
- Returns:
Morse code
- Return type:
- Raises:
NotImplementedError – if
langis not supported- Example:
>>> from pythainlp.util.morse import morse_encode >>> morse_encode("แมว", lang="th") '.-.- -- .--' >>> morse_encode("cat", lang="en") '-.-. .- -'
The pythainlp.util.morse.morse_encode function is convert text to Morse code.
- pythainlp.util.morse.morse_decode(morse_text: str, lang: str = 'th') str[source]
Convert Morse code to text.
Thai decoding may produce incorrect characters that can be fixed with a spell corrector.
- Parameters:
- Returns:
decoded text
- Return type:
- Raises:
NotImplementedError – if
langis not supported- Example:
>>> from pythainlp.util.morse import morse_decode >>> morse_decode(".-.- -- .--", lang="th") 'แมว' >>> morse_decode("-.-. .- -", lang="en") 'CAT'
The pythainlp.util.morse.morse_decode function is convert Morse code to text.