Go efficient text segmentation; support english, chinese, japanese and other.
GoApache License 2.0active
9 projectsNatural Language Processing › Tokenizers
Go efficient text segmentation; support english, chinese, japanese and other.
GoApache License 2.0active
This is a Go implementation of jieba which a Chinese word splitting algorithm.
GoMIT Licenseactive
Sentence tokenizer: converts text into a list of sentences.
GoMIT Licenseslowing
Go library for performing Unicode Text Segmentation as described in Unicode Standard Annex #29
GoApache License 2.0active
Go package for n-gram based text categorization, with support for utf-8 and raw text.
GoBSD 2-Clause "Simplified" Licenseslowing
This is a GO implementation of MMSEG which a Chinese word splitting algorithm.
GoOtherdormant
Stemmer packages for Go programming language. Includes English and German stemmers.
GoBSD 2-Clause "Simplified" Licensedormant
A tokenizer based on the dictionary and Bigram language models for Golang. (Now only support chinese segmentation)
GoApache License 2.0dormant
The shamoji is word filtering package written in Go.
GoMIT Licenseslowing