Go efficient text segmentation; support english, chinese, japanese and other.
Dictionary with double array trie (Double-Array Trie) to achieve, Sender algorithm is the shortest path based on word frequency plus dynamic programming, and DAG and HMM algorithm word segmentation.
Support common, search engine, full mode, precise mode and HMM mode multiple word segmentation modes, support user dictionary, POS tagging, run JSON RPC service.
Support HMM cut text use Viterbi algorithm.
Text Segmentation speed single thread 9.2MB/s,goroutines concurrent 26.8MB/s. HMM text segmentation single thread 3.2MB/s. (2core 4threads Macbook Pro).
gse-bind, binding JavaScript and other, support more language.
go get -u github.com/hhjpin/gse
go get -u github.com/go-ego/re
To create a new gse application
$ re gse my-gse
To run the application we just created, you can navigate to the application folder and execute:
$ cd my-gse && re run
package main
import (
"fmt"
"github.com/hhjpin/gse"
)
var (
text = "你好世界, Hello world."
seg gse.Segmenter
)
func cut() {
hmm := seg.Cut(text, true)
fmt.Println("cut use hmm: ", hmm)
hmm = seg.CutSearch(text, true)
fmt.Println("cut search use hmm: ", hmm)
hmm = seg.CutAll(text)
fmt.Println("cut all: ", hmm)
}
func segCut() {
// Text Segmentation
tb := []byte(text)
fmt.Println(seg.String(tb, true))
segments := seg.Segment(tb)
// Handle word segmentation results
// Support for normal mode and search mode two participle,
// see the comments in the code ToString function.
// The search mode is mainly used to provide search engines
// with as many keywords as possible
fmt.Println(gse.ToString(segments, true))
}
func main() {
// Loading the default dictionary
seg.LoadDict()
// Load the dictionary
// seg.LoadDict("your gopath"+"/src/github.com/hhjpin/gse/data/dict/dictionary.txt")
cut()
segCut()
}
Look at an custom dictionary example
package main
import (
"fmt"
"github.com/hhjpin/gse"
)
func main() {
var seg gse.Segmenter
seg.LoadDict("zh,testdata/test_dict.txt,testdata/test_dict1.txt")
text1 := []byte("你好世界, Hello world")
fmt.Println(seg.String(text1, true))
segments := seg.Segment(text1)
fmt.Println(gse.ToString(segments))
}
Gse is primarily distributed under the terms of both the MIT license and the Apache License (Version 2.0), thanks for sego and jieba.