You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
NGram distance returns wrong value for strings shorter than N.
fromstrsimpy.ngramimportNGramng=NGram()
ng.distance("abc", "abc") ==0.0ng.distance("a", "b") ==0.0# should be 1.0
Distance between a and b should be 1.0 as the strings are completely different. if N is two, then the code at https://github.com/luozhouyang/python-string-similarity/blob/master/strsimpy/ngram.py#L45 calculates cost of 0 and returns 1.0 * cost / max(sl, tl). This returns similarity (which actually is 0, because the strings are completely different). However the code is returning normalized distance, which should be maximum possible here.
NGram distance returns wrong value for strings shorter than N.
Distance between
a
andb
should be1.0
as the strings are completely different. if N is two, then the code at https://github.com/luozhouyang/python-string-similarity/blob/master/strsimpy/ngram.py#L45 calculates cost of 0 and returns1.0 * cost / max(sl, tl)
. This returns similarity (which actually is 0, because the strings are completely different). However the code is returning normalized distance, which should be maximum possible here.This issue seems to be at https://github.com/luozhouyang/python-string-similarity/blob/master/strsimpy/ngram.py#L49 where it does
however I think in this case it should return
I'll create PR to address this.
Thank you
The text was updated successfully, but these errors were encountered: