NGram issue with strings shorter than N #28

adamko147 · 2021-09-10T06:31:21Z

NGram distance returns wrong value for strings shorter than N.

from strsimpy.ngram import NGram
ng = NGram()
ng.distance("abc", "abc") == 0.0
ng.distance("a", "b") == 0.0 # should be 1.0

Distance between a and b should be 1.0 as the strings are completely different. if N is two, then the code at https://github.com/luozhouyang/python-string-similarity/blob/master/strsimpy/ngram.py#L45 calculates cost of 0 and returns 1.0 * cost / max(sl, tl). This returns similarity (which actually is 0, because the strings are completely different). However the code is returning normalized distance, which should be maximum possible here.

This issue seems to be at https://github.com/luozhouyang/python-string-similarity/blob/master/strsimpy/ngram.py#L49 where it does

return 1.0 * cost / max(sl, tl)

however I think in this case it should return

return 1.0 - cost / max(sl, tl)

I'll create PR to address this.

Thank you

The text was updated successfully, but these errors were encountered:

adamko147 · 2021-09-10T09:07:22Z

@luozhouyang any chance you can package a publish new version to pypi.org including this fix? Thanks a lot!

luozhouyang · 2021-09-10T09:15:46Z

@adamko147 just released v0.2.1! Thanks for you contribution!

adamko147 mentioned this issue Sep 10, 2021

fix ngram distance for strings shorter than N #29

Merged

luozhouyang closed this as completed in #29 Sep 10, 2021

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

NGram issue with strings shorter than N #28

NGram issue with strings shorter than N #28

adamko147 commented Sep 10, 2021 •

edited

Loading

adamko147 commented Sep 10, 2021

luozhouyang commented Sep 10, 2021

NGram issue with strings shorter than N #28

NGram issue with strings shorter than N #28

Comments

adamko147 commented Sep 10, 2021 • edited Loading

adamko147 commented Sep 10, 2021

luozhouyang commented Sep 10, 2021

adamko147 commented Sep 10, 2021 •

edited

Loading