Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

NGram issue with strings shorter than N #28

Closed
adamko147 opened this issue Sep 10, 2021 · 2 comments · Fixed by #29
Closed

NGram issue with strings shorter than N #28

adamko147 opened this issue Sep 10, 2021 · 2 comments · Fixed by #29

Comments

@adamko147
Copy link
Contributor

adamko147 commented Sep 10, 2021

NGram distance returns wrong value for strings shorter than N.

from strsimpy.ngram import NGram
ng = NGram()
ng.distance("abc", "abc") == 0.0
ng.distance("a", "b") == 0.0 # should be 1.0

Distance between a and b should be 1.0 as the strings are completely different. if N is two, then the code at https://github.com/luozhouyang/python-string-similarity/blob/master/strsimpy/ngram.py#L45 calculates cost of 0 and returns 1.0 * cost / max(sl, tl). This returns similarity (which actually is 0, because the strings are completely different). However the code is returning normalized distance, which should be maximum possible here.

This issue seems to be at https://github.com/luozhouyang/python-string-similarity/blob/master/strsimpy/ngram.py#L49 where it does

return 1.0 * cost / max(sl, tl)

however I think in this case it should return

return 1.0 - cost / max(sl, tl)

I'll create PR to address this.

Thank you

@adamko147
Copy link
Contributor Author

@luozhouyang any chance you can package a publish new version to pypi.org including this fix? Thanks a lot!

@luozhouyang
Copy link
Owner

@adamko147 just released v0.2.1! Thanks for you contribution!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
None yet
Projects
None yet
Development

Successfully merging a pull request may close this issue.

2 participants