logo
calendar30 Dekabr 2025
view14
Main language:Uzbek

Similarity detection based on Levenshtein distance in automatic bibliographic record linkage

Field of Science:Computer ScienceInformation Systems
pdf

63._Ishniyazov_O.O.__Ba....pdf

PDF

ARTICLE ANNOTATION

quote
This paper proposes and evaluates an approach based on the Levenshtein distance for determining string similarity in the process of automatic bibliographic record linkage. In bibliographic databases, spelling errors, transliteration differences, abbreviations, and formatting inconsistencies in fields such as author name, title, publisher, and publication year lead to “lost” links between records, reducing the completeness and accuracy of search results. Within the study, bibliographic record attributes are pre-normalized, after which a normalized similarity score based on the Levenshtein distance is calculated, and threshold values for match acceptance are experimentally evaluated. The experimental results demonstrate that threshold selection directly affects the balance between precision and recall metrics and confirm that the Levenshtein distance is an effective solution for candidate pair generation and the initial stage of record linkage in practical applications.

AUTHORS

O.Ishniyazov

"MUHAMMAD AL-XORAZMIY NOMIDAGI TOSHKENT AXBOROT TEXNOLOGIYALARI UNIVERSITETI" DAVLAT MUASSASASI

Tags

# precision# normalization# bibliographic record# record linkage# duplicate detection# Levenshtein distance# string similarity# threshold# recall# F1-score# electronic catalog# authority control

SIMILAR ARTICLES

OTHER ARTICLES IN THIS JOURNAL

Rate Article

0
0 ratings
5
4
3
2
1

Article Identifiers

References

Winkler W.E. Overview of Record Linkage and Current Research Directions. – U.S. Census Bureau, 2006.

Christen P. Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. – Springer, 2012.

Fellegi I.P., Sunter A.B. A Theory for Record Linkage // Journal of the American Statistical Association. – 1969. – Vol. 64, No. 328. – P. 1183–1210.

Hernández M.A., Stolfo S.J. Real-world data is dirty: Data cleansing and the merge/purge problem // Data Mining and Knowledge Discovery. – 1998. – Vol. 2, No. 1. – P. 9–37.

Hylton J.A. Identifying and Merging Related Bibliographic Records. – MIT Libraries, 2001.

Levenshtein V.I. Binary codes capable of correcting deletions, insertions, and reversals // Soviet Physics Doklady. – 1966. – Vol. 10. – P. 707–710.

Bilenko M., Mooney R.J. Adaptive Duplicate Detection Using Learnable String Similarity Measures // KDD. – 2003. – P. 39–48.

Christen P., Churches T. Febrl: A Freely Available Record Linkage System with a GUI. – Canberra, 2005.

Herzog T.N., Scheuren F.J., Winkler W.E. Data Quality and Record Linkage Techniques. – Springer, 2007.

Ishniyazov O. Linking model and algorithm of bibliographical database //AIP Conference Proceedings. – AIP Publishing LLC, 2024. – Т. 3147. – №. 1. – С. 030035.

Ishniyazov O. Levinstein's algorithm for comparing records // Multidiscipline Proceedings of “digital fashion conference”, 2024. – ISSN:2466-0744. – №.4(3). – P.11-13.