In data integration, entity resolution is an important technique to improve data quality. Existing researches typically assume that the target dataset only contain string-type data and use single similarity metric. For larger high-dimensional dataset, redundant information needs to be verified using traditional blocking or windowing techniques. In this work, we propose a novel ER-resolving method using a hybrid approach, including type-based multiblocks, varying window size, and more flexible similarity metrics. In our new ER workflow, we reduce the searching space for entity pairs by the constraint of redundant attributes and matching likelihood. We develop a reference implementation of our proposed approach and validate its performance using real-life dataset from one Internet of Things project. We evaluate the data processing system using five standard metrics including effectiveness, efficiency, accuracy, recall, and precision. Experimental results indicate that the proposed approach could be a promising alternative for entity resolution and could be feasibly applied in real-world data cleaning for large datasets.
from #AlexandrosSfakianakis via Alexandros G.Sfakianakis on Inoreader http://ift.tt/2Gq3pCS
via IFTTT
Εγγραφή σε:
Σχόλια ανάρτησης (Atom)
Δημοφιλείς αναρτήσεις
-
from #AlexandrosSfakianakis via Alexandros G.Sfakianakis on Inoreader http://ift.tt/2nL9dMr via IFTTT
-
Vol.30 from #AlexandrosSfakianakis via Alexandros G.Sfakianakis on Inoreader http://ift.tt/2nItCSB via IFTTT
-
Background Although pneumonia is a leading cause of death in New York City (NYC), limited data exist about the settings in which pneumonia ...
-
Summary We tested whether prophylactic droperidol and ondansetron, in combination with a moderate dose of dexamethasone, were equally effe...
-
by Demin Li, Carol Bentley, Jenna Yates, Maryam Salimi, Jenny Greig, Sarah Wiblin, Tasneem Hassanali, Alison H. Banham Therapeutic monoclon...
-
from #AlexandrosSfakianakis via Alexandros G.Sfakianakis on Inoreader http://ift.tt/1HDudvw via IFTTT
-
ACS Nano DOI: 10.1021/acsnano.6b08567 from #AlexandrosSfakianakis via Alexandros G.Sfakianakis on Inoreader http://ift.tt/2oNpdhD via...
-
Vol.69 No.3 from #AlexandrosSfakianakis via Alexandros G.Sfakianakis on Inoreader http://ift.tt/2ltDWNq via IFTTT
-
Abstract Dermoscopy has demonstrated clinical benefits in improving early melanoma diagnosis and reducing unnecessary biopsies. Despite th...
Δεν υπάρχουν σχόλια:
Δημοσίευση σχολίου