In data integration, entity resolution is an important technique to improve data quality. Existing researches typically assume that the target dataset only contain string-type data and use single similarity metric. For larger high-dimensional dataset, redundant information needs to be verified using traditional blocking or windowing techniques. In this work, we propose a novel ER-resolving method using a hybrid approach, including type-based multiblocks, varying window size, and more flexible similarity metrics. In our new ER workflow, we reduce the searching space for entity pairs by the constraint of redundant attributes and matching likelihood. We develop a reference implementation of our proposed approach and validate its performance using real-life dataset from one Internet of Things project. We evaluate the data processing system using five standard metrics including effectiveness, efficiency, accuracy, recall, and precision. Experimental results indicate that the proposed approach could be a promising alternative for entity resolution and could be feasibly applied in real-world data cleaning for large datasets.
from #AlexandrosSfakianakis via Alexandros G.Sfakianakis on Inoreader http://ift.tt/2Gq3pCS
via IFTTT
Εγγραφή σε:
Σχόλια ανάρτησης (Atom)
Δημοφιλείς αναρτήσεις
-
Treatment with a combination of ipilimumab and Coxsackievirus A21 led to durable responses in a number of patients with advanced melanoma, i...
-
3 TerTiary essay WriTing Essays are a common form of assessment in many tertiary-level disciplines. The ability to construct good essays inv...
-
What is a Critical Essay? A critical essay is a critique or review of another work, usually one which is arts related (. book, play, movie, ...
-
from #AlexandrosSfakianakis via Alexandros G.Sfakianakis on Inoreader http://ift.tt/2oUXfBR via IFTTT
-
Related Articles Extending the theoretical framework for curriculum integration in pre-clinical medical education. Perspect Med Educ....
-
bmj;357/apr04_10/j1651/FAF1faAfter registration, Alistair Peter Macdonald served with the Royal Army Medical Corps in Cyprus and Somaliland ...
-
Abstract Research on sex-related brain asymmetries has not yielded consistent results. Despite its importance to further understanding of n...
-
Exciting news from ecancer. We are now fully accredited medical education provider status by the EACCME.… https://t.co/DMfGvDyn7b from #Al...
-
The following details unlockables in Resident Evil 4. This is content players do not initially have access to. This does not include items h...
Δεν υπάρχουν σχόλια:
Δημοσίευση σχολίου