In data integration, entity resolution is an important technique to improve data quality. Existing researches typically assume that the target dataset only contain string-type data and use single similarity metric. For larger high-dimensional dataset, redundant information needs to be verified using traditional blocking or windowing techniques. In this work, we propose a novel ER-resolving method using a hybrid approach, including type-based multiblocks, varying window size, and more flexible similarity metrics. In our new ER workflow, we reduce the searching space for entity pairs by the constraint of redundant attributes and matching likelihood. We develop a reference implementation of our proposed approach and validate its performance using real-life dataset from one Internet of Things project. We evaluate the data processing system using five standard metrics including effectiveness, efficiency, accuracy, recall, and precision. Experimental results indicate that the proposed approach could be a promising alternative for entity resolution and could be feasibly applied in real-world data cleaning for large datasets.
from #AlexandrosSfakianakis via Alexandros G.Sfakianakis on Inoreader http://ift.tt/2Gq3pCS
via IFTTT
Εγγραφή σε:
Σχόλια ανάρτησης (Atom)
Δημοφιλείς αναρτήσεις
-
from #Medicine-SfakianakisAlexandros via o.lakala70 on Inoreader https://ift.tt/2Gchesc via IFTTT
-
Abstract Determining the cause of unexplained death in all age groups, including infants, is a priority in forensic medicine. The triple r...
-
from #Medicine-SfakianakisAlexandros via o.lakala70 on Inoreader https://ift.tt/2BeOBVJ via IFTTT
-
Abstract Layer-by-layer (LbL) dip coating, accompanying with the use of micelle structure, allows hydrophobic molecules to be coated on me...
-
from #Medicine-SfakianakisAlexandros via o.lakala70 on Inoreader https://ift.tt/2rxuJIO via IFTTT
-
Abstract In this paper we present the study of a skull belonging to a young male from the Italian Bronze Age showing three perimortem inju...
-
Find out more about the wide range of A Levels and full time courses available at Longley Park Sixth Form College, the only independent Sixt...
-
Abstract To measure integral doses in image-guided radiation therapy, we developed an integral condenser dosimeter comprising a disposable...
-
Objectives. To assess the association between short-term postoperative cognitive dysfuction (POCD) and inflammtory response in patients unde...
Δεν υπάρχουν σχόλια:
Δημοσίευση σχολίου