@schmidt2

Automatically incorporating new sources in keyword search-based data integration

, , and . Proceedings of the 2010 international conference on Management of data, page 387--398. New York, NY, USA, ACM, (2010)
DOI: 10.1145/1807167.1807211

Abstract

Scientific data offers some of the most interesting challenges in data integration today. Scientific fields evolve rapidly and accumulate masses of observational and experimental data that needs to be annotated, revised, interlinked, and made available to other scientists. From the perspective of the user, this can be a major headache as the data they seek may initially be spread across many databases in need of integration. Worse, even if users are given a solution that integrates the current state of the source databases, <i>new</i> data sources appear with new data items of interest to the user.</p> <p>Here we build upon recent ideas for creating integrated views over data sources using keyword search techniques, ranked answers, and user feedback 32 to investigate how to <i>automatically discover</i> when a new data source has content relevant to a user's view - in essence, performing <i>automatic data integration</i> for incoming data sets. The new architecture accommodates a variety of methods to discover related attributes, including <i>label propagation</i> algorithms from the machine learning community 2 and existing schema matchers 11. The user may provide <i>feedback</i> on the suggested new results, helping the system <i>repair</i> any bad alignments or <i>increase the cost</i> of including a new source that is not useful. We evaluate our approach on actual bioinformatics schemas and data, using state-of-the-art schema matchers as components. We also discuss how our architecture can be adapted to more traditional settings with a mediated schema.

Description

Automatically incorporating new sources in keyword search-based data integration

Links and resources

Tags

community

  • @schmidt2
  • @dblp
@schmidt2's tags highlighted