copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Textual Resource Acquisition and Engineering

J. Chu-Carroll, J. Fan, N. Schlaefer, and W. Zadrozny. IBM Journal of Research and Development, 56 (3/4): 4:1--4:11 (2012)
DOI: 10.1147/JRD.2012.2185901

Abstract

A key requirement for high-performing question-answering (QA) systems is access to high-quality reference corpora from which answers to questions can be hypothesized and evaluated. However, the topic of source acquisition and engineering has received very little attention so far. This is because most existing systems were developed under organized evaluation efforts that included reference corpora as part of the task specification. The task of answering Jeopardy! questions, on the other hand, does not come with such a well-circumscribed set of relevant resources. Therefore, it became part of the IBM Watson effort to develop a set of well-defined procedures to acquire high-quality resources that can effectively support a high-performing QA system. To this end, we developed three procedures, i.e., source acquisition, source transformation, and source expansion. Source acquisition is an iterative development process of acquiring new collections to cover salient topics deemed to be gaps in existing resources based on principled error analysis. Source transformation refers to the process in which information is extracted from existing sources, either as a whole or in part, and is represented in a form that the system can most easily use. Finally, source expansion attempts to increase the coverage in the content of each known topic by adding new information as well as lexical and syntactic variations of existing information extracted from external large collections. In this paper, we discuss the methodology that we developed for IBM Watson for performing acquisition, transformation, and expansion of textual resources. We demonstrate the effectiveness of each technique through its impact on candidate recall and on end-to-end QA performance.

Links and resources

BibTeX key: ChuCarrollFanEtAl12ibmjrd1
entry type: article
year: 2012
journal: IBM Journal of Research and Development
number: 3/4
pages: 4:1--4:11
volume: 56
file: IEEE Digital Library:2012/ChuCarrollFanEtAl12ibmjrd1.pdf:PDF
issn: 0018-8646
groups: public
intrahash: 5d8a81ea869f03e87004aef4912c62d3
DOI: 10.1147/JRD.2012.2185901
timestamp: 2012.05.04
username: flint63

@flint63's tags highlighted

Cite this publication

%0 Journal Article %1 ChuCarrollFanEtAl12ibmjrd1 %A Chu-Carroll, Jennifer %A Fan, James %A Schlaefer, Nico %A Zadrozny, Wlodek %D 2012 %J IBM Journal of Research and Development %K 01801 ieee paper ibm ai language processing information retrieval system engineering development tool zzz.iui %N 3/4 %P 4:1--4:11 %R 10.1147/JRD.2012.2185901 %T Textual Resource Acquisition and Engineering %V 56 %X A key requirement for high-performing question-answering (QA) systems is access to high-quality reference corpora from which answers to questions can be hypothesized and evaluated. However, the topic of source acquisition and engineering has received very little attention so far. This is because most existing systems were developed under organized evaluation efforts that included reference corpora as part of the task specification. The task of answering Jeopardy! questions, on the other hand, does not come with such a well-circumscribed set of relevant resources. Therefore, it became part of the IBM Watson effort to develop a set of well-defined procedures to acquire high-quality resources that can effectively support a high-performing QA system. To this end, we developed three procedures, i.e., source acquisition, source transformation, and source expansion. Source acquisition is an iterative development process of acquiring new collections to cover salient topics deemed to be gaps in existing resources based on principled error analysis. Source transformation refers to the process in which information is extracted from existing sources, either as a whole or in part, and is represented in a form that the system can most easily use. Finally, source expansion attempts to increase the coverage in the content of each known topic by adding new information as well as lexical and syntactic variations of existing information extracted from external large collections. In this paper, we discuss the methodology that we developed for IBM Watson for performing acquisition, transformation, and expansion of textual resources. We demonstrate the effectiveness of each technique through its impact on candidate recall and on end-to-end QA performance.

@article{ChuCarrollFanEtAl12ibmjrd1, abstract = {A key requirement for high-performing question-answering (QA) systems is access to high-quality reference corpora from which answers to questions can be hypothesized and evaluated. However, the topic of source acquisition and engineering has received very little attention so far. This is because most existing systems were developed under organized evaluation efforts that included reference corpora as part of the task specification. The task of answering Jeopardy! questions, on the other hand, does not come with such a well-circumscribed set of relevant resources. Therefore, it became part of the IBM Watson effort to develop a set of well-defined procedures to acquire high-quality resources that can effectively support a high-performing QA system. To this end, we developed three procedures, i.e., source acquisition, source transformation, and source expansion. Source acquisition is an iterative development process of acquiring new collections to cover salient topics deemed to be gaps in existing resources based on principled error analysis. Source transformation refers to the process in which information is extracted from existing sources, either as a whole or in part, and is represented in a form that the system can most easily use. Finally, source expansion attempts to increase the coverage in the content of each known topic by adding new information as well as lexical and syntactic variations of existing information extracted from external large collections. In this paper, we discuss the methodology that we developed for IBM Watson for performing acquisition, transformation, and expansion of textual resources. We demonstrate the effectiveness of each technique through its impact on candidate recall and on end-to-end QA performance.}, added-at = {2017-11-13T14:44:56.000+0100}, author = {Chu-Carroll, Jennifer and Fan, James and Schlaefer, Nico and Zadrozny, Wlodek}, biburl = {https://www.bibsonomy.org/bibtex/25d8a81ea869f03e87004aef4912c62d3/flint63}, doi = {10.1147/JRD.2012.2185901}, file = {IEEE Digital Library:2012/ChuCarrollFanEtAl12ibmjrd1.pdf:PDF}, groups = {public}, interhash = {02b69528073d6d5acaf44b59f5826c54}, intrahash = {5d8a81ea869f03e87004aef4912c62d3}, issn = {0018-8646}, journal = {IBM Journal of Research and Development}, keywords = {01801 ieee paper ibm ai language processing information retrieval system engineering development tool zzz.iui}, number = {3/4}, pages = {4:1--4:11}, timestamp = {2018-04-16T12:08:21.000+0200}, title = {Textual Resource Acquisition and Engineering}, username = {flint63}, volume = 56, year = 2012 }

BibSonomy

copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Textual Resource Acquisition and Engineering

Abstract

Links and resources

Tags

community

Cite this publication

More citation styles

search on

Meta data

Comments and Reviews
(0)

BibSonomy

copydeleteadd this publication to your clipboardcommunity posthistory of this postURLDOIBibTeXEndNoteAPAChicagoDIN 1505HarvardMSOffice XML Textual Resource Acquisition and Engineering

Abstract

Links and resources

Tags

community

Cite this publication

More citation styles

search on

Meta data

Comments and Reviews (0)

copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Textual Resource Acquisition and Engineering

Comments and Reviews
(0)