copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Mining time-changing data streams

G. Hulten, L. Spencer, and P. Domingos. KDD '01: Proceedings of the seventh ACM SIGKDD international conference on Knowledge discovery and data mining, page 97--106. New York, NY, USA, ACM, (2001)
DOI: http://doi.acm.org/10.1145/502512.502529

Abstract

Most statistical and machine-learning algorithms assume that the data is a random sample drawn from a stationary distribution. Unfortunately, most of the large databases available for mining today violate this assumption. They were gathered over months or years, and the underlying processes generating them changed during this time, sometimes radically. Although a number of algorithms have been proposed for learning time-changing concepts, they generally do not scale well to very large databases. In this paper we propose an efficient algorithm for mining decision trees from continuously-changing data streams, based on the ultra-fast VFDT decision tree learner. This algorithm, called CVFDT, stays current while making the most of old data by growing an alternative subtree whenever an old one becomes questionable, and replacing the old with the new when the new becomes more accurate. CVFDT learns a model which is similar in accuracy to the one that would be learned by reapplying VFDT to a moving window of examples every time a new example arrives, but with O(1) complexity per example, as opposed to O(w), where w is the size of the window. Experiments on a set of large time-changing data streams demonstrate the utility of this approach.

Description

Mining time-changing data streams

Links and resources

BibTeX key: 502529
entry type: inproceedings
address: New York, NY, USA
booktitle: KDD '01: Proceedings of the seventh ACM SIGKDD international conference on Knowledge discovery and data mining
year: 2001
pages: 97--106
publisher: ACM
location: San Francisco, California
isbn: 1-58113-391-X
DOI: http://doi.acm.org/10.1145/502512.502529
url: http://portal.acm.org/citation.cfm?id=502512.502529

@jamesh's tags highlighted

Cite this publication

@inproceedings{502529, abstract = {Most statistical and machine-learning algorithms assume that the data is a random sample drawn from a stationary distribution. Unfortunately, most of the large databases available for mining today violate this assumption. They were gathered over months or years, and the underlying processes generating them changed during this time, sometimes radically. Although a number of algorithms have been proposed for learning time-changing concepts, they generally do not scale well to very large databases. In this paper we propose an efficient algorithm for mining decision trees from continuously-changing data streams, based on the ultra-fast VFDT decision tree learner. This algorithm, called CVFDT, stays current while making the most of old data by growing an alternative subtree whenever an old one becomes questionable, and replacing the old with the new when the new becomes more accurate. CVFDT learns a model which is similar in accuracy to the one that would be learned by reapplying VFDT to a moving window of examples every time a new example arrives, but with O(1) complexity per example, as opposed to O(w), where w is the size of the window. Experiments on a set of large time-changing data streams demonstrate the utility of this approach.}, added-at = {2009-01-13T04:42:14.000+0100}, address = {New York, NY, USA}, author = {Hulten, Geoff and Spencer, Laurie and Domingos, Pedro}, biburl = {https://www.bibsonomy.org/bibtex/2a6733f71996409ce88594a6b9b61f728/jamesh}, booktitle = {KDD '01: Proceedings of the seventh ACM SIGKDD international conference on Knowledge discovery and data mining}, description = {Mining time-changing data streams}, doi = {http://doi.acm.org/10.1145/502512.502529}, interhash = {6f9e0b7ff453e9a74a64ce8e96b880e5}, intrahash = {a6733f71996409ce88594a6b9b61f728}, isbn = {1-58113-391-X}, keywords = {change classification}, location = {San Francisco, California}, pages = {97--106}, publisher = {ACM}, timestamp = {2009-01-13T04:42:14.000+0100}, title = {Mining time-changing data streams}, url = {http://portal.acm.org/citation.cfm?id=502512.502529}, year = 2001 }

BibSonomy

copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Mining time-changing data streams

Abstract

Description

Links and resources

Tags

community

Cite this publication

More citation styles

search on

Meta data

Comments and Reviews
(0)

BibSonomy

copydeleteadd this publication to your clipboardcommunity posthistory of this postURLDOIBibTeXEndNoteAPAChicagoDIN 1505HarvardMSOffice XML Mining time-changing data streams

Abstract

Description

Links and resources

Tags

community

Cite this publication

More citation styles

search on

Meta data

Comments and Reviews (0)

copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Mining time-changing data streams

Comments and Reviews
(0)