copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

IRS for Computer Character Sequences Filtration: a new software tool and algorithm to support the IRS at tokenization process

Q. Ahmad Al Badawi. International Journal of Advanced Computer Science and Applications(IJACSA), (2013)

Abstract

Tokenization is the task of chopping it up into pieces, called tokens, perhaps at the same time throwing away certain characters, such as punctuation. A token is an instance of token a sequence of characters in some particular document that are grouped together as a useful semantic unit for processing. New software tool and algorithm to support the IRS at tokenization process are presented. Our proposed tool will filter out the three computer character Sequences: IP-Addresses, Web URLs, Date, and Email Addresses. Our tool will use the pattern matching algorithms and filtration methods. After this process, the IRS can start a new tokenization process on the new retrieved text which will be free of these sequences.

Links and resources

BibTeX key: IJACSA.2013.040212
entry type: article
year: 2013
journal: International Journal of Advanced Computer Science and Applications(IJACSA)
number: 2
volume: 4
url: http://ijacsa.thesai.org/

BibSonomy

copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

IRS for Computer Character Sequences Filtration: a new software tool and algorithm to support the IRS at tokenization process

Abstract

Links and resources

Tags

Cite this publication

More citation styles

search on

Meta data

Comments and Reviews
(0)

BibSonomy

copydeleteadd this publication to your clipboardcommunity posthistory of this postURLDOIBibTeXEndNoteAPAChicagoDIN 1505HarvardMSOffice XML IRS for Computer Character Sequences Filtration: a new software tool and algorithm to support the IRS at tokenization process

Abstract

Links and resources

Tags

Cite this publication

More citation styles

search on

Meta data

Comments and Reviews (0)

copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

IRS for Computer Character Sequences Filtration: a new software tool and algorithm to support the IRS at tokenization process

Comments and Reviews
(0)