Article,

Classifying Amharic webnews

L. Asker, A. Argaw, B. Gambäck, S. Eyassu Asfeha, and L. Nigussie Habte.
Information Retrieval, 12 (3): 416--435 (Jun 1, 2009)
DOI: 10.1007/s10791-008-9080-x

Abstract

We present work aimed at compiling an Amharic corpus from the Web and automatically categorizing the texts. Amharic is the second most spoken Semitic language in the World (after Arabic) and used for countrywide communication in Ethiopia. It is highly inflectional and quite dialectally diversified. We discuss the issues of compiling and annotating a corpus of Amharic news articles from the Web. This corpus was then used in three sets of text classification experiments. Working with a less-researched language highlights a number of practical issues that might otherwise receive less attention or go unnoticed. The purpose of the experiments has not primarily been to develop a cutting-edge text classification system for Amharic, but rather to put the spotlight on some of these issues. The first two sets of experiments investigated the use of Self-Organizing Maps (SOMs) for document classification. Testing on small datasets, we first looked at classifying unseen data into 10 predefined categories of news items, and then at clustering it around query content, when taking 16 queries as class labels. The second set of experiments investigated the effect of operations such as stemming and part-of-speech tagging on text classification performance. We compared three representations while constructing classification models based on bagging of decision trees for the 10 predefined news categories. The best accuracy was achieved using the full text as representation. A representation using only the nouns performed almost equally well, confirming the assumption that most of the information required for distinguishing between various categories actually is contained in the nouns, while stemming did not have much effect on the performance of the classifier.

BibTeX key: Asker2009
entry type: article
year: 2009
month: jun
day: 01
journal: Information Retrieval
number: 3
pages: 416--435
volume: 12
issn: 1573-7659
DOI: 10.1007/s10791-008-9080-x
url: https://doi.org/10.1007/s10791-008-9080-x

Users

Comments and Reviewsshow / hide

Please log in to take part in the discussion (add own reviews or comments).

Cite this publication

%0 Journal Article %1 Asker2009 %A Asker, Lars %A Argaw, Atelach Alemu %A Gambäck, Björn %A Eyassu Asfeha, Samuel %A Nigussie Habte, Lemma %D 2009 %J Information Retrieval %K amharic classification corpus ethiopic %N 3 %P 416--435 %R 10.1007/s10791-008-9080-x %T Classifying Amharic webnews %U https://doi.org/10.1007/s10791-008-9080-x %V 12 %X We present work aimed at compiling an Amharic corpus from the Web and automatically categorizing the texts. Amharic is the second most spoken Semitic language in the World (after Arabic) and used for countrywide communication in Ethiopia. It is highly inflectional and quite dialectally diversified. We discuss the issues of compiling and annotating a corpus of Amharic news articles from the Web. This corpus was then used in three sets of text classification experiments. Working with a less-researched language highlights a number of practical issues that might otherwise receive less attention or go unnoticed. The purpose of the experiments has not primarily been to develop a cutting-edge text classification system for Amharic, but rather to put the spotlight on some of these issues. The first two sets of experiments investigated the use of Self-Organizing Maps (SOMs) for document classification. Testing on small datasets, we first looked at classifying unseen data into 10 predefined categories of news items, and then at clustering it around query content, when taking 16 queries as class labels. The second set of experiments investigated the effect of operations such as stemming and part-of-speech tagging on text classification performance. We compared three representations while constructing classification models based on bagging of decision trees for the 10 predefined news categories. The best accuracy was achieved using the full text as representation. A representation using only the nouns performed almost equally well, confirming the assumption that most of the information required for distinguishing between various categories actually is contained in the nouns, while stemming did not have much effect on the performance of the classifier.

@article{Asker2009, abstract = {We present work aimed at compiling an Amharic corpus from the Web and automatically categorizing the texts. Amharic is the second most spoken Semitic language in the World (after Arabic) and used for countrywide communication in Ethiopia. It is highly inflectional and quite dialectally diversified. We discuss the issues of compiling and annotating a corpus of Amharic news articles from the Web. This corpus was then used in three sets of text classification experiments. Working with a less-researched language highlights a number of practical issues that might otherwise receive less attention or go unnoticed. The purpose of the experiments has not primarily been to develop a cutting-edge text classification system for Amharic, but rather to put the spotlight on some of these issues. The first two sets of experiments investigated the use of Self-Organizing Maps (SOMs) for document classification. Testing on small datasets, we first looked at classifying unseen data into 10 predefined categories of news items, and then at clustering it around query content, when taking 16 queries as class labels. The second set of experiments investigated the effect of operations such as stemming and part-of-speech tagging on text classification performance. We compared three representations while constructing classification models based on bagging of decision trees for the 10 predefined news categories. The best accuracy was achieved using the full text as representation. A representation using only the nouns performed almost equally well, confirming the assumption that most of the information required for distinguishing between various categories actually is contained in the nouns, while stemming did not have much effect on the performance of the classifier.}, added-at = {2018-02-26T10:03:38.000+0100}, author = {Asker, Lars and Argaw, Atelach Alemu and Gamb{\"a}ck, Bj{\"o}rn and Eyassu Asfeha, Samuel and Nigussie Habte, Lemma}, biburl = {https://www.bibsonomy.org/bibtex/22b65a74e93e4c7d9dd2be4459e9c4cd6/asmelash}, day = 01, description = {Classifying Amharic webnews | SpringerLink}, doi = {10.1007/s10791-008-9080-x}, interhash = {9f82d9e5d5e87ea048233942b1a98b50}, intrahash = {2b65a74e93e4c7d9dd2be4459e9c4cd6}, issn = {1573-7659}, journal = {Information Retrieval}, keywords = {amharic classification corpus ethiopic}, month = jun, number = 3, pages = {416--435}, timestamp = {2018-02-26T10:03:38.000+0100}, title = {Classifying Amharic webnews}, url = {https://doi.org/10.1007/s10791-008-9080-x}, volume = 12, year = 2009 }

BibSonomy

Classifying Amharic webnews

Abstract

Tags

Users

Comments and Reviewsshow / hide

Cite this publication

More citation styles

search on