MuNPEx is a multi-lingual noun phrase (NP) extraction component developed for the GATE architecture, implemented in JAPE. It currently supports English, German, French, and Spanish (in beta).
MuNPEx requires a part-of-speech (POS) tagger to work and can additionally use detected named entities (NEs) to improve chunking performance. Please read the documentation (or source code) for more details.
Semantic MediaWiki (SMW) is a free extension of MediaWiki that helps to search, organise, tag, browse, evaluate, and share the wiki's content. While traditional wikis contain only texts which computers can neither understand nor evaluate, SMW adds semantic annotations that bring the power of the Semantic Web to the wiki.
Webstemmer is a web crawler and HTML layout analyzer that automatically extracts main text of a news site without having banners, ads and/or navigation links mixed up
Our goal is to develop a probabilistic knowledge base that mirrors the content of the web. We are developing a system that uses semi-supervised learning methods to learn to extract symbolic knowledge from unstructured text and HTML. We are exploring methods of continous learning, where our system runs 24x7, continuously learning to read better, and continuously extracting facts from the web.
M. Vargas-Vera, and D. Celjuska. WI '04: Proceedings of the 2004 IEEE/WIC/ACM International Conference on Web Intelligence, page 615--618. Washington, DC, USA, IEEE Computer Society, (2004)
R. Swan, and J. Allan. CIKM '99: Proceedings of the eighth international conference on Information and knowledge management, page 38--45. New York, NY, USA, ACM, (1999)
R. Grishman, and B. Sundheim. Proceedings of the 16th International Conference on Computational Linguistics (COLING), page 466--471. Kopenhagen, (1996)
Y. Li, R. Krishnamurthy, S. Raghavan, S. Vaithyanathan, and H. Jagadish. Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, page 21--30. Honolulu, Hawaii, Association for Computational Linguistics, (October 2008)
M. Mintz, S. Bills, R. Snow, and D. Jurafsky. Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP, page 1003--1011. Suntec, Singapore, Association for Computational Linguistics, (August 2009)
E. Riloff, and R. Jones. AAAI '99/IAAI '99: Proceedings of the sixteenth national conference on Artificial intelligence and the eleventh Innovative applications of artificial intelligence conference innovative applications of artificial intelligence, page 474--479. Menlo Park, CA, USA, American Association for Artificial Intelligence, (1999)
E. Riloff, C. Schafer, and D. Yarowsky. Proceedings of the 19th international conference on Computational linguistics, page 1--7. Morristown, NJ, USA, Association for Computational Linguistics, (2002)