University of Wollongong in Dubai - Papers

Evaluation of statistical part of speech tagging of Persian text

Samira Tasharoft, University of Tehran
Fahimeh Raja, University of Tehran
Farhad Oroumchian, University of Wollongong in DubaiFollow
Masoud Rahgozar, University of Tehran

RIS ID

23504

Publication Details

This conference paper was originally published as Tasharofi, S., Raja, F., Oroumchian, F. & Rahgozar, M. 2007, 'Evaluation of statistical part of speech tagging of persian text', Proceedings of the 9th International Symposium on Signal Processing and its Applications, IEEE, Piscataway, USA, pp. 1-4.

Abstract

Part of Speech (POS) tagging is an essential part of text processing applications. A POS tagger assigns a tag to each word of its input text specifying its grammatical properties. One of the popular POS taggers is TnT tagger which was shown to have high accuracy in English and some other languages. It is always interesting to see how a method in one language performs on another language because it would give us insight into the difference and similarities of the languages. In case of statistical methods such as TnT, this will have an added practical advantages also. This paper presents creation of a POS tagged corpus and evaluation of TnT tagger on Persian text. The results of experiments on Persian text show that TnT provides overall tagging accuracy of 96.64%, specifically, 97.01% on known words and 77.77% on unknown words.

Download

COinS

University of Wollongong in Dubai - Papers

Evaluation of statistical part of speech tagging of Persian text

RIS ID

Publication Details

Abstract

Search

Browse

Author Corner

Links

University of Wollongong in Dubai - Papers

Evaluation of statistical part of speech tagging of Persian text

Authors

RIS ID

Publication Details

Abstract

Share

Search

Browse

Author Corner

Links