Semantic Tagging for the Urdu Language - Research Portal

Home > Research > Publications & Outputs > Semantic Tagging for the Urdu Language

Associated organisational units

Electronic data

3582496 (1)
Accepted author manuscript, 560 KB, PDF document
Available under license: CC BY: Creative Commons Attribution 4.0 International License

Text available via DOI:

https://doi.org/10.1145/3582496
Final published version
Available under license: CC BY-NC: Creative Commons Attribution-NonCommercial 4.0 International License

View graph of relations

Semantic Tagging for the Urdu Language: Annotated Corpus and Multi-Target Classification Methods

Research output: Contribution to Journal/Magazine › Journal article › peer-review

Published

Jawad Shafi
Rao Muhammad Adeel Nawab
Paul Rayson

More...

Article number	175
<mark>Journal publication date</mark>	17/06/2023
<mark>Journal</mark>	ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP)
Issue number	6
Volume	22
Number of pages	32
Pages (from-to)	175:1-175:32
Publication Status	Published
Early online date	16/02/23
<mark>Original language</mark>	English

Abstract

Extracting and analysing meaning-related information from natural language data has attracted the attention of researchers in various fields, such as natural language processing, corpus linguistics, information retrieval, and data science. An important aspect of such automatic information extraction and analysis is the annotation of language data using semantic tagging tools. Different semantic tagging tools have been designed to carry out various levels of semantic analysis, for instance, named entity recognition and disambiguation, sentiment analysis, word sense disambiguation, content analysis, and semantic role labelling. Common to all of these tasks, in the supervised setting, is the requirement for a manually semantically annotated corpus, which acts as a knowledge base from which to train and test potential word and phrase-level sense annotations. Many benchmark corpora have been developed for various semantic tagging tasks, but most are for English and other European languages. There is a dearth of semantically annotated corpora for the Urdu language, which is widely spoken and used around the world. To fill this gap, this study presents a large benchmark corpus and methods for the semantic tagging task for the Urdu language. The proposed corpus contains 8,000 tokens in the following domains or genres: news, social media, Wikipedia, and historical text (each domain having 2K tokens). The corpus has been manually annotated with 21 major semantic fields and 232 sub-fields with the USAS (UCREL Semantic Analysis System) semantic taxonomy which provides a comprehensive set of semantic fields for coarse-grained annotation. Each word in our proposed corpus has been annotated with at least one and up to nine semantic field tags to provide a detailed semantic analysis of the language data, which allowed us to treat the problem of semantic tagging as a supervised multi-target classification task. To demonstrate how our proposed corpus can be used for the development and evaluation of Urdu semantic tagging methods, we extracted local, topical and semantic features from the proposed corpus and applied seven different supervised multi-target classifiers to them. Results show an accuracy of 94% on our proposed corpus which is free and publicly available to download.

Research

Associated organisational units

Electronic data

Links

Text available via DOI:

Semantic Tagging for the Urdu Language: Annotated Corpus and Multi-Target Classification Methods

Abstract

Quick Links

Connect With Us

Faculties & Depts

Contact Us