A Matrix-Based Heuristic Algorithm for Extracting Multiword Expressions from a Corpus

Linguistics and English Language

Electronic data

2022.mwe2022-1.7
Final published version, 1.02 MB, PDF document
Available under license: CC BY-NC: Creative Commons Attribution-NonCommercial 4.0 International License

View graph of relations

Research output: Contribution to conference - Without ISBN/ISSN › Conference paper › peer-review

Published

Orhan Bilgin

More...

Publication date	25/06/2022
Number of pages	12
Pages	37-48
<mark>Original language</mark>	English
Event	13th Language Resources and Evaluation Conference (LREC 2022): 18th Workshop on Multiword Expressions (MWE 2022) - Marseille, France Duration: 21/06/2022 → 25/06/2022 https://lrec2022.lrec-conf.org/en/

Conference

Conference	13th Language Resources and Evaluation Conference (LREC 2022)
Country/Territory	France
City	Marseille
Period	21/06/22 → 25/06/22
Internet address	https://lrec2022.lrec-conf.org/en/

Abstract

This paper describes an algorithm for automatically extracting multiword expressions (MWEs) from a corpus. The algorithm is node-based, ie extracts MWEs that contain the item specified by the user, using a fixed window-size around the node. The main idea is to detect the frequency anomalies that occur at the starting and ending points of an ngram that constitutes a MWE. This is achieved by locally comparing matrices of observed frequencies to matrices of expected frequencies, and determining, for each individual input, one or more sub-sequences that have the highest probability of being a MWE. Top-performing sub-sequences are then combined in a score-aggregation and ranking stage, thus producing a single list of score-ranked MWE candidates, without having to indiscriminately generate all possible sub-sequences of the input strings. The knowledge-poor and computationally efficient algorithm attempts to solve certain recurring problems in MWE extraction, such as the inability to deal with MWEs of arbitrary length, the repetitive counting of nested ngrams, and excessive sensitivity to frequency. Evaluation results show that the best-performing version generates top-50 precision values between 0.71 and 0.88 on Turkish and English data, and performs better than the baseline method even at n= 1000.

Research

Electronic data