ISSN 2996-8223
International Journal of Computer Science and Technology | Vol. 4, No. 4, April 2023 | pp. 1–8
DOI: 10.46882/2023/IJCST/000219
Article Type: Original Research Paper
Title: Enhancing Cross-Lingual Named Entity Recognition via Contextual Word Embedding Alignment
Names of Authors: Priya Nair¹, Hans-Jürgen Integrated²
Authors’ Affiliations: ¹Center for Computational Linguistics, Indian Institute of Technology, Chennai, India; ²Computational Linguistics Division, Munich Language Labs, Munich, Germany
Abstract: Named Entity Recognition (NER) models achieve high accuracy when trained on large, labeled English datasets but perform poorly in low-resource languages. Manually labeling text corpora across multiple languages is expensive and time-consuming. This paper introduces a cross-lingual NER framework that transfers entity knowledge from English to resource-scarce target languages without requiring manual translations. The framework uses a mathematical alignment layer that projects localized target-language text features into a shared semantic vector space alongside pre-trained English models. We incorporated a token-level attention module to handle variations in word order and syntax across languages. The model was validated on Spanish, Hindi, and Turkish text datasets. The experimental results show a mean F1-score improvement of 11.2% over existing unsupervised cross-lingual models. The framework proved highly effective at identifying complex entities like corporate names and geographic locations across different language syntax models. This alignment strategy provides an efficient method to scale multilingual text analytics tools.
Keywords: Named Entity Recognition, Cross-Lingual Transfer, Natural Language Processing, Word Embeddings, Low-Resource Languages, Semantic Alignment
Manuscript Timeline: Received: June 12, 2022; Revised: July 18, 2022; Accepted: August 08, 2022; Published: April 09, 2023
Subscribe to read the full article: https://internationalscholarsjournals.org/subscribe-to-read