International Journal of Computer Science and Technology

ISSN 2996-8223

International Journal of Computer Science and Technology | Vol. 4, No. 4, April 2023 | pp. 1–8

DOI: 10.46882/2023/IJCST/000219

Article Type: Original Research Paper

Title: Enhancing Cross-Lingual Named Entity Recognition via Contextual Word Embedding Alignment

Names of Authors: Priya Nair¹, Hans-Jürgen Integrated²

Authors’ Affiliations: ¹Center for Computational Linguistics, Indian Institute of Technology, Chennai, India; ²Computational Linguistics Division, Munich Language Labs, Munich, Germany

Abstract: Named Entity Recognition (NER) models achieve high accuracy when trained on large, labeled English datasets but perform poorly in low-resource languages. Manually labeling text corpora across multiple languages is expensive and time-consuming. This paper introduces a cross-lingual NER framework that transfers entity knowledge from English to resource-scarce target languages without requiring manual translations. The framework uses a mathematical alignment layer that projects localized target-language text features into a shared semantic vector space alongside pre-trained English models. We incorporated a token-level attention module to handle variations in word order and syntax across languages. The model was validated on Spanish, Hindi, and Turkish text datasets. The experimental results show a mean F1-score improvement of 11.2% over existing unsupervised cross-lingual models. The framework proved highly effective at identifying complex entities like corporate names and geographic locations across different language syntax models. This alignment strategy provides an efficient method to scale multilingual text analytics tools.

Keywords: Named Entity Recognition, Cross-Lingual Transfer, Natural Language Processing, Word Embeddings, Low-Resource Languages, Semantic Alignment

Manuscript Timeline: Received: June 12, 2022; Revised: July 18, 2022; Accepted: August 08, 2022; Published: April 09, 2023