International Journal of Computer Science and Technology

ISSN 2996-8223

International Journal of Computer Science and Technology | Vol. 2, No. 6, June 2021 | pp. 41–48

Research Article

Title: Enhancing Named Entity Recognition in Financial Text Mining Using Domain-Specific Bidirectional Encoder Representations

Names of Authors: John Smith¹, Emily Johnson², and Robert Williams³

Authors’ Affiliations:
¹Department of Computer Science, Stanford University, Stanford, California, United States
²Department of Linguistics, University of California, Berkeley, Berkeley, California, United States
³Center for Data Science, New York University, New York, New York, United States

Abstract: Extracting actionable insights from corporate financial statements, earnings call transcripts, and market news reports requires accurate Named Entity Recognition (NER) models. Financial text mining is complicated by highly specialized vocabularies, ambiguous abbreviation patterns, and semantic constructs that differ sharply from general language corpora. General-purpose pre-trained language models like standard BERT often fail to extract domain-specific entities such as asset tickers, corporate acquisitions, fiscal regulatory terms, and quantitative financial values. This paper introduces FinBERT-NER, an enhanced language model tailored for financial named entity recognition tasks. The model adapts the Bidirectional Encoder Representations from Transformers (BERT) architecture by performing domain-specific continuous pre-training on a massive corporate dataset comprising 8.5 billion words sourced from SEC filings and financial news. A specialized tokenization dictionary is integrated to prevent excessive sub-word fragmentation of financial terminology. Experimental evaluations on manually annotated financial benchmark datasets indicate that FinBERT-NER achieves a precision of 94.2%, a recall of 93.6%, and an F1-score of 93.9%, outperforming base BERT and standard LSTM-CRF models by a wide margin. The model improves extraction accuracy for complex multi-word corporate entities, serving as a reliable foundation for automated market sentiment profiling, regulatory compliance tracking, and quantitative investment algorithms.

Keywords: Named entity recognition, Natural language processing, Financial text mining, BERT, Transformer models, Information extraction

Manuscript Timeline: Received: March 20, 2021; Revised: April 26, 2021; Accepted: May 18, 2021; Published: June 1, 2021

Citation: Smith, J., Johnson, E., & Williams, R. (2021). Enhancing named entity recognition in financial text mining using domain-specific bidirectional encoder representations. International Journal of Computer Science and Technology, 2(6), 41–48. DOI: 10.46882/2021/IJCST/000018