An offline English optical character recognition and NER using LSTM and adaptive neuro-fuzzy inference system

Suganthi, M.; Arun Prakash, R.

doi:10.3233/JIFS-221486

An offline English optical character recognition and NER using LSTM and adaptive neuro-fuzzy inference system

Article type: Research Article

Authors: Suganthi, M.^{a; *} | Arun Prakash, R.^b

Affiliations: [a] Roever Engineering College, Perambalur, Tamilnadu, India | [b] University College of Engineering, Ariyalur, Tamilnadu, India

Correspondence: [*] Corresponding author. M. Suganthi, Assistant Professor, Roever Engineering College, Perambalur, Tamilnadu, India. E-mail: [email protected].

Abstract: Everything becomes smart in the modern era, for everything we need a better plan or arrangements. In the olden days, essential information was noted as a document with the help of paper and pen or printed texts. But the intelligent world needs a paperless environment by converting handwritten or printed text documents into software copies. This can be achieved by the electronic data conversion concept called Optical Character Recognition (OCR). OCR of some documents is complex because of different writing styles and quality of scanned image issues, which can be solved by adopting a deep learning technique for better accuracy. We employed Long Short Term Memory (LSTM) for English Optical Character Recognition for paperless and effortless data storage and fast access in this work. Still, the records may contain the entities like names, contact details, drug details, diseases, educational qualifications, dates, etc. These entities cannot be separated by employing OCR alone; we need an entity recognition framework for deeper and faster data analysis. For efficient Named Entity Recognition, we utilize the Adaptive Fuzzy Inference System (ANFIS) powered by the algorithms CRF and BERT to automatically label each entity by training the vast amount of unlabeled text data. The ANFIS model is equipped with both linguistic and numerical knowledge. It is more accurate than the ANN when it comes to identifying patterns and classification data. Also, it is more transparent to the user. Our proposed framework aims to improve the performance of the character recognition system by using a feed-forward network. One of the main issues that have been identified in the development of this system is noise. Through this network, we can provide a single input and one output layer. The main components of the system are the training and recognition sections. These two sections are mainly focused on image acquisition and feature extraction. Besides these, they also include training and simulation of the classifier. The first step in the process of image recognition is to extract the features from the normalized image matrix. We then train the network using a proposed training algorithm. Experimentation on medical records attains a higher accuracy value of 0.9637, recall value of 0.9627, and f1 score of 0.9627, respectively.

Keywords: Paperless environment, Optical Character Recognition, Named Entity Recognition, faster data analysis, accuracy

DOI: 10.3233/JIFS-221486

Journal: Journal of Intelligent & Fuzzy Systems, vol. 44, no. 3, pp. 3877-3890, 2023

Published: 09 March 2023

Price: EUR 27.50

North America

IOS Press, Inc.
6751 Tepper Drive
Clifton, VA 20124
USA

Tel: +1 703 830 6300
Fax: +1 703 830 2300
[email protected]

For editorial issues, like the status of your submitted paper or proposals, write to [email protected]

Europe

IOS Press
Nieuwe Hemweg 6B
1013 BG Amsterdam
The Netherlands

Tel: +31 20 688 3355
Fax: +31 20 687 0091
[email protected]

For editorial issues, permissions, book requests, submissions and proceedings, contact the Amsterdam office [email protected]

Asia

Inspirees International (China Office)
Ciyunsi Beili 207(CapitaLand), Bld 1, 7-901
100025, Beijing
China

Free service line: 400 661 8717
Fax: +86 10 8446 7947
[email protected]

For editorial issues, like the status of your submitted paper or proposals, write to [email protected]

如果您在出版方面需要帮助或有任何建, 件至: [email protected]

Share this:

North America

Europe

Asia