Recognition Of Printed Arabic-english Text
Husni Al-Muhtaseb, Lahouari Ghouti, محمد حسني نجيب يحيى
كلية علوم وهندسة الحاسب الآلي-جامعة الملك فهد للبترول والمعادن · السعودية
Based on its content, an image document in Optical Character Recognition (OCR) field can be either monolingual image document or multilingual image document. Typically, OCR systems are designed to address script of a specific language (e.g. English) and their recognition rate will decline if they are used directly (i.e. without modification) to address other languages (e.g. Chinese). Multilingual OCR systems arise to address challenges and difficulties of processing multilingual documents. In this research work, we address the challenges of developing a robust printed Arabic/English text recognition prototype using Hidden Markov Models (HMMs). Many research works have been carried out to investigate the technologies and approaches of monolingual, bilingual and multilingual text recognition. Our developed recognition prototype consists of four main modules: Data Preparation, feature extraction, classification, and post-processing. The lack of availability of bilingual text datasets, particularly Arabic/English datasets, motivates us to build a bilingual dataset consists of Arabic/English text images. The presented dataset consists of 10 sets; each contains 777 binary images of Arabic, English, and bilingual text lines and written with one of 10 fonts. In the feature extraction module, features based on statistical information (e.g. pixels density) were extracted from text line images using different structures using sliding window technique. In the classification module, HMMs with different settings and parameters were used in the recognition experiments of the bilingual single-font text images as well as the bilingual multi-font text images. Two types of images were used in the classification module; digitized images and scanned images. Recognition performance was evaluated using accuracy and correctness percentages. For the digitized single-font text images classification, the achieved highest correctness was 99.98% and the highest accuracy was 99.98%. The font that has the highest recognition rates was Tahoma. For the digitized multi-fonts text images classification, the achieved highest correctness was 98.06% and the highest accuracy was 97.70%. For the scanned single-font text images classification, the achieved highest correctness was 99.04% and the highest accuracy was 98.98%. The font that has the highest recognition rates was Tahoma. For the scanned multi-font text images classification, the achieved highest correctness was 97.07% and the highest accuracy was 96.62%. In addition to the classification experiments, several language identification experiments were conducted using the same modules of the recognition prototype. The achieved highest rate was 99.98% for Tahoma font. In the post-processing module, a post-processing methodology was developed to correct the recognition errors through voting mechanism, and spellchecking and correction. Using voting mechanism, the developed methodology showed an enhancement in the correctness rate by 1.16% and in the accuracy rate by 1.15%. Using spellchecking and correction, the developed methodology has corrected 68% of misrecognized words.