Paper
17 January 2005 Printed Arabic document recognition system
Author Affiliations +
Proceedings Volume 5676, Document Recognition and Retrieval XII; (2005) https://doi.org/10.1117/12.585711
Event: Electronic Imaging 2005, 2005, San Jose, California, United States
Abstract
As a cursive script, the characteristics of Arabic texts are different from Latin or Chinese greatly. For example, an Arabic character has up to four written forms and characters that can be joined are always joined on the baseline. Therefore, the methods used for Arabic document recognition are special, where character segmentation is the most critical problem. In this paper, a printed Arabic document recognition system is presented, which is composed of text line segmentation, word segmentation, character segmentation, character recognition and post-processing stages. In the beginning, a top-down and bottom-up hybrid method based on connected components classification is proposed to segment Arabic texts into lines and words. Subsequently, characters are segmented by analysis the word contour. At first the baseline position of a given word is estimated, and then a function denote the distance between contour and baseline is analyzed to find out all candidate segmentation points, at last structure rules are proposed to merge over-segmented characters. After character segmentation, both statistical features and structure features are used to do character recognition. Finally, lexicon is used to improve recognition results. Experiment shows that the recognition accuracy of the system has achieved 97.62%.
© (2005) COPYRIGHT Society of Photo-Optical Instrumentation Engineers (SPIE). Downloading of the abstract is permitted for personal use only.
Jianming Jin, Hua Wang, Xiaoqing Ding, and Liangrui Peng "Printed Arabic document recognition system", Proc. SPIE 5676, Document Recognition and Retrieval XII, (17 January 2005); https://doi.org/10.1117/12.585711
Lens.org Logo
CITATIONS
Cited by 22 scholarly publications.
Advertisement
Advertisement
RIGHTS & PERMISSIONS
Get copyright permission  Get copyright permission on Copyright Marketplace
KEYWORDS
Image segmentation

Optical character recognition

Astatine

Binary data

Intelligence systems

Statistical analysis

Computing systems

RELATED CONTENT

Intelligent word-based text recognition
Proceedings of SPIE (February 01 1991)
Music recognition system using ART-1 and GA
Proceedings of SPIE (March 06 2002)
Comparison of historical documents for writership
Proceedings of SPIE (January 18 2010)
Characteristics of digitized images of technical articles
Proceedings of SPIE (August 01 1992)

Back to Top