Article,

A conclusive methodology for rating OCR performance

, , and .
Journal of the American Society for Information Science and Technology, 56 (12): 1274--1287 (October 2005)
DOI: 10.1002/asi.20214

Abstract

One of the most challenging topics in the automatic document rating process is the development of a rating scheme for the image quality of documents. As part of the Department of Energy (DOE) document declassification program, we have developed a generalized rating system to predict the optical character recognition (OCR) accuracy level that is achieved when processing a document. The need for such a system emerged from the declassification of degraded, typewriter-era documents, which is currently a time-consuming manual process. This article presents the statistical analysis of the most influential document quality features affecting OCR accuracy, develops consistent predictive models for four currently used OCR engines, and studies the applicability of different OCR products to the DOE document declassification process. This study is expected to lead to an efficient and completely automated document declassification system.

Tags

Users

  • @cirrus

Comments and Reviews