Optical Character Recognition (OCR) is a fundamental component of document analysis pipelines, enabling downstream tasks such as information extraction, retrieval, and serving as a key enabler of automation. The recent emergence of multimodal Large Language Models (LLMs) has introduced alternative approaches that integrate implicit OCR capabilities within unified vision–language architectures. Despite their growing prominence, a systematic comparison between traditional OCR systems and multimodal LLMs is still lacking. This paper presents a comprehensive benchmarking study encompassing 16 OCR and multimodal systems, grouped into open-source OCR engines, commercial OCR services, commercial multimodal LLMs, and open-source multimodal LLMs. The evaluation spans four publicly available datasets representing printed, scanned, handwritten, English and Chinese dense-text multicolumn documents. Performance is assessed using character- and word-level accuracy, normalized edit distance, and flexible character accuracy, complemented by latency and cost analyses to provide a holistic view of each system’s operational suitability. The results show that no single paradigm excels universally: in our settings, multimodal LLMs achieve stronger performance on unstructured and handwritten inputs, while specialized OCR pipelines remain more reliable for structured and printed layouts. Lightweight domain adaptation through fine-tuning proves beneficial for open-source OCR models, whereas general-purpose LLMs offer competitive accuracy at higher computational costs. All assessment tools and configurations are publicly released to promote reproducibility and facilitate future OCR and LLM research within the document analysis community.

Caravani, V., De Cesaris, R., Merialdo, P. (2027). Assessing and comparing document OCR systems in the era of Large Language Models. INFORMATION PROCESSING & MANAGEMENT, 64(1) [10.1016/j.ipm.2026.105047].

Assessing and comparing document OCR systems in the era of Large Language Models

Valerio Caravani
;
Riccardo De Cesaris;Paolo Merialdo
2027-01-01

Abstract

Optical Character Recognition (OCR) is a fundamental component of document analysis pipelines, enabling downstream tasks such as information extraction, retrieval, and serving as a key enabler of automation. The recent emergence of multimodal Large Language Models (LLMs) has introduced alternative approaches that integrate implicit OCR capabilities within unified vision–language architectures. Despite their growing prominence, a systematic comparison between traditional OCR systems and multimodal LLMs is still lacking. This paper presents a comprehensive benchmarking study encompassing 16 OCR and multimodal systems, grouped into open-source OCR engines, commercial OCR services, commercial multimodal LLMs, and open-source multimodal LLMs. The evaluation spans four publicly available datasets representing printed, scanned, handwritten, English and Chinese dense-text multicolumn documents. Performance is assessed using character- and word-level accuracy, normalized edit distance, and flexible character accuracy, complemented by latency and cost analyses to provide a holistic view of each system’s operational suitability. The results show that no single paradigm excels universally: in our settings, multimodal LLMs achieve stronger performance on unstructured and handwritten inputs, while specialized OCR pipelines remain more reliable for structured and printed layouts. Lightweight domain adaptation through fine-tuning proves beneficial for open-source OCR models, whereas general-purpose LLMs offer competitive accuracy at higher computational costs. All assessment tools and configurations are publicly released to promote reproducibility and facilitate future OCR and LLM research within the document analysis community.
2027
Caravani, V., De Cesaris, R., Merialdo, P. (2027). Assessing and comparing document OCR systems in the era of Large Language Models. INFORMATION PROCESSING & MANAGEMENT, 64(1) [10.1016/j.ipm.2026.105047].
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11590/553740
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact