AlphaOCR: Integrated Deep Learning and Optical Character Recognition for Receipt Extraction
Abstract
Receipts are vital documents that validate transactions by capturing key details such as dates, items purchased, prices, and seller information. However, accurately documenting receipts is often compromised by issues like blurring, which can result from poor image quality, physical wear, or suboptimal scanning conditions. While the Processing Key Information Extraction from Documents Using Improved Graph Learning-Convolutional Networks (PICK) deep learning model is effective for structured information extraction, it struggles with processing blurred text, leading to inaccuracies in data retrieval. To overcome these challenges, this research introduces AlphaOCR, a web-based application that integrates the PICK model with advanced Optical Character Recognition (OCR) technology. The methodology integrates OCR-based line-item recognition with PICK-based structured-field extraction to extend the range of information recovered from receipts. The deep learning component was evaluated using mean Entity Precision (mEP), mean Entity Recall (mER), mean Entity F1-score (mEF), and mean Entity Accuracy (mEA), together with a paired t-test for the entity-level mEF results. PICK-2 increased the mEF from 83.32% to 83.88%, corresponding to a modest absolute gain of 0.56 percentage points. The complete AlphaOCR configuration combining PICK- 2 with EasyOCR achieved an overall accuracy of 87.96% under the evaluation setting used in this study. Among the OCR models evaluated, EasyOCR achieved a Character Accuracy Rate (CAR) of 92.04% and provided the most suitable output characteristics for integration with PICK-2 within the tested configuration. The reported results should be interpreted within the experimental scope of this study rather than as a direct performance ranking against previously published systems that used different datasets, tasks, extracted fields, and evaluation metrics. AlphaOCR provides a practical integrated workflow for receipt information extraction, although broader claims regarding robustness, scalability, and generalisability require evaluation on larger and more diverse receipt collections. This research contributes to receipt information extraction primarily through the system-level integration of graph-based key information extraction, OCR-based item recognition, output fusion, and a web-based verification workflow. Future work should evaluate the system on larger and more diverse datasets and investigate enhanced semantic understanding and additional document types.