A Lightweight CNN Framework for Efficient Strike-Out Removal in Handwritten Text Images
Abstract
Due to corrections and revisions, strike-out words are very common in handwritten documents. OCRs that process such words tend to cause recognition errors and low overall accuracy. Current approaches to managing strike-out text are often based on heuristic rules or large pretrained deep networks, which are either not robust or are computationally expensive. The present work proposes a lightweight convolutional neural network, termed StrikeCNN to detect strike-outs at the word level in handwritten documents. The method treats strike-out identification as a binary classification problem and operates directly on binarized word images which allows the model to learn relevant structural patterns without manual feature design. The proposed model achieves an accuracy of 98.8% for strike-out detection on Assamese handwritten words. The experimental results also show significant performance improvements in inference time, training time and model size compared to benchmark architectures. StrikeCNN is simple and efficient, which makes it applicable to the real world in OCR preprocessing pipelines.
Keywords
Handwritten document analysis, strike-out detection, word-level classification, binary classification, OCR preprocessing.