Indic Document Intelligence: Recognition, Correction, Translation and Grounded Question Answering for Indian-Language Documents on One Machine
B.E. Artificial Intelligence and Machine Learning, Adichunchanagiri Institute of Technology, Chikkamagaluru · TealSquare Technologies Private Limited
Abstract
A large amount of everyday paperwork in India exists only on paper or as a phone photograph, and much of it is written in a script that mainstream document tools handle poorly. This report presents an open pipeline that takes such a page from image to answer on a single machine: Tesseract recognises the text in English and nine Indian languages, a script-based detector identifies which languages are present, a language model reached through the Consciousnx API is used to repair recognition errors and to translate the text, and a retrieval step indexes the result so that questions can be answered from it with the supporting passages cited and questions that the documents do not cover are refused. Every component apart from the language model is open source and operates on the CPU of the machine that holds the documents.
We outline each stage and the choices behind it, then evaluate the pipeline on synthetic pages rendered from a parallel corpus in all ten languages under six image conditions, on real photographs, and on retrieval and answering tasks built from the same corpus. The experiments examine how recognition accuracy varies by script and image quality, whether image preparation and model-based correction help, how the language string supplied to the recogniser should be chosen, whether an English-only embedding model can search Indic text, and how well the refusal threshold separates answerable questions from the rest. During the evaluation, the new test suite identified a sign error in the deskew step that was present in the first release. The code, test suite, benchmark pages, and every raw result are released under the Apache 2.0 licence.
@techreport{bhovivaddar2026indic,
title = {Indic Document Intelligence: Recognition, Correction, Translation and Grounded Question Answering for {Indian}-Language Documents on One Machine},
author = {Bhovivaddar, Ganesh F. and {Manasa SR} and {Rishi A M}},
institution = {TealSquare Technologies Private Limited},
type = {Technical Report},
number = {TSQ-TR-2026-02},
year = {2026},
month = sep,
url = {https://tealsquare.in/papers/Indic-Document-Intelligence.pdf}
}
Bhovivaddar, G. F., Manasa SR, & Rishi A M. (2026). Indic document intelligence: Recognition, correction, translation and grounded question answering for Indian-language documents on one machine (Technical Report TSQ-TR-2026-02). TealSquare Technologies Private Limited. https://tealsquare.in/papers/Indic-Document-Intelligence.pdf
G. F. Bhovivaddar, Manasa SR, and Rishi A M, "Indic Document Intelligence: Recognition, Correction, Translation and Grounded Question Answering for Indian-Language Documents on One Machine," TealSquare Technologies Private Limited, Tech. Rep. TSQ-TR-2026-02, Sep. 2026. [Online]. Available: https://tealsquare.in/papers/Indic-Document-Intelligence.pdf