Research

Selected technical reports from TealSquare’s ongoing work in applied AI.

Technical Reports

First page of the Indic Document Intelligence report

TSQ-TR-2026-0222 September 202620 pages

Indic Document Intelligence: Recognition, Correction, Translation and Grounded Question Answering for Indian-Language Documents on One Machine

Ganesh F Bhovivaddar, Manasa SR, Rishi A M

B.E. Artificial Intelligence and Machine Learning, Adichunchanagiri Institute of Technology, Chikkamagaluru · TealSquare Technologies Private Limited

Abstract

A large amount of everyday paperwork in India exists only on paper or as a phone photograph, and much of it is written in a script that mainstream document tools handle poorly. This report presents an open pipeline that takes such a page from image to answer on a single machine: Tesseract recognises the text in English and nine Indian languages, a script-based detector identifies which languages are present, a language model reached through the Consciousnx API is used to repair recognition errors and to translate the text, and a retrieval step indexes the result so that questions can be answered from it with the supporting passages cited and questions that the documents do not cover are refused. Every component apart from the language model is open source and operates on the CPU of the machine that holds the documents.

We outline each stage and the choices behind it, then evaluate the pipeline on synthetic pages rendered from a parallel corpus in all ten languages under six image conditions, on real photographs, and on retrieval and answering tasks built from the same corpus. The experiments examine how recognition accuracy varies by script and image quality, whether image preparation and model-based correction help, how the language string supplied to the recogniser should be chosen, whether an English-only embedding model can search Indic text, and how well the refusal threshold separates answerable questions from the rest. During the evaluation, the new test suite identified a sign error in the deskew step that was present in the first release. The code, test suite, benchmark pages, and every raw result are released under the Apache 2.0 licence.

First page of the Rule-Based Context Shortening report

TSQ-TR-2026-0121 September 202611 pages

Less Is More: Rule-Based Context Shortening Improves Small-Model RAG Accuracy

Teja Velagala, Ganesh F Bhovivaddar

TealSquare Technologies Private Limited

Abstract

Retrieval-augmented generation (RAG) systems pass retrieved document chunks to a language model word for word. A substantial share of that text is grammatical scaffolding: articles, auxiliary verbs, prepositions, pronouns, and stock phrases such as “it is important to note that”. These words assist a human reader but carry little of the information needed to answer a question, and each one costs the model attention and time. Earlier work on prompt compression removes such material using a second neural model that scores or rewrites the context.

This report examines a simpler question: what happens if the retrieved context is shortened with a fixed word list and a handful of regular expressions, without a model at all? We describe a three-pass rule-based shortener and compare it with verbatim context on three question-answering benchmarks that differ in domain and reasoning type, using a small open-weight generator running on a laptop. We measure answer accuracy, the amount of context removed, and end-to-end latency, examine where the method helps and where it performs poorly, and discuss what this implies for small-model RAG deployments. The complete pipeline and results are released under the Apache 2.0 licence.