OCR is easy when the input is a clean page with one column of Latin text. It becomes a very different problem when the document is a long Indian-language book with mixed scripts, formulas, tables, and pages that do not follow the usual horizontal layout.
Recently, I used Surya OCR 2 from Datalab to process an 800+ page book written in Hindi and Sanskrit. The result was much more than a pile of recognized characters. With a cache-first processing pipeline, I was able to retain text, formulas, tables, page structure, and the intermediate images needed to inspect and reproduce the output.
Why this book was a hard OCR problem
The number of pages was only part of the challenge. The book combined several things that tend to expose weaknesses in document pipelines:
- Hindi and Sanskrit text, including Devanagari diacritics and unfamiliar vocabulary
- Formulas mixed into running prose
- Tables embedded in otherwise ordinary pages
- Tables whose reading direction or visual arrangement was vertical
- Different page layouts, margins, and scan quality across the book
- Enough pages that re-running an expensive inference step for one small fix was not practical
A basic image -> text loop would lose useful information. I needed the output to answer questions such as: Where did this text come from? Is this block a table or a formula? Which page image produced it? Can I retry one page without processing the entire book again?
What Surya OCR 2 brings to the pipeline
Surya is a 650M-parameter document OCR model. The current model card describes a single VLM-backed workflow for OCR, layout analysis, and table recognition. In addition to the recognized content, its OCR output can include:
- Reading order for blocks on a page
- Layout labels such as
Text,SectionHeader,Table,Equation, andPicture - HTML output for text, formulas, and tables
- Bounding boxes and polygons
- Per-block confidence values
- Error and skipped flags
That structure is important. Plain text is useful for searching, but structured output is what makes it possible to rebuild a document, render a table, inspect a formula, or locate a suspicious result on the source page.
Surya also supports a broad set of languages, and its multilingual performance was a good fit for the Hindi and Sanskrit content in this book. It was not necessary to split the document into separate language-specific jobs before inference.
The cache-first approach
The most important engineering decision was to treat each page as an independently cacheable unit. I did not make the final text file the only output of the job. For every page, the pipeline kept the inputs and intermediate results needed for a later retry:
- Render or load the page image and store it with a stable page identifier.
- Run OCR and layout inference only when that page has no valid cached result.
- Save the structured response, including blocks, geometry, labels, confidence, and errors.
- Save table and formula representations alongside the page result.
- Record processing metadata so failed pages can be identified and retried.
- Assemble the final book only after the page-level results are available.
In pseudocode, the control flow looked like this:
| |
The exact cache format is less important than the boundaries. A page is the unit of work, the image is an immutable input, and the structured inference result is an immutable artifact. If a post-processing rule changes, I can rebuild the final document without calling the model again.
One page, three stages
Page 83 from the scanned Surya Siddhanta book shows the complete path through the pipeline. The first image is the source scan. The second is the Markdown representation produced by the script, including the extracted Devanagari text and formula markup. The third is the DOCX rendering generated from that same result.
1. The scanned source

2. Markdown generated from the OCR output

3. DOCX generated from the OCR output

This comparison makes the value of structured OCR visible: the output is not just text copied from pixels. It is an intermediate document representation that can be rendered into different formats while retaining the page’s formulas, paragraphs, and illustration.
What worked particularly well
Text across Hindi and Sanskrit
The model handled the Devanagari content well enough to make the full book searchable and usable for downstream processing. OCR quality still needs review, especially around damaged scans and unusual glyph combinations, but the output was substantially more useful than treating the pages as unsearchable images.
Formulas as their own blocks
Formulas were not flattened into an undifferentiated paragraph. The block labels and HTML representation made it possible to keep formulas separate from surrounding prose and to apply a different renderer or cleanup rule to them later.
That separation matters for Sanskrit and Hindi books with scholarly or mathematical content: a formula should not be normalized like a sentence, and a text correction should not accidentally rewrite mathematical notation.
Tables and reading order
Tables were detected as layout objects rather than being read as a random sequence of words. Surya’s table-recognition output includes row and column geometry, cell relationships, and, when full prediction is enabled, HTML for the table.
This was especially useful for the book’s vertically arranged table pages. A conventional OCR pass often assumes that the page should be read from left to right and top to bottom. Here, the geometry and reading-order data gave the post-processing stage enough information to preserve the intended arrangement instead of blindly concatenating every line.
Reproducibility through cached images
Keeping the page images turned debugging from guesswork into a visual comparison:
- Open the source image for a page.
- Inspect the detected polygons and layout labels.
- Compare the extracted text, formula, or table with the scan.
- Retry only that page if the source image or inference settings need attention.
For an 800+ page run, this is not just a performance optimization. It is an audit trail.
A minimal starting point
The official package can be installed with:
| |
The command-line OCR entry point accepts an image, PDF, or directory and writes structured JSON:
| |
For a Python integration, the model card shows the inference manager and predictor pattern:
| |
For a production-sized book, I would wrap this in page-level checkpointing rather than call the predictor with the entire book and hope the process finishes uninterrupted. Tune the inference backend and batch size for the available GPU, but keep the cache independent of those runtime settings.
Lessons from processing a book this size
Do not overwrite good results. Write each page result atomically and keep failed pages visible. A missing result and an empty result are not the same thing.
Keep geometry with text. Bounding boxes and polygons are what make it possible to verify a questionable extraction and handle non-standard layouts.
Separate recognition from assembly. First produce reliable page artifacts. Then apply cleanup, table conversion, formula rendering, search indexing, and book-level ordering.
Expect a review pass. OCR is not a substitute for proofreading. Use confidence values and visual overlays to prioritize pages for inspection instead of manually reading all 800 pages again.
Measure the whole pipeline. Model throughput is only one part of the job. Image rendering, disk I/O, cache reads, retries, table conversion, and final export can dominate a long run.
Final take
Surya OCR 2 gave me a practical foundation for OCR-ing a difficult 800+ page Hindi and Sanskrit book. Its value was not only that it recognized text. It returned enough document structure to keep formulas and tables distinct, preserve reading order, and reason about vertical layouts.
The other half of the result came from the pipeline around the model: cached page images, page-level JSON artifacts, explicit retries, and a separate assembly step. That combination made a long OCR job resumable and inspectable instead of turning it into one fragile command.
If you are processing a multilingual book, especially one with mixed text and structured content, Surya is worth evaluating. Start with a representative sample of difficult pages, build caching in from the first run, and judge the complete structured output rather than only the extracted plain text.
Read the official Surya OCR 2 model card on Hugging Face for installation details, output fields, language coverage, and licensing information.