Mohammad Ahmed
All work
Document AI serviceMar 2026Coursework-scale service · demo hosted without the LLM

Offline Document Intelligence Studio

One FastAPI service that OCRs documents and runs local-LLM summarisation, extraction and question answering, with no cloud AI API.

01

Overview

A FastAPI app with a small HTML dashboard and six routers: OCR, chat, summarise, extract, predict and retrieval Q&A. LLM work is delegated to a local Ollama server, and the whole thing ships as a Docker image and a Windows executable.

02

Problem

Sensitive documents shouldn't have to leave the machine. The goal was OCR, summaries, field extraction and Q&A over uploaded files using only local models.

03

Architecture

  1. DashboardJinja2 + plain HTML/JS front end served by FastAPI.
  2. Routers/ocr /chat /summarize /extract /predict /rag, each a thin layer over a service.
  3. ServicesOCR, LLM, summary, extraction, prediction and retrieval logic.
  4. OllamaExternal local server called over HTTP, configurable through environment variables.
  5. PackagingDockerfile, compose file and a PyInstaller-aware path setup for the Windows executable.
04

AI pipeline

  1. 01

    Input

    Image or PDF upload.

  2. 02

    Preprocessing

    PDF pages at 300 dpi, then grayscale, resize, denoise, deskew and threshold with OpenCV.

  3. 03

    OCR

    Tesseract text extraction.

  4. 04

    Retrieval

    500-character chunks indexed with TF-IDF; top matches by cosine similarity.

  5. 05

    LLM

    Ollama answers only from the retrieved context, or summarises and extracts fields as JSON.

  6. 06

    Output

    Text, summary, extracted fields or a grounded answer in the dashboard.

  • OCR with OpenCV preprocessing (denoise, deskew, threshold)
  • Local LLM (Ollama, llama3.2) for chat, summaries and JSON extraction
  • TF-IDF retrieval for document Q&A
  • RandomForest classifier (Iris, as an assignment component)
05

Engineering

Router and service separation
Endpoints stay thin; OCR, LLM, retrieval and prediction each live in their own service module.
Environment-driven config
The Ollama URL, Poppler path and Tesseract command are environment variables, and a cloud-demo mode flag switches off what can't run in the hosted version.
Honest hosting write-up
The repo documents why Ollama can't run on a free Hugging Face Space, and what the hosted demo therefore omits.
Frozen-app support
Path handling accounts for PyInstaller's runtime layout so the same code runs as a script or a packaged exe.
06

Challenges

  • A local LLM doesn't fit the free Hugging Face Space limits, so the hosted demo runs without the model-backed features.
  • Packaging for Windows meant handling frozen-app resource paths.
07

Results & limits

Measured

  • No evaluation numbers are recorded. Testing was manual.

Known limits

  • Retrieval is lexical TF-IDF, not neural embeddings.
  • The hosted demo is English OCR only; the Docker image installs English language data.
  • Built to a course brief. The Iris classifier is a stand-in with no connection to the documents.
08

Stack

  • FastAPI
  • Python
  • Tesseract
  • OpenCV
  • scikit-learn
  • Ollama
  • Docker
  • Jinja2
10

Demo

Open the live demo

Hugging Face Space. It may take a minute to wake, and the Ollama-backed features are off in the hosted version.