Mohammad Ahmed
All work
RAG pipelineFeb 2026Working prototype

Multi-LLM RAG Assistant

One vector store, three model backends: ask the same grounded question to a local, a proprietary and a hosted open-source LLM.

Gradio interface of the multi-LLM RAG assistant answering a question about hot peppers using a local Ollama model
01

Overview

A PDF is chunked into ChromaDB once. A Gradio interface retrieves the top chunks for a question and sends the same prompt to the provider the user selects: Ollama (llama3.2), OpenAI (gpt-4o-mini) or Hugging Face Inference (Llama-3.1-8B-Instruct).

02

Problem

Which model should answer? Comparing local and cloud models is hard when each has its own retrieval setup, so the retrieval layer has to stay fixed while the model changes.

03

Architecture

  1. Ingestion scriptPDF loader → recursive splitter (350 / 100) → Chroma persistent collection.
  2. RetrieverChroma query returning the top 4 chunks.
  3. Prompt templateContext-only answering with a refusal when nothing is retrieved.
  4. Provider routerOne function per backend behind a single dispatch.
  5. Gradio UIProvider dropdown, question box, answer panel.
04

AI pipeline

  1. 01

    Input

    A question and a chosen provider.

  2. 02

    Preprocessing

    Documents were chunked at ingestion time.

  3. 03

    Retrieval

    Top-4 similarity search in ChromaDB.

  4. 04

    LLM

    The identical prompt goes to Ollama, OpenAI or Hugging Face.

  5. 05

    Output

    A grounded answer, or "I don't know" if nothing is retrieved.

  • Dense retrieval with ChromaDB
  • Prompt-level grounding with an "I don't know" fallback
  • Runtime LLM provider routing
05

Engineering

Retrieval fixed, model swappable
Only the generation step changes between providers, which keeps comparisons honest.
One function per provider
Ollama streams over REST and is parsed line by line; OpenAI and Hugging Face use their SDKs. A small router picks one.
Keys stay in the environment
API keys are read from environment variables and nothing secret is committed.
06

Challenges

  • No challenges are documented in the repository.
07

Results & limits

Measured

  • No evaluation has been run and no answer-quality numbers exist.

Known limits

  • Single source document, which is not included in the repo.
  • Embeddings use Chroma's default function rather than a chosen model.
  • No chat history, tests or evaluation set.
08

Stack

  • Python
  • ChromaDB
  • LangChain loaders
  • Gradio
  • Ollama
  • OpenAI API
  • Hugging Face Inference
10

Demo

No public demo for this one yet. The source repository has setup instructions.