AI Recruitment Assistant
A hiring platform that turns CVs and job posts into structured data, then ranks applicants with an explainable hybrid score.
Overview
Two role-based experiences on one backend. HR creates and publishes jobs and triggers screening runs; candidates register, upload a PDF résumé and apply. Claude extracts structured requirements and résumé data, embeddings and rules produce a weighted score, and every run is stored with the weights it used.
Problem
First-pass CV screening is repetitive and inconsistent. A reviewer needs structured candidate data, a ranking they can explain, and a record of what each screening run actually did.
Architecture
- Next.js appApp Router + TypeScript. Auth context and role guards for HR and Candidate areas.
- FastAPI routersauth, jobs, profile, applications, screening, results, chat, health.
- Services layerLLM, résumé parsing, matching and screening logic kept out of the routers.
- PostgreSQL 16Async SQLAlchemy models with five Alembic migrations.
- Model providersClaude for extraction and analysis; OpenAI for embeddings only.
AI pipeline
- 01
Input
Job description text and a candidate's PDF résumé.
- 02
Preprocessing
PDF text extraction with pypdf.
- 03
LLM extraction
Claude is forced to call a schema'd tool; output is validated with Pydantic and retried if no tool call comes back.
- 04
Retrieval / scoring
Score = 0.5 semantic similarity + 0.3 skill match + 0.2 experience. Falls back to token overlap if embeddings fail.
- 05
Analysis
Per-candidate analysis, strengths, gaps and interview questions.
- 06
Output
A persisted screening run, with the weights it used, that HR can reopen and question.
- Claude forced tool-use for schema-bound extraction
- OpenAI text-embedding-3-small similarity
- Hybrid scoring: semantic + skills + experience
- Grounded Q&A over stored screening results
Engineering
- Structured output, not parsed prose
- Extraction goes through a forced tool call with a JSON schema, then Pydantic validation, with a retry when the model returns no tool block. Downstream code never parses free text.
- Runs are auditable
- Each screening run snapshots its scoring weights and stores a success or failed status per candidate, so one bad résumé cannot sink a run and old results stay interpretable if the weights change.
- Graceful degradation
- If the embedding call fails, similarity falls back to a bag-of-words cosine instead of failing the whole screening.
- Tighter account model
- Public registration only ever creates a Candidate. HR accounts are provisioned through a CLI script.
- Chat explains, it doesn't re-rank
- The chat endpoint answers from stored results, so a conversation can never silently change a ranking.
Challenges
- Testing async SQLAlchemy with asyncpg: connections bind to an event loop, so the tests create a fresh engine per test.
- Docker Compose environment variables were being shadowed by the host shell, so the stack uses an APP_ prefix.
- The model occasionally returned no tool call, which led to the retry in the extraction path.
Results & limits
Measured
- No quantitative evaluation has been run. Ranking quality, latency and extraction accuracy are not measured yet.
- The backend has roughly 1,300 lines of pytest coverage across auth, jobs, applications, profile, screening and the database layer.
Known limits
- Skill matching is a simple substring check, so short skill names can false-match.
- The scoring weights are hand-chosen, not validated against labelled data.
- Not deployed, and no public demo.
Stack
- Next.js
- TypeScript
- Tailwind CSS
- FastAPI
- SQLAlchemy 2 (async)
- PostgreSQL 16
- Alembic
- Pydantic v2
- JWT
- Docker Compose
- pytest
Demo
No public demo for this one yet. The source repository has setup instructions.