- Python
- WebSockets
- TTS
AI Interview Platform
Real-time voice interviews run by an LLM, end to end.
An automated interview platform that conducts spoken interviews from a job description, a candidate profile and the live conversation. I built the real-time audio pipeline, the speech-to-text service and the LLM loop that decides the next question.
Real-time
streamed TTS audio to the browser
2 services
Node.js audio pipeline + Python STT
RBAC
multi-company recruiter workflows

01
Why I built it
Hiring teams lose days to first-round screening calls that ask the same questions. The goal: an interviewer that adapts to each candidate in real time, so a recruiter only reviews the people worth a human conversation. A conversation only feels natural if the gap between a candidate's answer and the next question is short, so latency drove every design decision.
02
The problem
- A voice conversation breaks if the system waits for a full LLM response before speaking.
- Microphone audio arrives in small chunks and has to be transcribed incrementally, not at the end.
- Follow-up questions must use the job requirements, the resume and everything said so far.
- Recruiters need company onboarding, scheduling, public interview links and dashboards around it.
03
System design
- 1
Browser
Captures microphone audio and plays streamed question audio.
- 2
Python WebSocket service
Receives audio incrementally and transcribes it with OpenAI speech-to-text.
- 3
Conversation loop
Combines job requirements, candidate profile and transcript to generate the next question.
- 4
Node.js audio pipeline
Turns the LLM question into TTS audio and streams it back in chunks.
- 5
Recruiter app
RBAC, company onboarding, interview creation and scheduling, resume–JD matching, HR dashboards.
04
Decisions & trade-offs
Stream audio in chunks instead of waiting for whole responses
Why: Playback starts as soon as the first chunk is ready, so the interviewer replies at conversational speed.
Trade-off: More moving parts: chunk ordering, buffering and cleanup when a session drops mid-stream.
A separate Python service for speech-to-text
Why: Audio processing lives in its own service with its own scaling and failure boundary, and the Node.js app stays focused on orchestration.
Trade-off: Two runtimes to deploy and monitor, plus a WebSocket contract between them.
05
Backend & frontend
Backend
- Node.js real-time audio pipeline for LLM-generated questions and streamed TTS.
- Python WebSocket service for incremental microphone audio and OpenAI speech-to-text.
- LLM conversation loop for context-aware follow-ups.
- RBAC, company onboarding, public interview links and resume–JD matching.
Frontend
- Browser audio capture and chunked playback for a live, spoken interview.
- Recruiter-facing dashboards, interview creation and scheduling flows.
06
Design
- Interview screen designed to feel like a calm call, not a form.
07
Results
- Interviews run end to end without a human interviewer on the first round.
What's next
- Measure end-to-end response latency per turn and set a budget for it.