Back to Projects

Real-time Transcription System

AI-powered speech-to-text pipeline using Silero VAD and Faster-Whisper for fully local, low-latency transcription

AI & Automation 2025

Overview

A Python pipeline that captures microphone audio in real time, uses Silero VAD to segment speech from silence, and passes each speech chunk to Faster-Whisper for transcription. The tool runs entirely on local hardware with no data sent to external APIs.

The architecture uses three concurrent threads:

  • Producer: captures audio at 16 kHz mono with sounddevice and pushes frames into a shared queue
  • VAD consumer: runs Silero VAD on each frame and accumulates speech until silence is detected for a configurable duration
  • Transcription worker: submits completed segments to Faster-Whisper and writes timestamped .wav and transcript files

Technical Implementation

Faster-Whisper is a CTranslate2-optimised port of OpenAI Whisper. GPU acceleration via CUDA is used when available, with a CPU fallback for standard hardware. The threading model avoids dropped frames even on modest hardware, as Silero VAD inference is lightweight enough to run well ahead of transcription without stalling.

Silence-threshold and minimum-segment-length parameters are configurable at startup, letting the system adapt to a noisy lecture hall or a quiet interview room without code changes.

Key Features

Voice Activity Detection

Silero VAD (a lightweight ONNX model) runs on every audio frame and classifies it as speech or non-speech, preventing Whisper from processing silent segments and dramatically reducing CPU/GPU load.

Automatic Transcription

Each speech segment is sent to Faster-Whisper for transcription. GPU acceleration via CUDA is used when available, with a CPU fallback for standard hardware.

Session Archive

Completed segments are saved as timestamped .wav files alongside their transcripts, building a searchable session log for later review or export.

Adjustable Sensitivity

Silence-threshold and minimum-segment-length parameters are configurable at startup, letting the system adapt to a noisy lecture hall or a quiet interview room.

Multi-threaded Architecture

Capture, VAD inference, and Whisper transcription run on separate threads with a queue between them, keeping the main thread non-blocking and the live display responsive.

Open for internships • CS @ Boston University • Java • Python • C++ • Arduino • Streamlit • Email: davidebonn.tn@gmail.com •