AI Voice Activity Detector (VAD)

Find where people speak and where it's silent, with timestamps, totals and an exportable segment list.

AI Private · runs on your device Free
Runs locally in your browser AI processing happens on your device. Your files are not uploaded for inference.
AI model
…
Download size
… ·
Device
Checking browser support…

    Where is the speech, where is the silence?

    Voice Activity Detection finds the parts of a recording that contain human speech. This tool uses Silero VAD (about 2 MB) to estimate the probability of speech every 32 milliseconds.

    Adjustable results

    • Speech threshold: set the sensitivity
    • Merge short pauses: keep breaths within a sentence in one segment
    • Ignore tiny blips: filter out clicks and coughs

    The timeline and segment list update instantly as you change settings, and total speech and silence times are shown.

    Use cases

    Podcast and interview editing, finding silent parts in long recordings, preparing for transcription and building speech datasets. Export segments as CSV or JSON. Your recordings are processed on your device.

    How to use AI Voice Activity Detector (VAD)

    1. Upload your audio or video file.
    2. The model calculates speech probability.
    3. Adjust the threshold and merge settings.
    4. Export the speech segments as CSV or JSON.

    Why use this tool?

    Removes the need to mark speech segments by hand in hours-long recordings.

    FAQ

    Does it work in every language?

    Yes. Voice activity detection looks at the acoustic features of speech rather than words, so it's language-independent.

    Does it tell speakers apart?

    No, it only detects whether there is speech, not who is speaking.

    Is music detected as speech?

    Song vocals can sometimes count as speech; raising the threshold reduces this.

    Related tools