Remove the noise.
Keep the voice.
Real-time noise removal and primary speaker extraction: every other voice is cut out, so only your caller is left.
Loud places.
Clear calls.
People call from cafés, shared offices, and busy streets. Clarity helps your voice agent focus on the person calling.
Clarity does two jobs, in real time. It removes background noise. And it extracts the primary speaker: every other voice, from the café crowd to the person at the next desk, is cut out, so only your caller comes through.
Your voice bot then understands your caller instead of the room: speech-to-text, turn detection and the model behind it all work from one clean voice.




Heard the difference?
Try Clarity on your own audio.
Built for real calls.
Your voice bot hears only the caller.
Clarity removes background noise and chatter before your voice bot hears them, so its turn detection and speech-to-text react only to the caller.
+50 msadded latency in pipelines with 240 ms chunks
Put to the test.
Against the best.
Target-speaker quality
Libri2Mix · DNSMOS, higher is better
| Metric | Clarity 1240 ms | StarTSE[1]560 ms |
|---|---|---|
| Speech quality SIG | 3.56 (best) | 3.54 |
| Background quality BAK | 4.03 (best) | 3.75 |
| Overall quality OVRL | 3.27 (best) | 3.12 |
Noise removal
DNS 2020 Challenge · DNSMOS, higher is better
| Metric | Clarity 1240 ms | UnprocessedOriginal audio | AI Acoustics15 ms | DeepFilterNet310 ms | GTCRN16 ms |
|---|---|---|---|---|---|
| Speech quality SIG | 3.62 (best) | 3.39 | 3.46 | 3.51 | 3.34 |
| Background quality BAK | 4.11 (best) | 2.62 | 3.60 | 4.11 (best) | 3.98 |
| Overall quality OVRL | 3.37 (best) | 2.48 | 2.98 | 3.25 | 3.05 |
Good to know.
What is Clarity?
Clarity is a speech enhancement model that does two jobs in real time. It removes background noise, and it extracts the primary speaker: every other voice, such as café chatter or colleagues nearby, is cut out so only the one voice that matters is left. Whoever is listening, a person or a voice agent, hears clear speech.
How much latency does Clarity add?
Clarity processes audio in 240 ms chunks and needs about 50 ms for each one. If your turn detection already reads audio in 240 ms chunks, Clarity adds only those 50 ms. With smaller input frames, a sound can come out up to 290 ms after it went in, depending on where it falls in its chunk.
How does it know which voice to keep?
Primary speaker extraction uses a short reference recording of the person you want to follow, and cuts out every other voice. Without a reference, it keeps the loudest speaker.
Does it help my voice bot understand callers?
Yes. Speech-to-text, turn detection and the language model behind your bot can only work with the audio they receive, so Clarity gives them one clean voice: your caller’s. Background talkers stop ending up in transcripts and stop setting off turn detection, and the bot responds to your caller instead of the room.
Does it wait for the caller to finish?
No. Clarity processes audio in a continuous stream, as it arrives. It is built for live conversations, rather than waiting for a completed recording.
Can I hear it with my own audio?
Yes. Sign up and run your own recordings through Clarity in the dashboard, free until October 30th. For a self-hosted deployment or a larger rollout, contact us and bring examples from your callers’ environments.

