Deepgram
Real-time AI speech-to-text, text-to-speech, and voice agent API
About Deepgram
Deepgram is an enterprise AI speech platform that provides world-class speech-to-text (STT), text-to-speech (TTS), and real-time voice agent APIs. Powered by end-to-end deep learning models like Nova-2 and Aura, Deepgram delivers industry-leading accuracy, sub-250ms latency, and high cost-efficiency for processing live conversational audio. Deepgram is the infrastructure backbone for next-generation conversational AI applications, customer service bots, meeting transcription engines, and autonomous voice agents. Its streaming WebSocket architecture allows models to transcribe noisy, multi-speaker phone audio in real time while simultaneously synthesizing natural, expressive voices with zero perceptible delay.
Deepgram replaced legacy heuristic speech pipelines with pure end-to-end deep learning neural networks trained on over 100,000 hours of diverse multi-lingual audio. Its Nova-2 STT model achieves lower Word Error Rates (WER) than traditional cloud providers while operating up to 40x faster and at a fraction of the cost. The platform offers specialized features including automatic language detection, smart formatting (punctuating numbers, dates, and acronyms), multichannel diarization, topic detection, and PII redaction. Deepgram also provides the Deepgram Voice Agent API, which bundles speech recognition, LLM reasoning, and ultra-low-latency voice synthesis into a single unified WebSocket connection.
Create an account and obtain your Deepgram API key with $200 in free credits.
Stream live audio via WebSocket or submit pre-recorded audio files via REST API.
Configure model parameters such as Nova-2, language, diarization, and smart formatting.
Receive real-time, timestamped transcriptions with word-level confidence scores.
Use Deepgram Aura to convert LLM text responses back into human-like audio in under 200ms.
Capabilities & Features
Common Use Cases
real-time-voice-agents
live-audio-transcription
conversational-ai-telephony
podcast-and-video-captioning
ultra-fast-text-to-speech
speech-sentiment-analysis
Frequently Asked Questions
How fast is Deepgram speech-to-text latency?
Deepgram streaming STT processes audio in real time with sub-250ms latency, making it ideal for live phone calls and voice agents.
What languages does Deepgram support?
Deepgram supports over 30 languages including English, Spanish, French, German, Japanese, Portuguese, and Mandarin with automatic language detection.
Can I self-host Deepgram on-premises?
Yes, Deepgram offers an on-premises and private cloud deployment option for enterprise customers requiring strict data sovereignty and compliance.
Free Plan
$200 free credit upon signup with full API access to Nova-2 and Aura voice models.
Paid Plan
Pay-as-you-go starting at $0.0043/min for speech-to-text and $0.015/1,000 chars for natural voice generation.
Direct link · Verified & reader-supported
Pros & Cons
Industry-leading Nova-2 model with highest transcription accuracy and lowest WER
Ultra-low sub-250ms streaming latency essential for conversational AI voice agents
Up to 40x faster and 3–5x cheaper than legacy cloud speech providers
Native multi-speaker diarization, smart formatting, and PII redaction
Unified Voice Agent API combining STT, LLM orchestration, and Aura TTS
API-first platform requiring developer integration
Advanced custom vocabulary tuning requires training on domain-specific datasets
Looking for the best alternatives to Deepgram?
Side-by-side feature matrix, pricing models, and decision frameworks.
Alternatives to Deepgram
Deep comparison hubElevenLabs
Ultra-realistic AI voices and voice cloning
A leading AI voice platform for generating ultra-realistic speech, voice cloning, and multilingual audio content.
Descript
Edit audio and video by editing text
An AI-powered video editor that works like a word processor. Transcribe your media and delete text to cut scenes or correct audio with AI cloning.
AssemblyAI
The API for speech AI
A developer-first platform for transcription, speaker diarization, and sentiment analysis.
Moshi by Kyutai
Real-time full-duplex conversational voice AI model with sub-200ms latency
Moshi is an open-source real-time conversational voice AI foundation model developed by Kyutai, the non-profit AI research lab based in Paris. Engineered to revolutionize human-AI verbal communication, Moshi operates on a full-duplex architecture capable of listening, thinking, and speaking simultaneously with sub-200ms end-to-end latency. Unlike traditional voice assistants that chain separate Speech-to-Text (STT), Large Language Model (LLM), and Text-to-Speech (TTS) pipelines together, Moshi processes raw multi-stream audio natively as continuous speech tokens. This allows Moshi to understand emotional nuances, interrupt and be interrupted naturally, chuckle, whisper, and express genuine conversational timing. Moshi is fully open-source with openly accessible weights, training recipes, and inference code, serving as a foundational milestone for research in real-time spoken language modeling and multi-modal conversational systems.
LiveKit Agents
Open-source real-time WebRTC infrastructure for building ultra-low-latency voice and multimodal AI agents
LiveKit Agents is an open-source real-time communication framework engineered to build conversational voice, video, and multimodal AI agents with sub-500ms latency. Leveraging WebRTC, it connects speech-to-text (Deepgram, Whisper), LLMs (OpenAI, Anthropic), and text-to-speech (Cartesia, ElevenLabs) in a tightly synchronized bidirectional stream. From customer service avatars to interactive language tutors and hands-free coding copilots, LiveKit Agents provides the enterprise infrastructure for real-time human-AI interaction.
Udio
The most musical AI track generator
An AI music creation platform known for its exceptional audio fidelity and sophisticated musical arrangements across every imaginable genre.
Compare Deepgram with Alternatives
Side-by-side feature, pricing, and pros & cons breakdowns
