Skip to content

AIExplore

How to Use Deepgram for Latency and cost model picks

Learn Deepgram latency and cost model picks with step by step workflows, realistic examples, and verified plan notes.

Deepgram works well for latency and cost model picks when you run it like production work: locked brief, SOURCE facts, then review before publish. Deepgram is an enterprise voice AI platform unifying speech-to-text, text-to-speech, and LLM orchestration for real-time conversational agents. Confirm live plans and limits on deepgram.com/pricing. Start at /explore/deepgram.

This guide focuses on latency and cost model picks in detail. Related Deepgram articles: /blog/how-to-use-deepgram-for-spend-alert-setups, /blog/how-to-use-deepgram-for-api-key-rotation-practices, /blog/how-to-use-deepgram-for-barge-in-agent-tests.

When this workflow is the right job

Use latency and cost model picks when the deliverable is specifically this Deepgram job. Switch to speech-to-text prototypes when that workflow already owns the asset.

Step by step workflow

1. Brief Latency and cost model picks

Write what must stay true for latency and cost model picks in Deepgram before settings or spend.

Brief: Latency and cost model picks
Keep: verified SOURCE facts only
Avoid: invented pricing or features
Success: one reviewable output

2. Open Deepgram for Latency and cost model picks

Use the Deepgram surface that owns latency and cost model picks. Do not mix a neighboring workflow in the same pass.

Surface: Latency and cost model picks
Start: pilot with one representative input
Plans: deepgram.com/pricing

3. Pilot Latency and cost model picks

Run a single latency and cost model picks pilot. Score clarity, grounding, and whether the output is reviewable.

Pilot: Latency and cost model picks
[ ] SOURCE facts match
[ ] Output reviewable
[ ] Settings logged

4. Refine Latency and cost model picks

Change one latency and cost model picks dimension only. Save a template from the best run.

Refine: Latency and cost model picks
Change: one control only
Keep: SOURCE and success criteria

Practical latency and cost model picks examples

Guardrail filter

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Guardrail filter".

Objective:
Produce a reviewable transcript or speech artifact for Guardrail filter with required features named.

Inputs:
- Audio/file reference for Guardrail filter
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Guardrail filter → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Guardrail filter matching requested features, ready for human edit.

API spend alert

Scenario:
A media team runs Latency and cost model picks in Deepgram for "API spend alert".

Objective:
Produce a reviewable transcript or speech artifact for API spend alert with required features named.

Inputs:
- Audio/file reference for API spend alert
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process API spend alert → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for API spend alert matching requested features, ready for human edit.

Key rotation

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Key rotation".

Objective:
Produce a reviewable transcript or speech artifact for Key rotation with required features named.

Inputs:
- Audio/file reference for Key rotation
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Key rotation → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Key rotation matching requested features, ready for human edit.

Latency pick

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Latency pick".

Objective:
Produce a reviewable transcript or speech artifact for Latency pick with required features named.

Inputs:
- Audio/file reference for Latency pick
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Latency pick → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Latency pick matching requested features, ready for human edit.

Word timestamps

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Word timestamps".

Objective:
Produce a reviewable transcript or speech artifact for Word timestamps with required features named.

Inputs:
- Audio/file reference for Word timestamps
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Word timestamps → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Word timestamps matching requested features, ready for human edit.

Custom vocab

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Custom vocab".

Objective:
Produce a reviewable transcript or speech artifact for Custom vocab with required features named.

Inputs:
- Audio/file reference for Custom vocab
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Custom vocab → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Custom vocab matching requested features, ready for human edit.

Noise note

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Noise note".

Objective:
Produce a reviewable transcript or speech artifact for Noise note with required features named.

Inputs:
- Audio/file reference for Noise note
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Noise note → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Noise note matching requested features, ready for human edit.

Chapter titles

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Chapter titles".

Objective:
Produce a reviewable transcript or speech artifact for Chapter titles with required features named.

Inputs:
- Audio/file reference for Chapter titles
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Chapter titles → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Chapter titles matching requested features, ready for human edit.

Action items audio

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Action items audio".

Objective:
Produce a reviewable transcript or speech artifact for Action items audio with required features named.

Inputs:
- Audio/file reference for Action items audio
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Action items audio → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Action items audio matching requested features, ready for human edit.

Multilingual flag

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Multilingual flag".

Objective:
Produce a reviewable transcript or speech artifact for Multilingual flag with required features named.

Inputs:
- Audio/file reference for Multilingual flag
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Multilingual flag → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Multilingual flag matching requested features, ready for human edit.

Webhook delivery

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Webhook delivery".

Objective:
Produce a reviewable transcript or speech artifact for Webhook delivery with required features named.

Inputs:
- Audio/file reference for Webhook delivery
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Webhook delivery → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Webhook delivery matching requested features, ready for human edit.

Retry policy

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Retry policy".

Objective:
Produce a reviewable transcript or speech artifact for Retry policy with required features named.

Inputs:
- Audio/file reference for Retry policy
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Retry policy → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Retry policy matching requested features, ready for human edit.

Content filter

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Content filter".

Objective:
Produce a reviewable transcript or speech artifact for Content filter with required features named.

Inputs:
- Audio/file reference for Content filter
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Content filter → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Content filter matching requested features, ready for human edit.

Short pilot clip

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Short pilot clip".

Objective:
Produce a reviewable transcript or speech artifact for Short pilot clip with required features named.

Inputs:
- Audio/file reference for Short pilot clip
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Short pilot clip → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Short pilot clip matching requested features, ready for human edit.

Speaker labels

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Speaker labels".

Objective:
Produce a reviewable transcript or speech artifact for Speaker labels with required features named.

Inputs:
- Audio/file reference for Speaker labels
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Speaker labels → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Speaker labels matching requested features, ready for human edit.

Auto chapters

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Auto chapters".

Objective:
Produce a reviewable transcript or speech artifact for Auto chapters with required features named.

Inputs:
- Audio/file reference for Auto chapters
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Auto chapters → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Auto chapters matching requested features, ready for human edit.

Topic summary

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Topic summary".

Objective:
Produce a reviewable transcript or speech artifact for Topic summary with required features named.

Inputs:
- Audio/file reference for Topic summary
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Topic summary → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Topic summary matching requested features, ready for human edit.

Streaming prototype

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Streaming prototype".

Objective:
Produce a reviewable transcript or speech artifact for Streaming prototype with required features named.

Inputs:
- Audio/file reference for Streaming prototype
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Streaming prototype → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Streaming prototype matching requested features, ready for human edit.

PII redaction

Scenario:
A media team runs Latency and cost model picks in Deepgram for "PII redaction".

Objective:
Produce a reviewable transcript or speech artifact for PII redaction with required features named.

Inputs:
- Audio/file reference for PII redaction
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process PII redaction → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for PII redaction matching requested features, ready for human edit.

Subtitle SRT

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Subtitle SRT".

Objective:
Produce a reviewable transcript or speech artifact for Subtitle SRT with required features named.

Inputs:
- Audio/file reference for Subtitle SRT
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Subtitle SRT → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Subtitle SRT matching requested features, ready for human edit.

Interview mp3

Scenario:
A media team runs Latency and cost model picks in Deepgram for "Interview mp3".

Objective:
Produce a reviewable transcript or speech artifact for Interview mp3 with required features named.

Inputs:
- Audio/file reference for Interview mp3
- Feature flags (speakers/chapters/etc.)
- Language locale
- PII handling rules

Workflow:
Upload or stream → Configure Latency and cost model picks → Process Interview mp3 → Review transcript → Export

Requirements:
- Stay within verified Deepgram capabilities; do not invent features.
- Confirm live plan notes on deepgram.com/pricing before promising volume.
- Change one variable between iterations.
- Human-review before external publish, send, billing, or clinical/legal use.
- Confirm API/plan limits on the official pricing page before batch jobs.

Expected output:
A transcript/artifact for Interview mp3 matching requested features, ready for human edit.

How to improve latency and cost model picks

Cut noise from latency and cost model picks by removing extra adjectives while preserving SOURCE facts in Deepgram.

Raise quality by insisting on a single success check before debating style.

Make review easier by labeling fields that must never change.

Speed iteration by cloning the last good run and altering only one control.

Stabilize outputs by pinning settings after the pilot is approved.

Reduce rework by rejecting drafts that invent claims.

Improve handoffs by recording which control produced the best result.

Harden the workflow by testing an incomplete input before trusting defaults.

Prompting and usage guidance

Name the latency and cost model picks job, audience, and success check before opening Deepgram.

Paste only verified facts under SOURCE so Deepgram cannot invent details.

Specify the deliverable shape up front.

Call out fixed details versus flexible style choices.

Ask Deepgram to flag unsupported claims before you accept the draft.

Limitations to respect

Check Deepgram plan gates for latency and cost model picks on deepgram.com/pricing before you promise timelines.

Keep drafts unpublished until a human confirms SOURCE facts.

Deepgram can be wrong. Treat latency and cost model picks as provisional until review.

If documentation is silent on a claim, leave it out rather than guessing.

Practical tips for this workflow

Pilot once before batching latency and cost model picks in Deepgram.

Keep a reusable template with variables for latency and cost model picks.

Separate creative instructions from SOURCE facts.

Log settings from the best run.

Common mistakes

  • Skipping the pilot run before scaling volume
  • Inventing pricing, quotas, or features not on official pages
  • Mixing unrelated workflows in one session
  • Publishing without a human review gate

Treat latency and cost model picks in Deepgram as a production workflow: brief, pilot, refine, then ship with review. Related reading: /blog/how-to-use-deepgram-for-spend-alert-setups, /blog/how-to-use-deepgram-for-api-key-rotation-practices, /blog/how-to-use-deepgram-for-barge-in-agent-tests.

Related articles