Deepgram Adds SageMaker AI Observability
- •Deepgram adds billing and usage metrics for SageMaker AI speech deployments in customer CloudWatch accounts
- •ConsumedUnits matches AWS Marketplace metered billing; usage stream tracks model tiers, methods, and features
- •Prometheus and OpenTelemetry support exposes per-GPU, host, and Deepgram engine metrics every 60 seconds
Deepgram said on August 27, 2026, that its self-hosted speech-to-text and text-to-speech models on Amazon SageMaker AI now provide Enhanced Metrics for billing and usage visibility, plus Prometheus and OpenTelemetry support for engine and GPU observability. The update applies to Deepgram SageMaker AI deployments, where audio and transcripts remain inside a customer’s own AWS account while SageMaker AI handles deployment, scaling, and monitoring. Deepgram said the capabilities are available today and are designed to show what customers are billed for, which features their traffic uses, and what the inference engine is doing on each GPU.
Deepgram’s speech models are sold as AWS Marketplace model packages and deployed as SageMaker AI real-time endpoints. SageMaker AI already provides endpoint metrics such as ConcurrentRequestsPerModel and FirstChunkLatency in CloudWatch, container logs in Amazon CloudWatch Logs, and alarm-driven automatic scaling. AWS Marketplace model packages run with network isolation, meaning the container cannot make outbound connections; Deepgram said the new metrics work within that limit because they do not require the container to open a network path and the data lands in the customer’s CloudWatch account.
Deepgram Enhanced Metrics use CloudWatch Embedded Metric Format, or EMF (structured logs converted into metrics), written by the Deepgram container to stdout. SageMaker AI forwards that output to the endpoint’s CloudWatch log group, and CloudWatch Logs extracts the records into metrics automatically. Deepgram said this path needs no agent, sidecar, collector, or additional IAM permissions beyond endpoint logging, and all dimensions are low-cardinality with no personally identifiable information, transcripts, TTS input, or per-request identifiers.
The Deepgram/SageMakerInference namespace emits one record per completed streaming session, pre-recorded request, and TTS request. Its ConsumedUnits metric carries the same billable-unit values used for AWS Marketplace metered billing, while AudioDurationSeconds tracks speech-to-text audio duration and CharCount tracks text-to-speech characters synthesized. Dimensions are published at [Category], [Category, Model], and [Category, Model, Transport] granularities, allowing customers to check streaming STT cost and model-specific use, including nova-3. SampleCount represents the number of billed requests.
A second stream, Deepgram/SelfHosted, reports raw usage by method, model tier, and enabled feature, independent of billing. Its metrics include AudioMs and Requests by Deployment and Method, TierAudioMs by Deployment and Tier, FeatureAudioMs and FeatureTokens by Deployment and Feature, and TtsCharacters, Tokens, and VoiceAgentMs by Deployment and Method. Deepgram cited example features such as diarize, smart_format, redact, and keyterm prompting. The usage stream is on by default and can be disabled with DEEPGRAM_API_01: emf.enabled=false, while the billing stream cannot be disabled because it is part of metering.
For deeper infrastructure visibility, Deepgram containers expose a Prometheus metrics endpoint, and SageMaker AI detailed observability runs an AWS managed OpenTelemetry Collector on each instance behind the endpoint. The collector exports GPU metrics such as DCGM_FI_DEV_GPU_UTIL and DCGM_FI_DEV_FB_USED, host metrics such as node_cpu_seconds_total and node_memory_MemTotal_bytes, and Deepgram engine metrics such as engine_active_requests{kind="stream"} and engine_estimated_stream_capacity. Deepgram said engine_estimated_stream_capacity estimates how many concurrent streams an instance can sustain, and comparing it with engine_active_requests gives a headroom signal for scaling decisions.
Every Prometheus and OpenTelemetry series carries SageMaker labels including aws.sagemaker.endpoint.name, variant name, and instance ID, so users can filter by endpoint, isolate one instance in a scaled-out fleet, or compare GPUs within an instance. Detailed observability is on by default for newly created endpoints and publishes every 60 seconds; older endpoints can enable it or change the publish frequency with MetricsConfig in a new endpoint configuration and update-endpoint, which Deepgram described as a blue/green deployment that keeps the endpoint in service. Metrics can be queried with PromQL from CloudWatch, Grafana, or any Prometheus-compatible tool.
Deepgram said Enhanced Metrics require no setup once a Deepgram Sag