AWS Uses Bedrock for Dashboard Validation
- •AWS Bedrock system cut dashboard failure detection from up to 72 hours to less than 1 hour
- •Visual validation detected 802 failures across 153,000 checks, with fewer than 1 percent reported by users
- •Numeric checks use LLM extraction and deterministic comparison for 50–70 data points per weekly cycle
An AWS team used Amazon Bedrock to build an automated system that scans dashboard content failures in AWS Insights, after finding that blank, stale, or incorrect dashboard elements can evade infrastructure monitoring even when servers, APIs, and data pipelines appear healthy. The team said its instrumentation showed such failures happen in fewer than 1 percent of cases, but they are often silent because users must notice and report them. The system reduced mean time to detection from up to 72 hours to less than 1 hour by running hourly validation cycles and alerting builders in real time.
The monitoring gap sits at the business intelligence presentation layer, where dashboard sections can appear blank, show stale data, display error states, or present wrong numbers despite upstream systems reporting normal status. In 30 days after automated monitoring was deployed, the visual validation mechanism detected 802 content failure instances, including row-level data permission errors, filters skipping records, and rendering issues; fewer than 1 percent had a corresponding user report. AWS said this risk increases when AI narrative systems consume dashboard data, because numeric errors can flow into insights used by business leaders.
The solution uses a five-stage serverless architecture built on AWS managed services. Amazon EventBridge triggers hourly visual checks, while weekly data refreshes trigger numeric validation cycles. A configuration registry in Amazon Redshift tracks section identifiers, owner assignments, and scheduling preferences. AWS Lambda orchestrates headless browser sessions for visual screenshots, while agentic browser automation handles numeric ground-truth capture when dashboards require navigation and filters.
Screenshots pass through a redaction step before storage. Amazon Rekognition detects text and numeric values through OCR (text recognition from images), then replaces text with redacted placeholders and numbers with synthetic values so sensitive data is not retained. The screenshots are stored in Amazon Simple Storage Service and served through Amazon CloudFront for low-latency access during analysis.
Stage 3 runs two parallel validation mechanisms on Amazon Bedrock. For visual content validation, Anthropic Claude models analyze screenshots with dashboard metadata to detect blank tiles, error states, missing visuals, and whether an empty section is legitimate or a content failure. AWS constrained model outputs to structured verdicts with confidence scores, routed ambiguous cases to human review, and kept screenshots redacted before analysis.
For numeric cross-source validation, the system checks whether the same metric is consistent across dashboards. AWS described the pattern as hybrid validation: LLMs handle semantic recognition of metrics with different labels or layouts, while deterministic code handles unit normalization, decimal precision, and final matched-or-mismatched verdicts. The article gave examples such as comparing $1.2B with $1,200M and 58.484 with 58.5.
In production, AWS said false positives were a major adoption risk because dashboard owners stop trusting alerts if notifications are wrong. The team also replaced an earlier two-layer LLM comparison design with deterministic code after production runs showed occasional comparison-logic errors around rounding, tolerance, and unit differences. After that change, subsequent model upgrades improved extraction and raised recall from 0.88 to 0.95 by reducing false alarms from misread dashboard values.
Over 30 days, the visual validation mechanism monitored hundreds of dashboards, performed 153,000 automated content checks, detected 802 content failures, and measured 0.52 percent failed checks, equivalent to 99.48 percent system content availability. The numeric validation mechanism has run weekly for more than 6 months and validates 50–70 data points per cycle. AWS said matched values are automatically approved through deterministic comparison, no data issue has bypassed human review as a false approval in production evaluation to date, and one production cycle found a systematic inconsistency affecting related metrics before the data reached downstream consumers.