Snowflake Previews Cortex AI Gateway
- •Snowflake put Cortex AI Gateway in public preview for AWS commercial regions on October 8, 2026.
- •Internal tests reported up to 3x token efficiency and roughly 25% fewer tokens with steady pull-request throughput.
- •Gateway offers model and MCP access controls, traces, guardrails, and budgets; tools feature remains in private preview.
Snowflake announced Cortex AI Gateway in public preview on October 8, 2026, in AWS commercial regions, excluding New Zealand, Malaysia and Spain. The service places a centralized control point between AI clients and enterprise systems, giving platform teams shared visibility, governance and oversight as organizations use models, agents and Model Context Protocol (MCP) servers. Snowflake says the gateway offers one endpoint for model inference, dynamic routing, MCP tool governance, observability, security policies and cost controls. Its stated purpose is to let platform teams manage model and tool access, attribute and control spending, and see actions taken by AI for users, while application teams build against a single endpoint.
For inference, the gateway supports the Chat Completions API for OpenAI, Grok, GLM, Gemini, Llama, Mistral and DeepSeek models, and the Messages API for Claude. Streaming, tool calling, structured output, prompt caching, image input and reasoning work as they do natively; existing clients can connect by changing the base URL and token. Snowflake says support for the OpenAI Responses API is coming soon. Administrators use Snowflake’s role-based access control (RBAC) to decide which models are available. The Snowflake CLI can configure an agent to route traffic through the gateway with the command `snow ai opencode`.
Dynamic model routing, which Snowflake said was coming to private preview soon, selects the least expensive available model it expects can handle each agent step. Snowflake says its research team fine-tuned the routing so repetitive, lower-complexity tasks go to efficient models and tasks needing deeper reasoning go to frontier models. In an internal dbt pipeline evaluation, the approach used up to 3x greater token efficiency than a frontier-model-only approach at comparable quality. In a separate coding test, engineering teams maintained pull-request throughput while using roughly 25% fewer tokens. Snowflake says only approved models are used, data residency settings are respected, and routing decisions are logged.
Snowflake also reported internal ADE-bench results using Snowflake CoCo as the agent harness: DeepSeek-V4-Flash scored 74.4%, ahead of the leading proprietary model it tested. GLM-5.3 is planned for private preview; its predecessor GLM-5.2 scored 66% on ADE-bench2 with the benchmark’s lowest token footprint. Snowflake says these models run inside its secure perimeter, next to governed data and under the same RBAC and audit trail. The gateway’s tools feature, in private preview, includes a curated catalog of more than 100 MCP servers with OAuth handled for users. Administrators can enable specific tools, all published tools, or current and future tools. Disabled tools are hidden from clients and cannot be invoked; calls can be logged and included in traces.
Snowflake cited a survey it conducted in which roughly one in three customers surveyed used custom-built dashboards to monitor agents. It also cited Cloud Security Alliance survey findings that 82% of organizations had discovered at least one previously unknown AI agent or autonomous workflow in the past year. The Cost of a Data Breach Report 2026 put average breach cost for organizations with high shadow AI use at $5.39 million and said shadow AI incidents affected 43% of breached organizations. Cortex AI Gateway records model and tool calls that pass through it as OpenTelemetry spans in Snowflake, including calls from local developer laptops and remote Kubernetes clusters, without requiring agent code changes. Trace records can include requested model, input and output tokens, cache activity, duration, status code, tools, client version and session. Prompt, response and tool-call content is captured only if an administrator enables payload capture.
Snowflake says its existing Cortex AI Guardrails, which protect its own AI products against prompt injection and jailbreak attempts, will soon extend to third-party agents routed through the gateway. Its cost controls include budgets set at gateway level or by resource and user tags, plus monthly and optional daily per-user quotas that enforce blocks within minutes. Gateway usage is recorded in AI_GATEWAY_USAGE_HISTORY with credits broken out by model and request IDs for tracing spending to conversations, workflows, tasks, requests and tool calls. Snowflake described Cortex AI Gateway as a starting point for enterprises to inspect agent