AIM Media House

Snowflake Adds Dynamic Model Routing to Cut Enterprise AI Costs

Snowflake Adds Dynamic Model Routing to Cut Enterprise AI Costs

Snowflake introduced dynamic model routing in Cortex AI Gateway, allowing enterprises to choose models based on cost, quality and performance.

Snowflake has introduced dynamic model routing in Cortex AI Gateway, allowing enterprises to automatically select AI models for individual tasks based on factors including cost, quality and latency.

The company announced the capability on August 18 as part of a push to improve what CEO Sridhar Ramaswamy calls “intelligence efficiency.” Snowflake said its initial benchmarks showed the routing approach could deliver better economics at a given quality level than relying on a single model.

The announcement comes as enterprises face growing costs from AI inference. As AI systems move into production, the economics of running models at scale are becoming a larger consideration, particularly as companies expand the number of AI agents and workloads they operate. AI infrastructure demand is already reshaping enterprise technology costs.

Snowflake Lets Enterprises Set Model Policies

Cortex AI Gateway allows customers to define which models they approve and the tradeoffs they want the system to consider. The gateway then evaluates each task against those policies and available cost and performance data before selecting a model.

Snowflake said the system is designed to adapt as models change. After one model completes a task, another model evaluates the result, creating a feedback loop that can inform future routing decisions.

The company said its internal testing found that dynamic routing delivered up to three times greater token efficiency on some workloads compared with relying on a frontier model alone while maintaining comparable quality. Snowflake also reported a 25% improvement in token efficiency at the same pull-request throughput in a separate coding evaluation. These figures come from Snowflake's own testing.

The routing capability is optional. Customers can continue to use a specific model or limit the system to an approved group of models.

Snowflake said it does not charge a separate fee for routing. Customers pay for model usage based on token consumption, meaning the choice of a lower-cost model can reduce the cost of a request.

The company is also expanding the number of models available through its platform. Snowflake said it supports models from Anthropic, Google, Mistral AI, OpenAI and SpaceXAI, and is adding GLM-5.3 and DeepSeek-V4-Flash 0731.

The expansion gives enterprises more options as model economics change. Snowflake's position is that companies should be able to switch between models rather than standardize permanently on one model as performance and pricing evolve. The growing focus on AI infrastructure economics is also changing how enterprises think about compute and inference.

Model Routing Becomes an Enterprise Infrastructure Layer

Snowflake is entering a market where other cloud and AI infrastructure providers are also building routing capabilities.

Microsoft's Azure AI Foundry includes a model router that can select models based on configured optimization goals, while Databricks offers AI Gateway capabilities for routing, governance and traffic management. NVIDIA has also introduced NeMo Switchyard, an open-source routing layer designed to dynamically direct requests between AI models.

Google Cloud has similarly developed AI Gateway capabilities that can route requests based on factors such as cost, latency and accuracy.

The competition reflects a change in how enterprises manage model choice. Instead of requiring developers to determine which model should handle every request, routing systems can make that decision at the infrastructure layer.

For Snowflake, that layer is tied to its existing data governance controls. The company's Cortex AI Gateway is positioned as a control point for model and agent traffic, while Snowflake's broader platform connects model access with enterprise data and permissions.

That approach also reflects the wider push to connect models, data and agents through governed infrastructure. The competition around enterprise AI orchestration is increasingly focused on that connective layer.

Snowflake's announcement ultimately shifts the model-selection question from choosing one model for an organization to deciding how models should be selected for individual tasks. As enterprises add more AI workloads, that distinction could make model routing an increasingly important part of managing the cost and performance of AI systems.

Key Takeaways

  • Snowflake's dynamic model routing improves AI cost efficiency by selecting models based on cost, quality, and performance.
  • Cortex AI Gateway allows enterprises to define model policies, optimizing task selection and resource usage.
  • Internal tests showed up to three times greater token efficiency compared to using a single AI model.
  • The dynamic routing system adapts over time, creating a feedback loop to enhance future model selection.
  • This new capability addresses the rising costs of AI inference as enterprises scale their AI operations.