Openrouter Model Routing

Implement intelligent model routing to optimize cost, quality, and latency on OpenRouter. Use when building multi-model systems or optimizing spend across task types. Triggers: 'openrouter routing', 'model routing', 'route to model', 'model selection openrouter'.

Beruf
Kategorien: Machine Learning

Overview

OpenRouter gives you access to 100+ models through one API. The key to cost efficiency is routing each request to the right model based on task complexity, required capabilities, cost budget, and latency requirements. This skill covers task-based routing, complexity classification, cost-aware selection, and OpenRouter's native routing features.

Task-Based Router

import os, re
from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
    default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
)

# Model tiers by cost and capability
MODELS = {
    "free":    "google/gemma-2-9b-it:free",          # $0/0 — testing only
    "budget":  "meta-llama/llama-3.1-8b-instruct",   # $0.06/$0.06 per 1M
    "mid":     "openai/gpt-4o-mini",                  # $0.15/$0.60 per 1M
    "standard":"anthropic/claude-3.5-sonnet",         # $3/$15 per 1M
    "premium": "openai/o1",                           # $15/$60 per 1M
}

TASK_ROUTING = {
    "classification":  "budget",   # Simple label assignment
    "translation":     "mid",      # Moderate quality needed
    "summarization":   "mid",      # Good quality, cost-effective
    "code_generation": "standard", # Needs high accuracy
    "code_review":     "standard", # Needs reasoning
    "analysis":        "standard", # Complex reasoning
    "creative_writing":"standard", # Quality matters
    "deep_reasoning":  "premium",  # Multi-step logic
    "simple_qa":       "budget",   # Basic questions
    "chat":            "mid",      # General conversation
}

def route_request(task_type: str, messages: list[dict], **kwargs) -> dict:
    """Route to appropriate model based on task type."""
    tier = TASK_ROUTING.get(task_type, "mid")
    model = MODELS[tier]

    response = client.chat.completions.create(
        model=model, messages=messages, **kwargs
    )
    return {
        "content": response.choices[0].message.content,
        "model": response.model,
        "tier": tier,
        "tokens": response.usage.prompt_tokens + response.usage.completion_tokens,
    }

Openrouter Model Routing

Beruf
Kategorien: Machine Learning

Overview

Task-Based Router

import os, re from openai import OpenAI client = OpenAI( base_url="https://openrouter.ai/api/v1", api_key=os.environ["OPENROUTER_API_KEY"], default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"}, ) # Model tiers by cost and capability MODELS = { "free": "google/gemma-2-9b-it:free", # $0/0 — testing only "budget": "meta-llama/llama-3.1-8b-instruct", # $0.06/$0.06 per 1M "mid": "openai/gpt-4o-mini", # $0.15/$0.60 per 1M "standard":"anthropic/claude-3.5-sonnet", # $3/$15 per 1M "premium": "openai/o1", # $15/$60 per 1M } TASK_ROUTING = { "classification": "budget", # Simple label assignment "translation": "mid", # Moderate quality needed "summarization": "mid", # Good quality, cost-effective "code_generation": "standard", # Needs high accuracy "code_review": "standard", # Needs reasoning "analysis": "standard", # Complex reasoning "creative_writing":"standard", # Quality matters "deep_reasoning": "premium", # Multi-step logic "simple_qa": "budget", # Basic questions "chat": "mid", # General conversation } def route_request(task_type: str, messages: list[dict], **kwargs) -> dict: """Route to appropriate model based on task type.""" tier = TASK_ROUTING.get(task_type, "mid") model = MODELS[tier] response = client.chat.completions.create( model=model, messages=messages, **kwargs ) return { "content": response.choices[0].message.content, "model": response.model, "tier": tier, "tokens": response.usage.prompt_tokens + response.usage.completion_tokens, }

Error	Cause	Fix
Wrong model selected	Classification too coarse	Add more task categories; test with diverse prompts
Model unavailable	Selected model temporarily down	Add fallback chain per tier
Cost overrun	Complex tasks routed to premium models	Set `max_tokens` and daily budget caps
Quality regression	Budget model can't handle task	Monitor output quality; escalate tier on poor results

Openrouter Model Routing

Overview

Task-Based Router

Openrouter Model Routing

Overview

Task-Based Router

Complexity-Based Auto-Router

OpenRouter Native Routing

Cost-Aware Router

Error Handling

Enterprise Considerations

References

Continuous Learning V2

Continuous Learning V2

Continuous Learning V2

Continuous Learning

Continuous Learning

Pytorch Patterns