Control Tower v0.2.1

Config file

Most setups are done in the console, but providers, models, fallbacks, MCP servers and alerting can also live in a config.yaml — reviewed in pull requests and applied at every start:

controltower --config config.yaml                 # or CT_CONFIG=config.yaml, or CONFIG_FILE_PATH
docker run -v $(pwd)/config.yaml:/app/config.yaml -p 4000:4000 ghcr.io/joshmaster2165/controltower --config /app/config.yaml
  • The file is applied at every start: edits and removals take effect on restart, and nothing is duplicated.
  • Rows it creates keep stable ids, so the map, history and gates keep pointing at the same models across restarts.
  • Providers, models, keys and gates added in the console or through the API are separate and are left alone. Models declared in the file can't be deleted through the API — remove them from the file.
  • A file that can't be read or parsed stops startup with the reason. A provider whose key isn't in the environment is skipped with a warning, and the server still starts.

Prefer a one-off import you can then edit in the console? Use Models → Import config with the same file. For zones and gates, see Policy as code and --policy.

Example

model_list:
  # One model on one provider
  - model_name: gpt-4o
    params:
      model: openai/gpt-4o
      api_key: os.environ/OPENAI_API_KEY

  # Two entries with the same name and different `order`: tried in order
  - model_name: claude-sonnet
    params:
      model: anthropic/claude-sonnet-4-5
      api_key: os.environ/ANTHROPIC_API_KEY
      order: 1
  - model_name: claude-sonnet
    params:
      model: bedrock/anthropic.claude-sonnet-4-5-20250929-v1:0
      aws_region_name: us-east-1
      order: 2

  # A local model, priced by you
  - model_name: local-llama
    params:
      model: ollama/llama3.2
      api_base: http://localhost:11434
    model_info:
      input_cost_per_token: 0
      output_cost_per_token: 0

  # Embeddings
  - model_name: text-embedding
    params:
      model: openai/text-embedding-3-small
      api_key: os.environ/OPENAI_API_KEY
    model_info:
      mode: embedding

  # Every Groq model, added the first time it is requested
  - model_name: "groq/*"
    params:
      model: "groq/*"
      api_key: os.environ/GROQ_API_KEY

router_settings:
  routing_strategy: simple-shuffle

settings:
  fallbacks: [{ "gpt-4o": ["claude-sonnet"] }]   # claude-sonnet only when gpt-4o fails

mcp_servers:
  github:
    url: https://api.githubcopilot.com/mcp/
    transport: http
    auth_type: bearer_token
    auth_value: os.environ/GITHUB_TOKEN

general_settings:
  master_key: os.environ/TOWER_ADMIN_KEY
  alerting: ["slack"]                      # needs SLACK_WEBHOOK_URL
  alert_types: ["llm_exceptions", "budget_alerts", "daily_reports"]

model_list

Each entry becomes a provider (one per distinct endpoint and credential) and a deployment.

FieldBecomes
model_nameThe name agents ask for
params.model<provider>/<model> — the provider and the upstream model
params.api_key, api_base, api_versionProvider credentials and endpoint. os.environ/NAME reads the environment; the usual provider variables (OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY, AWS_* …) are used when a key is left out
params.aws_access_key_id, aws_secret_access_key, aws_session_token, aws_region_name, aws_bedrock_runtime_endpointAWS Bedrock credentials (the session token is optional, for temporary credentials)
params.vertex_project, vertex_location, vertex_credentialsGoogle Vertex AI
params.credentialA named entry in credential_list (shared api_key / api_base values): credential_list: [{credential_name: shared-openai, credential_values: {api_key: os.environ/OPENAI_API_KEY}}]
params.orderPriority within a model group (lower first)
params.weight, rpm, tpmWeight within a model group; rpm and tpm are also the deployment's own limits
params.max_parallel_requestsThe deployment's limit on calls at a time
params.region_name (or aws_region_name, vertex_location)The deployment's region, for keys that must keep their data in a region
params.tagsReserved for tags (default serves untagged requests too)
params.timeout, stream_timeoutSeconds to wait for the provider's first byte
model_info.max_input_tokens (or max_tokens)The model's context window, when the price table doesn't know it
model_info.input_cost_per_token, output_cost_per_tokenA price override, in dollars per token
model_info.mode: embeddingAn embeddings model. chat (the default) and completion are chat models; other modes are skipped

Several entries with the same model_name become an alias that routes across them:

  • If every entry has a different order, they are tried in order, lowest first; the next is tried only when the one before fails.
  • Otherwise router_settings.routing_strategy picks how:
    • simple-shuffle (the default) → weighted: each request goes to one of the entries with the lowest order (or no order), picked at random in proportion to weight. Entries with a higher order are tried only if it fails.
    • latency-based-routing → fastest first.
    • cost-based-routing → cheapest first, by the price of a typical call (input price × 3 plus output price, per million tokens); entries without a known price go last.
    • least-busy, usage-based-routing and usage-based-routing-v2 are treated as weighted. Any other name is treated as weighted too, with a warning.

settings.fallbacks adds other models from the file as fallbacks: whatever the routing strategy, they are tried only after all of the model's own entries.

router_settings.model_group_alias: {gpt-4: gpt-4o} adds another name for a model in the file: agents asking for gpt-4 get gpt-4o's entries. An alias naming a model that isn't in the file, or a name already taken, is skipped with a warning.

These become the model's routing, in router_settings or settings:

FieldBecomes
context_window_fallbacks: [{gpt-4o-mini: [gpt-4.1]}]Prompt too long → try
content_policy_fallbacks: [{gpt-4o: [claude-sonnet]}]Content refused → try
default_fallbacks: [gpt-4.1], or fallbacks: [{"*": [...]}]Anything else → try, for every model
num_retries: 2Retries on the same deployment for rate limits, timeouts and server errors
retry_policy: {RateLimitErrorRetries: 3, TimeoutErrorRetries: 1, InternalServerErrorRetries: 2}Retries for each

A fallback may name a model in the file or one Control Tower already has.

Providers: openai, azure, azure_ai, anthropic, gemini, vertex_ai, bedrock, groq, mistral, together_ai, fireworks_ai, deepseek, xai, openrouter, perplexity, cerebras, deepinfra, nvidia_nim, sambanova, ollama, ollama_chat, hosted_vllm, lm_studio, and openai/ with any api_base (any OpenAI-compatible server).

Wildcards: model_name: "openai/*" with model: "openai/*" connects the provider and adds each model the first time it is requested. model_name: "*" does that for every provider whose key is set in the environment.

Secrets: values in the file are used as given; os.environ/NAME is read from the server's environment, or from the file's own top-level environment_variables: {NAME: value}. Control Tower's own CT_* variables are never read from a config file.

A file pasted into Import config or sent to the admin API (and a model sent to POST /model/new) is trusted less than the --config file the server starts with: its os.environ/ references can't read the server's own settings and secrets: CT_* and the other variables the server reads for itself, UI_*, DATABASE_URL, REDIS_URL, PG*, process variables such as PATH, HOME and NODE_*, and any name containing MASTER, PASSWORD, PASSWD, SESSION_SECRET or PRIVATE_KEY. So an admin session can't copy the server's database password or master key into a provider's settings and send it elsewhere. A --config file (or CT_CONFIG) is part of the deployment and may read them.

mcp_servers

mcp_servers:
  files:
    url: http://files-mcp:3001/mcp
    transport: http                  # Streamable HTTP
    auth_type: bearer_token          # bearer_token | api_key (x-api-key) | basic
    auth_value: os.environ/FILES_MCP_TOKEN
  crm:
    url: https://crm.example.com/mcp
    static_headers: { X-Team: support, X-Api-Key: os.environ/CRM_MCP_KEY }

authentication_token is another name for auth_value. static_headers are sent on their own or together with api_key or basic auth. Next to bearer_token they are ignored (the startup log warns): only the token is sent.

Each entry becomes an MCP server whose tools agents reach at /mcp as files__<tool>. Servers with only a command (stdio) are skipped with a warning.

general_settings

FieldBecomes
master_keyThe admin key, unless CT_ADMIN_KEY is set
alerting: ["slack"]A Slack alert channel from SLACK_WEBHOOK_URL
alert_typesAlert rules; without it: failed requests, slow requests, budgets and outages

Settings Control Tower manages itself — database, authentication, UI access and spend-log options — are listed as ignored in the startup log.

Not supported

include files, logging callbacks (Control Tower records every flight itself — use metrics or alert webhooks), and guardrails in the config (use inspect gates).

--model quick start

OPENAI_API_KEY=sk-… controltower --model openai/gpt-4.1-mini

Serves one model with credentials from the environment, with no config file. Agents ask for it by the name you passed — here openai/gpt-4.1-mini; after --model ollama/llama3.2 it is ollama/llama3.2. GET /v1/models lists it.