LLM Configuration & Prompt Overlays

What it is

The control center for how every model call is made — which provider / model, with what fallbacks, tools, sampling, limits, and system-prompt shaping, selected per call site.

Operator surface

  • Presets — llm_preset_list_v2, llm_preset_get_v2, llm_preset_create_v2, llm_preset_update, llm_preset_delete_v2, and llm_preset_simulate (dry-run resolution for a site).
  • Site attachment — llm_site_attach_preset, llm_site_detach_preset, llm_site_attached_preset, llm_preset_apply.
  • Modes (atomic bundles) — llm_mode_list / _get / _create / _update / _delete, llm_mode_preview, llm_mode_apply, llm_mode_history.
  • Providers / credentials / backends — llm_provider_list / _create / _test / _refresh_models, llm_credential_list / _create / _set_key / _quarantine / _set_coder_cap, llm_backend_list / _create, llm_config_set_fallback_chain, llm_config_set_max_turns, plus the legacy llm_config_* family and snapshots (llm_snapshot_list / _get, and llm_call_explain for one call's snapshot + every recorded attempt in one document — see Recipe 13 in docs/MCP-COOKBOOK.md).

REST: LlmConfigController, LlmConfigModesController, LlmValidationController, LlmAvailabilityController.

How it works

A preset (LlmPresetRecord) is a composable, optionally-inherited bundle of routing steps (LlmPresetStepRecord → RoutingChain / RoutingTarget with fallbacks), a tool policy, sampling and limits profiles, a prompt overlay, and MCP capability toggles (browser / emulator / windows). A prompt overlay (PromptOverlayRecord) injects or suppresses named experts and adds a custom system prefix / suffix.

Presets attach to call sites by writing a LLM.{SiteId}.Preset config key; at dispatch the ProviderResolver reads it (falling back to LLM.{Family}:Default.Preset when a specific site is unattached) and resolves the concrete provider / model. These catalogs live in the control plane; modes (LlmConfigModeRecord) apply a whole set of site attachments at once with full preview and history. This is the same resolution path the Agents planning sites and the Workflows call-site nodes use.

Reading what happened to a call

Configuration says how a call should have been made. Call telemetry says what actually happened, and it is readable without knowing the storage schema.

Every snapshot carries its own outcome. Alongside the materialized configuration, each snapshot row — in the list, in the detail, and in llm_preset_simulate output — reports disposition and dispositionAt (how and when the call ended), correlationId (which chain it belonged to), attemptCount and servedStepOrder (how many tries, and which step of the fallback chain finally served), errorCode, resolutionRule and resolutionId (which rule chose the preset), writer, totalLatencyMs, and estimatedCostUsd.

Dispositions are served, failed, walled, refused_preflight, synthetic, cancelled, timed_out and orphaned. An unset disposition means the call is still in flight — so an unfiltered list is not simply the sum of those states.

Filtering by outcome. GET /api/v1/llm/snapshots accepts disposition and correlationId filters, so "show me what went wrong on this site" is one query rather than a page-through. These filters, like siteId, are exact equality — there is no wildcard or prefix form.

The attempt history. GET /api/v1/llm/snapshots/{id}/attempts returns one row per attempt, in the order the runtime tried them: outcome, HTTP status, error detail, the provider-side request id to quote to a vendor, which credential and account were used, which rate-limit window was consumed and when it resets, first-token latency, cache-token counters, and start / completion times. A step the chain never reached has no row, so the rows themselves show how far the fallback chain got. A snapshot with no recorded attempts returns an empty list, which stays distinct from an unknown snapshot.

llm_call_explain answers the same question in a single call when you would otherwise chain a list, a detail read and an attempt read together.

Prompt text and the authorization boundary. The attempt history deliberately contains no prompt or response text — and that is exactly why it is available to any authorized caller: an operator can trace a failed call end to end without being exposed to conversation content. Prompt and response text remain on the snapshot detail route, behind the tenant-admin policy. The separation is deliberate: the fields needed to diagnose a failure are not the fields that carry customer data, so they do not have to share a permission level.