Agents

Audience: Genesis operators and Genesis agents alike. This page documents the agents subsystem end to end — what agents are and how you drive them (operator surface), then how the dispatch, scheduling, budgeting, and LLM-call internals actually work (architecture). It is a deep page; skim the operator half, read the architecture half when you need to reason about behaviour.


1. What an agent is

In Genesis, an agent is a long-lived autonomous worker that runs a plan-then-act loop on its own cadence. Unlike a one-shot LLM call, an agent:

  • persists as a row (agent_instances) in its tenant's database, with a lifecycle status, an active operating mode, and self-pacing scheduling state;
  • wakes on a schedule (a heartbeat "tick") and on external events (an operator message, an answered question, a mission state change);
  • reasons once per tick through a single provenance-recorded LLM call, then takes typed actions (propose a mission, escalate a decision, message the operator) strictly through governed tool calls;
  • remembers its own past turns (an append-only state ledger) and re-injects them as continuity context;
  • never self-applies anything risky — every consequential action is either governance-gated or routed to a human via the decision queue.

There are two broad families that share most of this machinery:

  1. Standing agents — always-on, tenant-scoped brains that run a continuous plan-then-act loop: the Strategic Orchestration Agent (SOA) and the Compliance agent. These are the focus of this page.
  2. Bursty / bound agents — agents that exist to do a finite job and then go quiet: the Prompt Engineer (one burst per call-site rework) and mission agents (one mission lifecycle each).

Separate from both is the lightweight in-memory agent registry — a coordination directory of external/worker agents (coders, reviewers, monitors) used for capability-based task routing. The same word "agent" covers both the heavyweight standing brains and the registry entries; the sections below keep them distinct.


2. Operator surface

2.1 Agent types

There are two complementary type systems.

(a) The persisted AgentType discriminator — the authoritative type of an agent_instances row. It selects the agent's policy, its planning preset/site, its operating-mode catalogue, and (for missions) the lifecycle pump.

AgentTypeProduct nameFamilyWhat it does
Soa (default)Strategic Orchestration AgentStanding, nativeLong-horizon tenant orchestration — surveys the tenant + platform health, proposes well-scoped strategic missions, escalates genuine forks to the operator.
ComplianceCompliance agentStanding, nativeStanding regulatory/compliance brain — sweeps posture, researches EU frameworks (CRA, DORA, NIS2, EU AI Act, PLD), authors dual-form regulatory dossiers, escalates compliance calls. Never self-applies (no mission-create tool).
PromptRedactorPrompt EngineerBursty, boundReworks the committed prompt layers for one model+provider couple at one call-site and proposes the result for human approval (MR-only; never self-approves). The enum member name / DB string stays PromptRedactor as a wire/DB contract; the C# types and UI were renamed to Prompt Engineer.
MissionFeatureFeature missionMissionGeneric code-producing mission (the decompose→code→review pipeline).
MissionTenantBootstrapFirst-run onboardingMissionFirst-run tenant onboarding.
MissionProjectBootstrapNew-project onboardingMissionReturning-user new-project onboarding.
MissionResearchResearch missionMissionRead-only investigation.
MissionReviewReview missionMissionReview-engine analysis.
MissionAspectOnboardingAspect onboardingMissionPer-project guided onboarding.
MissionBugInvestigationBug investigationMissionSupport-case investigation.
MissionBugFixBug fixMissionSupport-case fix.

The Mission* members mirror the mission "kind" one-for-one. A mission "runs as an agent-type" only at the policy + LLM-invocation seam — its lifecycle is still driven by the separate mission pump off the missions / mission_event_queue tables. Native standing agents (Soa, Compliance, PromptRedactor) run directly on agent_instances / agent_event_queue.

(b) The in-memory registry role — a lighter directory for coordination and routing. When you call agent_register, you supply a free-form role (Developer, Reviewer, Coordinator, Monitor, or Generic) plus capabilities. These registry agents are how coders and other workers advertise themselves so tasks can be routed to them by capability. Registry entries are ephemeral (in-process, tenant-partitioned, heartbeat-expired after 5 minutes); the agent_instances rows are durable.

2.2 Registration & discovery (the coordination registry)

These MCP tools manage the in-memory registry (they proxy Core's /api/v1/agents controller):

ToolPurpose
agent_registerRegister an agent: name, role, comma-separated capabilities, optional endpoint, lifecycleKind (manual/docker/process/aspire), transportKind (in-memory/rabbitmq/http), transportTarget, containerId, JSON labels.
agent_listList registered agents, filterable by role / capability / status / lifecycle / transport.
agent_list_availableList only active agents (heartbeat within the last 5 min, not offline) currently free for assignment.
agent_discoverScored discovery by capability — ranks by proficiency, language, and framework match (minProficiency 1–5). The intelligent-routing entry point.
agent_heartbeatKeep an agent alive (resets the 5-minute offline timer).
agent_deregisterRemove an agent from the registry.
agent_route_taskRoute a task to the best agent by capability match (capability/language/framework/priority) and dispatch over its transport, or force a targetAgentId.
agent_message_queue_depthInspect an agent's pending message count.
message_send / message_receivePoint-to-point or broadcast (*) messaging between registry agents.
task_assign / task_statusCreate/assign a task to a registry agent and query task status.

The standing SOA registers itself here automatically (role/type orchestration) so it appears on the /agents dashboard and can receive live turn pushes; that registration is a UI surfacing convenience, not the SOA's control plane.

2.3 Spawning and driving coders

Coders are containerised coding workers. Two layers:

On-demand coders (OnDemandCoderMcpTools, tenant-admin, proxy /api/v1/coders/on-demand):

ToolPurpose
coder_spawnStand up one coder for your tenant: tool (e.g. claude, copilot — must be fleet-enabled), variant (default standard), optional projectId (display/intent only), sessionId to resume a prior session (--resume), idleTtlMinutes (default 60). It survives queue-idle scale-to-zero and self-kills after the idle TTL.
coder_listList your tenant's on-demand coders (name, tool, variant, state, idle TTL, resume session id, project).
coder_killKill one coder by container name (your tenant only).

Coding tasks (CodingTaskMcpTools, tenant-admin) — the actual unit of work a coder executes:

ToolPurpose
coding_task_dispatchQueue a coding task: taskType (code/test/review/refactor), prompt, optional context, targetPath, branch, maxTurns (default 100, max 1000), enableBrowserMcp/enableEmulatorMcp/enableWindowsMcp (attach a resource pool to the task; default false), coderVariant (standard/ide image flavor — browser/android are retired, use the pool flags), coderTool (claude/copilot/codex/gemini/aider/octofriend/custom), model/provider/backend or a presetId, and optional conversationId+attachmentIds to deliver chat-uploaded files into the coder workspace. Async — poll status.
coding_task_statusCheck a task; returns the result when complete.
coding_task_listList pending + recently completed tasks.
coding_task_extendResume a task that recorded a coder session with extra turns (1–1000) and optional extra instructions; continues from that session. The session is the precondition, not the status — in practice TurnLimitReached. Failed is refused in practice (the harness routes any run that produced a session to TurnLimitReached, so a failed run died before one existed); a turn-limited run killed by the wall clock lacks one too. Both refusals name the case; dispatch a new task instead.
coding_task_session_infoTurns used / max turns / session id for resumption.

The model/provider/backend for a coding task resolves through the same preset machinery as everything else (see §3.7); presetId overrides the explicit model/provider/backend triple.

Windows lane (WindowsMcpTools, behind the windows-pool feature, attached with enableWindowsMcp) — a coder leases a Windows worker on first use (no lease verb; the session is held across calls until windows_release hands it back, the task settles, or the lease expires) and drives it through two shapes. The build shape: windows_run_spec (submit a script, get a task handle) + windows_task_status (poll). The desktop shape (ARDS-857), which drives real UI Automation on the leased machine the way the browser tools drive a page: windows_launch (start a program, get {pid, hwnd, title}), windows_windows (visible top-level windows with their hwnd and bounds), windows_focus / windows_close, windows_snapshot (the UI Automation tree as indented text, one node per line with a ref=e<N> that stays valid until the next snapshot), windows_click and windows_type (target by ref, else automationId / name / controlType, else a screen point; the returned method says whether the element's own pattern or a real mouse click / keystrokes did it), windows_key (chords such as Ctrl+S, Alt+F4), windows_wait_for (an element or a window title reaching exists / enabled / gone), windows_computer (raw mouse / keyboard in primary-screen pixels) and windows_screenshot (the primary screen, or cropped to an hwnd; its pixels are the coordinates windows_computer takes). windows_release hands the leased machine back to the pool at once: it ends the session (a windows_run_spec build still running on it is cancelled and the programs launched on it are closed), the next windows_* call leases a new one, and it is idempotent — with nothing held it answers released: false (outcome: not_held) and changes nothing. windows_type takes exactly one of text / secretRef: a secretRef names a secret the operator provisioned on the Windows pool (a credential, a licence key) — the pool types its value on the machine and the value never reaches the coder. A secretRef may only be typed into a password-masked control: aimed at anything else the guest refuses without typing (SECRET_TARGET_NOT_MASKED), and on a call that carried typed input the pool's error body reaches the coder projected to {error, reason, state, secretRef, host} — never the guest's free text. windows_screenshot with an hwnd is a pure read — the window is not brought to the front and a covered region shows what is on top; windows_focus first when it matters. When a coder's Windows session is released (by windows_release, expired, evicted, or the task settled), the pool has the guest close every process that session started with windows_launch before the slot is reused, so nothing a coder launched outlives its lease. windows_wait_for.timeoutMs is clamped to 10 minutes and windows_launch.waitForWindowMs to 55 s — the guest's own ceiling, ordered guest < pool < core (the pool's launch budget is the wait + 10 s), so a program that shows no window by then comes back as the guest's honest hwnd: null, never as UI_TIMEOUT (both clamped, not refused). Every error a Windows verb returns carries one of these codes. FEATURE_DISABLED: the windows-pool feature is off for this tenant — every verb, before any lease; an operator turns it on. VALIDATION_ERROR: windows_type was given neither or both of text / secretRef, or a secretRef that is not a bare name — fix the call, no machine was leased. A host with no interactive desktop answers every desktop verb DESKTOP_UNAVAILABLE with the pool's reason: session_0 is a compile-only install (no retry changes it); capture_probe_failed means the host's interactive session stopped rendering (its RDP client disconnected or was minimised) — an operator has to restore that window, after which the pool re-probes the desktop within 30 s, so retry only after the operator acts; ELEMENT_NOT_FOUND means re-snapshot, WINDOW_NOT_FOUND means that hwnd is gone (list again with windows_windows), WAIT_TIMEOUT carries the elapsed time, UI_TIMEOUT means the guest did not answer within the pool's per-call budget (a modal dialog may be blocking the desktop — windows_screenshot shows it; retryable), SECRET_NOT_FOUND names the missing secret for the operator, WINDOW_NOT_FOREGROUND means focus the target (or dismiss what covers it) and retry, INPUT_BLOCKED means a higher-integrity window or a locked desktop refuses input. Anything else is the pool's generic vocabulary, shared with the browser and emulator lanes: POOL_UNAVAILABLE (no HTTP answer — the pool is unreachable or the call outlived core's HTTP timeout; retry), POOL_ERROR (the pool answered 5xx; at capacity it names the current holder(s), so a busy pool reads differently from a leaked hold; retry), UNAUTHENTICATED / FORBIDDEN (the pool refused core's credential — an operator fixes the resource-pool-auth secret; no retry helps), NOT_FOUND with sessionEvicted: true (the pool no longer knows the session: the hold was dropped, the machine's state is gone, the next call leases a fresh one), CONFLICT (a 409 outside the desktop vocabulary, such as a windows_run_spec while a task is already running — its taskHandle is in poolDetail), RATE_LIMITED (429; retry later) and BAD_REQUEST (any other refusal — an unknown key name in windows_key, for instance). Each error's data carries httpStatus, retryable and the pool's own body as poolDetail.

2.4 Human-in-the-loop (HITL)

Agents never silently make irreversible calls. There are two HITL channels:

The decision queue (agent escalation). When a standing agent needs a human call it raises an AttentionRequest via soa_propose_decision (or, internally, soa_propose_mode_change for a mode switch). This:

  1. is constitution-gated before it is created (a violation is fed back to the agent, not thrown);
  2. appears on the operator's decision queue (the decision_* tools / dashboard) as a SOA-origin request carrying { source: "soa", soaInstanceId };
  3. pauses the relevant work until the operator answers (decision_respond).

When the operator resolves it, AgentControlService.EnqueueOperatorAnsweredEscalationAsync enqueues an OperatorAnsweredEscalation event back to the exact originating instance, carrying the operator's resolution + resolutionNotes. The agent resumes from that event on its next tick. The round-trip is idempotent (a unique partial index dedupes concurrent REST-then-MCP resolves). Critically, the agent never enqueues its own events — the control side owns every enqueue.

Operator chat. soa_send_message (agent→operator) and the operator's reply (which lands as an OperatorMessage event) form a non-blocking chat channel, persisted as the durable transcript and broadcast live to the agent's detail page.

Mission HITL (mission_ask_user). The mission family has its analog: a mission sub-agent can call mission_ask_user to pose a question that surfaces to the operator and feeds the answer back as refinement guidance the mission resumes from. It is the mission-side equivalent of the SOA's soa_propose_decision → OperatorAnsweredEscalation loop.

2.5 Operating modes, presets & call-sites

Operating modes shape what an agent focuses on and which tools it may use this episode. Each agent type has its own seeded mode catalogue (table-backed, in-code-seed fallback):

  • SOA (six modes): Strategic (default), Incident, Dialogue, Housekeeping, Idle, Discovery.
  • Compliance (three): Cadence (default), GateReview, Research.
  • Prompt Engineer (one): Redact.

The agent cannot switch its own mode (rail R3/R7). The mode is written only by a human:

  • the UI mode control (AgentControlService.SetActiveModeAsync / the per-type Compliance + redactor variants, source=ui), or
  • the approval of a soa_propose_mode_change AttentionRequest (source=approval).

Both paths funnel through one chokepoint and append an audit row to agent_mode_change_events. The agent only ever proposes a switch.

Presets & call-sites. Every LLM call in Genesis is attributed to a named call-site in the static LlmSiteRegistry. The agent-relevant sites:

Site idUsed by
Agent:DefaultFamily-default preset parent for native agents.
Agent:PlanningThe SOA's per-tick planning/reasoning call.
Agent:ComplianceThe Compliance agent's per-tick call (split from SOA on 2026-06-24 so it can be pinned independently).
PromptRedactor:Default / PromptRedactor:RedactThe Prompt Engineer's Document/Evaluate sub-agent calls.
Agent:PodThe remote agent-pod class gate (the /mcp/agent boundary).

An operator binds a preset (a model+provider+backend+overlay bundle) to a call-site (llm_site_attach_preset / _detach_preset). A leaf site inherits its parent's family-default binding until pinned (e.g. Agent:Planning → Agent:Default). This is how you pin Compliance to a different model than the SOA without touching code. See §3.7 for resolution mechanics.

2.6 Dashboard pages

RoutePageShows
/soaSOA dashboardThe Strategic Orchestration Agent: status, start/stop/reset, operator chat, live turn timeline, current mode + mode control, mode-change audit.
/agentsAgents dashboardRoster of agents across the tenant (the standing brains + registry workers).
/agents/{kind}/{id}Agent detailPer-instance live detail: turn ledger, reasoning summaries, served-model + cost per turn (SignalR live push), chat.
/agents/complianceCompliance screenThe Compliance agent: posture, dossiers, start/stop, mode (Cadence/GateReview/Research).
/agents/redactorPrompt Engineer screenThe bound prompt-engineer instance per call-site, its proposals, mode.
/agents/{id}Registry detailA registry agent's detail.
/fleetCoder fleetOn-demand coders + fleet config; /coders is now a tab here.

3. Architecture

3.1 The data model

Standing agents live entirely on per-tenant tables.

agent_instances — one row per agent instance (the SOA analog of a mission row):

ColumnMeaning
idPK (Guid).
agent_typeThe AgentType discriminator (text NOT NULL DEFAULT 'Soa', stored as string).
bound_entity_type / bound_entity_id / bound_entity_keyBinding seam for bound agents. The Prompt Engineer binds to a call-site via (PromptRedactor, "CallSite", siteId) — bound_entity_key is the string sibling, with a partial unique index guaranteeing one instance per couple@call-site. Null for SOA.
labelDisplay label.
statusAgentInstanceStatus: Active (pumped), Paused (parked, resumable), Closed (terminal). Only Active instances are ticked.
current_phaseFree-form phase label the agent stamps each turn.
next_event_sequence_numberPer-instance monotonic counter sourcing the event queue's sequence.
lease_holder_session_id / lease_expires_atCAS lease columns (multi-replica safety).
next_tick_atWhen the next self-paced tick is due (audit; the queue's visible_at is authoritative).
agent_string_idThe registry record id surfacing this instance on /agents.
consecutive_idle_ticksB1 backoff counter (idle ticks since inputs last changed).
last_plan_input_hashHash of the planner's external inputs from the last tick (the no-op short-circuit key).
consecutive_blocked_turnsSTUCK-guard counter (consecutive PhaseBlocked ticks).
active_modeThe operating mode name (nullable; default applied in code). Written only by human paths.
is_building_starting_workset / workset_deadline_utcFoundational-phase flag + cost-safety deadline (see §3.6).
created_at / updated_atTimestamps.

agent_event_queue — the per-instance event log. Each row is one event with an event_kind, an optional JSON payload, a sequence_number, a visible_at (the authoritative due-gate), and a consumed marker (events are consumed, not deleted, so resume-dedup can see history). The AgentEventKind discriminator:

  • AgentInitialized — first event for a fresh instance (drives the plan-then-greet opener).
  • AgentTick — the self-scheduled heartbeat.
  • OperatorMessage — an operator chat message to read next tick.
  • OperatorAnsweredEscalation — an operator answered a decision the agent raised.
  • MissionStateChanged / DecisionResolved / DependencyCompleted — orchestration signals (a steered mission changed status, a decision resolved, a dependency unblocked).
  • ExternalSignalReceived — a generic agent-type-agnostic signal whose JSON payload carries an AgentSignalKind (NewCouple / GuidanceUpdated / RegressionDetected / HumanRequest). Used by the Prompt Engineer (Step 6 wires only HumanRequest); the SOA never receives it.

agent_state_ledger (+ agent_state_ledger_actions) — the append-only turn record. Each turn row carries the triggering event kind/payload, the AgentDecisionKind, the structured decision payload ({actions, waitCondition, reasoningSummary, activeMode}), the wait-condition, the reasoning summary, and a back-link to the LLM call snapshot (served model + cost). Inserts are lease-enforced. Per-action child rows normalise the actions the turn produced.

agent_modes — the per-type mode catalogue (name, description, prompt fragment, toolset-subset JSON, optional preset, optional budget override JSON, system-seed flag). agent_mode_change_events — the mode-change audit trail. agent_messages — the operator↔agent chat transcript (with direction).

AgentDecisionKind classifies each turn: Initialized, Heartbeat, MessageAcknowledged, EscalationResumed, NoOp (legacy), plus the planning-loop kinds Planned, ActedWithTool, WaitingForEvent, Escalated. It is additive-only (persisted as string; historical rows must keep classifying).

3.2 The dispatch model

The AgentDispatcher is a hosted BackgroundService that pumps standing agents. Its shape:

  1. Gated by the standing-agents feature, at two levels. The dispatcher is hosted whenever the deploy level is on (services.core.features.standingAgents.enabled, historically the AgentDispatcher:Enabled env; deploy default on), but it only ticks a tenant whose tenant level is on — and the tenant default is off: each tenant opts in from Settings → Features (see Feature flags). On start it enumerates active tenants and fans out one polling loop per tenant. Newly added tenants need a pod restart to be picked up.
  2. Per-tenant bootstrap. When AutoBootstrap=true, each tenant gets one singleton SOA instance ensured at startup (the schema does not enforce the singleton; the "one SOA per tenant" invariant lives in AgentControlService / the bootstrap caller). With StandingAgentsIdleByDefault=true it is created Active and seeded an AgentInitialized event (so it runs its plan-then-greet opener immediately, then idles); otherwise it is created Paused, awaiting an operator start. On first creation the instance also enters the foundational phase (§3.6).
  3. Polling cycle. Every PollingInterval (default 5 s) the dispatcher opens a per-cycle DI scope under an ambient tenant scope, lists instances "needing attention", and processes up to MaxConcurrentInstancesPerTenant (default 4) of them.
  4. Per-instance processing is lease-bounded and runs through the shared tick supervisor (next section).
  5. Error isolation at cycle and instance level — one bad instance or cycle never stops the loop.

The body to run is resolved per instance by AgentType via keyed DI (AgentType.Soa → AgentAgent; AgentType.Compliance → also AgentAgent but with the Compliance policy; PromptRedactor → PromptRedactorAgent). This is how one dispatcher serves every standing type.

Two execution regimes, one code path. A single config flag, AgentDispatcher:EphemeralPod, picks the regime:

  • Warm (default, EphemeralPod=false) — the in-process poll loop ticks instances directly. KeepWarm == !EphemeralPod.
  • Ephemeral pod (EphemeralPod=true) — the in-core pump stops ticking instances in-proc; KEDA-spawned agent pods pull ticks over the /mcp/agent endpoint via agent_tick_claim / agent_tick_complete. Bootstrap and event-seeding still run in-core; only per-instance ticking moves to the pods.

Both regimes drive ticks through the same AgentTickSupervisor primitives, so persistence, heartbeat, and STUCK logic are byte-identical whichever runs.

3.3 Tick scheduling: claim → run → complete

AgentTickSupervisor factors the lease + queue + passivation primitives so the warm dispatcher and the pod runner share them exactly.

ClaimTickAsync:

  1. Acquire the per-instance CAS lease for a fresh sessionId (skip if held by another session → null).
  2. Dequeue the next due event (visible_at <= now). If none is due, seed a heartbeat AgentTick if none is pending (first beat / chain recovery), release the lease, and return null.
  3. Read the pre-tick instance snapshot (the idle/STUCK counters key on it) and return a TickClaim.

CompleteTickAsync is the single passivation writer — exactly one place writes phase/next-tick/counters per tick:

  1. Session-guard the write — re-read the instance and confirm the lease is still held by this session (a pod whose lease expired mid-run must not clobber). If lost, skip the write and return.
  2. Derive the decision kind/phase (warm path passes the agent's already-decided kind/phase straight through, so the persisted value is byte-identical; the pod path lets the supervisor re-derive from the raw outcome via the shared AgentDecisionMapper).
  3. Idle detection — a tick is idle when it was an AgentTick whose PlanInputHash matches the prior tick's. consecutive_idle_ticks increments on idle, resets otherwise.
  4. STUCK guard — a genuine PhaseBlocked (the planner returned blocked, mapped to WaitingForEvent with phase idle:blocked) increments consecutive_blocked_turns. The counter is always persisted; the halt only fires when enforceStuckGuard=true (the pod path). On halt at MaxConsecutiveBlockedTurns (default 3) the instance is set Closed, the heartbeat is suppressed (so the queue drains and KEDA scales to zero), and an operator AttentionRequest is surfaced. The warm path never halts (preserving legacy behaviour).
  5. Maintain exactly one heartbeat — unless halting, enqueue a single future AgentTick at next_tick_at if none is pending.
  6. Surface the instance on /agents (warm only), write the single passivation update, push the turn live over SignalR (warm only), and release the lease in finally.

ComputeNextHeartbeatDelay sets the cadence:

  • When StandingAgentIdleHeartbeatEnabled=true (default), the next heartbeat is a flat StandingAgentHeartbeatInterval (default 24 h) — a slow daily check-in that is immune to telemetry jitter. So an agent goes quiet right after its opening greeting and stays quiet until a real event.
  • Otherwise it falls back to geometric idle backoff (ComputeTickDelay): base · BackoffFactor^priorIdle, capped at MaxBackoffInterval (default 30 min). This only ever engages on unchanged inputs.

Either way, real events bypass the cadence entirely — operator messages and answered escalations enqueue with visible_at = now and wake the agent on the next cycle.

This is the "standing vs bursty" distinction in practice: standing agents (SOA, Compliance) settle to the slow flat heartbeat and react to events; bursty agents (Prompt Engineer) are never dispatcher-bootstrapped — they are created Paused on a binding/signal and ticked only once activated.

3.4 The per-tick planning loop (AgentAgent)

AgentAgent is the shared plan-then-act body for every standing planning agent. Per tick it does not touch leases or the event queue (the dispatcher owns those); it only re-injects, plans, and records. The sequence:

  1. Re-inject — load the recent turn tail (last 12 turns) for continuity across ticks / pod restarts.
  2. Ops digest — probe "ARDS-self" operational health (DB deadlocks/rollbacks, error-rate spikes, firing alerts) via IOpsDigestProbe, bucketed into the plan-input hash so a real incident flips the bucket and forces a re-plan.
  3. Load the instance — needed for both the short-circuit and the mode resolution. If the read fails, the tick aborts cleanly (it cannot safely resolve mode/type) and retries next tick.
  4. B1 no-op short-circuit — compute a SHA-256 PlanInputHash over the external drivers only (event kind + payload + bucketed ops digest; the agent's own ever-changing reasoning tail is deliberately excluded). If the hash matches the last tick's and a forced re-plan isn't due, skip the LLM call entirely and idle. A safety valve (ForcedReplanEveryNIdleTicks, default 20) forces a real re-plan periodically to catch drift. This is the lever that stops idle agents re-planning every 30 s and burning spend.
  5. Per-event preamble — steer by what triggered the tick: AgentInitialized → survey + post a short grounded greeting then idle; OperatorMessage → read and reply; OperatorAnsweredEscalation → resume incorporating the answer; AgentTick → baseline.
  6. Open-items roster — the missions this agent already created + the escalations it raised (linked via a JSONB soaInstanceId metadata helper, index-served). Injected at the top of the objective so the model grounds against in-flight work before proposing — the direct lever on duplicate proposals.
  7. Resolve the active mode — read active_mode (default per type), look it up in the seeded catalogue, and assert its toolset ⊆ the §3 superset fail-closed (the resolver throws if a mode would grant an out-of-membership tool — there is no [Authorize] backstop on the warm path). The resolved mode's toolset becomes both the invoker's authorized tools and the objective's advertised action menu; its prompt fragment is appended to the invariant base.
  8. (Compliance only) cold-sweep corpus discovery — resolve the tenant's compliance project id(s) so the objective can name the corpus to query.
  9. Credit gate (B-iii) — resolve the tenant's weekly agent-planning credit window (default $60/wk). When CreditEnforcementEnabled and the pool is exhausted, skip the LLM call and idle until regen; otherwise meter the drawdown.
  10. Plan — one provenance-recorded LLM call through ISubAgentInvoker (§3.5), with the mode's toolset as the authorized tool surface and the resolved per-tick budget (§3.6).
  11. Map the outcome — histogram-driven precedence (raised an escalation → Escalated; took any side-effecting action → ActedWithTool; else clean completion → Planned/MessageAcknowledged/EscalationResumed by event; else blocked → WaitingForEvent). The ladder lives in the shared AgentDecisionMapper so warm and pod map identically.
  12. Record the append-only ledger turn (with the back-linked provenance snapshot) and return the AgentDecision — the dispatcher persists the phase, schedules the next tick, and releases the lease.

A "burst" is therefore a sequence of these ticks: the first (AgentInitialized) greets and idles; subsequent ticks fire on the flat heartbeat or on injected events; within a single tick the invoker runs a bounded multi-turn agentic loop (the "episode") that may query memory/telemetry then emit one or two proposals before terminating.

3.5 How agents do LLM calls (the in-process pipeline)

Agents reason through SubAgentInvoker — the codebase's manual agentic loop, deliberately not the SDK auto-invoke path, because it must observe the loop boundary every turn (budget enforcement, named-tool termination, per-turn provenance, dispatch-failure policy).

Each invocation takes a SubAgentRequest carrying the parent entity id (the agent instance), the AgentType, a phase kind (the agent uses a findings-only Research phase so no mission proposal batch is opened), the preset site key (Agent:Planning / Agent:Compliance, resolved from the agent's IPlanningAgentPolicy), the objective payload, the budget envelope, the authorized MCP tools (the active mode's toolset), an invocation id, and the mode's system-prompt suffix. It also carries cost-attribution hooks (OwnerAgentInstanceId, QuotaWindowId) and the foundational per-turn output-token override.

The loop body: seed messages → while not terminated: check budget → resolve the (model, provider, backend) couple from the preset/site (see §3.7) → call the LLM via a Microsoft.Extensions.AI IChatClient → write an LlmCallSnapshot (tagged AgentPlanning, keyed by invocation id) → update consumption → dispatch each FunctionCallContent through its typed MCP handler → react. The load-bearing rule (§1.1): the tool call is the commit; response text is reasoning trace and is never parsed for state. Termination is by named control-plane tool (phase_complete / phase_blocked), not by interpreting prose; if a turn emits zero tool calls the loop appends a synthetic nudge to terminate.

Fallback & provenance. The invoker drives the preset's fallback chain — within-provider credential failover first (e.g. an OAuth-Max 429 → the lower-priority API key), then provider advance — emitting OpenTelemetry spans and counters (subagent.llm_call.fallback_advance, …all_providers_exhausted, etc.). Each attempt records a CallAttempt against the snapshot, with served-vs-requested model and cost computed Core-side through the canonical pricing path (the agent never self-reports cost). AgentAgent then back-links the latest snapshot id onto the ledger turn so the dashboard shows served model + cost per turn.

Overlays are part of the resolved preset — a prompt overlay (and routing-chain / tool-policy / sampling / limits profiles) bound to the site. The mode's prompt fragment is appended to the invariant base system prompt inside the invoker (base + "\n\n" + fragment), keeping the base mode-invariant.

Out-of-process parity (pod path). When EphemeralPod=true the programmatic agent runner has no in-proc invoker, so AgentPodMcpTools ports the exact two-step provenance onto the agent surface: agent_record_call_attempt mints the AgentPlanning snapshot + terminal attempt so the dashboards light up identically. The pod tick itself runs agent_tick_claim → (substrate turn) → agent_tick_complete, with the supervisor doing the single passivation write server-side.

3.6 Per-tick budgets

AgentTickBudgetResolver is the single source of truth for a standing agent's per-tick budget envelope. Each axis reads the DB config layer first (the AgentDispatcher:Tick* keys in system_configs), falling back to the AgentDispatcherOptions value (appsettings/env, itself defaulting to the code default). The DB-first read is what makes the caps operator-tunable live — a write via the Settings UI / config_set takes effect on the very next tick, no restart.

The envelope (SOA defaults, all config-overridable):

AxisDefaultConfig key
Max tokens (in+out)24,000AgentDispatcher:TickMaxTokens
Max tool calls12…:TickMaxToolCalls
Max wall-clock300 s…:TickMaxWallClockSeconds
Max cost$2.00…:TickMaxCostUsd
Max turns (hard anti-runaway)12…:TickMaxTurns

AgentAgent.ResolveTickBudget layers three sources, in precedence:

  1. Foundational phase first. While is_building_starting_workset is set and within workset_deadline_utc, every axis is lifted to Unlimited and the per-turn output cap rises to FoundationalMaxOutputTokensPerTurn (default 16,384) — so a brand-new agent can build its one-time starting workset (its baseline knowledge / dossier corpus) unthrottled. The agent ends the phase by calling the agent_workset_complete control terminator; the deadline (default now + 24 h) self-heals it if the terminator is never called. The weekly dollar ceiling stays in force throughout as the runaway-cost backstop. The exemption is universal — it lives on the shared AgentAgent seam, so it applies to every standing type.
  2. Mode override. A mode may carry a BudgetJson override (e.g. Compliance Research widens to 48k tokens / 24 tool-calls / 600 s / $4.00 / 24 turns for whole-chapter dossier authoring) — applied only to that mode, never widening the shared default.
  3. Policy default — the IPlanningAgentPolicy.TickBudget (read live from options each access).

3.7 Resolving the (model, provider) couple

Every agent LLM call is attributed to a call-site in the static LlmSiteRegistry (each LlmCallSite carries a SiteId, display name, a default model, supported backends, default allowed/denied tools, and the catalog axes — category, the track boundary it belongs to, dispatch mode, default routing-chain/tool-policy/overlay ids). The agent sites are Agent:Default, Agent:Planning, Agent:Compliance, PromptRedactor:Default, PromptRedactor:Redact, and Agent:Pod.

Resolution at the invoker:

  1. The agent's policy supplies the preset site key (SOA → Agent:Planning; Compliance → Agent:Compliance).
  2. The provider resolver looks up the site-attached preset (llm_site_attach_preset binding). If the leaf site has none, it walks the parent chain (GetParentSiteId: Agent:Planning → Agent:Default; Agent:Compliance → Agent:Default; PromptRedactor:Redact → PromptRedactor:Default) and uses the family-default binding (LLM.Agent:Default.Preset). A default seeder seeds an initial binding so resolution lands a real preset rather than the legacy [synthetic, anthropic] tail.
  3. The resolved preset yields the concrete (model, provider, backend) couple plus its fallback chain and any attached overlays/profiles. The invoker tries each couple in order (credential failover within a provider before advancing).

So "which model does the SOA think with" is answered by: the preset bound to Agent:Planning (or inherited from Agent:Default), pinned live by an operator via llm_site_attach_preset — no code change, effective on the next tick. The current code default for the agent family is claude-opus-4-8.

The Agent:Pod deny-list on that site is advisory only (a tools/list trim). The authoritative boundary for the remote agent-pod surface is the class gate: AgentPodMcpTools is the only class an agent token can reach (it satisfies AgentPodPolicy and nothing else), and every dangerous tool (decision_respond, *_approve, config_set, llm_* mutations, constitution_*, project_delete, gitlab_* writes, coder_spawn, modification_apply, browser_*) lives in a class gated by a policy the agent token fails. The deny-list enumerates each member explicitly because tool filtering matches exact names (no glob).

3.8 The tool/governance rail

The set of tools a standing agent can ever reach on the warm path is the §3 tool superset (AgentModeCatalog.ToolSuperset), and a mode's toolset is always a strict subset (asserted fail-closed). Every superset member is either read-only (soa_query_telemetry, agent_query_memory, intelligence_knowledge_query, intelligence_code_search, compliance_query_posture, compliance_knowledge_query, orchestrator_research_get, compliance_dossier_list) or artifact-record (compliance_dossier_record / _translate, compliance_knowledge_ingest, orchestrator_research_create) or a human-gated escalation/control verb (soa_propose_decision, soa_send_message, soa_propose_mode_change, agent_workset_complete).

The one self-applying tenant mutation — soa_propose_task (governance-gated mission-create) — is in the superset for the SOA only. The machine-guarded SelfApplyingTools deny-list asserts that no Compliance (or Prompt Engineer) mode toolset ever contains it: those agents can escalate to a human but can never self-apply a change. This "never self-applies" rail is enforced in code (a test asserts the superset contains no self-applying tool beyond the SOA's, and that Compliance modes exclude it), not just documented.

The action handlers themselves keep their governance gates intact regardless of surface (warm in-process or pod): soa_propose_task runs constitution + boundary checks before persisting; soa_propose_decision / soa_send_message / soa_propose_mode_change run constitution on their content. A violation is block-but-feed-back (returned to the model as a structured { blockedReason, violations }, never thrown), so the agent can reconsider within the same episode.

3.9 Lifecycle & control

AgentControlService owns every operator action and every enqueue (the agent never enqueues):

  • Start / Stop / Reset (SOA) and StartCompliance / StopCompliance — flip status Active↔Paused, seeding an AgentInitialized event on start. Reset additionally clears the ledger memory, drains stale events, and re-seeds a fresh opener (refusing if a tick is mid-flight, to protect the single-passivation-writer invariant).
  • Operator chat — persist the message (durable transcript), enqueue an OperatorMessage event, broadcast live.
  • Mode control — the UI set (source=ui) and the approval chokepoint (source=approval), each validating the target against the agent's own catalogue and auditing the change.
  • Escalation round-trip — on an operator resolving a SOA-origin AttentionRequest, apply any approved mode change, then enqueue the OperatorAnsweredEscalation resume (idempotent via the resume-dedup index).
  • Prompt Engineer signals — RaiseRedactorSignalAsync ensure-creates the Paused call-site-bound redactor instance and enqueues an ExternalSignalReceived event (ships dark — durably queued but not ticked until the redactor body is activated).

The agent registry (AgentRegistryService) is the separate in-memory coordination directory: tenant-partitioned, thread-safe, heartbeat-expired (5 min), surfacing registry agents (including the SOA's self-registration) on /agents. It is ephemeral — re-populated after a pod restart — and strictly tenant-isolated.


4. Putting it together: a standing agent's life

  1. Bootstrap. The dispatcher ensures one SOA per tenant. With idle-by-default on, it is created Active, enters the foundational phase (unthrottled per-tick, weekly $ ceiling still enforced), and is seeded AgentInitialized.
  2. Opener. The first tick surveys context and posts a short grounded greeting (soa_send_message), then schedules the next heartbeat ~24 h out and goes quiet.
  3. Foundation. Over the next ticks (unthrottled) it builds its starting workset, then calls agent_workset_complete to revert to the normal per-tick budget (or the deadline self-heals it).
  4. Steady state. It idles on the flat daily heartbeat. Each heartbeat: probe ops health, hash the external inputs, and short-circuit without an LLM call if nothing changed.
  5. React. A real event — an operator message, an answered escalation, a mission state change, an incident flipping the ops bucket — enqueues with visible_at = now, wakes the agent immediately, and forces a re-plan. Within that tick the invoker runs a bounded episode: query memory/telemetry to ground, then take at most one or two governed actions (propose a mission, escalate a decision, send a message), never duplicating something already on the open-items roster, then phase_complete.
  6. Human gate. Any genuine fork becomes an AttentionRequest on the decision queue; the operator's answer comes back as OperatorAnsweredEscalation and the agent resumes.
  7. Modes. A human (UI or approval) can switch the agent's focus — Strategic/Incident/Discovery/… for the SOA, Cadence/GateReview/Research for Compliance — narrowing or widening its tool surface and budget per episode. The agent can only propose a switch.
  8. Safety. If a pod-path agent gets wedged (MaxConsecutiveBlockedTurns consecutive blocks) it halts to Closed, drains its queue (KEDA scales to zero), and surfaces an operator AttentionRequest.