Ghost Evaluation

Ghost Evaluation is ARDS's "would a different model have done as well?" machinery. Useful when you're deciding which model to pin to a call site and want evidence, not vibes.

What it does

Ghost re-runs LLM calls against one or more ghost models and compares each ghost's answer with the response the product actually used (the authoritative response). Two things feed it:

  • Shadow evaluations, automatically. Live calls from the platform's own call sites — dialogue, research, critic, decomposition, classification, coder — are re-run in the background against the configured ghost models. You don't start these; they accumulate on their own.
  • On-demand evaluations. You can queue one yourself from the dashboard (see below).

Ghost runs are ghost: they never change product state. The original conversation is untouched; results live on the Ghost Evaluations dashboard, scored and comparable. Each evaluation compares one ghost model side by side against the authoritative response — there is no N-model matrix.

When it's useful

  • "Could a cheaper model hold the planning call site? The leaderboard has weeks of evidence."
  • "Does an open-weight model answer dialogue as well as the default, at a fraction of the cost?"
  • "We changed the default model — did the character of the answers actually move?"

The dashboard

Two main tabs:

  • Stats — filter by source type, read the counters (Total / Completed / Failed / Timeout) and the Model Leaderboard: per ghost model, its completed evaluations, average concordance, judge score, latency, tokens and failures. Rows highlighted green are promotion candidates — models with ≥ 90% concordance over ≥ 50 evaluations.
  • Evaluations — the individual runs, filterable by source, status and model. Click a row to open the detail: the concordance breakdown, judge scores when enabled, a latency and token comparison, and a Side-by-Side Comparison of the ghost response against the authoritative response. A Re-run button queues the same prompt again.

Running one yourself

Click New Evaluation on the dashboard:

  1. On the Manual tab, type a prompt (plus an optional system prompt) — or switch to From History and pick a single past conversation message (with options to include its context and system prompt).
  2. Optionally tick Enable LLM judge scoring.
  3. Click Run Evaluation. The prompt is queued against every configured ghost model.

You don't pick models in the dialog. Which ghost models run — and which model judges — is tenant-wide configuration resolved server-side from LLM Config (the ghost-models call site takes a comma-separated model list).

Scoring

Every completed evaluation gets a concordance score computed without an LLM: length ratio, keyword overlap and structural similarity against the authoritative response.

If judging is enabled, one judge model also compares the two answers on fixed dimensions — semantic similarity, factual accuracy, completeness — plus an overall score and a short written reasoning. The dimensions are fixed; there is no per-run rubric and no multi-judge mode. (The judge's instructions can be overridden tenant-wide on the LLM Config Prompts tab, like other system prompts.) The judge is itself an LLM call with biases of its own — read the reasoning, not just the number.

Cost considerations

Each evaluated prompt runs once per configured ghost model, and the judge adds one more LLM call per evaluation. Ghost spend scales with the size of the ghost-model list — keep it short. There is no projected-cost display; watch the Costs dashboard if you're unsure what Ghost is spending.

What Ghost doesn't do

  • Doesn't replay tool calls. When the original work involved file edits and commands, Ghost can't re-execute the side effects. It re-runs the LLM-level exchange only.
  • Doesn't change anything live. Even when a ghost model outscores the authoritative one, nothing swaps in automatically. Use the evidence to update LLM Config, then run new work.