Models, providers & costs
ARDS doesn't run its own models. You bring API keys and ARDS routes work to those providers. This page covers the model lineup, what each is good for, and how billing works.
Provider families
- Anthropic Claude — the recommended default. Per-token billing through your Anthropic account. Four model tiers (Fable / Opus / Sonnet / Haiku) trading intelligence for cost.
- OpenAI & OpenAI-compatible — supported via the same
sk-…API key shape. Use this for OpenAI directly or for any provider that exposes an OpenAI-style API (e.g. Azure OpenAI, OpenRouter). - Synthetic.new — flat-rate subscription, not per-token. Hosts a curated catalogue of strong open-weight models (Qwen, Kimi, GLM, MiniMax and others). Good for high-volume work where token costs would otherwise pile up. Models carry the
hf:prefix in their IDs. - JetBrains — also selectable in the add-provider dialog on the LLM Config page.
- Ollama (local) — supported but only when you're running ARDS-adjacent infrastructure yourself; not relevant for hosted testers.
You may also meet a GitHub Copilot provider entry: a bring-your-own-key kind, backed by a fine-grained GitHub PAT with the Copilot Requests scope.
You can connect more than one provider. ARDS uses fallback chains so that, e.g., a primary Claude call that fails on credits silently retries on Synthetic. Configure chains on the LLM Config page.
Picking a model
For your first session, leave the defaults alone. The default conversation model is Opus; coder tasks follow the platform's presets, which pick a model per kind of work.
When you're comfortable, the rough rules of thumb:
- Planning / decomposition → Sonnet. It's fast and good at structure.
- Implementation tasks → Opus for harder tasks, Sonnet for routine.
- Test scaffolding / boilerplate → Haiku. A fifth of Opus's price per token and entirely capable for repetitive work.
- Cost-controlled high volume → Synthetic.new (e.g. Qwen3 Coder 480B). Flat-rate so you can leave it running.
Costs you'll see
The Costs dashboard tracks every dollar spent. Its two main panels are Costs by Source (which part of the platform spent it) and Costs by Model (where your spend concentrates). Filters above them slice by provider and date range, and switch the cost basis between Billed, API-equivalent (what subscription-covered usage would have cost at per-token rates), and Both. The page also shows subscription-usage windows, rate-limit events, and a table of recent cost records.
Anthropic prices are set per model family (per million tokens, current as of v1):
| Family | Input | Output | Context |
|---|---|---|---|
| Claude Fable | $10 | $50 | 1M tokens |
| Claude Opus | $5 | $25 | 200K tokens |
| Claude Sonnet | $3 | $15 | 200K tokens |
| Claude Haiku | $1 | $5 | 200K tokens |
Newly discovered versions inherit their family's figures — Opus 4.8 and Opus 4.7 both bill at the Opus rates. Fable is the top tier, with a 1-million-token context window for work that needs a very large view. The full current lineup is in the models reference.
Synthetic.new is flat-rate by subscription — no per-token math; pick a plan and pour as much work through as you like.
Honest expectation: a "fix this bug" task costs cents on Sonnet. A "build a new microservice end to end" mission spanning several Opus tasks can run several dollars. The Costs page tells you in real time.
What happens when a key runs out
- Out of credits → you get an error in the mission, and (if a fallback chain is configured) ARDS retries on the next provider in the chain.
- Rate-limited → ARDS backs off and retries. Long-running rate limits surface as an attention request.
- Key revoked or invalid → mission fails. Re-add the key on the LLM Config page.
See Troubleshooting for the specific error shapes.