Usage & Cost
What Usage & Cost shows
Every LLM call AI Partner makes is logged with its token counts and estimated cost. The Usage & Cost panel surfaces this as a dashboard so you can see exactly where your API spend is going and set alerts before it gets out of hand.
Go to sidebar → Usage & Cost.
Summary metrics
The top row shows totals for your selected period (daily / weekly / monthly):
| Metric | Description |
|---|---|
| Total calls | Number of LLM API calls made |
| Prompt tokens | Tokens sent to the model (your context + tool results) |
| Completion tokens | Tokens generated by the model (responses) |
| Total tokens | Prompt + completion |
| Total cost | Estimated USD cost across all providers |
| Avg latency | Average response time per call |
Switch between Daily, Weekly, and Monthly views using the period selector at the top.
Per-provider breakdown
A table showing usage and cost broken down by provider and model:
| Provider | Model | Calls | Prompt tokens | Completion tokens | Cost | Avg latency |
|---|---|---|---|---|---|---|
| Anthropic | claude-3-5-haiku-20241022 | 142 | 284,000 | 47,000 | $0.38 | 1.2s |
| Anthropic | claude-sonnet-4-5 | 28 | 112,000 | 18,000 | $0.87 | 2.1s |
| Groq | llama-3.3-70b-versatile | 89 | 178,000 | 31,000 | $0.04 | 0.4s |
| OpenRouter | google/gemini-2.0-flash | 15 | 45,000 | 9,000 | $0.01 | 0.8s |
This breakdown tells you:
- Which models are being used most
- Which models cost the most per call
- Where the smart router is sending which task types
Daily cost tracker
A separate card shows today's spend vs. your configured alert threshold:
Today's cost: $0.42
Alert at: $10.00 [Edit threshold]
████░░░░░░░░░░░░░░ 4.2% of daily budget
When daily cost exceeds the threshold, AI Partner:
- Sends a notification to your Telegram (if configured)
- Shows a warning banner in the UI
- Optionally pauses new goal executions until you acknowledge (configurable)
Setting a cost alert
Click Edit threshold next to the daily cost card, or go to Settings → Usage:
Daily cost alert: $10.00
Action on threshold:
○ Notify only
● Notify + pause new goals
○ Notify + pause all LLM calls
Cost estimates are based on each provider's published per-token pricing. Actual billing may differ slightly due to rounding, minimum charges, or promotional rates. Always check your provider dashboard for authoritative billing figures.
Cost-saving tips
The model router sends classification tasks to Groq (Llama-3.3-70b at $0.05/M tokens) instead of Claude or GPT-4o. Check Settings → Model Routing to confirm this is configured.
Lower the maximum steps per goal in settings. Most goals finish in 5–15 steps, so a cap of 20 removes the runaway case without affecting normal work.
The conversation compressor summarizes history when it exceeds 30 messages — reducing the prompt tokens on every subsequent call. Configurable via CONVERSATION_COMPRESS_THRESHOLD.
If you use Claude heavily, turn on prompt caching in settings. The parts of every request that don't change — the agent's identity, your workspace files — are then charged at a fraction of the usual rate.
Per-user attribution (multi-user)
Every model call is attributed to whoever's chat or goal caused it. On a shared instance, administrators can see a per-user breakdown — calls, tokens, cost and average latency per member — for the day, the week or the month.
Attribution is honest about its own gaps: anything that genuinely can't be traced back to a person is shown as unattributed rather than being spread around or quietly dropped.