Skip to main content

Usage & Cost

What Usage & Cost shows​

Every LLM call AI Partner makes is logged with its token counts and estimated cost. The Usage & Cost panel surfaces this as a dashboard so you can see exactly where your API spend is going and set alerts before it gets out of hand.

Go to sidebar → Usage & Cost.


Summary metrics​

The top row shows totals for your selected period (daily / weekly / monthly):

MetricDescription
Total callsNumber of LLM API calls made
Prompt tokensTokens sent to the model (your context + tool results)
Completion tokensTokens generated by the model (responses)
Total tokensPrompt + completion
Total costEstimated USD cost across all providers
Avg latencyAverage response time per call

Switch between Daily, Weekly, and Monthly views using the period selector at the top.


Per-provider breakdown​

A table showing usage and cost broken down by provider and model:

ProviderModelCallsPrompt tokensCompletion tokensCostAvg latency
Anthropicclaude-3-5-haiku-20241022142284,00047,000$0.381.2s
Anthropicclaude-sonnet-4-528112,00018,000$0.872.1s
Groqllama-3.3-70b-versatile89178,00031,000$0.040.4s
OpenRoutergoogle/gemini-2.0-flash1545,0009,000$0.010.8s

This breakdown tells you:

  • Which models are being used most
  • Which models cost the most per call
  • Where the smart router is sending which task types

Daily cost tracker​

A separate card shows today's spend vs. your configured alert threshold:

Today's cost: $0.42
Alert at: $10.00 [Edit threshold]

████░░░░░░░░░░░░░░ 4.2% of daily budget

When daily cost exceeds the threshold, AI Partner:

  1. Sends a notification to your Telegram (if configured)
  2. Shows a warning banner in the UI
  3. Optionally pauses new goal executions until you acknowledge (configurable)

Setting a cost alert​

Click Edit threshold next to the daily cost card, or go to Settings → Usage:

Daily cost alert: $10.00
Action on threshold:
○ Notify only
● Notify + pause new goals
○ Notify + pause all LLM calls

Cost estimates are based on each provider's published per-token pricing. Actual billing may differ slightly due to rounding, minimum charges, or promotional rates. Always check your provider dashboard for authoritative billing figures.


Cost-saving tips​

Use fast models for classification

The model router sends classification tasks to Groq (Llama-3.3-70b at $0.05/M tokens) instead of Claude or GPT-4o. Check Settings → Model Routing to confirm this is configured.

Limit goal iterations

Lower the maximum steps per goal in settings. Most goals finish in 5–15 steps, so a cap of 20 removes the runaway case without affecting normal work.

Use conversation compression

The conversation compressor summarizes history when it exceeds 30 messages — reducing the prompt tokens on every subsequent call. Configurable via CONVERSATION_COMPRESS_THRESHOLD.

Cache with Anthropic

If you use Claude heavily, turn on prompt caching in settings. The parts of every request that don't change — the agent's identity, your workspace files — are then charged at a fraction of the usual rate.


Per-user attribution (multi-user)​

Every model call is attributed to whoever's chat or goal caused it. On a shared instance, administrators can see a per-user breakdown — calls, tokens, cost and average latency per member — for the day, the week or the month.

Attribution is honest about its own gaps: anything that genuinely can't be traced back to a person is shown as unattributed rather than being spread around or quietly dropped.