Tracks resource usage for agent runs.
Usage tracking helps monitor costs and performance across agent executions. You can aggregate usage from multiple agent runs to track total consumption.
Examples
iex> usage = Usage.new()
iex> usage = Usage.inc_requests(usage)
iex> usage = Usage.add_tokens(usage, input: 100, output: 50)
iex> {usage.requests, usage.total_tokens}
{1, 150}Every counter is additive, so runs aggregate with add/2:
iex> a = Usage.add_tokens(Usage.new(), input: 100, output: 50)
iex> b = Usage.add_tokens(Usage.new(), input: 10, output: 5)
iex> Usage.add(a, b).total_tokens
165Cost
cost/2 prices a usage struct against Nous.Usage.Pricing. Cost is
deliberately not a struct field: it is derived from a price table that
changes independently of the run (vendors reprice, operators override rates),
so a number computed once and persisted inside %Usage{} would silently
become wrong while looking authoritative. Counters are facts about the run;
money is a function of those facts and a rate card. Compute it on demand.
Summary
Functions
Add two usage trackers together.
Add token counts from options.
Price a usage struct against Nous.Usage.Pricing, in USD.
Create usage from OpenAI API usage format.
Increment request count by 1.
Increment tool call count.
Create a new empty usage tracker.
Types
@type cost() :: %{ input: float(), output: float(), cache_read: float(), cache_write: float(), total: float() }
A cost breakdown in USD. :total is the sum of the four component costs.
@type t() :: %Nous.Usage{ cache_creation_input_tokens: non_neg_integer(), cache_read_input_tokens: non_neg_integer(), input_tokens: non_neg_integer(), output_tokens: non_neg_integer(), requests: non_neg_integer(), tool_calls: non_neg_integer(), total_tokens: non_neg_integer() }
Functions
Add two usage trackers together.
Useful for aggregating usage across multiple agent runs.
Examples
iex> usage1 = %Usage{requests: 1, total_tokens: 100}
iex> usage2 = %Usage{requests: 2, total_tokens: 200}
iex> total = Usage.add(usage1, usage2)
iex> {total.requests, total.total_tokens}
{3, 300}
Add token counts from options.
Options
:input- Number of input tokens (default: 0):output- Number of output tokens (default: 0)
Examples
iex> usage = Usage.add_tokens(Usage.new(), input: 50, output: 30)
iex> {usage.input_tokens, usage.output_tokens, usage.total_tokens}
{50, 30, 80}Omitted counts default to zero, so partial updates are safe:
iex> usage = Usage.add_tokens(Usage.new(), output: 12)
iex> {usage.input_tokens, usage.total_tokens}
{0, 12}
@spec cost(t(), Nous.Model.t() | String.t()) :: {:ok, cost()} | {:error, :unknown_model}
Price a usage struct against Nous.Usage.Pricing, in USD.
Accepts a %Nous.Model{} or a "provider:model" string (parsed with
Nous.Model.parse/2). Returns {:error, :unknown_model} — never a guess,
never a raise — when the model has no price entry, or when the string is not
a valid "provider:model" spec.
What gets billed at which rate
Providers do not agree on whether the input count they report already
includes cached tokens, and this matters: naively adding input_tokens and
cache_read_input_tokens double-counts every cached token on most providers.
What the code in this repo actually produces:
- Anthropic (
Nous.Messages.Anthropic.parse_usage/1,lib/nous/messages/anthropic.ex:281-292) copies Anthropic'sinput_tokens,cache_creation_input_tokensandcache_read_input_tokensthrough unchanged. Anthropic reports those three as disjoint counts — its docs defineinput_tokensas the tokens "not read from or used to create a cache" and givetotal = cache_read + cache_creation + input_tokens— so each counter is billed at its own rate and nothing is subtracted. - Gemini / Vertex AI (
Nous.Messages.Gemini.parse_usage/1,lib/nous/messages/gemini.ex:606-612) mapsinput_tokensfrompromptTokenCountandcache_read_input_tokensfromcachedContentTokenCount. Google documentspromptTokenCountas "the total effective prompt size meaning this includes the number of tokens in the cached content", so cached tokens are subtracted from the input count before the input rate is applied. - OpenAI and OpenAI-compatible providers
(
Nous.Messages.OpenAI.parse_usage/1,lib/nous/messages/openai.ex:265-269, andfrom_openai/1above) mapinput_tokensfromprompt_tokensand never populate either cache counter, so cached reads are invisible here and are billed at the full input rate.prompt_tokenslikewise includesprompt_tokens_details.cached_tokens, so if you populatecache_read_input_tokensyourself, the same subtraction applies and stays correct.
In short: cache reads are subtracted from input_tokens for every provider
except Anthropic, which is the only one that reports them separately.
Anthropic cache writes are also reported outside input_tokens, and no
other provider populates cache_creation_input_tokens, so cache-write
tokens are always billed on top.
Examples
iex> {:ok, cost} = Usage.cost(%Usage{input_tokens: 1_000}, "ollama:llama3.3")
iex> cost.total
0.0
iex> Usage.cost(%Usage{}, "openai:no-such-model-exists")
{:error, :unknown_model}
Create usage from OpenAI API usage format.
Converts the usage object from OpenAI responses to our format.
Examples
Keys are read as atoms. A JSON-decoded provider payload has string keys, so convert it before calling this, or build the map yourself:
iex> usage = Usage.from_openai(%{prompt_tokens: 100, completion_tokens: 50, total_tokens: 150})
iex> {usage.requests, usage.input_tokens, usage.output_tokens, usage.total_tokens}
{1, 100, 50, 150}Missing keys count as zero:
iex> usage = Usage.from_openai(%{})
iex> {usage.requests, usage.total_tokens}
{1, 0}
Increment request count by 1.
Examples
iex> usage = Usage.new() |> Usage.inc_requests()
iex> usage.requests
1
@spec inc_tool_calls(t(), non_neg_integer()) :: t()
Increment tool call count.
Examples
iex> usage = Usage.new() |> Usage.inc_tool_calls(3)
iex> usage.tool_calls
3
iex> usage = Usage.new() |> Usage.inc_tool_calls()
iex> usage.tool_calls
1
@spec new() :: t()
Create a new empty usage tracker.
Examples
iex> Usage.new()
%Usage{}