Integrations
Every recipe on this page is the same three lines wearing different clothes:
after the AI call, hand track the provider, the response-reported
model, and the response's usage object pasted as-is, plus your
fields. Marginal detects each provider's usage dialect
server-side — you never convert token counts yourself. The only part that
needs care is streaming, where usage arrives at the end of the stream and
each provider hands it over differently.
All snippets assume the quickstart setup and follow your connection choice — TypeScript SDK, Python SDK, or the raw HTTP API. Streaming recipes are SDK code by nature; in HTTP mode the POST is the same, only the usage extraction shown in the SDK tabs is yours to reproduce.
Let your AI assistant do it#
The fastest integration is the one your coding assistant writes. Paste this into Claude Code, Cursor, or any agent with access to your codebase:
Read https://marginalhq.com/llms.txt and instrument every AI/LLM
call site in this codebase with Marginal event tracking, following that
guide exactly. My Marginal API key is in the MARGINAL_API_KEY env var.
Before writing code, list the call sites you found and the fields
(customer / feature) you plan to attach to each, and ask me
to confirm./llms.txt is this entire page plus the wire contract in one
plain-text document, written for machines. The recipes below are the same
material, written for you.
OpenAI#
Chat Completions — the usage object carries the cached/uncached split; paste it whole:
const completion = await openai.chat.completions.create({ model: "gpt-4o", messages });
marginal.track({
provider: "openai",
model: completion.model,
usage: completion.usage,
fields: { customer: "acme", feature: "chat" },
});completion = client.chat.completions.create(model="gpt-4o", messages=messages)
marginal.track(
provider="openai",
model=completion.model,
usage=completion.usage.model_dump(),
fields={"customer": "acme", "feature": "chat"},
)curl -X POST https://api.marginalhq.com/v1/events \
-H "Authorization: Bearer $MARGINAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "events": [{
"provider": "openai", "model": "gpt-4o-2024-08-06",
"usage": { "prompt_tokens": 2006, "completion_tokens": 300,
"prompt_tokens_details": { "cached_tokens": 1920 } },
"fields": { "customer": "acme", "feature": "chat" } }] }'The Responses API is the same recipe — response.model and
response.usage — and its different usage shape is detected automatically.
When streaming, OpenAI sends no usage at all unless you opt in with
stream_options: { include_usage: true }. The final chunk
carries it.
const stream = await openai.chat.completions.create({
model: "gpt-4o",
messages,
stream: true,
stream_options: { include_usage: true }, // without this, no usage arrives
});
let model = "";
let usage;
for await (const chunk of stream) {
if (chunk.usage) ({ model, usage } = chunk); // final chunk only
// …handle chunk.choices[0]?.delta as usual
}
marginal.track({ provider: "openai", model, usage, fields: { customer: "acme" } });stream = client.chat.completions.create(
model="gpt-4o",
messages=messages,
stream=True,
stream_options={"include_usage": True}, # without this, no usage arrives
)
model, usage = "", None
for chunk in stream:
if chunk.usage: # final chunk only
model, usage = chunk.model, chunk.usage.model_dump()
# …handle chunk.choices[0].delta as usual
marginal.track(provider="openai", model=model, usage=usage, fields={"customer": "acme"})Anthropic#
message.usage already includes the cache fields
(cache_read_input_tokens, cache_creation_input_tokens) — paste it whole:
const message = await anthropic.messages.create({
model: "claude-sonnet-5",
max_tokens: 1024,
messages,
});
marginal.track({
provider: "anthropic",
model: message.model,
usage: message.usage,
fields: { customer: "acme", feature: "chat" },
});message = client.messages.create(model="claude-sonnet-5", max_tokens=1024, messages=messages)
marginal.track(
provider="anthropic",
model=message.model,
usage=message.usage.model_dump(),
fields={"customer": "acme", "feature": "chat"},
)curl -X POST https://api.marginalhq.com/v1/events \
-H "Authorization: Bearer $MARGINAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "events": [{
"provider": "anthropic", "model": "claude-sonnet-5",
"usage": { "input_tokens": 810, "output_tokens": 512,
"cache_read_input_tokens": 5100 },
"fields": { "customer": "acme", "feature": "chat" } }] }'Streaming: Anthropic splits usage across stream events (input counts in
message_start, output counts in message_delta). Don't reassemble it —
the SDK's stream helper does that for you:
const stream = anthropic.messages.stream({ model: "claude-sonnet-5", max_tokens: 1024, messages });
for await (const text of stream.textStream) {
// …forward text
}
const message = await stream.finalMessage(); // fully-merged usage
marginal.track({ provider: "anthropic", model: message.model, usage: message.usage, fields: { customer: "acme" } });with client.messages.stream(model="claude-sonnet-5", max_tokens=1024, messages=messages) as stream:
for text in stream.text_stream:
... # forward text
message = stream.get_final_message()
marginal.track(provider="anthropic", model=message.model, usage=message.usage.model_dump(), fields={"customer": "acme"})Vercel AI SDK#
The AI SDK normalizes usage across all its providers, so one small mapper
covers every model you route through it. Its noCacheTokens /
cacheReadTokens / cacheWriteTokens split maps directly onto Marginal's
canonical shape:
import { generateText, streamText, type LanguageModelUsage } from "ai";
// AI SDK (v7) usage → Marginal's canonical usage shape.
function toMarginalUsage(usage: LanguageModelUsage) {
return {
input_tokens: usage.inputTokenDetails.noCacheTokens ?? usage.inputTokens ?? 0,
cache_read_tokens: usage.inputTokenDetails.cacheReadTokens ?? 0,
cache_write_tokens: usage.inputTokenDetails.cacheWriteTokens ?? 0,
output_tokens: usage.outputTokens ?? 0,
};
}
const result = await generateText({ model: openai("gpt-4o"), prompt });
marginal.track({
provider: "openai", // whichever provider you routed to
model: result.response.modelId,
usage: toMarginalUsage(result.usage), // summed across tool-loop steps
fields: { customer: "acme", feature: "chat" },
});Streaming is the same mapper — usage and response resolve once the
stream finishes:
const result = streamText({ model: openai("gpt-4o"), prompt });
for await (const text of result.textStream) {
// …forward text
}
marginal.track({
provider: "openai",
model: (await result.response).modelId,
usage: toMarginalUsage(await result.usage),
fields: { customer: "acme", feature: "chat" },
});On AI SDK 6 and earlier the usage object is flat — map
input_tokens: usage.inputTokens - (usage.cachedInputTokens ?? 0),
cache_read_tokens: usage.cachedInputTokens ?? 0, and
output_tokens: usage.outputTokens instead.
Google Gemini#
The JS SDK's usageMetadata is already the wire shape — paste it whole
(thoughtsTokenCount and cachedContentTokenCount are billed correctly
server-side). The Python SDK renames the same counts to snake_case
attributes — send them under their wire (camelCase) names:
const response = await ai.models.generateContent({ model: "gemini-2.5-flash", contents });
marginal.track({
provider: "gemini",
model: response.modelVersion ?? "gemini-2.5-flash",
usage: response.usageMetadata,
fields: { customer: "acme", feature: "chat" },
});response = client.models.generate_content(model="gemini-2.5-flash", contents=contents)
um = response.usage_metadata
marginal.track(
provider="gemini",
model=response.model_version or "gemini-2.5-flash",
usage={
"promptTokenCount": um.prompt_token_count or 0,
"candidatesTokenCount": um.candidates_token_count or 0,
"thoughtsTokenCount": um.thoughts_token_count or 0,
"cachedContentTokenCount": um.cached_content_token_count or 0,
},
fields={"customer": "acme", "feature": "chat"},
)curl -X POST https://api.marginalhq.com/v1/events \
-H "Authorization: Bearer $MARGINAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "events": [{
"provider": "gemini", "model": "gemini-2.5-flash",
"usage": { "promptTokenCount": 1400, "candidatesTokenCount": 300,
"thoughtsTokenCount": 180, "cachedContentTokenCount": 1024 },
"fields": { "customer": "acme", "feature": "chat" } }] }'When streaming, the last chunk's usageMetadata carries the totals.
AWS Bedrock#
Converse responses don't echo the model back, so send the modelId you
requested; the usage object is pasted as-is:
import { ConverseCommand } from "@aws-sdk/client-bedrock-runtime";
const modelId = "anthropic.claude-sonnet-5";
const response = await client.send(new ConverseCommand({ modelId, messages }));
marginal.track({
provider: "bedrock",
model: modelId,
usage: response.usage, // inputTokens / outputTokens / cache fields
fields: { customer: "acme", feature: "chat" },
});model_id = "anthropic.claude-sonnet-5"
response = client.converse(modelId=model_id, messages=messages)
marginal.track(
provider="bedrock",
model=model_id,
usage=response["usage"], # inputTokens / outputTokens / cache fields
fields={"customer": "acme", "feature": "chat"},
)curl -X POST https://api.marginalhq.com/v1/events \
-H "Authorization: Bearer $MARGINAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "events": [{
"provider": "bedrock", "model": "anthropic.claude-sonnet-5",
"usage": { "inputTokens": 812, "outputTokens": 204 },
"fields": { "customer": "acme", "feature": "chat" } }] }'InvokeModel responses carry the model vendor's native usage instead (an
Anthropic-shaped object for Claude models) — paste that as-is too; the
dialect is detected from its shape, not from the provider string.
Azure OpenAI and OpenAI-compatible providers#
Groq, Mistral, xAI, DeepSeek, Together, Fireworks, Ollama, and Azure OpenAI
all return OpenAI-shaped responses — the OpenAI recipes above work verbatim.
The only change is the provider string, which namespaces the price lookup:
marginal.track({ provider: "groq", model: completion.model, usage: completion.usage, fields: { customer: "acme" } });marginal.track(provider="groq", model=completion.model, usage=completion.usage.model_dump(), fields={"customer": "acme"})Use azure, groq, mistral, deepseek, xai, together_ai, … —
pricing resolves provider/model first, then the bare model string, so
OpenAI-compatible hosts price correctly either way.
Everything else#
For a self-hosted or custom model, send Marginal's canonical usage shape —
input_tokens is the uncached input count — and add a custom price on
the Models page so it prices:
marginal.track({
provider: "self-hosted",
model: "my-fine-tune",
usage: { input_tokens: 812, output_tokens: 204 },
fields: { customer: "acme", feature: "chat" },
});marginal.track(
provider="self-hosted",
model="my-fine-tune",
usage={"input_tokens": 812, "output_tokens": 204},
fields={"customer": "acme", "feature": "chat"},
)curl -X POST https://api.marginalhq.com/v1/events \
-H "Authorization: Bearer $MARGINAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "events": [{
"provider": "self-hosted", "model": "my-fine-tune",
"usage": { "input_tokens": 812, "output_tokens": 204 },
"fields": { "customer": "acme", "feature": "chat" } }] }'And for spend that isn't token-shaped at all — voice minutes, image generation, a cost you already computed — send the dollars directly:
marginal.track({ cost: 0.05, fields: { customer: "acme", feature: "voice" } });marginal.track(cost=0.05, fields={"customer": "acme", "feature": "voice"})curl -X POST https://api.marginalhq.com/v1/events \
-H "Authorization: Bearer $MARGINAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "events": [{ "cost": 0.05, "fields": { "customer": "acme", "feature": "voice" } }] }'Check it landed#
Send one request through each instrumented call site, then open the
project's Ingest log — every API request appears there, tagged
success, warn, or error, with rejected events (including the payload
you sent), stripped field keys, and unpriced models spelled out per
request. Accepted events land in the Explorer with their model, cost,
and fields. The same warnings also reach your code through the SDK's
onError / on_error: ignored-fields means a field key isn't
registered (it was stripped, not silently kept), and
unpriced-models means the model needs a custom price on the Models
page. Nothing fails silently — a success row in the Ingest log plus the
event in the Explorer means the integration is done.