aliteq.

Stripe token billing for AI apps: how the markup works, what it costs, and what to use while it's in preview

Stripe will now keep your model price list and add your markup to every token, but it runs on Metronome, it's in public preview, and your app still has to report usage and stop users at zero. Here is the markup math for four models and the fees on top, from Stripe's own docs read on 2 Oct 2026.

LedgerUpdated 1h ago11 min readWeb story
Flat vector illustration of a person pressing a button on a tall black vending machine that pours glowing gold tokens into its tray, a gold price tag with a plus sign on the machine, on a saturated indigo-violet background
Share

If your AI app resells model calls, you have two pricing jobs. You need a markup on every token, and you need to keep that markup right when OpenAI, Anthropic or Google change their prices. Stripe's LLM token billing does the second job for you and makes the first one a setting.

This piece is about that one feature: what it is, how usage gets into it, what a markup actually charges per model, and what it costs. If you're still choosing between credits, metering and a flat plan, or comparing billing platforms, start with our usage-based billing guide for AI apps instead. We read Stripe's and Metronome's docs on 2 October 2026. aliteq hasn't used token billing and isn't paid by Stripe or Metronome. The markup math below is worked arithmetic on stated assumptions, not a forecast of what you'll earn.

What Stripe token billing is

It's a managed price list plus a markup. Stripe tracks what each model provider charges per token, Metronome turns those prices into your rates at the markup you set, and Stripe collects payment on the resulting invoices.

Stripe's docs describe two uses. "Resell LLM tokens" is for platforms and gateways that "want to set their own markup on token costs". "Charge your app's users" is for apps billing their own users for AI usage. Both run through Metronome, which Stripe bought in January 2026.

The useful part is the sync. In Stripe's words: "Stripe syncs token prices for OpenAI, Anthropic, and Google models… When providers update their pricing or release new models we automatically apply new prices to all customers." Metronome's guide adds that a price change updates your rate card "while preserving the markup", and new models join at your default markup.

Pricing shapes are flexible. Stripe lists pure usage, a fixed fee with included usage, "AI credit packs and top-ups" in a custom unit such as an "AI Credit", and hybrids of those. One limit: Metronome says non-USD currencies aren't supported for token billing, "as provider prices are denominated in USD".

Preview status and who can use it

It's a public preview. Stripe's page says: "Billing for LLM tokens is in public preview and available to all users with a Metronome account." Metronome's own guide adds: "We will continue to build on the product and there may be quick changes!"

So there's no waitlist on the current docs. You sign up for Metronome and build there. Neither page gives a launch date or a date for general availability, so we don't either.

Earlier material tells a different story, which is why some guides still say you need an invite. Stripe's npm packages @stripe/ai-sdk and @stripe/token-meter, first published in October 2025, still say they're "only available to organizations participating in the Billing for LLM Tokens Private Preview". That was the first version of the feature, built on Stripe Billing itself. The current docs point new integrations at Metronome.

How usage gets captured: proxy, SDK or events

On the current docs, your app sends the token counts itself. After each model call, you post one event to Metronome's ingest API with the customer, model, provider and token counts. Stripe says to "send usage events directly to Metronome from your own integration."

The event carries model, provider, input_tokens, cached_input_tokens, output_tokens and cached_write_tokens. Metronome notes cache-write tokens are supported "for Anthropic models and OpenAI GPT-5.6+ models only". Use the counts the provider returns in its response, not your own estimate.

Five-step flow of Stripe token billing as documented on 2 Oct 2026. Setup: in Metronome, create a rate card that charges on AI provider pricing, pick models and set a default markup. Always on: Stripe syncs provider token prices, and price changes or new models update rates while keeping the markup. Every call: the app sends one usage event to Metronome's ingest API with customer, model, provider and token counts. Every call: for prepaid credits the app checks a local copy of the balance, because billing does not block calls. Period end: Metronome invoices and pushes the invoice to Stripe, which charges the card and calculates tax. Callout: public preview for anyone with a Metronome account; Checkout needs custom integration.
Stripe owns the price list and the payment. Metronome owns the rating and the invoice. Your app owns the usage events and the balance check. · aliteq research

The "LLM proxy" you may have read about belongs to the first version. Stripe's @stripe/ai-sdk package describes a provider that "routes requests through Stripe's LLM proxy at llm.stripe.com", plus a wrapper that meters any Vercel AI SDK model. @stripe/token-meter wraps native OpenAI, Anthropic and Gemini SDK responses. Stripe's Chipp.ai case study describes the same setup: usage "routes them through Stripe's AI gateway to record tokens consumed".

We found no proxy or SDK setup in the current Stripe docs; the pages we tried for one returned 404. We can't tell you whether the proxy still accepts new users. What we can say is that llm.stripe.com still serves the model catalog that Metronome's setup skill reads, and the documented path for a new app is plain event ingestion.

Which models it prices

Far more than the three providers in Stripe's headline. On 2 Oct 2026 Stripe's rendered price table held 389 rows covering 224 models from 32 providers, including Anthropic, OpenAI, Azure, Google, Amazon Bedrock, Groq, Together AI, DeepInfra, Mistral AI, xAI and DeepSeek.

Stripe separates the publisher, who made the model, from the provider, who serves it and "sets the rate used for billing". So the same model can carry two prices. On the day we read it, OpenAI's gpt-5.6-sol had cached input at $0.40 per million tokens direct from OpenAI and $0.50 through Azure. Bill the provider you actually call.

For the four models in our worked example, Stripe's synced prices matched the providers' own pages exactly:

gpt-6.1-sol

Provider
OpenAI
Input per 1M
$2.00
Output per 1M
$10.00
Cached input per 1M
$0.10

Claude Sonnet 5.5

Provider
Anthropic
Input per 1M
$2.00
Output per 1M
$10.00
Cached input per 1M
$0.20

Claude Haiku 4.5

Provider
Anthropic
Input per 1M
$1.00
Output per 1M
$5.00
Cached input per 1M
$0.10

gpt-6-luna

Provider
OpenAI
Input per 1M
$0.10
Output per 1M
$0.50
Cached input per 1M
$0.01

New to tokens? Our plain-English token explainer shows how text turns into the numbers being billed here.

The markup math, per model

The markup multiplies the list price. Metronome's own formula is the token price times (1 + markup), applied to input and output alike. So a 20% markup on $2 and $10 per million gives $2.40 and $12.

Our worked request is 2,500 input tokens and 500 output tokens, the same size we use in our usage-based billing guide. That's an assumption, not measured traffic. No caching.

Table of token billing markups per model, 2 Oct 2026. gpt-6.1-sol and Claude Sonnet 5.5: list $2 input and $10 output per million tokens, $2.40 and $12 at plus 20 percent, $3 and $15 at plus 50 percent; one request of 2,500 input and 500 output tokens costs $0.0100, sells for $0.0120 or $0.0150; minimum markup to cover fees 4.3 percent. Claude Haiku 4.5: $1 and $5, $1.20 and $6, $1.50 and $7.50; request $0.0050, $0.0060, $0.0075; minimum 4.7 percent. gpt-6-luna: $0.10 and $0.50, $0.12 and $0.60, $0.15 and $0.75; request $0.0005, $0.0006, $0.00075; minimum 12.1 percent because the per-event fee is large next to a cheap request. A 20 percent markup is a 16.7 percent gross margin; 50 percent is 33.3 percent.
Same markup, very different cents. The cheaper the model, the more the fixed per-event fee matters. · aliteq research

gpt-6.1-sol

Cost of one request
$0.0100
Price at +20%
$0.0120
Price at +50%
$0.0150

Claude Sonnet 5.5

Cost of one request
$0.0100
Price at +20%
$0.0120
Price at +50%
$0.0150

Claude Haiku 4.5

Cost of one request
$0.0050
Price at +20%
$0.0060
Price at +50%
$0.0075

gpt-6-luna

Cost of one request
$0.0005
Price at +20%
$0.0006
Price at +50%
$0.00075

How the cost column is built: Sonnet 5.5 is 2,500 × $2 ÷ 1M = $0.005 of input plus 500 × $10 ÷ 1M = $0.005 of output. Haiku 4.5 is half that. gpt-6-luna is a twentieth of Sonnet.

Two things the table hides:

  • Markup is not margin. Margin is markup ÷ (1 + markup). A 20% markup keeps 16.7% of what the user pays. A 50% markup keeps 33.3%. Fees come out of that.
  • The per-event fee is flat. Metronome charges $0.04 per 1,000 events, which is $0.00004 per request if you send one event per call. That's 0.3% of a $0.0120 Sonnet request but 6.7% of a $0.0006 luna request.

Neither Stripe nor Metronome suggests a markup. Metronome's setup skill tells AI agents to "never recommend a percentage or describe one as typical". For reference only: Metronome's fictional example charges 10%, and Stripe's Chipp.ai case study describes a 30% markup. Your number depends on what else the price has to cover.

One cost trap sits outside the markup. Anthropic says its Claude 4.7 and newer models use "approximately 30% more tokens for the same text". Sonnet 5.5 is on that newer tokenizer; Haiku 4.5 isn't. Token billing charges the counts the provider reports, so your users pay for those extra tokens. That's correct billing, but it can surprise a user comparing models by price per million.

The fees on top

Token billing has no separate fee on the pricing pages we read. You pay Metronome's Startup rate, 0.8% of billing volume plus $0.04 per 1,000 ingested events, and card processing on each invoice, 2.9% + 30¢ for a US card on plain Stripe.

One thing we couldn't confirm: whether Stripe also charges Billing's 0.7% on invoices Metronome pushes into Stripe. The pages we read don't say. Ask Stripe before you model a fee stack.

Here is one customer worked through, on our assumptions: 4,000 Sonnet 5.5 requests in a month, $40 of token cost, one invoice paid by card.

Invoice

+20% markup
$48.00
+50% markup
$60.00

Metronome 0.8%

+20% markup
$0.38
+50% markup
$0.48

Events, 4,000 × $0.04 per 1,000

+20% markup
$0.16
+50% markup
$0.16

Card, 2.9% + 30¢

+20% markup
$1.69
+50% markup
$2.04

Token cost

+20% markup
$40.00
+50% markup
$40.00

**Left after fees and tokens**

+20% markup
$5.76 (12.0%)
+50% markup
$17.32 (28.9%)

If Stripe's 0.7% also applied, take another $0.34 or $0.42 off. Hosting, tax, refunds and failed payments aren't in this.

The same customer on gpt-6-luna shows why cheap models are different. Their $2.00 of tokens bills at $2.40 with a 20% markup, and the fixed 30¢ card fee plus 16 cents of events leaves you about 15 cents short. At 50% you keep about 43 cents. A small invoice for cheap-model usage barely covers its own fees. Fold that usage into a subscription or a credit pack so it rides on a bigger charge.

The minimum markup that covers the percentage fees and the event fee, before the 30¢, is about 4.3% on Sonnet 5.5 or gpt-6.1-sol, 4.7% on Haiku 4.5 and 12.1% on gpt-6-luna. Below that, every request loses money.

To see how card, billing and tax fees stack on a single subscription charge, try the calculator:

What each option takes from one sale

Cheapest per sale

Stripe + Billing

$1.02 a sale · you file the tax

Cheapest that files tax for you

Creem

$1.18 a sale · 5.9%

Price of not filing yourself

$16/mo

Creem vs Stripe + Billing, at 100 sales a month

OptionPer saleTake rateFiles tax
Stripe + Billing$1.025.1%You
Stripe + Billing + Tax BasicTax Basic only calculates where you're registered$1.125.6%You
Creem$1.185.9%Provider
Dodo Payments$1.306.5%Provider
Paddle$1.507.5%Provider
Polar (free plan)$1.507.5%Provider
Lemon Squeezy$1.608.0%Provider
Stripe Managed Payments + Billing3.5% is charged on the total including tax$1.728.6%Provider

US sales tax: you only register in a state once your sales there pass its threshold ($100,000 in most states, $500,000 in California, Texas and New York), and not every state taxes software.

Rates from Stripe (US and Denmark pricing pages), Paddle, Lemon Squeezy, Polar, Dodo Payments and Creem, and Lovable's payments docs, all read 27 Sep 2026. Card payments; fee on the amount charged, before disputes, payouts and currency conversion. Stripe Denmark's 1.80 kr fixed fee converted at the ECB rate of 25 Sep 2026. “You keep” is before VAT or sales tax and before income tax. Not tax advice.

If you'd rather not handle sales tax yourself, Stripe's own merchant-of-record option adds 3.5% on top; our Stripe Managed Payments fee breakdown covers it. Its docs require subscriptions created through Checkout or Payment Links, and Metronome invoices don't go through Checkout, so check the combination with Stripe before you plan on it.

What it won't do for you

Token billing counts, prices and invoices. It doesn't stop a user from spending past their balance, and it doesn't plug into Stripe Checkout out of the box.

  • Stopping at zero is your job. Metronome's setup skill says that for a prepaid hard stop you "keep a low-latency projection" of the balance in your app and "gate an LLM request on that local state". The check runs on your server before each model call. Chipp did the same with Stripe's billing alerts: when a customer hits the limit, its system "freezes their LLM access until they purchase more".
  • Checkout needs custom work. Stripe's docs say Metronome invoices use automatic charging or the hosted invoice page, not Checkout, and Checkout with Metronome "requires custom integration". You can keep an existing Stripe subscription and let Metronome send separate usage invoices to the same customer.
  • No Stripe Dashboard or Connect support on Metronome. Stripe's comparison page lists both as unsupported for Metronome, which matters if you run a marketplace.
  • USD only. Provider prices are in dollars, so token billing rate cards are too. You can show customers a custom credit unit with a dollar conversion.

If an AI coding tool is writing your billing, and with vibe coding it often is, ask it where the balance check runs. A check in the browser is a suggestion. Ask how it handles two requests arriving at once with one credit left.

What to use if the preview doesn't fit

Three options are generally available today. Stripe Billing meters if you want the cheapest Stripe-native path, Polar if you want the sales tax handled, and Lago if you'd rather run billing yourself. None of them keeps model prices in sync for you.

Scorecard of four ways to bill AI app users for tokens, 2 Oct 2026. Stripe token billing on Metronome: public preview, 0.8 percent plus 4 cents per 1,000 events, model prices synced by Stripe, does not block usage at zero, not a merchant of record. Stripe Billing meters: generally available, 0.7 percent with 100 million events a month included, you maintain model prices, credits settle at invoice time, not a merchant of record. Polar: usage billing is a new feature, 5 percent plus 50 cents per sale, you send token counts, does not block at zero, merchant of record. Lago open source: self-hosted, free under AGPLv3, you maintain prices, real-time balance is a premium feature, not a merchant of record.
Only the preview syncs model prices. Only Polar does your sales tax. None of them blocks a user at zero for you. · aliteq research
  • Stripe Billing meters. 0.7% of billing volume with up to 100 million meter events a month included. It works with Checkout, Connect and the Dashboard. You meter tokens yourself and price them with your own rate per unit, so a provider price change means you edit your prices. Credits only settle at invoice time. Stripe calls it "fully supported for existing integrations" while recommending Metronome for new ones.
  • Polar. 5% + 50¢ per sale on the free plan, and it's the merchant of record, so it handles sales tax and VAT. It has meters, metered prices with caps, credits, and a helper that wraps an AI SDK model to report prompt and completion tokens. Its docs call usage billing "a new feature" and say Polar "doesn't block usage if the customer exceeds their balance".
  • Lago. The self-hosted edition is free under AGPLv3 and does usage, prepaid and hybrid billing. Lago Cloud is quote-only with a five-figure annual minimum, and a real-time wallet balance is a premium feature.

Our usage-based billing comparison covers these and Orb and Chargebee in more depth, with the margin math for a $20 plan. For the tax side of selling software in the US, see our business payments hub.

Should you use it now

Use it if you resell many models and their prices change faster than you'd update them by hand. Skip it for now if you bill one or two models, rely on Checkout, or need Connect.

Many models, prices that move: token billing on Metronome saves you maintaining a price table. Budget 0.8% plus 4 cents per 1,000 events, and expect preview changes.

Set the markup from your costs, not a rule of thumb. Check it clears the fee floor: about 4.3% on Sonnet 5.5 or gpt-6.1-sol, 12.1% on gpt-6-luna, before the 30-cent card fee.

One or two models: Stripe Billing meters at 0.7% with your own per-token price is simpler, and you only update it when a provider changes prices.

Whatever you choose, gate every model call on the customer's balance on your server, and send the token counts the provider returns.

Want the tax handled or to own the stack: Polar as merchant of record at 5% + 50 cents, or self-hosted Lago for free if you can run a database.

Quick answers

What is Stripe token billing?
A Stripe feature, run on Metronome, that bills your customers for LLM tokens at the model provider's price plus a markup you set. Stripe keeps the model prices in sync, so when a provider changes a price or releases a model, your rates update and keep your markup.
Is Stripe LLM token billing still in preview?
Yes. On 2 Oct 2026 Stripe's docs said billing for LLM tokens is in public preview and available to all users with a Metronome account. Older Stripe npm packages still mention a private preview; the current docs send new integrations to Metronome.
Does Stripe token billing use a proxy?
The first version did: Stripe's AI SDK package routed calls through an LLM proxy at llm.stripe.com. The current docs have your app send one usage event per model call to Metronome's ingest API with the model, provider and token counts.
How much does Stripe token billing cost?
We found no separate fee for it. You pay Metronome's Startup rate of 0.8% of billing volume plus 4 cents per 1,000 events, plus card processing of 2.9% + 30 cents for a US card. Whether Stripe Billing's 0.7% also applies to Metronome invoices isn't stated on the pages we read.
What markup should I charge on LLM tokens?
Stripe and Metronome don't recommend one. Work it out from your costs. A 20% markup is a 16.7% gross margin and a 50% markup is 33.3%, before fees. On cheap models, the per-event fee means you need about 12% just to break even on gpt-6-luna.
Will token billing stop a user when their credits run out?
No. Metronome's own guidance is to keep a copy of the customer's balance in your app and check it before each model call. Billing counts usage and invoices it; blocking is up to your server.

Use this in your own page

Teaching this? Paste the live version into your course, blog or answer. Free, no sign-up; the credit line links back here.

Embed
Cite

Found this useful? Share it

Share
Ledger

Payments & Deals Editor

Ledger

I'm US-based. I don't believe a discount until I've seen the price history, and I don't believe a fee table until I've done the math at three price points. I cover what it really costs to get paid — processors, merchants of record, VAT — and the tech deals that are actually deals.

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading