Stripe will now keep your model price list and add your markup to every token, but it runs on Metronome, it's in public preview, and your app still has to report usage and stop users at zero. Here is the markup math for four models and the fees on top, from Stripe's own docs read on 2 Oct 2026.
If your AI app resells model calls, you have two pricing jobs. You need a markup on every token, and you need to keep that markup right when OpenAI, Anthropic or Google change their prices. Stripe's LLM token billing does the second job for you and makes the first one a setting.
This piece is about that one feature: what it is, how usage gets into it, what a markup actually charges per model, and what it costs. If you're still choosing between credits, metering and a flat plan, or comparing billing platforms, start with our usage-based billing guide for AI apps instead. We read Stripe's and Metronome's docs on 2 October 2026. aliteq hasn't used token billing and isn't paid by Stripe or Metronome. The markup math below is worked arithmetic on stated assumptions, not a forecast of what you'll earn.
What Stripe token billing is
It's a managed price list plus a markup. Stripe tracks what each model provider charges per token, Metronome turns those prices into your rates at the markup you set, and Stripe collects payment on the resulting invoices.
Stripe's docs describe two uses. "Resell LLM tokens" is for platforms and gateways that "want to set their own markup on token costs". "Charge your app's users" is for apps billing their own users for AI usage. Both run through Metronome, which Stripe bought in January 2026.
The useful part is the sync. In Stripe's words: "Stripe syncs token prices for OpenAI, Anthropic, and Google models… When providers update their pricing or release new models we automatically apply new prices to all customers." Metronome's guide adds that a price change updates your rate card "while preserving the markup", and new models join at your default markup.
Pricing shapes are flexible. Stripe lists pure usage, a fixed fee with included usage, "AI credit packs and top-ups" in a custom unit such as an "AI Credit", and hybrids of those. One limit: Metronome says non-USD currencies aren't supported for token billing, "as provider prices are denominated in USD".
Preview status and who can use it
It's a public preview. Stripe's page says: "Billing for LLM tokens is in public preview and available to all users with a Metronome account." Metronome's own guide adds: "We will continue to build on the product and there may be quick changes!"
So there's no waitlist on the current docs. You sign up for Metronome and build there. Neither page gives a launch date or a date for general availability, so we don't either.
Earlier material tells a different story, which is why some guides still say you need an invite. Stripe's npm packages @stripe/ai-sdk and @stripe/token-meter, first published in October 2025, still say they're "only available to organizations participating in the Billing for LLM Tokens Private Preview". That was the first version of the feature, built on Stripe Billing itself. The current docs point new integrations at Metronome.
How usage gets captured: proxy, SDK or events
On the current docs, your app sends the token counts itself. After each model call, you post one event to Metronome's ingest API with the customer, model, provider and token counts. Stripe says to "send usage events directly to Metronome from your own integration."
The event carries model, provider, input_tokens, cached_input_tokens, output_tokens and cached_write_tokens. Metronome notes cache-write tokens are supported "for Anthropic models and OpenAI GPT-5.6+ models only". Use the counts the provider returns in its response, not your own estimate.
Stripe owns the price list and the payment. Metronome owns the rating and the invoice. Your app owns the usage events and the balance check. · aliteq research
The "LLM proxy" you may have read about belongs to the first version. Stripe's @stripe/ai-sdk package describes a provider that "routes requests through Stripe's LLM proxy at llm.stripe.com", plus a wrapper that meters any Vercel AI SDK model. @stripe/token-meter wraps native OpenAI, Anthropic and Gemini SDK responses. Stripe's Chipp.ai case study describes the same setup: usage "routes them through Stripe's AI gateway to record tokens consumed".
We found no proxy or SDK setup in the current Stripe docs; the pages we tried for one returned 404. We can't tell you whether the proxy still accepts new users. What we can say is that llm.stripe.com still serves the model catalog that Metronome's setup skill reads, and the documented path for a new app is plain event ingestion.
Which models it prices
Far more than the three providers in Stripe's headline. On 2 Oct 2026 Stripe's rendered price table held 389 rows covering 224 models from 32 providers, including Anthropic, OpenAI, Azure, Google, Amazon Bedrock, Groq, Together AI, DeepInfra, Mistral AI, xAI and DeepSeek.
Stripe separates the publisher, who made the model, from the provider, who serves it and "sets the rate used for billing". So the same model can carry two prices. On the day we read it, OpenAI's gpt-5.6-sol had cached input at $0.40 per million tokens direct from OpenAI and $0.50 through Azure. Bill the provider you actually call.
For the four models in our worked example, Stripe's synced prices matched the providers' own pages exactly:
The markup multiplies the list price. Metronome's own formula is the token price times (1 + markup), applied to input and output alike. So a 20% markup on $2 and $10 per million gives $2.40 and $12.
Our worked request is 2,500 input tokens and 500 output tokens, the same size we use in our usage-based billing guide. That's an assumption, not measured traffic. No caching.
Same markup, very different cents. The cheaper the model, the more the fixed per-event fee matters. · aliteq research
gpt-6.1-sol
Cost of one request
$0.0100
Price at +20%
$0.0120
Price at +50%
$0.0150
Claude Sonnet 5.5
Cost of one request
$0.0100
Price at +20%
$0.0120
Price at +50%
$0.0150
Claude Haiku 4.5
Cost of one request
$0.0050
Price at +20%
$0.0060
Price at +50%
$0.0075
gpt-6-luna
Cost of one request
$0.0005
Price at +20%
$0.0006
Price at +50%
$0.00075
Cost of one request
Price at +20%
Price at +50%
gpt-6.1-sol
$0.0100
$0.0120
$0.0150
Claude Sonnet 5.5
$0.0100
$0.0120
$0.0150
Claude Haiku 4.5
$0.0050
$0.0060
$0.0075
gpt-6-luna
$0.0005
$0.0006
$0.00075
How the cost column is built: Sonnet 5.5 is 2,500 × $2 ÷ 1M = $0.005 of input plus 500 × $10 ÷ 1M = $0.005 of output. Haiku 4.5 is half that. gpt-6-luna is a twentieth of Sonnet.
Two things the table hides:
Markup is not margin. Margin is markup ÷ (1 + markup). A 20% markup keeps 16.7% of what the user pays. A 50% markup keeps 33.3%. Fees come out of that.
The per-event fee is flat. Metronome charges $0.04 per 1,000 events, which is $0.00004 per request if you send one event per call. That's 0.3% of a $0.0120 Sonnet request but 6.7% of a $0.0006 luna request.
Neither Stripe nor Metronome suggests a markup. Metronome's setup skill tells AI agents to "never recommend a percentage or describe one as typical". For reference only: Metronome's fictional example charges 10%, and Stripe's Chipp.ai case study describes a 30% markup. Your number depends on what else the price has to cover.
One cost trap sits outside the markup. Anthropic says its Claude 4.7 and newer models use "approximately 30% more tokens for the same text". Sonnet 5.5 is on that newer tokenizer; Haiku 4.5 isn't. Token billing charges the counts the provider reports, so your users pay for those extra tokens. That's correct billing, but it can surprise a user comparing models by price per million.
The fees on top
Token billing has no separate fee on the pricing pages we read. You pay Metronome's Startup rate, 0.8% of billing volume plus $0.04 per 1,000 ingested events, and card processing on each invoice, 2.9% + 30¢ for a US card on plain Stripe.
One thing we couldn't confirm: whether Stripe also charges Billing's 0.7% on invoices Metronome pushes into Stripe. The pages we read don't say. Ask Stripe before you model a fee stack.
Here is one customer worked through, on our assumptions: 4,000 Sonnet 5.5 requests in a month, $40 of token cost, one invoice paid by card.
Invoice
+20% markup
$48.00
+50% markup
$60.00
Metronome 0.8%
+20% markup
$0.38
+50% markup
$0.48
Events, 4,000 × $0.04 per 1,000
+20% markup
$0.16
+50% markup
$0.16
Card, 2.9% + 30¢
+20% markup
$1.69
+50% markup
$2.04
Token cost
+20% markup
$40.00
+50% markup
$40.00
**Left after fees and tokens**
+20% markup
$5.76 (12.0%)
+50% markup
$17.32 (28.9%)
+20% markup
+50% markup
Invoice
$48.00
$60.00
Metronome 0.8%
$0.38
$0.48
Events, 4,000 × $0.04 per 1,000
$0.16
$0.16
Card, 2.9% + 30¢
$1.69
$2.04
Token cost
$40.00
$40.00
**Left after fees and tokens**
$5.76 (12.0%)
$17.32 (28.9%)
If Stripe's 0.7% also applied, take another $0.34 or $0.42 off. Hosting, tax, refunds and failed payments aren't in this.
The same customer on gpt-6-luna shows why cheap models are different. Their $2.00 of tokens bills at $2.40 with a 20% markup, and the fixed 30¢ card fee plus 16 cents of events leaves you about 15 cents short. At 50% you keep about 43 cents. A small invoice for cheap-model usage barely covers its own fees. Fold that usage into a subscription or a credit pack so it rides on a bigger charge.
The minimum markup that covers the percentage fees and the event fee, before the 30¢, is about 4.3% on Sonnet 5.5 or gpt-6.1-sol, 4.7% on Haiku 4.5 and 12.1% on gpt-6-luna. Below that, every request loses money.
To see how card, billing and tax fees stack on a single subscription charge, try the calculator:
What each option takes from one sale
Cheapest per sale
Stripe + Billing
$1.02 a sale · you file the tax
Cheapest that files tax for you
Creem
$1.18 a sale · 5.9%
Price of not filing yourself
$16/mo
Creem vs Stripe + Billing, at 100 sales a month
Option
Per sale
Take rate
You keep a month
Files tax
Stripe + Billing
$1.02
5.1%
$1,898
You
Stripe + Billing + Tax BasicTax Basic only calculates where you're registered
$1.12
5.6%
$1,888
You
Creem
$1.18
5.9%
$1,882
Provider
Dodo Payments
$1.30
6.5%
$1,870
Provider
Paddle
$1.50
7.5%
$1,850
Provider
Polar (free plan)
$1.50
7.5%
$1,850
Provider
Lemon Squeezy
$1.60
8.0%
$1,840
Provider
Stripe Managed Payments + Billing3.5% is charged on the total including tax
$1.72
8.6%
$1,828
Provider
US sales tax: you only register in a state once your sales there pass its threshold ($100,000 in most states, $500,000 in California, Texas and New York), and not every state taxes software.
Rates from Stripe (US and Denmark pricing pages), Paddle, Lemon Squeezy, Polar, Dodo Payments and Creem, and Lovable's payments docs, all read 27 Sep 2026. Card payments; fee on the amount charged, before disputes, payouts and currency conversion. Stripe Denmark's 1.80 kr fixed fee converted at the ECB rate of 25 Sep 2026. “You keep” is before VAT or sales tax and before income tax. Not tax advice.
If you'd rather not handle sales tax yourself, Stripe's own merchant-of-record option adds 3.5% on top; our Stripe Managed Payments fee breakdown covers it. Its docs require subscriptions created through Checkout or Payment Links, and Metronome invoices don't go through Checkout, so check the combination with Stripe before you plan on it.
What it won't do for you
Token billing counts, prices and invoices. It doesn't stop a user from spending past their balance, and it doesn't plug into Stripe Checkout out of the box.
Stopping at zero is your job. Metronome's setup skill says that for a prepaid hard stop you "keep a low-latency projection" of the balance in your app and "gate an LLM request on that local state". The check runs on your server before each model call. Chipp did the same with Stripe's billing alerts: when a customer hits the limit, its system "freezes their LLM access until they purchase more".
Checkout needs custom work. Stripe's docs say Metronome invoices use automatic charging or the hosted invoice page, not Checkout, and Checkout with Metronome "requires custom integration". You can keep an existing Stripe subscription and let Metronome send separate usage invoices to the same customer.
No Stripe Dashboard or Connect support on Metronome. Stripe's comparison page lists both as unsupported for Metronome, which matters if you run a marketplace.
USD only. Provider prices are in dollars, so token billing rate cards are too. You can show customers a custom credit unit with a dollar conversion.
If an AI coding tool is writing your billing, and with vibe coding it often is, ask it where the balance check runs. A check in the browser is a suggestion. Ask how it handles two requests arriving at once with one credit left.
What to use if the preview doesn't fit
Three options are generally available today. Stripe Billing meters if you want the cheapest Stripe-native path, Polar if you want the sales tax handled, and Lago if you'd rather run billing yourself. None of them keeps model prices in sync for you.
Only the preview syncs model prices. Only Polar does your sales tax. None of them blocks a user at zero for you. · aliteq research
Stripe Billing meters. 0.7% of billing volume with up to 100 million meter events a month included. It works with Checkout, Connect and the Dashboard. You meter tokens yourself and price them with your own rate per unit, so a provider price change means you edit your prices. Credits only settle at invoice time. Stripe calls it "fully supported for existing integrations" while recommending Metronome for new ones.
Polar. 5% + 50¢ per sale on the free plan, and it's the merchant of record, so it handles sales tax and VAT. It has meters, metered prices with caps, credits, and a helper that wraps an AI SDK model to report prompt and completion tokens. Its docs call usage billing "a new feature" and say Polar "doesn't block usage if the customer exceeds their balance".
Lago. The self-hosted edition is free under AGPLv3 and does usage, prepaid and hybrid billing. Lago Cloud is quote-only with a five-figure annual minimum, and a real-time wallet balance is a premium feature.
Use it if you resell many models and their prices change faster than you'd update them by hand. Skip it for now if you bill one or two models, rely on Checkout, or need Connect.
Many models, prices that move: token billing on Metronome saves you maintaining a price table. Budget 0.8% plus 4 cents per 1,000 events, and expect preview changes.
Set the markup from your costs, not a rule of thumb. Check it clears the fee floor: about 4.3% on Sonnet 5.5 or gpt-6.1-sol, 12.1% on gpt-6-luna, before the 30-cent card fee.
One or two models: Stripe Billing meters at 0.7% with your own per-token price is simpler, and you only update it when a provider changes prices.
Whatever you choose, gate every model call on the customer's balance on your server, and send the token counts the provider returns.
Want the tax handled or to own the stack: Polar as merchant of record at 5% + 50 cents, or self-hosted Lago for free if you can run a database.
Quick answers
What is Stripe token billing?
A Stripe feature, run on Metronome, that bills your customers for LLM tokens at the model provider's price plus a markup you set. Stripe keeps the model prices in sync, so when a provider changes a price or releases a model, your rates update and keep your markup.
Is Stripe LLM token billing still in preview?
Yes. On 2 Oct 2026 Stripe's docs said billing for LLM tokens is in public preview and available to all users with a Metronome account. Older Stripe npm packages still mention a private preview; the current docs send new integrations to Metronome.
Does Stripe token billing use a proxy?
The first version did: Stripe's AI SDK package routed calls through an LLM proxy at llm.stripe.com. The current docs have your app send one usage event per model call to Metronome's ingest API with the model, provider and token counts.
How much does Stripe token billing cost?
We found no separate fee for it. You pay Metronome's Startup rate of 0.8% of billing volume plus 4 cents per 1,000 events, plus card processing of 2.9% + 30 cents for a US card. Whether Stripe Billing's 0.7% also applies to Metronome invoices isn't stated on the pages we read.
What markup should I charge on LLM tokens?
Stripe and Metronome don't recommend one. Work it out from your costs. A 20% markup is a 16.7% gross margin and a 50% markup is 33.3%, before fees. On cheap models, the per-event fee means you need about 12% just to break even on gpt-6-luna.
Will token billing stop a user when their credits run out?
No. Metronome's own guidance is to keep a copy of the customer's balance in your app and check it before each model call. Billing counts usage and invoices it; blocking is up to your server.
Use this in your own page
Teaching this? Paste the live version into your course, blog or answer. Free, no sign-up; the credit line links back here.