When you broker AI for your member installs — they use your provider account instead of holding their own API keys — you need to know what each member costs you, and you need a ceiling that stops a runaway install before it runs up your bill. This page covers both halves: what a brokered call costs and how you price it, and how much AI each plan includes.
Not the same as AI Usage & Costs. That page reports your own spend at your provider, read from their billing API for your whole organization. This system reports what each member spent on your key, priced by your rate card, in real time.
The problem with counting tokens
The obvious meter is tokens, and it's the wrong one. A token of a flagship model costs several times a token of a fast one, an image is billed per picture rather than per token, and an embedding is two orders of magnitude below chat. A single token allowance therefore charges every member a different real amount depending on which model they happened to pick — and your margin moves with their choices instead of staying where you set it.
Pricing every call in money fixes that. The meter and the ceiling end up in the same unit your provider bill arrives in.
The rate card
Dashboard → Memberships → AI Rate Card (admin only) is where you set what each model costs you, per provider.
- Token rates are entered in dollars per million tokens — the same unit every provider publishes. Copy the numbers straight off their pricing page.
- Image rates are entered in dollars per picture, for models billed that way.
- Leave a field blank to use the shipped rate. A blank field means "no opinion", never "free".
- Every shipped model arrives with a rate already filled in, checked against each provider's published price list. Those are the official rates, and updates to them (shipped with CMS releases) reach every install automatically — only fields where you typed something different are stored as yours.
- An official rate update outranks your earlier edit. Your entry stands until the shipped figure for that field moves; when it does, the new official rate takes over and your stale entry is dropped. If your negotiated figure still applies after the change, re-enter it — it can never silently block a rate push.
The page shows which models your fleet actually used in the last 30 days, and what they cost, so you can see at a glance which ones are worth pricing carefully.
Which providers can be brokered
Four text dialects can carry brokered inference — Claude (Anthropic), ChatGPT (OpenAI), Gemini (Google), and DeepSeek — each toggled independently on the Managed AI card (Memberships → Settings) with its own house key. Members then pick their provider per conversation from whatever you broker, and every lane meters through the same rate card. Two shipped-rate notes worth knowing:
- Gemini 3.7 Flash ships at Google's introductory price ($0.75 / $3.75 per million tokens), which Google has announced doubles on January 1, 2027 — the shipped card will be updated then, but if you're reading this after that date and haven't updated, enter the current rate yourself. Gemini 3.1 Pro reprices whole requests past 200K input tokens (2× input, 1.5× output), and the card models that automatically.
- DeepSeek's shipped rates are the peak-hours, cache-miss price (their schedule discounts off-peak hours and repeated-context cache hits). The meter can't see DeepSeek's clock or cache from here, so the shipped figure deliberately never under-bills; if you want to pass the blended discount on, enter your own rate.
Gemini and DeepSeek are text lanes only — images and embeddings keep their own upstreams. (fal.ai is deliberately not offered as a brokered lane.)
Unpriced models are billed high, never free
A model with no rate — a new one a member requested, or one you removed a rate for — is billed at a deliberately high fallback estimate rather than at zero. Zero would be a silent hole: your key gets spent, nothing is measured, and you find out when the provider bills you.
The rate card and the member drill-down both flag any model in use with no rate set, and an unrecognised model that shows up in usage is added to the rate-card form automatically so you can price it. Set its real rate and the estimate stops applying to future calls.
Long prompts can reprice a whole request
Some providers charge more once a request's input passes a threshold — and they reprice the entire request, input and output alike, not just the tokens past the line. OpenAI's current models do this above 272,000 input tokens (2× input, 1.5× output).
The rate card carries that as a multiplier against each model's own base rate, so it stays correct when you override the base rate rather than reverting to a shipped figure. Without it, the very largest calls — the ones that matter most to your margin — would be billed at roughly half what they cost.
If you compare our shipped OpenAI rates against a third-party pricing summary and they look low, this is usually why: several aggregators quote the long-context figure as if it were the standard rate.
Images: per picture or per token
Providers differ. Some bill a flat price per picture; others (including OpenAI's current image model) bill images as tokens, where the generated picture is a few thousand output tokens.
- Set a per-image price and the model is billed per picture, times the number requested.
- Leave per-image blank and set token rates and the model is billed on its tokens.
- A token-billed image whose provider returns no usage figures falls back to the per-picture estimate rather than costing nothing — the one case where token pricing would otherwise give a picture away.
Image cost also swings several-fold with size and quality, so a per-image rate can be a table keyed "{size}/{quality}" with a default fallback. The dashboard form edits a single flat price; a tiered table authored directly in the setting survives form saves while its field stays blank — such rows show "Tiered" in the field — and typing a flat price is a deliberate edit that replaces the table.
Entitlement: who is served at all
Brokered AI is a paid add-on, and being an actively-licensed member is not the same as having bought it. A member reaches the house keys only when one of three things is true:
- A per-member credit is set by hand (
ai_credit_centson the member edit page). Zero is a real value here: entitled, capped at nothing. - They are comped — the owner's own installs, capped by the fleet-wide comped credit.
- An active line item on a plan that grants an AI credit — i.e. a plan with a margin set. The AI add-on plan is exactly this: a $10/mo plan with a $5 margin grants $5 of AI a month.
Everyone else gets no managed_ai block in their signed claims, and the AI proxy refuses them with ai_not_enabled, pointing them at the add-on or at adding their own provider key.
Both gates matter, and the proxy's is the load-bearing one. Withholding the block only stops an install learning about the lanes; the proxy sees nothing but a license key, so an install that ingested a block before the subscription lapsed — or one hand-crafting calls — would keep spending the house account. This mirrors how image entitlement already worked.
The bug this closes. Entitlement and the ceiling are different questions, and for a while only the second was asked.
monthlyCreditMicrosFor()returns null for "the owner has not capped this member", which the allowance check reads as unmetered — correct for a comp with no ceiling set, catastrophic for a software-plan subscriber who bought no AI. Every active member on a plan with no margin could spend the house keys without limit. Locked byManagedAiProxyTestandManagedAiAllowanceTest.
A note on plan shape. The add-on has to sit alongside a plan that unlocks the license: both the proxy and the signed block still require an active license-unlocking line item, so an AI plan bought on its own does not currently work. If the add-on should be sellable to free-tier members, that gate is what has to change — not the entitlement rule above.
Setting what a plan includes: your margin, not their credit
Each plan carries your AI margin — what you keep from that plan's price. The member's monthly AI credit is whatever the price leaves over:
credit = what the member pays − your margin
So a $20 plan with a $10 margin includes $10 of AI. Move the same member to $40 and they get $30, at $60 they get $50 — and your $10 never moves. You set one number per plan, on Dashboard → Memberships → Plans, and the editor shows the resulting credit as you type.
The reason it's stored this way round is that the margin is the thing you actually decide. Store the credit instead and every price change needs a matching credit edit; the day someone forgets, your margin has silently changed and nothing says so.
Details worth knowing:
- The margin is per month; a yearly price is divided down first. A $240/yr plan with a $10 margin grants $10/mo of credit ($20 monthly-equivalent − $10) — not $230/mo off the annual figure. The plan editor's preview shows the monthly equivalent for yearly plans.
- The member's own price wins. Credit comes out of what that member actually pays, so someone on a negotiated $25 against a $40 list plan gets $15 of credit, not $30.
- Several plans stack. A member paying for two plans gets both plans' credit.
- Blank margin = no AI credit from that plan, which is what every plan starts as — and a member with no margin set on any of their plans is not served brokered AI at all (see Entitlement below). It used to mean unmetered, which was the same sentence answering the wrong question.
- A margin at or above the price grants zero credit — not unlimited. "All margin, no AI" is a real configuration and is enforced as one.
- A custom-pricing plan with no price grants nothing, since there's no price for the margin to come out of.
- Comped members don't go through this at all — comping is an exemption with no subscription, so there's no price to take a margin out of. They get a flat ceiling instead; see below.
Comped installs: a ceiling, and the alert comes to you
A comped install is is_exempt with no Stripe subscription, so it holds no line items and the margin arithmetic above has nothing to work from. That used to mean unmetered — reasonable while comps were your own machines, and wrong the moment a comped install can spend your house key without a ceiling. Comping grants free features; it was never meant to grant a free hand with your provider bill.
Two controls:
- Monthly AI credit for comped installs (Memberships → Settings → Managed AI) — one number that every comped install gets. It ships set to $5/month, seeded as a real stored value so it shows filled in on that page rather than hiding as a fallback you can't see. $5 is a working allowance, not a token gesture; change it to whatever suits your fleet. Clearing the field means unmetered — the only way to say "no ceiling", since 0 means "no AI".
- Monthly AI credit on any member's edit page — a per-member override that wins over everything, including the comped default and a paying member's derivation. For the one install the general rule prices wrongly.
Once metered, a comped install behaves like any other: the same per-surface shares, the same total-is-runaway-protection rule, the same meter on its own dashboard. Two things differ:
- The refusal offers different remedies. It has no plan to move up from and no Plans page to buy credit on, so it's told to add its own provider key or ask you — not to go looking for controls that aren't there.
- The 80/95/100% emails come to you, not to the site. They go to the Fleet alert address (Fleet → Settings), falling back to every admin on this server, and they name the install's host — you run many, and the figures alone identify none of them. This is the point of metering a comp: the spend is on your bill and you're the only one who can act on it.
Zero is a real ceiling in both fields ("this install gets no AI"), distinct from blank ("no ceiling"). Clearing a field never means zero — a blank that read as zero would black out every comp at once. The seed only ever fills an absent setting, so a figure you have already chosen — including a deliberate blank — is never overwritten by an update.
What members see when they run out
Staff-facing features — the editing assistant, content generation, image generation — stop with a message naming the actual figures ("$10.00 of $10.00 this month") and pointing at two ways forward: add your own provider API key, or move up a plan.
Visitor-facing features behave differently on purpose. A per-feature share stops any one feature quietly eating the whole credit, but going over that share only reports for anything a visitor sees — your client's site chatbot keeps answering, because a dead chatbot in front of their customer costs more than the tokens it would have spent. Only exhausting the member's whole credit stops everything; that ceiling is runaway protection and outranks the rule.
Members get emailed at 80%, 95% and 100% of their credit, once per threshold per month, so the first sight of the cap is a heads-up rather than a mid-task failure. Two refinements: a member on a zero-credit ("all margin") plan is never emailed — percent-of-nothing thresholds are meaningless — and the 100% email checks their top-up balance first, so someone whose bought credit is still covering usage is told "nothing is paused, $X remaining" rather than the paused wording their situation contradicts. A comped install's alerts are addressed to you instead of to the site (see above).
And they can act on it themselves: the Plans page in their account area (/members/upgrade) shows their credit meter alongside every other plan in their category and what each one's credit is worth, so "move up a plan" is a decision they can make with the figures in front of them rather than a support conversation. See Memberships for how the switch itself works.
Image generation is Super-only until you open it up
Pictures are the most expensive thing a bundled credit can be spent on — a handful of hero images empties a month's allowance — so managed image generation is limited to an install's Super users until it is opened to everyone. Text, translation, alt text and search are unaffected.
This used to work by withholding the lane entirely, and that gated the wrong person: an agency owner is a Super on every install in their fleet, so keeping a client's staff from spending the credit took the owner's own generate buttons away with it. The lane now follows AI entitlement, and the switch decides who may use it.
Two things open it to everyone:
- Buying AI credit. When a member's top-up payment clears, the webhook flips the switch as part of minting the credit. An owner-granted comp does not — that is goodwill, not a signal they are paying for pictures — and a replayed webhook mints nothing and flips nothing.
- By hand. The member's edit page has a Let all users generate AI images switch beside their monthly credit; the members list shows an AI images: all users badge for anyone who has it.
The role check runs on the install, and it has to. The AI proxy authenticates a license key, never a person, so nothing on the mothership can tell a Super's request from an admin's — a check there could only restate what the install already decided. AiImageGenerator::available() is the gate every Generate with AI control and every spend path reads (page builder image fields and row backgrounds, the media library's Generate Image button, the featured-image sparkles, the logo / dark-logo / favicon generators, the AI site generator and the blog agent). The member's monthly allowance remains the ceiling that bounds what any of this can cost.
An unattended run — the scheduled blog agent, a queued site generation — is held to the same bar as an admin, since no Super is present to authorise the spend. The blog agent judges the person who requested the task, so a Super's own scheduled post still gets its image.
One version rule: an install older than the release that added the role gate reads a bare "images enabled" as show the buttons to everyone, so a Super-only member is not offered the lane there at all until it updates. Nothing changes for a member who has already opened it up.
The install learns about a change on its next license check (every six hours, or immediately with php artisan cms:check-membership). Settings → AI says so when the lane is Super-only, so an admin who picks Managed isn't left with missing buttons and no explanation; an install with its own OpenAI / Google / fal key is unaffected by any of this — that key is the site's own, so it is never role-gated.
Top-ups: credit that doesn't wait for next month
Plan credit resets every month. A top-up is a one-off block of credit that persists until it's spent — bought by the member, or granted by you.
Members buy their own from the Plans page in their account area, in fixed denominations. It's a one-off payment that never touches their subscription or their monthly bill. Only metered members are offered (or can complete) a purchase — the drawdown only ever runs against a plan credit, so selling a block to an unmetered member would take money for a balance nothing draws.
You can grant credit from Dashboard → Memberships → Members → (a member): an amount and a reason. Granted credit is money you're giving away, so who granted it and why is recorded in the activity log, and the reason is required.
How it behaves:
- Plan credit is spent first. A top-up is only drawn on once the month's included credit is gone, so a member who stays under their plan never touches what they bought.
- It doesn't expire. The column exists to add an expiry later without a migration, but the shipped policy is that bought credit keeps.
- Holding a balance lifts the per-feature shares. Those shares exist to stop one feature quietly eating the included allowance; someone who has paid for extra has already said "let it run", and a purchase that left a feature blocked anyway would be money taken for nothing.
- It shows up everywhere the meter does — the member's own install shows it under the monthly bar as a reserve, the account area shows it beside their meter, and the member drill-down shows the balance plus the last ten top-ups with how much of each is left.
- Pay-as-you-go is a real configuration. A plan whose margin equals its price includes no monthly credit, and its members' only spending power is the top-ups they buy. Metered-at-zero is not unmetered: their install shows the top-up balance in place of the monthly bar (and a pointer to buy credit when it is empty), so bought credit is never invisible.
Why credit is only ever created by a webhook
A paid checkout sends the customer's browser to a success URL. That URL can be reopened, shared, or never reached at all on a closed tab — so treating the redirect as proof of payment either mints credit twice or loses it entirely. Credit is created only when Stripe's webhook confirms the session, and only when the session actually reports as paid, since a session can complete unpaid.
Stripe retries webhooks, so the checkout session id is stored as a unique key: a duplicate delivery credits nothing.
Delayed payment methods are covered. A method like ACH debit completes the checkout session
unpaid and settles days later; the credit is minted when Stripe delivers
checkout.session.async_payment_succeeded for the now-paid session — same idempotent path, so
however many events arrive, one payment mints one block of credit.
The meter on the member's own install
An install using your brokered AI shows its own credit meter on Settings → API Keys, inside the managed-AI callout: "$2.65 of $10.00 AI credit used this month" with a progress bar that turns amber at 75% and red at 90%.
The figures ride inside the signed license response, so they can't be tampered with by anything sitting between the install and you. They refresh on the 6-hourly license check rather than live, which is why the meter labels itself "as of" a time rather than pretending to be current. Unmetered members see no meter at all — an absent figure, never a 0-of-0 bar.
Installs on a release older than this feature show no meter until they update. The old meter counted tokens, and there is no honest way to express a dollar credit in that unit, so the figure is withheld rather than mistranslated. Their cap is still enforced and the refusal message still explains itself.
Fleet managers get all of this too
Everything here works on any install that issues license keys, not just the mothership — including an Enterprise fleet server brokering AI on its own house key. Same rate card, same derived credits, same cutoff, same meter, and the same 80/95/100% alert emails to its own members.
That last one was a real gap: the alert sweep used to run only on the mothership, while the proxy it warns about runs anywhere license keys are issued. A fleet owner's members were metered, capped, and shown a meter — and never emailed, so they met the cap with no warning.
Where per-member cost shows up
Dashboard → Memberships → Members → (a member) shows their last 30 days of brokered AI:
- A total cost for the window, beside the request and token counts.
- Per-feature badges — chatbot, site assistant, content generation, translation, images, alt text, semantic search — showing tokens and cost, ordered by cost. A feature running an expensive model can sit near the bottom of a token ranking while being the biggest line on your bill, so the ordering follows the money.
- A per-day table with a Cost column. Rows using an unpriced model are marked with an asterisk.
The one number on the dashboard
An AI Credits card sits in the dashboard's stat grid on any install that actually brokers managed AI: the credit the whole fleet has burned so far this calendar month, with the number of installs it came from underneath, linking straight to the rate card. It is there because the alternative was opening that page and adding up a per-model table to answer "is my margin holding this month?".
It is deliberately the burn alone, not burn-against-credited. Working out what you credited means walking every member's active line items through the plan derivation — a table's worth of work, on every dashboard render, to fill a single sub-line. The breakdown belongs on the rate card, one click away.
The card is admin-only, matching the rate-card page's own gate, and invisible on client installs.
How a call gets priced
Pricing happens at the single point where usage is recorded, not at the call sites. A caller can't record a call as free by forgetting to price it — the recorder prices it, and only needs to be told what the token counts can't say: which capability lane the call used, and for images how many pictures at what size and quality.
Consequences worth knowing:
- Dated model ids price as their base model. Providers ship dated snapshot ids that cost the same as the base model, and installs genuinely store them. They fold onto the same rate rather than reading as unpriced.
- Cost is stored in micro-dollars (millionths of a dollar). A single cheap call is worth tens of them, so cents would round most calls to zero.
- Historical rows read as zero cost, not zero spend. Usage recorded before this feature existed has no cost figure; back-filling it would mean pricing old traffic against today's rates and presenting the guess as a bill.
- A rate change is not retroactive. Cost is computed when the call happens, so editing the rate card changes what future calls cost and leaves recorded history alone.
- Only served calls bill. A call the provider refused (a content-policy rejection, a rate
limit, an outage) meters nothing — a failed image request in particular must not fall back to
the per-picture estimate for pictures that were never produced, and the billable picture count
is capped at the provider's own maximum however large an
nthe request claims.
Works for any fleet manager
None of this is specific to the mothership. Any install issuing license keys — including an Enterprise fleet server running its own license channel — gets its own rate card against its own house key, and prices its own members. The rate card is per-install configuration, not a shipped constant.
Setup
- Turn on Managed AI at Dashboard → Settings → Memberships and add your house provider key.
- Open Dashboard → Memberships → AI Rate Card and check the pre-filled rates against what your own account is actually charged. Override anything that differs.
- Watch the member drill-down for models flagged as unpriced, and price them as they appear.
Rates were last checked against the providers' published price lists in August 2026. Providers change pricing without notice and nothing here re-checks it for you, so it's worth a look whenever your provider bill moves unexpectedly.