Share
This Pricing Model Spotlight series is a recurring feature where we tear down the packaging mechanics, infrastructure choices, and unit economics of companies on the leading edge of monetization. Each post deconstructs how modern platforms are navigating the transition to usage-based, hybrid, and agentic billing frameworks.
When we launched our Pricing Model Index and analyzed over 50 leading AI companies, we saw a central tension started taking shape:
How do you give users the freedom to experiment without collapsing gross margins under the weight of raw compute costs?
On a relative scale, LLM tokens are cheap in text-based AI, but in generative AI video and imagery, just one click can cost orders of magnitude more. If a creative runs a complex prompt sequence or task that doesn’t meet their artistic intent, the vendor still pays 100% of that rendering cost, whether or not they charge for it. To survive, generative platforms have to build a billing architecture as sophisticated as their AI model architecture, one that acts as a buffer between hardware bills and user experience. Today, we’re looking at a consumer-to-enterprise generative media company that has mastered this balance: Leonardo.ai.
With a sophisticated multi-tiered token architecture that’s complete with automated token rollover banks, dynamic task-weighting, and strict structural boundaries between first-party and third-party models, Leonardo.ai has built a monetization model for the modern, multi-model creative suite.
The core innovation: Multi-currency token banking
Unlike early SaaS players that tried to cram generative AI into standard, rigid seat-based tiers, Leonardo has built its entire monetization framework around an abstracted currency. Tokens act as a buffer between fluctuating infrastructure costs (raw GPU compute) and user-facing value. To encourage high product stickiness while establishing reliable margin guardrails, they engineered a triple-currency hybrid token system:
- Fast Tokens
These are the primary currency included in every monthly subscription tier (ranging from 8,500/mo on the Essential tier to 60,000/mo on Ultimate). Fast tokens grant priority access to premium, high-speed GPU queues. - Rollover Tokens
A common friction point in pure consumption models is the use-it-or-lose-it anxiety that can actually result in churn. To put users at ease, Leonardo offers an automated rollover bank. Unused Fast Tokens drop into a secondary bank that’s capped at 3x the monthly tier capacity. This generous rollover setup ensures that users can keep pulling from the unused excess into the future, without worrying about losing out during lower-activity months. - Top-Up Tokens
Top-ups can be purchased whenever a user burns through their monthly allocations. These don’t expire but do require that the user has an active subscription. This setup safeguards predictable recurring revenue (ARR) for the company, while still offering good value to the user.
By decoupling usage into distinct prioritizations and structural banks, Leonardo offers the financial flexibility creative professionals need while retaining a predictable subscription floor.
Dynamic token pricing: Metering the GPU load, not the asset
The fatal mistake in early generative AI packaging was flat-rate asset pricing (think "$0.05 per image"). In reality, generating a 512x512 thumbnail using an older, lighter model takes a fraction of the compute required to output a 4K asset using advanced, multi-model prompt styling, background removal, and motion upscaling.
Leonardo addresses this with variable token metering. A token doesn't correspond to an image; it maps directly to computational intensity. This ensures that as users demand heavier work from Leonardo's infrastructure, the billing meter scales alongside underlying API or data center costs. Leonardo protects its gross margins on hyper-complex renderings while keeping light, exploratory prompt generation highly accessible.
The "relaxed mode" moat: Segmenting native vs. third-party models
On its highest tiers ($30/mo Premium and $60/mo Ultimate), Leonardo offers a feature that seems financially terrifying on paper: unlimited Relaxed Generation. Even if a user drains their priority Fast Token balance, they still aren’t blocked from creating. Instead, their tasks move to a lower-priority queue that processes assets during off-peak GPU capacity. Interesting enough on its own, but look closer at the structural guardrails. Leonardo has drawn a hard monetization line between first-party models and third-party models:
It seems straightforward as it’s laid out here, but this is a masterclass in AI packaging. By keeping third-party capabilities walled behind hard token consumption, Leonardo de-risks themselves from paying third-party API bills for non-paying or over-consuming users. At the same time, they’re using their proprietary, first-party models as a retention tool, ensuring users never feel cut off.
Scaling into the enterprise: From solo banks to shared pools
As Leonardo moves from individual hobbyists into studio and enterprise environments with their Team and API tiers, the monetization architecture adapts seamlessly from individual token banks to Shared Token capacity.
On the Starter and Growth team tiers, companies pay a per-seat monthly fee to buy into a massive, macro pool of tokens (up to 540,000 bank capacity). This shift transforms tokens from a personal usage counter into an enterprise resource planning tool. Studios can allocate token budgets across various projects, while Leonardo captures predictable platform premiums on features like Private Team Generations and collaborative Realtime Canvases.
Key takeaways
- Abstract costs with a tailored currency.
If your underlying infrastructure or API costs vary wildly depending on user inputs, it’s best to avoid using flat-rate asset pricing. Build an abstracted credit or token layer that dynamically maps to computational weight so it can flex with usage and scale. - Mitigate consumption anxiety with rollover banks.
Usage-based billing should feel fair instead of punitive. Giving users a structured safety net to carry over unused value minimizes end-of-month churn risks while supporting a steady billing cadence. - Never underwrite a third-party API fee on an unlimited tier.
If you don't control the infrastructure or the margins of a partner model, keep it strictly behind a metered, pay-as-you-go firewall to protect your own costs. From there, use your proprietary features or data layers to anchor any "unlimited" value propositions.
Leonardo.ai proves that successful AI monetization is about engineering a flexible financial system that moves in response with your infrastructure costs. By abstracting resource consumption into a tiered, multi-currency token economy, they’ve managed to encourage creative exploration while de-risking their gross margins. As the line between application SaaS and infrastructure consumption continues to blur, this hybrid, margin-protected setup is one way the next generation of software leaders can build sustainable, multi-model platforms.
Explore more Pricing Model Spotlights with a look inside Hubspot's move to outcome-based AI credits and find out why Clay cut data prices by +50% to win the long game.











.png)

