How to Calculate Data Product Margins Honestly
A data product can show attractive gross margins right up until a large customer asks for fresh data, a custom workflow, an audit trail, and an answer from a human who understands the system. Then the spreadsheet reveals what the demo concealed: the company is not selling software at software economics. It is selling a partially automated service with unusually expensive infrastructure.
To calculate data product margins correctly, start by treating every recurring cost required to deliver a usable customer outcome as real. Not just cloud spend. Not just API calls. The work, reliability burden, data rights, and customer-specific exceptions count too. If revenue grows while those costs grow almost as fast, calling the product scalable does not make it so.
The margin formula is simple. The inputs are not.
At the account level, gross margin is:
Gross margin = (Recurring revenue - cost to serve) / recurring revenue
The arithmetic is not where teams fail. They fail by defining cost to serve so narrowly that the answer becomes a fundraising artifact rather than a management tool.
For a conventional software product, direct costs often include hosting, third-party infrastructure, payment processing, and customer support. A data product has more moving parts. It may ingest third-party data, pay usage-based model or enrichment fees, run expensive transformations, maintain data quality controls, and absorb significant implementation work before a customer gets value.
The right question is not, “What does it cost to keep the application online?” It is, “What does it cost us to reliably produce the outcome this customer renews for?” Those are frequently different numbers.
A company selling real-time supply-chain intelligence, for example, may have modest dashboard hosting costs but meaningful expenses for source feeds, normalization, exception handling, model inference, and analyst review when confidence drops. Excluding all but the dashboard bill produces a beautiful margin and a useless operating model.
Calculate data product margins by cost behavior
Do not lump every expense into one opaque bucket. Classify costs based on what causes them to rise. That distinction tells you whether scale will improve the business or expose it.
Direct variable costs
These costs increase with customer usage, volume, or required output. They include per-record data licensing, inference and API charges, compute tied to processing volume, storage and egress, identity verification, and transaction-level services.
Variable cost is not inherently bad. A product can have healthy margins with usage costs if pricing captures the value and the unit economics improve at volume. But a usage-based cost base paired with flat pricing is a quiet margin leak. It often looks fine during pilots, when usage is constrained and the team is still flattered that anyone is using the product.
Measure variable cost per unit that maps to customer value: per monitored asset, decision, verified record, workflow, or active seat. “Per customer” is usually too blunt to diagnose the problem.
Direct fixed and semi-fixed costs
Some delivery costs do not move cleanly with usage but are still necessary to run the product: baseline cloud commitments, data platform tooling, security monitoring, data quality infrastructure, and the on-call function supporting production systems.
Allocate these costs carefully. A young company should not pretend these costs do not exist because one finance category calls them overhead. At the same time, it should not allocate them so aggressively that early-stage gross margin becomes meaningless. Use two views: contribution margin before shared delivery infrastructure, and fully loaded gross margin after it. Investors and operators need both.
Contribution margin answers whether an additional customer is economically attractive. Fully loaded gross margin answers whether the business can support the platform it claims to have built.
Human delivery cost
This is where many AI and data companies start lying to themselves without intending to. Implementation engineers, analysts, data operations staff, customer success specialists, and technical support personnel are often booked below gross margin because they sit in a different department. Accounting conventions can be reasonable. Commercial denial is not.
If a customer needs recurring analyst review to trust the output, that review is part of delivery. If every enterprise deployment requires two months of bespoke data mapping, the implementation effort belongs in the account economics, whether it is amortized over the contract term or tracked separately as onboarding cost.
There is a legitimate judgment call here. Early customer work can be product discovery, not a permanent service obligation. The test is whether the work disappears as the product matures. If the same custom work repeats across accounts, it is not learning. It is the product wearing a services costume.
Build an account-level cost-to-serve model
Portfolio averages are useful for board reporting and dangerous for diagnosis. One high-usage customer, one difficult integration, or one generously priced legacy contract can distort the picture. Build the model at the account level first, then aggregate.
For each account, capture contracted recurring revenue, usage volume, direct infrastructure cost, third-party data and model costs, allocated support, ongoing operations, and amortized implementation. If the customer demands a dedicated environment, specialized compliance controls, or a unique data source, assign those costs directly rather than hiding them in a shared pool.
Then calculate three measures:
Account contribution margin shows revenue less directly variable delivery cost.
Account gross margin adds the account’s share of recurring support, operations, and delivery infrastructure.
Customer lifetime contribution considers gross profit over the expected retention period, less acquisition and onboarding expense.
These measures prevent a familiar mistake: celebrating a large enterprise logo whose contract is profitable only if nobody counts the people required to keep it alive. Revenue is not validation when the cost-to-serve profile gets worse after signature.
Price the expensive truth, not the cheap demo
Once costs are visible, pricing becomes a strategy question rather than an act of wishful thinking. The market may not tolerate pricing that covers an unlimited volume of costly outputs. That is not a pricing-page problem. It may be evidence that the underlying product architecture or target customer is wrong.
Use pricing metrics that track value and cost closely enough to preserve margin. If each additional monitored entity produces material data and inference expense, price by entity, volume band, or workflow. If value is created through high-stakes decisions rather than raw usage, a platform fee plus outcome-linked capacity can be more rational.
Avoid charging only by seat when the cost driver is transactions, data refreshes, or model calls. Seat pricing is familiar, which is why it is so often used to conceal an economic mismatch. Familiarity does not pay cloud invoices.
There are cases where you deliberately accept lower margins. A design partner may justify subsidized onboarding if it produces reusable product capabilities, defensible referenceability, or a repeatable vertical playbook. Make that subsidy explicit, time-bound, and approved as an investment. Do not label it normal gross margin and hope later customers behave better.
Stress-test the model before the market does
A credible margin model does not rely on average usage, perfect data quality, or customers behaving exactly as procurement promised. Run downside scenarios.
What happens if the customer doubles volume without upgrading? What happens when a source provider raises rates, a model provider changes pricing, or regulatory requirements force longer data retention? What happens when the system’s confidence drops and human review increases? If a few plausible operating changes erase gross profit, the company has not found product-market fit. It has found a temporary accounting arrangement.
Also test margin by customer segment. Smaller customers may be less demanding but carry disproportionate support load. Large enterprises can generate excellent contract value while imposing security, integration, and governance requirements that consume engineering capacity. Neither segment is automatically better. The answer depends on whether the delivery model can be standardized.
For investors, margin diligence should connect directly to retention. A high-margin data product with weak renewals is still fragile. A lower-margin product with strong retention may be worth improving if the company can identify the costs that decline through automation, better data contracts, or narrower positioning. The question is whether the path to improvement is operationally credible, not whether a deck contains a future-state margin chart.
Treat margins as a product decision
Data product margins are not owned by finance alone. They are shaped by product scope, architecture, data sourcing, automation thresholds, implementation design, and the willingness to say no to customer requests that convert a platform into an agency.
The strongest companies make margin visible to product and commercial teams before a contract is signed. They know which features create recurring delivery burden, which integrations can be repeated, and which customers should pay more or be declined. That discipline is less glamorous than a dazzling demo. It is also how deep technical capability becomes trusted, revenue-generating infrastructure.
If the model only works when costs are excluded, usage stays unusually low, and the customer never asks for help, do not call it a high-margin data product. Call it what it is: an untested assumption waiting for deployment.
Where is your leadership effective, and where is it costing the company?
Most of the problems this blog covers trace back to how the founder runs the company. The Trellis Leadership Diagnostic maps that in 24 behavior-anchored items across six dimensions: about 12 minutes, instant results, free to take self-serve.
Take the Leadership Diagnostic →Exploring a fractional or advisory engagement instead? Book a discovery call →
