Digital-asset & online-business income
Data and API Licensing
Selling access to a dataset or a live API under a licence agreement, with customers paying per call, per volume tier, or a flat annual fee for defined rights of use.
Data and API licensing turns a maintained dataset into recurring income by selling defined rights of access rather than selling the data outright. Customers pay per API call, on volume tiers, or as a flat enterprise licence, and the agreement specifies exactly what fields they get, whether they may cache or redistribute, and for how long. The product is really the licence plus the freshness pipeline behind it, which is why the maintenance obligation never ends.
Royalties and licensing Semi-passive
The labels above place this income type before you read a word. The first is the mechanism — how the money actually reaches you, whether by lending it out, owning a slice of something, renting an asset, licensing a right, selling an option or owning a business somebody else runs. The second says whether the income keeps arriving on its own once it is running, or whether it needs work from you to keep coming. Both are descriptions of how the thing is built, not verdicts on it.
How it works
The licensor assembles a dataset that is expensive or awkward for a customer to reproduce: collected observations, aggregated public records, proprietary transaction data, curated reference data, or a cleaned and normalised version of scattered public sources. What is actually being sold is not the data itself but a defined right to use it. Delivery takes several forms, from a live REST or streaming API to scheduled bulk files dropped into cloud storage, a shared table in a data-warehouse marketplace, or a direct database share.
Distribution runs through self-serve API marketplaces such as RapidAPI, cloud data marketplaces like AWS Data Exchange, Snowflake Marketplace and Databricks Marketplace, or direct enterprise contracts negotiated one at a time. The licence agreement is the real product: it specifies permitted fields, seats or applications covered, whether the data may be cached and for how long, whether derivative works can be kept after termination, redistribution rights, attribution requirements and audit rights.
Term structure mirrors enterprise software: an annual subscription with auto-renewal, defined notice periods, and a mechanism for price increases. None of it works without provenance — the licensor must actually hold the right to license the underlying data, which depends on how it was collected, what the source's own terms allowed, and whether personal data is involved.
As customers scale up, service commitments move from courtesy to contract: uptime, latency, update frequency and support response times, backed by credits when missed. The newest large buyer category is AI model training, where labs increasingly license text, image, audio and structured corpora under negotiated agreements rather than scraping them, creating a market specifically for rights-cleared data.
What it pays
Pricing runs per call, tiered by monthly volume, per seat, per application, or as a flat annual enterprise licence — often the same underlying dataset sold under several of these models to different sizes of customer at once. The practical ceiling on price is what it would cost the customer to build or buy equivalent coverage themselves, plus the value of not waiting to build it.
Exclusivity is priced separately from access. A customer buying exclusive rights in a field or territory pays a large multiple of the non-exclusive rate, and the licensor permanently gives up sales to that market. Freshness and coverage completeness are what customers actually renew for, so update frequency drives both the price and the retention rate more than any single feature.
Under accrual accounting, revenue is recognised across the licence term rather than at signature, so a large annual prepayment sits on the books as deferred revenue rather than as an immediate windfall. Self-serve API tiers tend to generate many small, volatile subscriptions, while enterprise licences generate a few large, sticky ones with long sales cycles — most licensing businesses end up running both ends of that barbell. Marketplace channels take a commission on brokered transactions in exchange for discovery and billing infrastructure.
Costs and taxes
The collection and refresh pipeline is the permanent cost center. Scrapers break, source formats change without notice, upstream providers alter their terms, and coverage decays the moment continuous work stops. Infrastructure costs sit alongside it: storage, processing compute, API serving with rate limiting and authentication, plus monitoring and on-call coverage to meet the uptime commitments made in enterprise contracts.
Where inputs are purchased or licensed in, data acquisition costs and any minimum volume commitments owed upstream add another layer. Legal costs run unusually high relative to revenue at this scale — licence drafting, negotiating customer redlines, data-processing agreements, and provenance review on every new source before it enters the pipeline.
Compliance is substantive rather than boilerplate: US state privacy laws and the GDPR restrict the sale and sharing of personal data, several states require data-broker registration, and consumer opt-out or deletion requests have to be honored all the way downstream, including at customers who already received the data.
US tax treatment is ordinary business income reported on Schedule C or an entity return, with self-employment tax on net profit for a sole proprietor. Sales tax on information or data-processing services varies by state, and whether a given feed counts as data, software or a service is a fact-specific question. Cross-border licensing payments raise royalty withholding questions that depend on treaty terms and how the payment is characterised.
Liquidity and time commitment
A data business sells cleanly when contracts are assignable and the pipeline is documented, and sells badly when the whole operation depends on one person's undocumented scrapers. Buyers examine the contracts before anything else: assignment clauses, change-of-control provisions, notice periods, and whether major customers can walk away the moment ownership changes.
Provenance diligence is usually the deal-breaker. A buyer will not acquire a dataset whose right to be licensed cannot be established, and this scrutiny intensifies sharply where personal data or scraped sources are involved.
Day to day, the commitment is genuinely continuous rather than passive: the entire value proposition is freshness, so the pipeline has to run and be watched every day, and enterprise customers require real relationship management for support and renewal rather than an automated notice. Cash flow is smooth where annual enterprise licences bill upfront, and considerably lumpier where usage-based tiers move with customers' own transaction volumes.
How it goes wrong
The core failure is provenance: licensing data the licensor never had the right to license, whether through scraping in breach of a site's terms, misuse of an upstream source's own licence, or collecting personal data without a lawful basis. Legal exposure from the collection method itself follows close behind — Computer Fraud and Abuse Act theories, breach-of-contract claims over terms of service, and copyright or database-rights claims depending on jurisdiction.
Privacy regulation can restrict or prohibit selling entire categories of personal data outright, with data-broker registration and consumer deletion rights that must propagate to every downstream licensee, not just the primary sale. Source dependency is a structural risk: an upstream provider that raises its price, restricts redistribution or cuts off access can end the product overnight, and customer concentration compounds it, since losing one anchor enterprise customer can remove most of the revenue at once.
Customers sometimes reverse-engineer the dataset themselves after a year of seeing its structure, which is precisely why cache limits and derivative-works restrictions exist in the licence in the first place. Freshness decay is a quieter risk — it can go unnoticed until customers discover stale records, and by then the renewal is already lost. Commoditisation closes the loop: a competitor offering comparable coverage free, or a public source publishing the same data openly, removes pricing power regardless of how good the pipeline is.
What to remember
- The product being sold is the licence and the freshness pipeline behind it, not the data itself — maintenance never stops.
- Pricing follows what it would cost the customer to reproduce the coverage, plus a large premium for exclusivity where it is granted.
- Provenance — the actual legal right to license the underlying data — is the precondition for the whole business and the first thing a buyer or regulator checks.
- Customer concentration and single-source dependency are the two most common causes of sudden revenue loss.
- This is semi-passive work: continuous pipeline operation, contract negotiation and enterprise support are required, not optional.
- AI model training has become a major new buyer category, specifically for corpora that can prove clean rights of use.
This page explains how the income type works, which does not change from week to week, so it deliberately carries no rate and no price. The links below go to the pages that hold the current figures for it, each one stamped with the date the data was pulled. Read the mechanism here first: the numbers there are far easier to judge once you know what they are measuring.
See the live numbers: Digital Income, Royalties.
Frequently asked
What is actually being sold in a data licence?
Can I license data I scraped from public websites?
What is the AI-training licensing market?
Why do data licences include audit rights?
How is this different from running a SaaS business?
Written for information only. Nothing here is investment, tax or legal advice, and no page on this site recommends buying or selling anything. Rules and tax treatment change; verify anything that matters with a professional who knows your situation. This explainer was drafted by a language model (claude-sonnet-5) from an editor-approved outline and fact sheet, under the rules set out in our editorial policy, and carries no market figures. Last updated Jul 29, 2026.