nightlydata

What Is Short-Term Rental Data? Metrics, Sources, Aggregation

By Daniel Carrow (pen name) analyse
What Is Short-Term Rental Data? Metrics, Sources, Aggregation - cover image

Short-term rental data is the set of numbers a lodging operator uses to decide what to charge, where to buy or sign, and whether a unit is even legal to rent. At its core it answers four operator questions: how full will this listing be, what will it earn per night, how far ahead do bookings land, and what do the local rules allow. If you run 5 to 50 doors as a business, this data is the line between pricing on instinct and pricing on evidence.

This is a definitional page. It explains the metrics, the source categories, and the reading errors that cost operators money. It does not publish benchmark numbers: for live figures by city, use Nightlydata’s regulatory tracker and quarterly reports.

TL;DR: Short-term rental (STR) data breaks into three families: performance metrics (occupancy, ADR, RevPAR, booking pace, minimum stay), supply-and-demand signals (listing counts, seasonality), and regulatory status (caps, licenses, primary-residence rules). RevPAR = ADR x occupancy, so those metrics move together. Sources fall into four buckets: OTA and platform data (your own bookings), open data such as Inside Airbnb, paid market-data tools, and official regulatory sources. Aggregation into a market figure happens by scrape-and-model, by panel supply, or by official register, and it breaks on normalization: duplicates across platforms, active versus total listings in the denominator, unit mix, and vendor-drawn market boundaries. Aggregating your own portfolio across channels is the harder version of the same problem, and it turns on four choices: unit of analysis, time basis, revenue definition, and cancellation handling. Two rules keep you out of trouble: know whether a number is measured or estimated, and never treat one market’s benchmark as another’s.

The core performance metrics, defined

These come from hotel benchmarking and translate cleanly to STR if you swap “rooms” for “available listing-nights.”

Occupancy rate. The percentage of available nights that were booked over a period. CoStar/STR defines occupancy as the percentage of rooms occupied in a property, segment, or area for a given period, one of the industry’s foundational metrics alongside ADR and RevPAR (source: CoStar/STR, retrieved July 2026). For a single listing it is booked nights divided by available nights.

ADR (average daily rate). Average revenue per occupied night: room revenue divided by nights sold. It reflects the price charged per booked night, not per available night (source: CoStar/STR, retrieved July 2026). A high ADR paired with low occupancy can earn less than a moderate ADR that stays full.

RevPAR (revenue per available room, or per available night for STR). Total guest revenue divided by total available nights. CoStar/STR defines RevPAR as a function of occupancy rate and ADR (source: CoStar/STR, retrieved July 2026). The identity follows directly: revenue equals ADR times nights sold, and dividing by available nights gives RevPAR = ADR x occupancy. RevPAR is the one number that catches the tradeoff ADR and occupancy hide on their own.

Booking pace and lead time. Booking pace is how fast reservations accumulate for a future date; lead time is the gap between when a guest books and when they check in. These are standard revenue-management concepts rather than a single externally standardized statistic, so treat them as directional. They tell you whether to hold rate or discount as a date approaches.

Minimum stay. The shortest reservation a listing accepts. It is a pricing lever and a compliance lever at once: minimum-night settings shape yield, and in several cities the regulatory cap is itself expressed in continuous nights (see below).

Regulatory status is a data field, not a footnote

For an STR operator the legal status of an address is as much a number as its occupancy. It sets the ceiling on nights you can sell and whether a platform will even process the booking. Treat it as structured data with a source and a date. It reads very differently by market. All figures current as of July 2026:

  • New York City: Local Law 18 requires hosts to register with the Mayor’s Office of Special Enforcement and bars platforms from processing transactions for unregistered short-term rentals. Stays of 30 consecutive days or more are exempt (source: NYC OSE, retrieved July 2026).
  • Paris: a primary residence can be let as a furnished tourist rental for up to 90 days per year, a registration number must appear on every listing, and failure to register carries a fine of up to 5,000 euros (source: Ville de Paris, retrieved July 2026).
  • London: under the Deregulation Act 2015, a property can be used as short-term accommodation for a maximum of 90 nights per calendar year without planning permission (source: Greater London Authority, retrieved July 2026).
  • Seattle: operators need a business license tax certificate plus a short-term rental regulatory license (75 dollars per unit, valid one year) and may run up to two units they own, and if they run two, one must be their primary residence (source: City of Seattle, retrieved July 2026). Detail in the Seattle regulations guide.
  • Washington, DC: a vacation rental (entire home, host absent) is capped at 90 nights per calendar year, while a host-present short-term rental has no annual cap but limits each stay to 30 or fewer continuous nights, and both require the property to be the owner’s primary residence (source: DC DLCP, retrieved July 2026). Detail in the furnished DC rentals guide.

Same metric, “nights you may legally sell,” five different rules. That is why regulatory status belongs in the dataset, dated and sourced, not in a blog post you read once. Rules change; verify each figure with the cited authority before acting.

The four source categories

1. OTA and platform data (your own). Your booking records from the major platforms are the most accurate data you own: real, measured, first-party. Occupancy, ADR, lead time, cancellation rate, and channel mix all come out of your own reservations. The limit is scope: you see your own listings, not the market around them. Spreading those bookings across channels is its own lever, covered in OTA distribution beyond Airbnb.

2. Open data (Inside Airbnb and similar). Inside Airbnb is an independent, non-commercial project that compiles public information from the Airbnb site (each listing’s 365-day availability calendar and its reviews), then cleanses and aggregates it. Its site states it “is not associated with or endorsed by Airbnb” (source: Inside Airbnb, retrieved July 2026). It is released under a CC BY 4.0 license, so you can reuse it with attribution (source: Inside Airbnb, retrieved July 2026). The caveat that matters: its occupancy figures are estimates, not measured bookings. The San Francisco Model converts reviews to bookings at a 50% review rate, applies an average length of stay, and caps estimated occupancy at 70% (source: Inside Airbnb, retrieved July 2026). Useful for market shape, wrong tool for your own precise occupancy.

3. Paid market-data tools. A category of subscription platforms estimates neighborhood-level occupancy, ADR, and seasonality from scraped and modeled listing data. They fill the blind spot your own bookings cannot: what the market around you is doing. Depth, freshness, and geographic coverage vary, and the outputs are modeled estimates, not measured bookings. How to weigh these tools against free alternatives is covered in STR market-research tools compared.

4. Official regulatory data. Municipal and national government pages are the primary source for legal status: caps, license requirements, primary-residence tests, tax rules. Always cite the official page with a retrieval date, because these move (the Paris guidance page was last updated in April 2026). Secondary summaries go stale fast.

How short-term rental data gets aggregated

A single listing’s calendar is a fact. “Occupancy in this market is 61%” is a construction. Aggregation is the step in between, and it is where most of the error enters. Four mechanics produce almost every market figure you will see.

1. Scrape and model. Public listing pages, their 365-day availability calendars, and their reviews are collected, cleansed, and aggregated. Bookings are never observed directly, so they are inferred from review counts. Inside Airbnb documents its San Francisco Model precisely: reviews are converted to bookings at a 50% review rate (chosen as a middle ground between a 72% rate it judges unreliable and a 30.5% rate derived from New York Attorney General data), multiplied by an average length of stay (5.5 nights where Airbnb reported that figure for San Francisco, otherwise 3 nights, or the listing’s minimum stay if that is higher), with the resulting occupancy capped at 70% (source: Inside Airbnb, retrieved July 2026). Every assumption in that chain is a lever on the output.

2. Panel and first-party supply. The opposite construction. Instead of inferring bookings from public traces, the aggregator receives measured reservations from participants, usually through property-management-system integrations. The numbers are real bookings rather than estimates, which removes the modeling error. It does not remove bias: coverage now depends on who joined the panel rather than on what is publicly visible. Different bias, not a smaller one.

3. Official registration datasets. Where a city runs a registration regime, the register itself is aggregated data, and it is the one route where nothing is modeled. New York’s Office of Special Enforcement publishes its registration dataset with Registration Number, STRR Status (“Registered”, “Expired”, or “Terminated”), Street Address, Unit Number, Borough, Zip Code, Building Identification Number, and the associated booking service and listing URL. It is published under Local Law 18 and, as such, excludes the applicant’s name; the current file is dated January 7, 2026 (source: NYC OSE, retrieved July 2026). This tells you legal supply at address level, which no scrape can.

4. Normalization, where aggregation actually breaks. The arithmetic is easy; the bookkeeping is not.

  • Cross-platform duplicates. The same unit listed on Airbnb, Vrbo, and Booking.com is one property and up to three rows. Undeduplicated, supply is overstated and per-listing revenue is understated.
  • Active versus total listings. A listing that has taken no booking in a year still sits in the raw count. Whether it is in the denominator changes occupancy materially. Inside Airbnb’s own guidance is to combine the “Only highly available” and “Only recent and frequently booked” filters to isolate listings reviewed in the last six months and booked regularly (source: Inside Airbnb, retrieved July 2026).
  • Unit mix. Entire homes and private rooms earn differently. A market average that blends them describes no actual property.
  • Market boundaries. A “market” is a polygon someone drew. It is rarely a census boundary, and two vendors rarely draw the same one, which is enough on its own to make their figures disagree.

Before acting on any aggregate, make it answer four questions: is this measured or modeled, what is in the denominator, how were duplicates handled, and what date is the snapshot. A figure that cannot answer all four is a directional hint, not an input to a pro forma.

One caution that is easy to underrate: a single platform is not the market. A study of six online rental listing platforms across five US metropolitan areas found that they “target different audiences and offer distinct information on units within those market segments, resulting in markedly different estimates of local rental costs and unit size distribution depending on the platform,” concluding that one platform “may not sufficiently represent the current state of the rental stock” without adjustment (source: Costa et al., Cityscape 23(2), 2021, HUD PD&R, retrieved August 2026). That study looked at long-term rental platforms, not Airbnb and Vrbo, so treat it as a structural warning rather than a measurement of short-term rental supply. The mechanism transfers directly: an aggregate built on one platform inherits that platform’s audience, not the market’s shape.

Aggregating your own portfolio, which is the harder problem

Everything above is about reading someone else’s aggregate. The aggregation most operators actually get wrong is their own, when they roll five channels and a PMS into a single portfolio number. Nobody checks this one, because it feels like arithmetic rather than methodology. Four decisions determine whether the result means anything.

Pick a unit of analysis and hold it. A listing is not a unit and a unit is not a property. One apartment can be three listings across three OTAs; one building can be eight units under one address; one listing can be a whole villa that you also sell room by room in the shoulder season. Occupancy per listing, per unit, and per property are three different numbers, and portfolios routinely mix them without noticing. Decide which one your dashboard means, write it down, and make every channel report into it.

Fix the time basis before you sum anything. A reservation has at least three dates that matter: when it was booked, when the guest stayed, and when the money landed. Revenue aggregated on booking date tells you how sales are pacing. On stay date, it tells you what the asset produced. On payout date, it tells you what the bank saw, which is neither. All three are legitimate; mixing them inside one total is not, and it is the single most common way a portfolio revenue figure ends up unreconcilable with the accounts. Channels do not default to the same basis, so this is a choice you have to impose rather than inherit.

Decide what “revenue” includes, once. Gross booking value, net of channel commission, with or without the cleaning fee, with or without occupancy taxes you merely collect and remit. Each channel reports a different default, and a portfolio total that silently blends them will overstate the properties on whichever channel reports gross. Pick the definition, then check each channel’s export against it rather than assuming the column header means what it says.

Handle cancellations and modifications explicitly. A booking that was made, modified twice, and then cancelled can appear as one row, three rows, or none, depending on the export. If your channel manager restates history in place, a total you computed last month will not reproduce this month, and you will not know why.

The test for your own aggregate is the same one you apply to a vendor’s: could someone else, given your raw exports and your rules, land on the identical number. If the answer depends on which analyst ran it, you have a report rather than a metric.

Market occupancy is not your occupancy

“What is occupancy in this market” has three answers, and they disagree by construction rather than by error.

Your own platform data. Measured, first-party, exact. Covers only the units you operate.

Open data estimates. Modeled from public traces, and hard-capped. Because the San Francisco Model caps estimated occupancy at 70%, a market where competent operators genuinely run at 80% cannot read above 70 in that dataset (source: Inside Airbnb, retrieved July 2026). The ceiling is a deliberate conservatism, not a measurement, and it compresses exactly the top of the distribution you would most want to see.

Paid market-data tools. Modeled from scraped listings, with the vendor’s tracked set as the denominator and the vendor’s polygon as the market.

Two rules follow. Use market occupancy comparatively, to rank one market against another under a consistent method, not absolutely, to forecast what your unit will earn. And never mix the registers: setting your measured 74% against a market’s modeled 61% compares two different quantities and will make an ordinary market look like an opportunity. If you need an absolute occupancy number for a pro forma, the defensible one comes from your own bookings in a comparable unit, or from a stated, sourced assumption you are willing to defend, and it belongs in a break-even calculation rather than a revenue projection. Our break-even occupancy guide works that direction: the occupancy you need, rather than the occupancy you hope for.

How to read the numbers without fooling yourself

Measured versus estimated. The most expensive mistake is treating an estimate as a measurement. Your platform bookings are measured. Inside Airbnb occupancy and market-tool figures are modeled. Use estimates to size a market, use your own data to run it.

RevPAR over ADR alone. ADR flatters a listing that sits empty behind a high headline price. Because RevPAR = ADR x occupancy, RevPAR is the honest scoreboard when you compare units, dates, or a price change.

Match the window and the geography. A metric is only comparable within the same period and the same market. Peak-season occupancy in one city says nothing about shoulder-season demand in another. Do not carry a benchmark across a border, a season, or a property type.

Anchor to the source and the date. Every number worth acting on has a source and an as-of date. A 90-night cap that was 120 nights last year is a different business. This is why Nightlydata dates every figure and links the official source.

Where Nightlydata fits

Nightlydata does not sell a data feed to replace your own bookings. It maintains the reference layer around them: a per-city tracker of regulatory status with official sources and dates (for example Paris, London, and New York), quarterly market reports with disclosed methodology, and a directory of the tools and services that produce or consume this data. Use your platform data for what you operate, and these assets for the market and rules around it.

Key facts

  • STR data covers three families: performance metrics, supply-and-demand signals, and regulatory status.
  • The three core metrics are occupancy (booked nights divided by available nights), ADR (revenue divided by booked nights), and RevPAR (revenue divided by available nights).
  • RevPAR = ADR x occupancy, so it is the single metric that captures the price-versus-fill tradeoff (source: CoStar/STR, retrieved July 2026).
  • Booking pace and lead time are standard revenue-management concepts, directional rather than externally standardized.
  • Inside Airbnb occupancy is an estimate (reviews-to-bookings model, capped at 70%), not measured bookings, and is CC BY 4.0 licensed (source: Inside Airbnb, retrieved July 2026).
  • Regulatory caps differ sharply by market: NYC exempts stays of 30-plus consecutive days, Paris and London cap primary residences at 90 days per year, DC caps host-absent rentals at 90 nights, Seattle requires a per-unit license (source: official municipal pages, retrieved July 2026).
  • Market figures are aggregated three ways: scrape-and-model (bookings inferred from reviews), panel supply (measured reservations from participants), and official registers such as New York’s Local Law 18 dataset, which models nothing (source: NYC OSE, retrieved July 2026).
  • Aggregation fails on normalization, not arithmetic: cross-platform duplicates, active versus total listings in the denominator, entire-home versus private-room mix, and vendor-drawn market boundaries.
  • Aggregating your own portfolio turns on four choices made once and held: unit of analysis (listing, unit, or property), time basis (booking, stay, or payout date), revenue definition (gross, net of commission, fees in or out), and how cancellations are counted. Mixing any of them is why portfolio totals fail to reconcile with the accounts.
  • A single platform is not the market: across six online rental platforms in five US metros, estimates of rental cost and unit-size distribution differed markedly by platform (source: Costa et al., Cityscape 23(2), 2021; long-term rental platforms, so a structural warning rather than an STR measurement).
  • Market occupancy and your occupancy are different quantities. The San Francisco Model’s 70% cap means a market truly running at 80% cannot read above 70 in that data, so use market occupancy to rank markets, not to forecast your revenue.
  • Two reading rules: know whether a number is measured or estimated, and never carry a benchmark across market, season, or property type.

Information current as of July 2026; verify regulatory details against the official source before acting.

Frequently asked questions

What does short-term rental data actually include?
It spans three families. Performance metrics (occupancy, ADR, RevPAR, booking pace, minimum stay) tell you how a listing earns. Supply-and-demand signals (listing counts, seasonality) describe the market around it. Regulatory status (nightly caps, license and primary-residence rules) sets what you can legally sell. An operator running 5 to 50 doors uses all three to price, place, and stay compliant.
How is RevPAR different from ADR and occupancy?
ADR is revenue per booked night; occupancy is the share of available nights that sell. RevPAR is revenue per available night, and CoStar/STR describes it as a function of occupancy and ADR, which gives the identity RevPAR = ADR x occupancy. ADR alone can flatter a listing that sits empty at a high price. RevPAR combines rate and fill into one number, so it is the fairer scoreboard when you compare units, dates, or the effect of a price change.
Can I trust Inside Airbnb occupancy for my own listing?
Use it for market shape, not for your own precise occupancy. Inside Airbnb is public-listing data compiled independently and released under CC BY 4.0, but its occupancy figures are estimates, not measured bookings: the San Francisco Model converts reviews to bookings at a 50% review rate, applies an average length of stay, and caps estimated occupancy at 70%. Your own platform booking records are the measured source for how full your units actually run.
Why treat regulatory status as data rather than background?
Because it sets a hard ceiling on the nights you can sell and whether a platform will process the booking, and it varies sharply by market. New York exempts stays of 30-plus consecutive days, Paris and London cap primary residences at 90 days or nights per year, DC caps host-absent rentals at 90 nights, and Seattle requires a per-unit license. A rule that shifts changes the economics, so it belongs in your dataset with a source and a date, verified against the official page before acting.
How is short-term rental data aggregated?
Four mechanics produce nearly every market figure. Scrape-and-model collects public listing pages, 365-day calendars, and reviews, then infers bookings: Inside Airbnb's San Francisco Model converts reviews to bookings at a 50% review rate, applies an average length of stay (5.5 nights where reported for San Francisco, otherwise 3 nights or the listing minimum if higher), and caps occupancy at 70%. Panel supply works the reverse way, taking measured reservations from participating operators, usually via PMS integrations. Official registration datasets, such as New York's Local Law 18 register, aggregate legal status at address level with nothing modeled. The fourth step, normalization, is where aggregation breaks: cross-platform duplicates, active versus total listings in the denominator, entire-home versus private-room mix, and vendor-drawn market boundaries.
How do I aggregate my own short-term rental data across channels?
Four decisions do the work, and getting them wrong is why portfolio numbers fail to reconcile with the accounts. First, pick a unit of analysis and hold it: a listing is not a unit and a unit is not a property, so occupancy per listing, per unit and per property are three different numbers. Second, fix the time basis. A reservation has a booking date, a stay date and a payout date; revenue on booking date shows sales pacing, on stay date shows what the asset produced, on payout date shows what the bank saw. All three are valid, mixing them in one total is not. Third, define revenue once: gross, net of channel commission, cleaning fee in or out, taxes you merely collect excluded. Each channel exports a different default. Fourth, decide how cancellations and modifications are counted, or last month's total will not reproduce. The test: could someone else, given your exports and rules, reach the identical number.
Where does market occupancy data come from, and can I trust it?
It comes from three places that disagree by construction. Your platform bookings are measured but cover only your units. Open data estimates are modeled from public traces and hard-capped: because the San Francisco Model caps estimated occupancy at 70%, a market genuinely running at 80% cannot read above 70 there, so the top of the distribution is compressed by design. Paid tools are modeled from scraped listings using the vendor's tracked set as the denominator and the vendor's own polygon as the market. Use market occupancy comparatively, to rank markets under one consistent method, not absolutely to forecast your revenue. Never set your measured occupancy against a modeled market figure: those are two different quantities, and the comparison flatters an ordinary market.
Do 5 to 50 door operators need to pay for market-data tools?
Not always. Your own platform bookings already measure occupancy, ADR, and lead time for what you operate, and free open data like Inside Airbnb covers rough market shape. Paid market-data tools earn their cost mainly when you are evaluating a new market you have no bookings in, since they estimate neighborhood-level demand your own data cannot see. Treat that as a due-diligence spend during evaluation rather than a permanent subscription.