Before the First Spot Airs: An Instrumentation Checklist for Closed-Loop Streaming and TV Attribution - Enterprise Digital Marketing
Most television and streaming flights cannot be measured afterward because nothing was put in place beforehand. The measurement an awareness flight needs is cheap, but it has to exist before the flight starts: a six-month branded-search baseline, a call log with dynamic number insertion, a desk tag, a flight schedule with edges, and a held-out market. A checklist for multi-location operators, with the figures from our own measured awareness flight as the benchmarks each step should be able to produce.
The reason most TV and streaming flights end in an argument about whether they worked is that the measurement was designed after the spend. By the time the operator asks for proof, the flight has been on for a quarter, the baseline is gone, the schedule had no edges, every market got the same weight, and the only data anyone has is the vendor's matched-household report. At that point the honest answer is that the flight is unmeasurable, and the practical answer is that the vendor's number wins by default.
Everything below costs little or nothing. All of it has to be in place before the first spot airs. The benchmarks cited are from a continuous awareness flight we measured across a four-location group in the book, documented in the TV and streaming attribution paper and the lift valuation paper.
1. Search Console ownership and a clean brand term
What: Owner-level access to the Search Console property for every location's site, or the single domain property if the locations share one. Confirm that a query-by-date export works and that it reaches back at least six months.
Why: Branded organic impressions are the primary demand signal for an awareness flight, and Search Console is the only instrument that counts them independent of ad spend. History is retained for sixteen months and cannot be reconstructed later.
Check the brand term. Pull the query list and confirm that a substring match on the brand name classifies cleanly. A brand that shares a word with a place, a person, or a dictionary term needs a curated query list instead. The group in our case had 157 distinct branded queries on a single unambiguous substring. A brand campaign in another account in the book, labeled "Brand" by its previous agency, was 1.0 percent branded when classified by actual query text, so never classify by campaign name.
Benchmark: With six months of baseline the group's flight produced a weekly branded series with a Welch t of 6.95 against its pre-period. With six weeks of baseline the same lift would have been indistinguishable from noise.
2. A six-month pre-flight baseline, pulled and saved
What: Export branded and generic impressions and clicks, weekly, for the six months before launch. Save the file outside the reporting platform.
Why: The baseline is what the flight will be measured against, and it needs to be frozen before anyone has a reason to move it. Exporting it also forces the question of whether the brand is already trending, which changes the test design.
Benchmark: The group's pre-period ran 444 branded impressions per week against 11,870 generic. Those two numbers, and their ratio, are the whole baseline.
3. Generic impressions, or a held-out market, as the control
What: Decide before launch what the control series is. For a single-market operator, it is generic organic impressions for the same site. For a multi-market operator, it is the branded series in markets where the flight does not run.
Why: A branded lift with no control is a before-and-after comparison, and before-and-after comparisons are confounded by season, algorithm changes, and anything else that launched the same month. Generic impressions absorb site-wide effects. A held-out market absorbs brand-specific effects too (press coverage, a competitor closing), which makes it the stronger control where it is available.
Benchmark: The group's generic control fell 8.9 percent across the window while branded rose 41.4 percent, for a difference-in-differences of 50.3 points. Without the control, a reader could have argued that all of search was up. With it, that argument is closed.
4. A flight schedule with edges
What: Buy the flight in blocks with dark periods between them: two weeks on, two weeks off, or four on, two off, for at least three cycles. Write the schedule down with dates and keep it with the baseline.
Why: A continuous flight has one edge, the start, and whatever else launched that month is confounded with it forever. Every on-off edge is an independent test. The expected signature is a branded rise beginning one to two weeks after a block starts and a decay beginning one to three weeks after it ends. Three clean cycles settle the question in a way a year of always-on data cannot.
Benchmark: Our measured flight was always-on, launched in the same month as the rest of the group's program, and the paper says so: the 41 percent is attributable to the program as a whole, with the social flight the only component that plausibly moves branded demand. An edged schedule would have let us say more.
5. Call tracking with dynamic number insertion on the site
What: A call-tracking layer that logs every inbound call with source, and dynamic number insertion so that calls from branded-search sessions are separable from calls from everything else.
Why: Inbound call volume is the operator's own record and is not subject to vendor matching. It is the second demand signal, and with DNI it can be joined to the first: the branded-session call rate, which our valuation had to carry as a 15 to 35 percent scenario range, collapses to a measured point.
Benchmark: The group's six-month call log held 11,867 inbound calls, which with the ledger gave $273 of completed revenue per inbound call. That figure is the price list for the whole valuation and it came from the call log, not from any ad platform.
6. The front-desk source tag, with an "ad on TV" option
What: A required source field on every new-customer record, with an explicit option for the flight (the platform, or simply "saw our ad"). Staff trained on it before launch, not after.
Why: The desk tag is the only route that produces a ticket-level record. It undercounts awareness channels severely, because people do not remember spots, but a tag count that rises during on-blocks and falls during dark periods is corroborating evidence for the aggregate signals. The best-instrumented location in the book reached a 9.8 percent tag ratio on first-time callers; a flight-specific tag will run lower than that, and should still show the edges.
7. Revenue per call and a first-time floor from the ledger
What: Before launch, compute two numbers from the operator's own data: blended completed revenue per inbound call (ledger revenue over call-log volume, same window) and a first-time-caller floor (measured new-customer booking rate times new-customer average ticket).
Why: These are what turn an impression count into a revenue figure. Computing them before the flight removes any temptation to pick the window that makes the flight look best afterward.
Benchmark: $273 blended and $107 floor at the group. At a 25 percent click-to-call rate those priced the group's 127 incremental branded clicks per month at $3,400 to $8,700, against roughly $3,000 of monthly flight spend. A streaming flight at five to ten times that spend should produce a proportionally larger signal or be questioned.
8. The vendor report, pre-negotiated
What: Before signing, get written answers to five questions about the vendor's attribution report: the share of credited transactions that are existing customers, the credit window, the control group and how it is chosen, whether lift is stated against the control or against zero, and whether credited revenue will be reconciled to the operator's ledger.
Why: Household-matched TV attribution inherits the inflation mechanism of direct-mail matchback: "exposed, then transacted" with no causal requirement. Asked afterward, these questions become a dispute. Asked beforehand, they are a specification. A vendor that will not specify a control has not offered a measurement.
9. The reporting template, built empty
What: A four-line report, built before launch with the baseline filled in and the flight columns blank: branded impressions and clicks weekly with the control overlaid; call volume weekly with the flight schedule overlaid; the vendor's figure, labeled vendor-reported; and the modeled revenue line with the call rate shown as a range and the per-call basis stated.
Why: A report built after the fact is built to a conclusion. A report built empty is built to a question. It also makes the first weekly review, two weeks into the flight, a matter of filling in a column rather than deciding what to look at.
What this costs
Steps 1 through 3, 7, and 9 are analyst hours against data the operator already owns. Step 4 is a media-buying decision that costs nothing and usually improves frequency. Step 6 is a training session. Steps 5 and 8 are the only items with an external cost, and most multi-location operators already run call tracking for other reasons.
The alternative is a flight that cannot be evaluated, renewed or cancelled on the strength of a vendor's matched-household number that, on the mail evidence in the book, can overstate a channel's contribution by a factor of six.
Methodology notes
- Benchmarks from a four-location group: Search Console query-by-date, 479 days ending late July 2026; pre period 2025 weeks 15 through 39, post period 2025 week 43 through 2026 week 30, three ramp weeks excluded; 11,867 inbound calls over six months; completed revenue from the point-of-sale per-ticket export for the three locations with ledger access.
- Desk-tag ratio from a separate single location in the book: 181 first-time Google-tagged tickets against 1,848 first-time callers from the tracked line.
- Matchback comparison from a separate single location: vendor-credited $970,173 against $141,770 by front-desk tag (53 percent versus 7.8 percent of ledger revenue).
- All figures anonymized across the locations we manage.