Auditing Direct Mail ROI: Matchback Reports, Penetration Reports, and the Tests That Catch Inflation - Enterprise Digital Marketing
How direct mail vendor reports inflate returns, demonstrated on two real vendor reports from an instrumented auto service book: a matchback report claiming 53 percent of a shop's entire ledger against desk tags showing 7.8 percent, and a penetration report whose growth story decomposed into a counting change. Includes the reusable audit tests and a four-lens ROI method.
Direct mail is usually the largest marketing line item at an independent auto service location, and it is the least audited, because the audit evidence arrives from the vendor being audited. The mail company prints the pieces, runs the matchback, and writes the ROI report. No other channel in the building gets to grade its own homework at that scale.
We audited two of these vendor reports against the shops' own systems: point-of-sale ledgers, desk-tag source attributions, call tracking, and customer lists. The two reports represent the two dominant genres. One was a matchback ROI report (vendor matches its mailing list against the customer file and claims the revenue of every match). One was a year-over-year penetration report (vendor reports customer counts and spend by carrier route across two time windows and claims growth). Both inflated, through different mechanisms, and both mechanisms are detectable from the report's own arithmetic plus the shop's ledger.
This paper documents the mechanisms and the tests. A companion paper covers what direct mail is actually for inside a digital program once the measurement is honest.
Case 1: the matchback report
The vendor's report claimed 1,256 repair orders worth $970,173 for the year to date. The shop's entire ledger for the identical period was 2,384 repair orders worth $1,823,423.
The claim amounted to 53 percent of every dollar the shop wrote. The shop's own desk tags (the front counter asking "how did you hear about us" and recording it on the repair order) put the mail channel at 7.8 percent of the same ledger. That is a 6 to 7x gap between the vendor's number and the shop's number, and the shop's number is the one a buyer should start from.
Three mechanisms produced the gap:
1. Population matchback with no behavioral test. The vendor matched its mailing footprint against the customer file and claimed every match's revenue. But a saturation mail footprint covers most of the shop's natural trade area, so most customers match whether or not mail influenced anything. Matching is not attribution; it is a census.
2. The suppression paradox. This vendor runs real, systematic suppression: before each drop it removes customers with a service visit in the prior eleven months from the mailing list (we verified the suppression files: 1,231 to 1,518 addresses per drop, 90 percent matching the shop's customer list by address). Good practice. But the matchback base for the penetration claims included 1,113 of the 1,255 suppressed households, households with $2.1M in lifetime spend at the shop. The vendor claimed revenue from active customers it had itself excluded from the mail. That single cross-check is close to a complete audit on its own.
3. Response-rate claims on the shop's own customer list. Two "targeted" campaigns mailed roughly 800 pieces to the shop's existing customers and claimed 26 to 34 percent response at $130 to $221 returned per dollar. Mailing your own active customers and claiming their next visit as response is the purest form of the paradox above.
The confession in the vendor's own data: when one campaign was configured as prospect-only counting, its claimed return collapsed from $30.02 to $7.12 per dollar, a 76 percent drop with a near-identical mailing footprint. Nothing about the mail changed. The counting did.
Case 2: the penetration report
The second report contained no spend and no campaign detail: 488 carrier routes, customer counts and revenue in two 14.5-month windows, presented as growth. Its inflation mechanisms were structural, and each is testable from the report itself:
The overlapping-windows test. The two comparison windows overlapped by 2.5 months, double-counting roughly $830K of revenue into both sides of a "growth" comparison. Check the window boundaries before anything else.
The lifetime-spend self-proof. Divide each route's reported revenue by that route's reported average ticket and sum the implied visit count. In window one, the implied count (2,759) matched the report's own lifetime-visits column (2,764): the "period revenue" was lifetime revenue relabeled. Window two's claim of $6.95M came to 1.5 to 1.7x the shop's total revenue from all sources in that window, which is arithmetically impossible for one channel. A channel claim should always be divided by the whole ledger; anything near or above 100 percent is self-refuting.
The counting-switch test. Window one's customer counts fit new-customers-in-footprint. Window two's fit all-active-customers-in-footprint. We verified this by joining the vendor's ZIP-level counts to the shop's POS actives (the vendor counts tracked POS actives roughly 1:1, with an 87 percent name match). The reported "growth" was a definition change. The shop's actual ledger between fairly matched windows: active customers down 6.9 percent, new customers down 16.8 percent, revenue down 14.2 percent. The report showed growth over a period in which the shop shrank.
This is the same failure class as the conversion-scope changes documented in the revenue-lift measurement paper: a definition moved inside a trend and was reported as performance.
The four-lens ROI method
A single mail ROI number hides the choice of counting rule, so we compute four and publish all of them. From the matchback shop:
| Lens | Counting rule | Result (gross-profit multiple) |
|---|---|---|
| Vendor claim | All matchback matches | 14.0x (not credible) |
| Desk-tag | All ROs the counter tagged to mail | 2.0x |
| New-customer-only | First-visit customers tagged to mail | 0.53x |
| Lifetime | Acquisition cost vs full customer LTV | 1.6x |
The spread is the finding. As a prospecting channel this program loses money (0.53x: roughly $41K of annual spend produced 136 new-customer repair orders at a $268 average ticket). As a retention touch it roughly breaks even to modestly pays (1.6 to 2.0x). The vendor's 14.0x describes neither.
The pattern repeated at the second shop from the customer-list side: mail-sourced customers were the lowest-lifetime-value paid cohort in the ledger ($1,126 LTV at a $542 average ticket, against $2,840 and $1,206 for search-sourced customers at the same shop). Whatever mail is doing, it is not sourcing the customers the growth case depends on.
The audit checklist
Runnable by any operator with the vendor report and their own POS export:
- Divide the channel's claimed revenue by the whole ledger for the same period. Anything above roughly 15 percent for mail deserves hostility; above 50 percent is self-refuting.
- Compare the claim to desk tags. Desk tags undercount somewhat, but a 6x gap is not undercounting.
- Check window boundaries for overlap and for equal lengths.
- Run the lifetime-spend self-proof: implied visits from the report's own columns, and the whole-ledger ceiling.
- Ask for the counting definition in each period, in writing. If a "growth" report cannot state one definition that holds across both windows, the growth is the definition.
- Cross-check the matchback base against the suppression files. Suppressed households appearing in the claimed-revenue base is dispositive.
- Ask the two questions vendors resist: what share of claimed responders were existing customers, and what is the response rate on holdout routes that received no mail.
None of this requires bad faith on the vendor's part, only defaults that all point the flattering direction. The checklist exists so the defaults get examined before the budget renews.
Methodology note: both reports are real vendor deliverables audited in 2026 against the shops' live systems (POS ledger, desk-tag source attribution, customer list, suppression files supplied by the vendor). Shops, vendors, and markets are anonymized. Joins are deterministic (address, ZIP-level counts, exact name) with match rates stated inline; figures are floors or exact ledger values, not models.