Metric Discrepancies & Attribution

Self-Reported Attribution vs. Your Pixel: What "How Did You Hear About Us?" Is Actually Good For

By Chinmay Raibagkar·August 28, 2026·9 min read·Some SQL

The 60-second version

A post-purchase survey will never reconcile with your pixel, and it was never supposed to. Here is the question it does answer, and how to weight the answer without fooling yourself.

  • What happened, in one line
  • What to do about it this week
  • What you can safely ignore

Somewhere in your post-purchase flow there is a dropdown that says "How did you hear about us?" Roughly 40% of customers answer it. Of those, a large share pick "Google" when they came from Instagram, "a friend" when they came from a podcast, and "Instagram" when what they actually mean is "somewhere on my phone".

Then you compare it to your pixel data and the two disagree wildly, and someone concludes the survey is useless.

The survey is not useless. It is answering a completely different question from your pixel, and the mistake is expecting the two to reconcile. They cannot, they never will, and the reason is more interesting than the disagreement.


Two instruments, two questions

What Each Instrument Actually Measures

Data Journey
Stage 1Mechanical
The Pixel

Records the last technically observable digital touchpoint before checkout. Precise about what it sees, blind to everything it cannot.

Answers: what was the last click?
Stage 2Recalled
The Survey

Records what the customer remembers as the reason they know you. Imprecise, unreliable on detail, and covering channels no pixel touches.

Answers: what made you aware of us?
Stage 3Signal
The Gap Between

Not error. The systematic difference between the two is itself the most useful output — it locates demand your pixel structurally cannot see.

Answers: what is untracked?

Your pixel answers a question about mechanism: which URL, carrying which parameters, was loaded immediately before this purchase. It is extremely accurate about that and completely silent about everything else.

The survey answers a question about memory: what does this person believe made them aware of you. It is unreliable about the specific channel and uniquely capable of surfacing things no tracking system will ever see — a friend's recommendation, a podcast mention, a physical billboard, a comment thread, a WhatsApp group.

Neither is a degraded version of the other.

The reconciliation trap. Teams try to score the survey against the pixel: "63% of people who said Google actually came from Meta, so the survey is 37% accurate." That framing assumes the pixel is ground truth for a question the pixel does not answer. The customer who said "Google" may have discovered you on a podcast, searched your brand name on Google, and clicked a paid brand ad. All three statements are true. Only one is in your pixel.


Why the survey is systematically wrong in predictable ways

Understanding the specific biases is what makes the data usable. These are not random noise; they lean in known directions.

Recency bias

People report the most recent touchpoint they consciously noticed — which is rarely the one that created the demand. Someone who saw your ad twelve times over three weeks and finally searched your name reports "Google search", because that is the action they took deliberately.

Brand-name substitution

"Google" frequently means "the internet". "Instagram" often means "my phone". "Facebook" is used by a whole demographic to mean any social feed. Options in your dropdown are being used as rough categories, not precise sources.

Options bias

Whatever you list gets picked. Adding "TikTok" to your dropdown produces TikTok attribution from people who have never opened it, purely because the option is visible and plausible. This is the strongest and least-appreciated effect: your dropdown shapes its own results.

Social desirability and the escape hatch

"A friend recommended you" is a flattering answer that costs nothing to give. Meanwhile, "Other" and "I don't remember" absorb everything the respondent cannot be bothered to classify — which is why leaving those options out does not improve the data, it just redistributes the same confusion into your named channels.

Non-response bias

The 40% who answer are not a random 40%. They skew towards more engaged, more deliberate buyers — the ones with a story about how they found you. Impulse purchasers, who are disproportionately paid-social-driven, skip it.

Where the two instruments genuinely disagree

Typical pattern
Pixel saysPaid search 34%Mostly branded terms — people searching a name they already knew
Survey saysPaid search 11%Few people experience a branded search as 'how they heard'
Pixel saysPodcast 0.4%Only the fraction who used a vanity URL or promo code
Survey saysPodcast 14%The channel that created awareness, invisible to the pixel by construction
Neither instrument is broken. Branded paid search harvests demand the podcast created — the pixel sees the harvest, the survey sees the planting. Cutting podcast spend because 'the pixel shows 0.4%' would cut the thing generating the searches the pixel is taking credit for.

The four things a survey is genuinely good at

1. Finding channels your pixel cannot see

This is the primary job. Podcasts, out-of-home, PR, word of mouth, offline events, WhatsApp and Telegram forwards, private communities — none of these produce a trackable click for most of the people they influence. If 14% of respondents name a channel your pixel scores at under 1%, you have located spend that is working invisibly. That is the highest-value output of the whole exercise.

2. Tracking changes rather than levels

The absolute percentages are unreliable. The trend is much less so, because most of the biases are constant. If "a friend recommended you" moves from 8% to 19% over two quarters with your dropdown unchanged, word of mouth genuinely increased. The bias did not move; the signal did.

3. Cross-checking a suspicious channel

When a channel shows implausible pixel performance — a 14x ROAS on branded search, say — the survey is a cheap sanity check. If almost nobody names it unprompted, you are very likely harvesting demand rather than creating it, and the ROAS is a measurement artefact rather than a result.

4. Segmenting by what customers are worth

Join survey responses to actual order data and the question changes from "where did they come from" to "which discovery paths produce good customers". This is where the data stops being a curiosity:

Survey Response Joined to Realised Customer Value

Show query

This routinely produces the most actionable finding of the whole exercise: word-of-mouth and community-sourced customers usually show materially higher repeat rates than paid-social ones, which changes what you are willing to pay for each.


Designing a survey that produces usable data

Five Design Rules

Process Flow
1

Ask after the purchase, never during

A field in the checkout flow costs you conversions. Put it on the confirmation page or in the first post-purchase email, where abandoning has no cost.

2

Keep the options stable for at least two quarters

Every change to the option list breaks the time series. Since trend is the reliable output, changing options destroys the main value.

3

Make it optional and single-question

Required fields produce compliance answers — whichever option is first. A skippable question yields a smaller, more honest sample.

4

Include 'Other' and 'I don't remember'

Without an escape hatch, that uncertainty gets redistributed into your real channels and quietly corrupts them.

5

Store the response against the customer, not the order

Joined to the customer record, it becomes a segmentation dimension for life. Stored on the order, it is a one-off statistic.

Do not put the survey in your checkout. Every additional field in a checkout flow reduces completion. A post-purchase survey with a 40% response rate and no conversion cost beats an in-checkout survey with a 90% response rate and a 2% conversion cost — that trade is never worth it.


How to weight it against the pixel

The honest answer is: do not blend them into one number. There is no defensible weight, and any weighted average is a number with no owner and no meaning.

What works instead is using each instrument for the decision it is competent to make:

Which Instrument Answers Which Decision

Reporting Hierarchy
Tier 1
Bid and creative decisions

Pixel only. The platform's optimisation needs high-frequency mechanical signal, and survey data is neither fast enough nor granular enough to inform a bid.

Platform-reported conversions
Tier 2
Channel budget envelope

Blended CAC and MER, with the survey used to decide whether an untracked channel deserves a line in the budget at all.

Blended CAC + survey trend
Tier 3
Untracked channel investment

Survey only, because no pixel exists. A podcast or PR line is justified by survey mentions and by blended CAC moving in the right direction while it runs.

Survey share + blended CAC delta
Tier 4
Causal truth

Neither. If the answer materially changes your budget, run a holdout — geo, audience, or a spend pause. Both instruments are correlational.

Incrementality test

The one place they combine usefully is as a triangulation check. If a channel's survey share is rising, its pixel-attributed conversions are rising, and blended CAC is falling, you have three independent instruments agreeing. That is as close to confidence as marketing measurement gets without an experiment.


Frequently asked questions

What response rate should I expect?

Around 30–50% for an optional single question on a confirmation page; 15–25% in a post-purchase email. Rates above 80% usually mean the field is required, which means you are collecting compliance rather than information.

Should I offer an incentive to answer?

No. Incentives raise response rate and lower response quality, because they recruit people optimising for the incentive rather than answering thoughtfully. They also skew the sample towards discount-sensitive buyers.

Free text or dropdown?

Dropdown with an "Other, please specify" free-text field. Pure free text is genuinely richer but needs categorisation before it is usable, and in practice that categorisation never gets done. The free-text box on "Other" is where you discover the channel you should add to the dropdown next year — not this quarter, because changing options mid-series breaks the trend.

Can I use survey data to set channel budgets directly?

Not directly. Use it to decide whether an untracked channel gets a budget line at all, and to size that line roughly. Use blended CAC and, where it matters enough, an incrementality test to decide the amount.

My survey says 30% "word of mouth". Should I stop paying for ads?

No — and this is the most common misreading. Word of mouth is downstream of the ads. Customers acquired through paid channels tell their friends. The survey shows the last link in a chain whose first link was very likely something you paid for. The only way to know is to reduce spend in one geography and watch whether word-of-mouth volume follows.


The summary

  • The pixel measures the last mechanical touch. The survey measures remembered awareness. Different questions, so a disagreement is not an error.
  • Survey bias is systematic and predictable: recency, brand-name substitution, options bias, and non-response. Because it is systematic, trend is trustworthy even when level is not.
  • Its unique value is visibility into channels no pixel can see, and segmentation of customer value by discovery path.
  • Do not blend the two into a single attributed number. Route each to the decisions it is competent to make.
  • When the answer genuinely changes your budget, stop triangulating and run a holdout.
Free tool

Blended CAC Calculator

Total spend across every channel, divided by total new customers — the acquisition cost number that reconciles with what you actually spent.

CR

Chinmay Raibagkar

About author →

Founder of DataLens AI. He helps non-technical teams read their ad and database numbers with confidence — which number to trust, what to do next, and what to ignore.