Field note · opportunity

What Evidence Should Product Teams Collect About Feedback Contradictions

A fixed 40-review audit shows why ratings alone create false contradictions. Use this evidence matrix before turning opposing feedback into a roadmap decision.

15 minute read
  • customer feedback
  • product research
  • evidence
Illustration of a product team comparing opposing feedback records with missing context fields

When two customers appear to want opposite things, I don't start with a vote. I check whether the records describe the same product area, user, workflow stage, and outcome.

In a fixed sample of 40 public Netflix Google Play reviews, a simple text-only baseline produced 36 opposite-rating candidate pairs. A bounded audit of 12 candidates confirmed none as a genuinely opposed requirement.

That gap is the useful finding. Product teams should collect enough evidence to explain a contradiction before they turn it into a roadmap item.

Observed result: In the fixed sample, 36 text-only candidates became 8 false positives and 4 context-unknown or rating-text cases in the audit. No audited pair justified a product-wide “build one thing, reject the other” decision. The sample is not representative, and the result is not a prevalence estimate.

What evidence should a product team preserve for every feedback item?

Preserve provenance and decision context together. A quote without its source, timing, customer situation, and outcome is not enough to resolve a contradiction.

At minimum, keep this record for every item:

FieldWhat to captureWhy it matters when feedback conflicts
Stable source IDReview ID, ticket ID, interview ID, or stable URLLets a reviewer reopen the raw evidence instead of trusting a summary
Source and collection dateChannel, storefront, export date, and original timestampSeparates a recent regression from an old preference and preserves provenance
Raw text and ratingOriginal wording, title, star score, and edits if availablePrevents a synthesis from hiding qualifiers or a rating-text mismatch
Product area and decision objectPlayback, billing, onboarding, catalog, search, or the exact featureStops two different problems from being called one contradiction
Request or problem typeAdd, remove, restore, fix, simplify, speed up, or understandDistinguishes opposing requirements from requests that point in the same direction
Segment or use-case proxyPlan, role, device, region, account type, job, or declared contextTests whether different customers are optimizing for different outcomes
Workflow stageDiscovery, setup, first use, repeat use, recovery, or cancellationA positive first-use report can coexist with a severe recovery failure
Severity and impactFrequency, task blockage, workaround, revenue or retention exposurePrevents a loud but low-impact preference from outweighing a rare blocker
Behavioral or outcome signalUsage, completion, repeat use, retention, cancellation, conversion, or support escalationTests whether stated preference changes what people actually do
Opposing evidenceThe exact item or group that appears to disagreeKeeps both sides visible in the decision record
Unresolved questionThe missing fact that could change the decisionCreates a concrete next evidence task instead of a vague “needs research” note

Atlassian's product-feedback guidance recommends tracking source, segment, product area, request type, frequency, severity, business impact, related work, and next step. Research on continuous software engineering makes the same distinction from another angle: teams combine qualitative feedback with quantitative signals, but interpreting reviews often requires additional user context. The study's full text is especially useful here because it treats access to vaguely defined user groups and metric selection as practical problems, not as fields an analyst can safely invent.

The rule is simple: if a field could change which product action you choose, preserve it before summarizing the feedback.

When I teach product managers to move from writing specifications to building and shipping, the recurring failure is rarely a missing model trick. It is that nobody can say what “done” means or which evidence would change the decision. The same problem appears here. A contradiction ledger is useful because it makes the missing decision fields visible.

/blog/what-evidence-should-product-teams-collect-about-customer-feedback-contradictions-feedback-contradiction-ledger.webp

What did the fixed public-review sample actually show?

The sample showed that coarse polarity creates candidates faster than it creates understanding.

I fixed the corpus before drafting: rows 0 through 39 of the public splevine/netflix-google-play-reviews-sample train split, accessed on 2026-08-24. The dataset page exposes real review text, UUID review IDs, timestamps, ratings, app versions where present, platform, country, language, helpful-vote counts, developer-response fields, and collection timestamps. The selected rows are dated 2025-06-26, with some earlier row dates in the dataset's stated range. The public dataset viewer is the provenance for the raw records.

I mapped ratings of 4 or 5 to positive, 3 to neutral, and 1 or 2 to negative. I then assigned one primary product area using a fixed keyword order and paired items in the same area when their rating polarities were opposite.

MeasureResultWhat it does not prove
Review rows40Not a representative sample of Netflix users
Positive, neutral, negative ratings22, 1, 17Ratings are not the same as requirements
Rows assigned to a known area27The baseline leaves 13 vague or uncoded rows
Playback or reliability rows6Nine candidate pairs, not nine conflicts
Content or catalog rows10Twenty-four candidate pairs, not twenty-four opposing needs
Pricing or billing rows6No opposite-rating pair in this slice
Account or access rows4Three candidate pairs, all needing context
Interface rows1No opposite-rating pair in this slice
Candidate pairs36A triage output from a weak baseline
Audited candidate links12A bounded audit, not a model-accuracy benchmark
Confirmed opposing requirements0 of 12No claim about the full 5,000-row dataset

The most important column is the last one. A text-only contradiction detector can tell you where to look. It cannot tell you whether two people are making the same decision under the same conditions.

How should a team define a contradiction before detecting one?

Define a contradiction as semantically related feedback that points to opposing judgments or requirements about the same decision object, then keep separate labels for the reasons the texts differ.

For this analysis, an item pair became a candidate only when it met all three conditions:

  1. Both items matched the same primary product area.
  2. One rating was positive and the other was negative.
  3. Both contained a product judgment, request, or problem statement.

The audit then asked a harder question: are the customers actually asking the product team to move in opposite directions?

Use these explanations instead of a single “conflict” label:

ExplanationMeaningEvidence needed
Segment differenceThe same feature serves groups with different prioritiesDeclared or verified segment, role, plan, device, region, or account context
Workflow stageThe same area behaves differently during setup, first use, repeat use, or recoveryEvent stage, app state, version, and task step
Use-case differencePeople need the feature for different jobsJob-to-be-done, desired outcome, and actual workflow
Severity differenceThe same problem affects people with different consequencesBlockage, workaround, frequency, business or user impact
Preference conflictThe same people, context, and decision object genuinely prefer opposite optionsComparable context and a stated trade-off
Same directionRatings differ, but both records point toward the same improvementRequest type and exact object
Rating-text mismatchThe star rating and written text disagree inside one recordRaw rating, raw text, and any edit history
Context unknownThe record lacks enough information to classify the differenceA named missing field and a collection plan

The mobile-review comparator supplied in the topic package uses semantic similarity, sentiment distributions, topic modeling, antonym and negation handling, ratings, and frequency to identify apparent feature conflicts. Its publisher record is a useful comparator, but its detection method should not be mistaken for a product decision standard. Your team still needs to know what the feedback means in context.

What does a transparent contradiction baseline look like?

Start with a baseline that is easy to inspect and easy to disprove. A team should see why an item entered the candidate set.

My dated configuration used lowercase text, a fixed keyword family, the rating polarity rule above, and all within-area opposite-polarity pairs. The area families were:

pricing/billing: payment, recharge, price, pricing, plan, pay, budget, greedy
account/access: household, login, log in, password, kicked, account, profile, watching
interface: ui, user interface, pane, scroll, navigate, navigation, screen
playback/reliability: load, loading, play, playing, open, working, crash, stuck, resume
content/catalog: movie, movies, show, shows, series, content, squid, interactive, game, games, selection

The first matching family became the primary area. A row that matched none became unknown. Candidate generation was the equivalent of:

rows = load_dataset(
    "splevine/netflix-google-play-reviews-sample",
    split="train",
).select(range(40))

for row in rows:
    text = row["content"].lower()
    row["sentiment"] = (
        "positive" if row["rating"] >= 4
        else "neutral" if row["rating"] == 3
        else "negative"
    )
    row["product_area"] = first_matching_keyword_family(text) or "unknown"

candidates = [
    (a["review_id"], b["review_id"])
    for a, b in combinations(rows, 2)
    if a["product_area"] == b["product_area"] != "unknown"
    and {a["sentiment"], b["sentiment"]} == {"positive", "negative"}
]

This baseline is not meant to win a benchmark. It is meant to make a team's first failure legible. It cannot understand synonyms, mixed-area reviews, translated meaning, or whether “good but needs more movies” is support for the current product or a request to change it. Those weaknesses are exactly why the audit exists.

/blog/what-evidence-should-product-teams-collect-about-customer-feedback-contradictions-feedback-baseline-candidate-pairs.webp

What should the contradiction ledger contain?

The ledger should make every candidate answerable without reopening a dozen disconnected tools.

Use one row per feedback item, then link rows into a contradiction group. Do not collapse the raw records into one synthetic summary before the decision is made.

Review IDDate and ratingAreaRequest or problemContext proxyOutcome signalOpposing evidenceMissing question
1bb85164-2946-48d9-98e1-4c55c59ad5312025-06-26, 5content/catalogWants more moviesUS, English, app version missingHelpful votes 0Selection complaintsWhich titles, region, plan, and viewing behavior?
dd6106a0-0fb6-4ed0-854a-a754dfddf4172025-06-26, 1content/catalogSmall selection and cancellationsUS, English, version presentHelpful votes 0Broad catalog praiseDid the person cancel, stop watching, or only dislike the current titles?
b330cbbe-fa71-476a-bb82-9d28359d5aa62025-06-26, 1content/homeRemove games and short clips from the main surfaceUS, English, version presentHelpful votes 38Strategy that adds more surfacesTV or mobile, discovery behavior, and task completion
c04516ac-b6df-4cd0-a745-f6effcaa26602025-06-26, 5content/catalogBroad praiseUS, English, version presentHelpful votes 0Negative catalog-specific recordsWhat did the reviewer use and value?
ac1625c3-4e7e-492d-baed-e47260f016212025-06-26, 1content/catalogRestore interactive showsUS, English, version presentHelpful votes 0General catalog approvalIs interactive content a meaningful use case for a defined segment?
6077f080-67d8-401b-abc9-e2bcf57880482025-06-26, 5content/catalogGeneral approval for moviesUS, English, version missingHelpful votes 0Requests for more or different contentActive use, plan, region, and renewal behavior

Two fields deserve special attention:

  • Opposing evidence keeps the counterexample beside the claim. A synthesis that says “customers want more content” without showing the rows that reject, qualify, or redirect that reading is not auditable.
  • Missing question turns uncertainty into work. “Need more context” is weak. “Which titles were available in the reviewer's region, and did the person cancel?” is testable.

This is also where a reviewable AI workflow should stop. The parent guide on testing AI with contradictory feedback can help you implement the workflow, but the workflow should not invent the missing fields. If a model cannot distinguish a segment difference from a use-case difference, its output should say so.

What changes when context is added to the roadmap decision?

Context changed the worked answer from “decide” to “collect more evidence.”

Decision question: Should the team make a broad roadmap commitment to expand the catalog based on this review slice?

Decision fieldText-only viewContext-added view
Visible evidenceTen content or catalog rows, six positive and four negative; several mention more or different contentThe rows mix broad praise, title-specific requests, catalog breadth complaints, cancellations, interactive content, and home-surface clutter
ActionDecide: expand the catalogCollect more evidence: no broad commitment yet
WhyThe same area contains clear demand and clear dissatisfactionThe records do not identify segment, region, plan, title availability, workflow stage, actual viewing behavior, or retention impact
Evidence that would change the callMore reviews in the same themeExact titles and region, search-to-play behavior, completion, repeat use, cancellation or renewal reason, and a follow-up interview sample
Owner and next checkProduct manager after another synthesis passProduct manager with analytics and research after a dated context-enrichment pass
ReversibilityA roadmap commitment creates opportunity costA short evidence-collection task keeps the decision open

The text-only column is a deliberate trap, not a claim about Netflix's actual roadmap. It is what a majority-style synthesis might do when positive ratings, praise, and requests are treated as one kind of demand.

The context-added column is the artifact I would want in a product review: both sides remain visible, the unknowns are named, and the next collection step is small enough to run. A team can still decide later. It just has not earned that decision from this public slice.

When is contradictory feedback strong enough to decide?

Decide only when the evidence shows that the items refer to the same decision object and the difference survives context enrichment.

Use this rule:

Evidence stateActionMinimum record
Same area, same decision object, comparable context, opposing preferences, meaningful impactDecideRaw IDs, context fields, severity, opposing evidence, owner, and expected outcome
Same area, but different segment, workflow stage, use case, or severitySplit the decisionSeparate groups and a decision for each group or stage
Same area and opposite polarity, but missing contextCollect more evidenceNamed missing field, collection method, owner, and date
Weak text, unclear area, no request or outcomeDeferPreserve the raw record and exclude it from roadmap counting
Rating and text disagreeRead the text and preserve bothRating-text mismatch flag, raw text, version, and any edit history

The exception is a safety or access blocker. A single well-proven severe failure may justify immediate containment even when it does not establish broad demand. That is not a contradiction-resolution shortcut. It is a severity rule, and the evidence record should say so.

Productboard's older survey report is useful as a historical warning: it describes a gap between teams believing they are customer-driven and the completeness of their feedback-capture processes, and it identifies segmentation as one of the less-used prioritization dimensions. The 2020 report should not be read as a current benchmark, but its implication fits this artifact: a team cannot claim to understand contradictory demand if it never recorded who, when, where, or under what workflow conditions.

What can public reviews tell you, and what can they not tell you?

Public reviews are useful for finding language, product areas, visible pain, and candidate contradictions. They are weak evidence for segment-level prioritization unless the source includes the missing context or you link it to first-party behavioral data.

In this sample, every selected row carried us and en as country and language fields. Some rows carried app versions. Helpful votes were present, but a helpful vote tells you that other people engaged with the review, not that the underlying issue caused a failed task, churn, or lost revenue. Developer-response fields were null in the selected rows, so the sample cannot show whether a complaint was acknowledged or fixed.

The limits matter because a text-only contradiction can be real at one level and irrelevant at another. “Add more movies” and “remove games from the home screen” may both be negative reviews, but they do not necessarily describe a single product decision. “The app is great” and “the catalog is too small” may coexist because a customer values the service while still requesting an improvement.

Do not solve that ambiguity by asking an AI model to sound more certain. Add the field that would change the decision, then collect it.

How does this evidence standard prepare an AI feedback workflow?

Use the ledger as the input contract and the decision record as the output contract.

Before testing an AI workflow, require it to:

  1. preserve every source ID and raw review link;
  2. show the records it grouped together;
  3. separate rating, sentiment, request type, and product judgment;
  4. mark segment, workflow stage, use case, severity, and outcome as unknown when the source does not provide them;
  5. label a candidate as same-direction, context-unknown, or confirmed conflict;
  6. show the evidence that supports each label;
  7. produce a decide, split, defer, or collect-more-evidence recommendation with an owner and next evidence task.

The related guide on product-feedback synthesis evidence is the natural next read if your team needs a broader commitment gate. For model comparison, use a blind comparison test for AI feedback synthesis. For implementation, the parent workflow guide covers the reviewable system around this evidence contract.

The article can answer the query without a model. The model becomes useful only after the product team has decided which evidence must remain visible.

If you want your team to own this kind of evidence work rather than outsource the judgment, Marius Manolachi's AI learning and consulting work is the relevant next step. The artifact above is complete without that step.

/blog/what-evidence-should-product-teams-collect-about-customer-feedback-contradictions-feedback-decision-record.webp