Check whether the claims really conflict.
One report says a service is gaining users; another says fewer people are using it. Before choosing which to trust, check what each counts. The first may count everyone who has ever registered, while the second counts people active this month. Both can be right. A summary that says “the sources disagree” would miss the useful finding: registrations are growing while current use is falling.
Write each claim beside the passage that supports it. Include who or what was studied, when, what was measured, and what it was compared with. For a claim about remote-work productivity, for example, ask which jobs and workers were studied and how productivity was measured. Compare the actual findings before comparing the headlines.
Also check how far the conclusion extends beyond the study. Yarkoni describes how researchers can make broad claims that their statistical analysis does not support, particularly when they intend to generalize beyond the tasks or settings studied.[4] A finding about one group may be useful without answering the question for every group.
First check whether two reports are answering the same question.
Keep the original findings beside your summary.
When you combine several reports, it is tempting to give them the same wording. That makes a summary easier to read, but can hide an important difference. Keep the original passage, number, unit, and context beside your shortened version so you can see what changed.
A triangulation protocol is a structured way to compare findings from different methods. It distinguishes agreement, partial agreement, silence, and disagreement.[1] “Silence” means one set of findings does not address something the other covers. A survey that never asked about waiting times does not contradict an interview that describes a long wait. The interview may add a part of the experience the survey missed.
Put the findings side by side before writing your conclusion. O’Cathain and colleagues describe tables that help researchers see these relationships and examine cases that do not fit an emerging explanation.[2] Add a question to your comparison: could both findings be true, and under what conditions? This gives you something specific to investigate.
Compare like with like, and trace where the evidence came from.
Check definitions and calculations first. Does “users” mean accounts, people, or sessions? Does a percentage use everyone invited or only those who replied? Then compare dates, locations, product versions, study entry rules, and the time allowed for an outcome to occur. Convert units where appropriate, but do not make different measures look equivalent by giving them the same label.
Consider what each source can establish. A forecast estimates what may happen; a final report records what happened. One person’s account may document a failure without telling you how common it is. Administrative records count reported cases but may miss people who never contacted the service. These sources can inform the same question while answering different parts of it.
Trace repeated claims back to their evidence. Three reports may all quote one survey. That is three publications, but still one measurement. Mark shared surveys, datasets, and samples, and distinguish a new analysis from newly collected evidence. Repetition alone does not make a finding independently confirmed.
When three reports cite the same original record
An illustrative example: three publications, one measurement
Three reports quoting one survey still rely on one measurement.
Name the difference you need to investigate.
Use a plain description of the remaining problem. Perhaps the reports define success differently, study different groups, cover different periods, or measure the outcome in different ways. Perhaps they observe the same result but disagree about its cause. Or they agree on the facts but differ on whether the result is acceptable.
The description should help you choose the next check. Different definitions call for a comparison of what was counted. Different periods call for matched dates. A possible measurement error calls for checking the data or using another measure. If two explanations predict different outcomes, look for a way to test those predictions. If the dispute is about an acceptable trade-off, make that judgment explicit; another search may not settle it.
Several differences can matter at once. A later study may use both a different group and a stricter definition of success. Keep both possibilities open until you can check them. Also record when a written policy differs from what people experienced in practice. Update your explanation as you learn, keeping a short note of what each check ruled out.
Keep the original passages beside this comparison. Revise your explanation when a check reveals something new.
Ask what an average would hide.
Combining results can help when the studies are comparable. But first look at their direction, size, uncertainty, and the people and conditions studied. An average improvement may be of little comfort if the service works worse for the group you need to support.
Cochrane’s guidance warns that a random-effects meta-analysis does not explain away differences between studies. This method allows for varying effects when estimating an average, but that average can mislead when studies point in different directions.[3] A side-by-side comparison may be more useful than one combined number. If you notice a possible subgroup explanation only after seeing the results, treat it as a hypothesis to check, not an established rule.
Ask what your decision needs to know. A service team may need to understand why completion falls on older phones or in one language, even if the overall rate looks good. Report the conditions where the result appears to hold and where it changes. Say when there are too few observations to tell whether a difference is reliable.
Find a check that could separate the explanations.
Suppose one team thinks people abandon an application because the instructions are unclear. Another thinks the wait is too long. Both explanations fit a low completion rate. They make different predictions if you simplify the instructions while keeping the wait the same, or shorten the wait without changing the wording. Write down those predictions before examining new results.
Researchers sometimes design such a test together even when they favor competing theories. This is called adversarial collaboration. A 2025 study of two theories of consciousness agreed on predictions and tests in advance, using several laboratories and methods. Its results supported some predictions and challenged important claims of both theories.[5] The shared test clarified the disagreement without declaring a simple winner.
In everyday research, ask a well-informed critic to help improve both explanations and choose a useful check. That might be a record, a field visit, a comparison case, or a small intervention. Agree on what result would weaken each explanation. If no feasible check can distinguish them, say so and consider an action that can be revised as more evidence becomes available.
What would you expect to observe if each explanation were right?
Use AI to find differences, then verify them.
AI can help compare passages, spot changes in terminology, trace repeated citations, and suggest possible explanations. Ask it to return the exact claims, supporting passages, suspected difference, and what still needs checking. Treat that output as a list of leads to review against the sources.
One research method, Chain-of-Verification, drafts an answer, creates checking questions, answers them separately, and then revises the response. It reduced hallucinations on the tasks tested.[6] A useful adaptation for research is to separate mapping the disagreement from checking it. In the checking pass, return to original dates, quotations, samples, and calculations rather than asking only whether the first summary sounds reasonable.
Debate between models has improved reasoning and factual answers in some evaluations.[7] Other experiments found that one deliberately misleading agent could persuade a group toward a wrong answer; more agents or debate rounds did not reliably fix this.[8] Agreement between models is therefore not enough. Verify important conclusions with source material or an appropriate real-world check, and keep unresolved objections visible.
Give the reader the answer you can support.
“The evidence is mixed” tells the reader too little. Explain which findings are supported, which apply only under particular conditions, and which remain unresolved. Distinguish a disagreement about facts from a disagreement about what is acceptable. Mention missing groups or outcomes when they limit the answer.
Connect that answer to the decision. What can reasonably be done now? What should be easy to reverse? What needs a further check, and who will do it by when? If you recommend acting before the uncertainty is resolved, explain the cost of waiting and the possible cost of being wrong. Urgency does not make the evidence stronger.
The useful result may be a narrower answer than you hoped for: a service works in one setting, two reports count different things, or the current evidence cannot distinguish two causes. Explain how you reached that answer and what could change it. A disagreement becomes useful when it shows the reader what to investigate next.

