(01)

Start with a claim you can check.

You ask an AI tool to research a question. It returns a polished report with links, clear headings, and a confident conclusion. Then someone asks, “Which source supports this sentence?” You open the link and find that it supports only part of what the report says.

The problem is easy to miss when you review a whole paragraph at once. One sentence may contain several claims about different people, dates, or outcomes. SimpleQA, a benchmark for factual questions, keeps answers short so they are easier to assess.[1] The same idea is useful when checking a report: separate its important claims and check them one at a time.

For each claim, keep the source and the exact passage or observation that supports it. Also note anything that disagrees with it. You can then explain why you trust a conclusion without having to repeat the entire search.

A convincing sentence still needs a source that supports it.
(02)

Write down what you need to find out.

A folder full of papers tells you what you have collected. It does not tell you what you know. Before searching in depth, write down the claims you may need to test. These are questions to investigate, even if you first write them as statements.

Suppose you are studying a service that uses AI to decide who qualifies for support. You might need to check whether people understand a refusal, know how to appeal it, and can get an error corrected without unreasonable effort. The notice shows what the service told them. Interviews show what they understood. Case records show whether an appeal changed the outcome. You need these different sources because they answer different parts of the question.

Be specific enough that you could discover you are wrong. “AI improves research” leaves too much open. A claim such as “For this project, AI helped us find more relevant sources within two hours than our usual search method” gives you something to compare. You would still need to check the sources before including them. Name the people or tasks involved, the comparison, the outcome, and the limits of the claim.

(03)

Ask what each source can actually tell you.

A source can be reliable and still be the wrong evidence for your claim. A policy document tells you what the rules say; it does not tell you how people experience those rules. A product announcement records what a company says it released; it does not independently establish how well the product works. An interview tells you about one person’s account, not how common that experience is.

Look for the kind of evidence your question needs. To check a measured effect, read the original study. To understand a process, examine records or observe it. Use review articles to find the main findings and disagreements in a field, and commentary to find leads worth following. A forum post, for example, may reveal a failure worth investigating without telling you how often it happens.

AI can help with this search. Anthropic describes a research system that divides a question among agents, which search different parts of it in parallel.[2] That can bring more candidate sources into view. Each candidate still needs to be opened and checked before you use it to support a claim.

(04)

Keep a small evidence table.

A citation at the end of a paragraph may support one sentence and say nothing about the next. A simple table makes that gap easier to see: give each important claim its own row. This is sometimes called a claim ledger. A spreadsheet is enough.

Start with the claim, a source link, and the passage or observation you are relying on. Record when the source was published and which people or setting it covers. Add who checked it, how confident they are, and any conflicting evidence. A status such as proposed, supported, contested, rejected, or unresolved tells the next reader where the work stands.

Use this table as you write. If two sources seem to agree, check that they support the same statement. If AI produces a new summary, compare its claims with the relevant rows. If a source no longer supports a sentence in the draft, revise or remove that sentence. The table is useful only if it stays connected to the report.

(05)

Work out why sources disagree.

Imagine a study says a tool saves time, while interviews say it creates extra work. Both accounts may be accurate. The study may cover a short task under controlled conditions, while the interviews include the time spent correcting mistakes. Before choosing one account, check whether they cover the same people, dates, tasks, and measures.

Keep both claims and their strongest evidence in your notes. Write down a possible reason for the difference and what you could check next. You may conclude that the tool saves time in one setting but not another. If you cannot explain the difference, leave it unresolved rather than blending the findings into a reassuring average.

An AI agent also needs a way to stop and reconsider its approach. If an early assumption is wrong, more searches can take it further in the wrong direction. Anthropic’s guidance on agents recommends checking results against the environment and setting stopping conditions.[3] In a research task, that means checking what the sources actually say and returning to a person when the next step needs judgment.

Before choosing between two findings, check whether they answer the same question.
(06)

Give AI work you can inspect.

Ask AI to suggest search terms, find related terminology, group similar findings, compare tables, or flag sentences that may lack support. It can also suggest evidence that would challenge your answer. These are concrete tasks whose results you can inspect, and they can reduce the time spent sorting material.

A fluent explanation is not a record of what happened. Check the source rather than treating the model’s account as evidence. Even providing the original document does not guarantee the model will use the relevant passage. The study “Lost in the Middle” found that performance could change with the position of relevant information in a long input.[4]

Keep the inputs and outputs so you can check or repeat the work. The researcher still decides which question to ask, which evidence to accept, how to handle ambiguity, and how confident the conclusion should be. If AI proposes one of these judgments, ask for its basis and review it before using it.

(07)

Decide when another search is worth doing.

After a while, new search results may repeat what you already have. Ten reports that quote the same original study do not give you ten independent confirmations. Ask whether another search is likely to add evidence that could change the answer, rather than simply making the source list longer.

The amount of checking should reflect the cost of being wrong. For some factual claims, an authoritative original record and an independent confirmation may be enough. A safety-critical claim may need replication, specialist review, and testing where the result will be used. An early exploratory finding can remain provisional if you say so. Decide what your claim requires before the search grows.

Sometimes reading more is not the check you need. PaperBench asks AI agents to reproduce research from papers, and even its strongest evaluated systems completed only part of the work.[5] Describing a result and reproducing it are different achievements. If your decision depends on whether a finding holds in a new setting, try the relevant test there or observe the process directly.

(08)

Tell the reader how far the answer reaches.

End with an answer the reader can use: what you found, the evidence that matters most, and how confident you are. Include the strongest conflicting finding and say which people or settings you did not examine. Then name the next check that would matter if the reader wants to act.

For example, “The service was easier to use in these interviews, but we have not tested the appeal process” gives a team a useful next step. It supports a limited conclusion without suggesting the whole service has been validated. Saying what remains unknown helps the reader avoid using a finding for a decision it cannot support.

When someone asks why a conclusion is in the report, you should be able to show the claim, the evidence, and the reasoning that connects them. That record also makes the research easier to update: a new source can change a specific judgment without forcing you to start again.