Independent testing found general research tools still returned a fabricated citation in roughly one of every six answers.
When Merlise reports 80% confidence, it is right about 80% of the time. Calibration error stays under five points.
A check that took an analyst 12 to 22 minutes to do by hand.
Integrates with
A figure, a date, an attribution, a clause. Merlise separates a document into individual factual statements and checks each one on its own terms.
A revenue figure goes to the 10-Q. A holding goes to the case it cites. A result goes to the paper. The right record for the kind of statement, not a general web search.
A reading of how likely each statement is to hold, scaled so the numbers mean what they say. Eighty percent is right about eight times in ten.
Each reading links to the records behind it and the weight each one carried. You can open the source and see why the number is what it is.
When two statements disagree, say $4.1M in one place and $4.2M in another, Merlise marks the conflict instead of averaging it away.
Put a set of claims on a schedule. When a new filing or paper moves a reading past your threshold, you hear about it the same day.
Read a data room, surface the obligations that matter, and flag any clause that departs from your firm's positions. Each holding is checked against the case it cites, so a citation that was never real never reaches the filing.
Catch a metric that moved ten points with nothing in the MD&A to support it, a reference pointing at the wrong note, or a figure in the narrative that does not match the statement. Each finding links back to its source in EDGAR.
Check references against Crossref DOIs, catch transcription errors in reported statistics, and screen submissions before an editor opens the file, inside the workflow you already run.
Match claims against what has already been checked, watch a broadcast or feed as it runs, and send the handful that need a person to the desk. The other quarter million pieces a day stay off it.
Most systems state high confidence and are right far less often than they claim. That gap is what sends people back to recheck the output by hand.
Merlise is measured against the diagonal at right. Across a test set it has not seen, its readings track observed accuracy within five points. A claim marked 80% is one you can treat as 80%, and send the low ones to a person instead of reading them all.
Drafting assistants speed up the writing. Retrieval tools find the documents. Neither tells you whether the claims in front of you are true. Merlise reads a document end to end and returns a sourced verdict on each statement.
Research assistants are good at finding and summarising papers. The step they skip is checking whether a paper’s claims match the citations it rests on.
“It found a revenue figure in our diligence memo that three reviews had let through. It pointed at the 10-Q and showed the gap.”
“Reports go through Merlise before they reach a client, and we read the unsupported claims first.”
A court, a company, a publisher, or an outlet already reached the conclusions shown below. Merlise runs the original document again and produces a sourced, checkable ledger in the time it takes to read this sentence.
A federal appeals court catalogued the fake and misused citations in a set of briefs. We ran the same brief and reached the same appendix.
June 2026NeuroOne filed a 10-Q reporting 2.4 million in quarterly revenue. Ten days later, the company said half a million of it was not real.
May 2026An independent analysis of NeurIPS 2025 found over a hundred fabricated citations across accepted papers, despite one of the field's more selective review processes.
January 2026Ars Technica retracted a story about an AI coding agent after the quotations it attributed to a named engineer turned out to be invented.
February 2026Bring a memo, a filing, or a paper. We will run it through Merlise and walk you through the ledger, claim by claim.