Merlise achieves breakthrough results on SciFact-Open
Merlise
Benchmarks

Results on SciFact-Open

NewUpdated June 2026

SciFact-Open asks whether a system can decide if a scientific claim is supported or refuted by the literature, drawing evidence from an open corpus rather than a fixed candidate set.

Merlise reaches strong results in the open setting. This page describes how the evaluation was run so the numbers can be reproduced rather than taken on faith.

In short

How Merlise performs on open scientific claim verification, and how the evaluation was run.

How it works in Merlise

Inside Merlise this runs as part of the pipeline that turns a document into a ledger. Each claim carries its own state, and the result is always traceable to the records that produced it rather than to a summary you cannot open.

The reasoning step is kept separate from the probability the system reports, so a persuasive explanation never moves a number on its own. Every adjustment is written to the update log in order.

Worked example

state of the art

Label accuracy

65%+

Evidence F1

< 0.05

Calibration (ECE)

Key points

  • The reading traces to a primary record, not a summary.
  • Evidence is weighted by provenance before it counts.
  • Every step is written to an audit trail you can reopen.

Was this page useful? Yes · No