Merlise achieves breakthrough results on SciFact-Open
Merlise
Benchmarks

Results on SciFact-Open, in context

Merlise ResearchJune 10, 20267 min read

SciFact-Open asks a system to decide whether a scientific claim is supported or refuted by the literature, drawing evidence from an open corpus rather than a fixed shortlist. It is closer to the real task than the closed setting, and harder.

Open retrieval is the realistic setting: the evidence is somewhere in the literature, not handed to you.

What the open setting measures

In the closed setting a system chooses among a small set of candidate abstracts. In the open setting it has to find the evidence first, which means retrieval errors and label errors compound.

We report label accuracy alongside evidence selection, because getting the verdict right for the wrong reason is not a result worth trusting.

How the evaluation was run

The corpus, the retrieval depth, and the scoring are described so the numbers can be reproduced rather than taken on faith. Calibration is reported next to accuracy, so the confidence behind each verdict is legible.

Where the hard cases remain

Claims that depend on a specific table value, or on a figure that the source reports only in a chart, are still the frontier. We would rather flag those for a person than guess, and the meta state records that the evidence was thin.

Takeaways

  • Open retrieval finds evidence rather than choosing from a shortlist.
  • We report calibration next to accuracy.
  • Thin evidence is flagged, not guessed.