Results
What the numbers say, and how closely they were reproduced
Data
An imbalanced task
Label balance
SUPPORTS dominates and DISPUTED is rare in both splits. A model can score 44% on dev by always answering "supports".
Training set · 1,228 claims
- Supports42% · 519
- Refutes16% · 199
- Not enough info31% · 386
- Disputed10% · 124
Dev set · 154 claims
- Supports44% · 68
- Refutes18% · 27
- Not enough info27% · 41
- Disputed12% · 18
Evidence passage lengths
Most passages are one Wikipedia sentence of 11–50 words. The 13,578 passages of five words or fewer are easy to match spuriously on cosine similarity, as the report notes.
| Words | Passages (recomputed) | Report |
|---|---|---|
| ≤ 5 | 13,578 | 13,578 ✓ |
| 6–10 | 180,771 | 180,771 ✓ |
| 11–20 | 547,481 | 547,481 ✓ |
| 21–50 | 451,498 | 451,498 ✓ |
| 51–100 | 15,217 | 15,217 ✓ |
| 101–200 | 270 | 270 ✓ |
| > 200 | 12 | 12 ✓ |
Retrieval
Why retrieval was the bottleneck
The retrieval threshold trade-off
One setting at a time is swept on the dev set while the others keep their 2024 values (dotted line). With the minimum cosine similarity, the best dev F is 0.0520 at 0.30, against 0.0430 for the original setting. No setting lifts F much above 0.05: the passages are simply not being found.
Share of claims that fall back to “top n by score”
Classification
Transformer vs LSTM, then and now
Training curves: Transformer vs LSTM
The retrain (solid) tracks the curves printed in the 2024 notebook (dashed) closely. The Transformer learns from the first epoch; the LSTM sits at the majority-class rate until epoch 6. Validation here uses gold evidence, which is easier than the retrieved evidence used at test time.
Confusion matrix on the dev set
The model only ever predicts SUPPORTS or NOT_ENOUGH_INFO. Refuted and disputed claims are always missed, whichever evidence it is given.
Accuracy 38.3% on 154 dev claims. Dev claims predicted in batches of 16, in file order, from the 2024 retrieved evidence, as the notebook does.
| gold ↓ / predicted → | Supports | Refutes | Not enough info | Disputed |
|---|---|---|---|---|
| Supports | 42 | 0 | 26 | 0 |
| Refutes | 20 | 0 | 7 | 0 |
| Not enough info | 24 | 0 | 17 | 0 |
| Disputed | 17 | 0 | 1 | 0 |
Outlined cells are correct predictions. Stronger blue means more claims. S = supports, R = refutes, NEI = not enough info, D = disputed.
Parity
Reproduction checklist
Dev evidence F-score (submitted retrieval)
ExactReport, Table 2
- 2024
- 0.04299
- 2026 re-run
- 0.04299
Dev evidence F-score, notebook's final cell
ExactNotebook cell 58
- 2024
- 0.011205
- 2026 re-run
- 0.011205
eval.py harmonic mean (TypeScript port)
ExactNotebook cell 58
- 2024
- 0.021717
- 2026 re-run
- 0.021717 from the same F and accuracy
Retrieved passages for claim-752
ExactNotebook cell 26
- 2024
- 4 passages
- 2026 re-run
- same 4, same order
Saved 2024 retrieval lists re-derived
CloseTeam's saved output files
- 2024
- 307 dev + test claims
- 2026 re-run
- 303 with identical scores (244 identical passage sets); 4 not explained
Claim tags and keyword extraction
ExactNotebook cells 10, 16; saved tags
- 2024
- printed examples + 1.19M tagged passages
- 2026 re-run
- identical (TS port, numpy tie order included)
Passage-length table
ExactReport, Table 1
- 2024
- 1,208,827 passages in 7 buckets
- 2026 re-run
- every bucket identical
Transformer validation accuracy (gold evidence)
CloseReport §3.3.1, cell 43
- 2024
- 57.14% final epoch · 60.39% best (epoch 8)
- 2026 re-run
- 54.55% final epoch · 57.79% best (epoch 5)
Dev accuracy on retrieved evidence
DiffersReport, Table 2
- 2024
- 55.84%
- 2026 re-run
- 38.31% (retrained)
| Check | 2024 source | 2024 value | 2026 re-run | Status |
|---|---|---|---|---|
| Dev evidence F-score (submitted retrieval) | Report, Table 2 | 0.04299 | 0.04299 | Exact |
| Dev evidence F-score, notebook's final cell | Notebook cell 58 | 0.011205 | 0.011205 | Exact |
| eval.py harmonic mean (TypeScript port) | Notebook cell 58 | 0.021717 | 0.021717 from the same F and accuracy | Exact |
| Retrieved passages for claim-752 | Notebook cell 26 | 4 passages | same 4, same order | Exact |
| Saved 2024 retrieval lists re-derived | Team's saved output files | 307 dev + test claims | 303 with identical scores (244 identical passage sets); 4 not explained | Close |
| Claim tags and keyword extraction | Notebook cells 10, 16; saved tags | printed examples + 1.19M tagged passages | identical (TS port, numpy tie order included) | Exact |
| Passage-length table | Report, Table 1 | 1,208,827 passages in 7 buckets | every bucket identical | Exact |
| Transformer validation accuracy (gold evidence) | Report §3.3.1, cell 43 | 57.14% final epoch · 60.39% best (epoch 8) | 54.55% final epoch · 57.79% best (epoch 5) | Close |
| Dev accuracy on retrieved evidence | Report, Table 2 | 55.84% | 38.31% (retrained) | Differs |
Why the accuracy differs. The notebook never saved the trained weights, so the classifier was retrained with the same code, data and hyper-parameters (seed 42, on a CPU instead of Colab). Its training curves match the 2024 log closely. On retrieved evidence, though, it scores 38.3% rather than the reported 55.8%. The notebook itself printed 35.1% in its final cell, and its closing note says the report numbers came from a model the team was still tweaking. With 154 claims, each one is worth 0.65 percentage points, so small differences between trained models move accuracy visibly.
Why the re-run does not match every saved list. pandas’ default quicksort is not stable, so when several passages have exactly the same score, which of them makes the cut depends on the sort’s internals (library version, row layout) rather than on the data. The re-run returns passages with the same scores as the saved lists for 303 of 307 claims. 244 of those are the very same passages; the other 59 differ only in which tied passage was kept. The dev F-score is identical. The remaining 4 saved lists (claim-1160, claim-540, claim-2329, claim-1582) are not explained by the rule: they look like top-six fallback lists even though passages pass the filter, which suggests some extra 2024 logic that the notebook does not contain. Their selection path is therefore shown as “not recorded”.
Why the training curves differ. The 2024 log peaked at 60.39% in epoch 8 and ended at 57.14%, the figure the report quotes. The retrain peaked at 57.79% in epoch 5 and ended at 54.55%, so it trails the 2024 run by about 2.6 percentage points either way. The site uses the best epoch, as the notebook’s model selection did.