Skip to content
Climate Claim Checker

Results

What the numbers say, and how closely they were reproduced

All charts are computed from the original code and data by the build scripts. Nothing here is re-estimated or tuned. Where the 2026 re-run differs from 2024, the table at the bottom says so and why.

Data

An imbalanced task

Label balance

SUPPORTS dominates and DISPUTED is rare in both splits. A model can score 44% on dev by always answering "supports".

Training set · 1,228 claims

  • Supports42% · 519
  • Refutes16% · 199
  • Not enough info31% · 386
  • Disputed10% · 124

Dev set · 154 claims

  • Supports44% · 68
  • Refutes18% · 27
  • Not enough info27% · 41
  • Disputed12% · 18
Gold labels from the course's train-claims.json and dev-claims.json.

Evidence passage lengths

Most passages are one Wikipedia sentence of 11–50 words. The 13,578 passages of five words or fewer are easy to match spuriously on cosine similarity, as the report notes.

Evidence passages by length in words
WordsPassages (recomputed)Report
≤ 5
13,578
13,578 ✓
6–10
180,771
180,771 ✓
11–20
547,481
547,481 ✓
21–50
451,498
451,498 ✓
51–100
15,217
15,217 ✓
101–200
270
270 ✓
> 200
12
12 ✓
Recomputed from all 1,208,827 passages (words = whitespace-separated tokens) and checked against the report's Table 1.

Retrieval

Why retrieval was the bottleneck

Retrieval compares tags, not meaning. A passage that shares the stems “sea” and “level” with the claim scores well whether or not it is about sea-level rise. The tags also carry noise: the Arabic word محمد sits in 167,592 of 1,192,697 passages’ tags because of how numpy breaks ties.

The retrieval threshold trade-off

One setting at a time is swept on the dev set while the others keep their 2024 values (dotted line). With the minimum cosine similarity, the best dev F is 0.0520 at 0.30, against 0.0430 for the original setting. No setting lifts F much above 0.05: the passages are simply not being found.

Vary
Rule
PrecisionRecallF-score

Share of claims that fall back to “top n by score”

Computed by scripts/build_retrieval.py over all 1.19M passages for the 154 dev claims, scored with the course's eval.py definitions.

Classification

Transformer vs LSTM, then and now

Training curves: Transformer vs LSTM

The retrain (solid) tracks the curves printed in the 2024 notebook (dashed) closely. The Transformer learns from the first epoch; the LSTM sits at the majority-class rate until epoch 6. Validation here uses gold evidence, which is easier than the retrieved evidence used at test time.

Metric
Transformer · 2026 retrainTransformer · 2024 notebookLSTM · 2026 retrainLSTM · 2024 notebook
Source: training logs printed by notebook cells 43–44 (2024) and scripts/train_classifier.py (2026: same code and hyper-parameters, seed 42, CPU). 10 epochs, batch 16, Adam lr 1e-4.

Confusion matrix on the dev set

The model only ever predicts SUPPORTS or NOT_ENOUGH_INFO. Refuted and disputed claims are always missed, whichever evidence it is given.

Predictions

Accuracy 38.3% on 154 dev claims. Dev claims predicted in batches of 16, in file order, from the 2024 retrieved evidence, as the notebook does.

Confusion matrix: rows are gold labels, columns are predicted labels
gold ↓ / predicted →SupportsRefutesNot enough infoDisputed
Supports420260
Refutes20070
Not enough info240170
Disputed17010

Outlined cells are correct predictions. Stronger blue means more claims. S = supports, R = refutes, NEI = not enough info, D = disputed.

Retrained Transformer (best epoch by validation accuracy). Rows: annotators' labels; columns: model verdicts.

Parity

Reproduction checklist

Each row is enforced by an automated test in the repository (vitest for the TypeScript ports, assertions in the uv build scripts for the Python re-runs).
  • Dev evidence F-score (submitted retrieval)

    Exact

    Report, Table 2

    2024
    0.04299
    2026 re-run
    0.04299
  • Dev evidence F-score, notebook's final cell

    Exact

    Notebook cell 58

    2024
    0.011205
    2026 re-run
    0.011205
  • eval.py harmonic mean (TypeScript port)

    Exact

    Notebook cell 58

    2024
    0.021717
    2026 re-run
    0.021717 from the same F and accuracy
  • Retrieved passages for claim-752

    Exact

    Notebook cell 26

    2024
    4 passages
    2026 re-run
    same 4, same order
  • Saved 2024 retrieval lists re-derived

    Close

    Team's saved output files

    2024
    307 dev + test claims
    2026 re-run
    303 with identical scores (244 identical passage sets); 4 not explained
  • Claim tags and keyword extraction

    Exact

    Notebook cells 10, 16; saved tags

    2024
    printed examples + 1.19M tagged passages
    2026 re-run
    identical (TS port, numpy tie order included)
  • Passage-length table

    Exact

    Report, Table 1

    2024
    1,208,827 passages in 7 buckets
    2026 re-run
    every bucket identical
  • Transformer validation accuracy (gold evidence)

    Close

    Report §3.3.1, cell 43

    2024
    57.14% final epoch · 60.39% best (epoch 8)
    2026 re-run
    54.55% final epoch · 57.79% best (epoch 5)
  • Dev accuracy on retrieved evidence

    Differs

    Report, Table 2

    2024
    55.84%
    2026 re-run
    38.31% (retrained)

Why the accuracy differs. The notebook never saved the trained weights, so the classifier was retrained with the same code, data and hyper-parameters (seed 42, on a CPU instead of Colab). Its training curves match the 2024 log closely. On retrieved evidence, though, it scores 38.3% rather than the reported 55.8%. The notebook itself printed 35.1% in its final cell, and its closing note says the report numbers came from a model the team was still tweaking. With 154 claims, each one is worth 0.65 percentage points, so small differences between trained models move accuracy visibly.

Why the re-run does not match every saved list. pandas’ default quicksort is not stable, so when several passages have exactly the same score, which of them makes the cut depends on the sort’s internals (library version, row layout) rather than on the data. The re-run returns passages with the same scores as the saved lists for 303 of 307 claims. 244 of those are the very same passages; the other 59 differ only in which tied passage was kept. The dev F-score is identical. The remaining 4 saved lists (claim-1160, claim-540, claim-2329, claim-1582) are not explained by the rule: they look like top-six fallback lists even though passages pass the filter, which suggests some extra 2024 logic that the notebook does not contain. Their selection path is therefore shown as “not recorded”.

Why the training curves differ. The 2024 log peaked at 60.39% in epoch 8 and ended at 57.14%, the figure the report quotes. The retrain peaked at 57.79% in epoch 5 and ended at 54.55%, so it trails the 2024 run by about 2.6 percentage points either way. The site uses the best epoch, as the notebook’s model selection did.