Skip to content
Climate Claim Checker

COMP90042 Natural Language Processing · The University of Melbourne · 2024, Semester 1

Can a student system from 2024 fact-check a climate claim?

Type a statement about climate science. The original pipeline searches Wikipedia passages for evidence using TF-IDF, then a Transformer trained from scratch labels the claim supported, refuted, disputed or not enough info. It is re-run faithfully here, weak scores included.

Dev claim 752Open
“[South Australia] has the most expensive electricity in the world.”

Carl Albert Unbehaun (1851 -- 5 February 1924) was an electrical engineer in South Australia.

[citation needed] South Australia has the highest retail price for electricity in the country.✓ gold

"South Australia has the highest power prices in the world".✓ gold

The class are the first electric trains to operate in South Australia.

Annotators:SupportsModel:Supportsevidence F 0.67 (4 retrieved, 2 of 2 gold)

The assignment

What the coursework asked for

The task, paraphrased: build a system that, given a claim about climate science, finds supporting or contradicting passages in a knowledge source of about 1.2 million Wikipedia sentences and classifies the claim. Pretrained language models and embeddings were not allowed. Everything had to be trained from scratch on 1,228 labelled claims.

1 · Retrieve

Return a small set of evidence passages for each claim, from a corpus of 1,208,827 passages.

2 · Classify

Label the claim SUPPORTS, REFUTES, NOT_ENOUGH_INFO or DISPUTED given that evidence.

3 · Be scored

Evidence F-score against human-annotated passages, label accuracy, and their harmonic mean as the headline metric.

What the team built

A two-stage pipeline, kept exactly as it was

Every step below is a TypeScript port of the 2024 notebook. Each was checked against the original Python, down to NLTK’s stemmer and the order in which numpy breaks ties.
  1. 01

    Clean

    Expand contractions, lowercase, strip punctuation, tokenise, drop stopwords and Porter-stem every word.

  2. 02

    Tag

    A claim's tags are its sorted stems. Each passage was tagged in advance with its ten highest-scoring TF-IDF keywords.

  3. 03

    Score

    Compare tags with a 1,000-term TF-IDF vectorizer (cosine similarity) plus the share of tags in common.

  4. 04

    Select

    Keep passages with cosine > 0.55 and overlap > 0.5 that share the most tags (up to six). If none qualify, take the top six by score.

  5. 05

    Classify

    A Transformer trained from scratch reads claim + evidence stems (128 tokens) and picks one of four verdicts.

Walk through each stage with a real claim on the method page.

Key results

Honest numbers: retrieval was the weak link

The team’s report scored the system on the validation (dev) set and on the hidden test set of the course leaderboard. Retrieval found the right passages for only 13 of 154 dev claims.
Results reported in the team’s 2024 report (Table 2)
Reported in 2024Evidence FAccuracyHarmonic mean
Validation (154 dev claims)0.042990.558440.07984
Test (course leaderboard)0.033100.407900.06120

Reproduced exactly

Evidence F = 0.04299

Re-running the retrieval rule over all 1.19M passages gives the reported dev F-score to every printed digit, and the notebook's own F = 0.0112 too.

Retrained, weights were never saved

Accuracy 38.3% (report: 55.8%)

Same code, data and hyper-parameters. The training curves track the 2024 log closely, but on retrieved evidence the model lands below the 44.2% always-"supports" baseline. With gold evidence it reaches 57.8%.

Matches the report

Corpus statistics

The passage-length table in the report is recomputed from the corpus count for count (1,208,827 passages).

Revisiting the code in 2026

Four things the re-run revealed

Finding 1

The Transformer never attended across words

The encoder was built without batch_first=True but fed (batch, sequence) tensors. So self-attention mixed the 16 claims in a mini-batch, never the words within one claim. Run on a single claim, the whole network reduces exactly to a lookup table of per-word scores. That makes every prediction on this site fully explainable.

Finding 2

The committed notebook is not the submitted code

The saved 2024 results come from a rule scoring similarity + overlap. The notebook on GitHub adds the similarity twice, and its last evaluation cell passes the raw claim text instead of its stems. That is why it printed F = 0.011 while the report says 0.043. All three variants are reproduced here.

Finding 3

A stray Arabic word in 14% of evidence tags

Keyword extraction took numpy’s top-10 TF-IDF features. For short passages, numpy pads that list with zero-weight features in its own tie order. The last word of the 20,000-term vocabulary, محمد, is alphabetic, so it landed in 167,592 passages’ tags.

Finding 4

Two of the four verdicts are never predicted

With 42% of training claims labelled SUPPORTS and only 10% DISPUTED, the retrained model only ever answers Supports or Not enough info. That is the class-imbalance problem the report’s discussion warned about.

About this project

COMP90042 Natural Language Processing

The University of Melbourne, 2024, Semester 1. Group project by Wed5PM Group 1.

Team

  • Sunchuangyu (Rin) Huang

    System design, preprocessing, TF-IDF evidence retrieval, report and presentation

  • Wei Zhao

    Transformer and LSTM classifiers: design, training, evaluation and model selection

  • Xuan Wang

    Retrieval testing and debugging, literature review, report and presentation

The 2026 revival (this website, the reproducible build scripts and the TypeScript ports) was made by Rin Huang. The original notebook, report and code are kept unchanged in the repository for reference. The assignment’s datasets and specification belong to the subject and are not hosted here.

Stack, then and now

2024
Python in Google Colab: pandas, NLTK, contractions, scikit-learn TF-IDF, PyTorch (Transformer and LSTM)
2026
Next.js 16 and TypeScript ports of every step, a read-only SQLite index (39,666 passages), uv scripts that re-run the original Python, and vitest parity tests
View the code on GitHub