Skip to content
Climate Claim Checker

Method

One claim, through every stage of the pipeline

Following dev claim 752 from raw text to verdict. Every value on this page is computed by the TypeScript ports when the site is built, the same code the Try-it page runs.
“[South Australia] has the most expensive electricity in the world.”

Preprocess

preprocess_and_tokenize · notebook cell 9

Contractions are expanded with the contractions package. Text is lowercased, and every ASCII punctuation mark becomes a space (so sentence.abc splits in two). NLTK tokenises, English stopwords and non-alphabetic tokens are dropped, and the Porter stemmer reduces what is left.

tokens
south australia has the most expensive electricity in the world
kept
south australia expensive electricity world
stems
south australia expens electr world

Tag the claim and the evidence

cells 10, 14–19

A claim’s tags are simply its stems, sorted: australia electr expens south world. The evidence was tagged once, offline, across all 1.2M passages: each passage keeps its ten highest-weighted features under a TF-IDF model with 20,000 uni- to tri-gram features, and only single words survive. Here is the first gold passage for this claim:

[citation needed] South Australia has the highest retail price for electricity in the country.

  • citat 0.41
  • retail 0.36
  • price 0.33
  • highest 0.30
  • need 0.30
  • electr 0.30
  • australia 0.27
  • countri 0.24
  • south 0.22

A quirk: numpy’s argsort pads short passages’ top ten with zero-weight features, in an order that depends on its unstable quicksort. The last vocabulary entry, محمد, is alphabetic, so it ended up in 167,592 passages’ tags (14.1%). For “Their number was increased to forty in 1855.” the tags are forti, increa, number, محمد (weight 0). The port reimplements numpy’s introsort so these tags come out identical.

Score every passage

find_top_evidence · cell 24

Claim tags and passage tags are vectorised by a second, 1,000-term TF-IDF model (the tags are stemmed again on the way in). Each passage then gets three numbers:

  • cosine

    cosine similarity of the two tag vectors

  • overlap

    shared tags ÷ the smaller tag set

  • score

    cosine + overlap (the notebook adds the cosine twice)

Scores of the passages selected for claim 752
Selected passagecosineoverlapscoreshared
Carl Albert Unbehaun (1851 -- 5 February 1924) was an electrical engineer in South Australia.0.6740.6001.2743
[citation needed] South Australia has the highest retail price for electricity in the country.gold ✓0.6050.6001.2053
"South Australia has the highest power prices in the world".gold ✓0.6000.6001.2003
The class are the first electric trains to operate in South Australia.0.5950.6001.1953

Select

cell 24

Keep passages with cosine > 0.55 and overlap > 0.5. Sort them by score and keep only those sharing the most tags with the claim, at most 6. For this claim 8 passages passed and the 4 sharing 3 tags were kept. They include 2 of the 2 gold passages (F = 0.67). When nothing passes the filter, the top six by score are returned instead. That happened for 47% of dev claims.

Classify

cells 33–37, 51

The claim’s stems and the stems of every retrieved passage are joined into one string. The string is wrapped in [CLS] … [SEP], mapped to a 5,886-word vocabulary built from the training claims and their gold evidence, and cut or padded to 128 tokens. This claim uses 37 real tokens and 91 [PAD]s.

model input
south australia expens electr world carl albert unbehaun februari electr engin south australia citat need south australia highest retail price electr countri south australia highest power price world class first electr train oper south australia
Embedding
256-d
Encoder
6 layers × 8 heads
Feed-forward
512
Training
10 epochs, Adam lr 1e-4

The quirk that makes it explainable. PyTorch’s nn.TransformerEncoder expects (sequence, batch, features) unless batch_first=True. The notebook passes (batch, sequence, features), so attention ran across the 16 claims of a mini-batch at each position, never across the words of a claim. The positional encoding likewise marks a claim’s place in the batch, not a word’s place in the sentence. Given one claim on its own, every position is transformed independently and the mean-pooled output becomes

logits = bias + (1/128) × Σt g[tokent]

Here g is a 5,886 × 4 table exported from the retrained weights. It matches PyTorch to within 7.7e-7, so the site needs a 94 KB table, not a 4.7M-parameter network, and can show each word’s exact share of the verdict. The notebook evaluated dev claims in batches of 16, so its predictions also depend on their neighbours in the file. The Explore page shows both protocols.

Data & provenance

  • Claims come from the subject’s public project repository (1,228 train, 154 dev, 153 unlabelled test). The test labels were never released, so test-set results exist only as the leaderboard score in the report.
  • Evidence is a pruned index of 39,666 Wikipedia sentences out of the course’s 1,208,827: every gold passage for the train and dev claims (3,443), every passage either rule could select for any train, dev or test claim (12,350), every passage it could select for the Try-it examples (129), and a seeded random sample (25,000). The sets overlap. The full corpus is not redistributed.
  • What the pruning costs. For every dataset claim and example the pruned index returns the same passages as the full corpus (tested). For new text it may not: on 60 hand-written climate claims that played no part in building the index, it picked the same passages as the full 1.19M search for 39 under the submission rule (35 of 52 on the filtered path, 4 of 8 on the fallback path) and 40 under the notebook rule. The claims are listed in scripts/free_text_claims.py.
  • Model artefacts (vectorizer vocabularies, the token table, retrieval runs and predictions) are derived by the build scripts in the repository and stored in a read-only SQLite file queried on the server.
  • The assignment specification is paraphrased, not reproduced. The team’s own report and notebook are kept in the GitHub repository.