Skip to content
Climate Claim Checker

Methods & decisions

Where every number comes from, and how far to trust it

This page sets out where the data comes from and how the system was evaluated. It states what the evaluation assumes and where it falls short, explains how AI is used on the site, and links to the decision records behind it, weak results included.

Data provenance

  • Claims come from the subject’s public COMP90042_2024 repository: 1,228 training, 154 dev and 153 test claims. The test labels were never released.
  • Evidence is a corpus of 1,208,827 English Wikipedia sentences. The site keeps a pruned index of 39,666 of them: every gold passage for the train and dev claims, every passage either rule could select for a dataset claim or a Try-it example, and a seeded random sample of 25,000 (seed 2024).
  • Model artefacts (vectoriser vocabularies, the classifier’s token table, retrieval runs and predictions) are rebuilt from the original code and data by the uv scripts in scripts/. The build is deterministic and writes one read-only SQLite file.
  • The classifier was retrained in 2026 with the notebook’s code and hyper-parameters (seed 42, PyTorch 2.3.0), because the 2024 weights were never saved.
  • The site displays the dev claims’ text. Train and test claim texts and the full corpus are not redistributed. The data statement has the details.

Method

Two stages, kept exactly as the team built them. Retrieval tags each passage with its top TF-IDF keywords and scores it against the claim’s stems by cosine similarity plus tag overlap. Passages above both thresholds that share the most tags are kept, up to six, with a fallback to the six best scores. A Transformer encoder trained from scratch then reads the claim and evidence stems and picks one of four verdicts.

Every step runs as a TypeScript port that is tested against the original Python, down to NLTK’s stemmer and numpy’s tie-breaking. The pipeline walkthrough follows one claim through every stage.

Evaluation design

  • Unit and sample. The unit is the claim. All results use the 154 dev claims, because the test labels were never released. Each claim moves accuracy by 0.65 percentage points.
  • Metrics. The course’s own metrics come first: evidence precision, recall and F-score per claim, label accuracy, and the harmonic mean of mean F and accuracy. Macro-F1 and per-label recall are added because the labels are imbalanced.
  • Uncertainty. Every performance estimate carries a 95% interval. Numbers quoted only to show that the re-run matches the 2024 report are given as reported. Proportions use the Wilson score interval. Means, F1 and the harmonic mean use a percentile bootstrap that resamples whole claims (10,000 resamples, seed 2026), so a claim’s retrieval and verdict move together.
  • Comparisons are paired. Two systems are compared on the same claims. The accuracy difference gets a paired bootstrap interval, McNemar’s exact test uses only the claims where the two disagree, and Cohen’s h gives the size of the gap.
  • Training-seed spread. The Transformer and the LSTM are each retrained under five seeds with the notebook’s code, so the run-to-run spread sits next to the single-run numbers. The seed-42 run is the site’s model and reproduces its stored predictions exactly.
  • Baselines. Always answering “supports” (44.2% on dev) is the floor. The same classifier with gold evidence (57.8%) shows the cost of retrieval.
  • Checked twice. scripts/stats_reference.py recomputes every interval with numpy, scikit-learn and statsmodels, using a Python port of the same random number generator. The test suite requires the two to agree to twelve decimal places.
  • The LLM harness. A seeded random sample of dev claims (default N = 20, seed 42) goes to the visitor’s model with the same retrieved passages the classifier read, and again with the gold passages. Calls that fail for infrastructure reasons are excluded and counted. Unusable answers are scored as wrong. Macro-F1 averages both systems over one label set: the labels in the gold set or in either system’s verdicts, re-derived on each resample. Citation validity is the share of answers whose cited ids were all among the passages shown. In evidence F, cited ids that were not among the passages shown count as wrong predictions. The classifier’s gold-evidence verdicts come from the epoch chosen on the same dev claims, so the gold condition flatters it. Its bootstrap uses 5,000 resamples, seed 2026.

The results are on the results page and the comparison is on the LLM evaluation page.

Assumptions

  • The dev claims are a fair sample of the claims the system would meet, drawn independently. The bootstrap, the Wilson interval and McNemar’s test all rest on that.
  • The annotators’ labels and gold passages are correct and complete enough. A relevant passage that is not on the gold list counts as a miss.
  • One retrain stands in for the 2024 model. Its training curves track the 2024 log closely, but its accuracy on retrieved evidence differs from the report’s. Retraining under five seeds shows how far another run could move it.
  • For the LLM, the instruction to judge only from the passages is followed. Nothing can enforce that, and a model may still use what it already knows.

Limitations

  • The dev set chose the best training epoch and then reported the results, so the classifier’s dev numbers are optimistic.
  • With 154 claims the intervals are wide, and for the rarer labels (18 to 68 claims) wider still.
  • The retrained classifier scores 38.3% on retrieved evidence where the report gave 55.8%. The report’s model cannot be recovered, so the gap cannot be explained further.
  • Free-text claims search a 39,666-passage subset. On 60 held-out sentences it matched the full search 39 times.
  • No LLM results ship with the site. They depend on the visitor’s model, the date and the sample, they vary between runs, and the public dataset may be in a model’s training data.

What I'd change

  • Evaluate retrieval on its own first, with recall at k and a BM25 baseline, before building a classifier on top of it.
  • Hold out part of the training claims for model selection and use the dev set once.
  • Fit a TF-IDF and logistic regression baseline, weight the classes, and keep only models that beat the baseline on a paired test.
  • Report every number with an interval from the start, and retrain over several seeds so one lucky run cannot decide a comparison.

AI use statement

AI is optional on this site. Every page works without it, and nothing is sent to any AI provider unless you add your own API key in AI settings (the key icon in the header).

What AI does here

  • Second opinion. On a claim page or a Try-it result, an LLM can read the claim and the evidence passages and return a verdict, the ids of the passages it relied on and a one-line rationale.
  • Evaluation harness. The LLM evaluation page runs the same task on a sample of dev claims and compares the answers with the 2024 classifier’s.

What AI never does here

  • It never changes the 2024 system’s results or any number in the database.
  • It never runs without your key and your click.
  • It is never presented as a fact-check of a real-world claim.
  • Its output is never shown without an “AI-generated” label.

What is sent, and where

Your browser sends the claim text, the evidence passages with their ids and a fixed instruction straight to the provider you chose: Anthropic (claude-haiku-4-5 by default, or claude-sonnet-5-5) or OpenAI (gpt-5-mini by default). Your key goes only in that request’s headers. It is kept in this tab’s sessionStorage, or in localStorage if you tick “remember on this device”, and “forget keys” deletes it. It is never sent to this site’s server, never logged and never written to the audit log. The provider’s own terms and data retention apply to what you send.

Human in the loop and audit trail

Every call, including failed ones and calls you stop while they are in flight, is recorded in the AI audit log in your browser. A record holds the prompt, the output, the model requested and the model id the provider reported, the latency, the token usage the provider reported, any retries after a rate limit or overload, and your decision on the output: accepted, edited or rejected. You can export it as JSON or CSV and clear it at any time. Answers are checked against a schema, and cited passage ids are checked against the passages the model was shown.

The design is informed by the Australian Government’s policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It is a personal project and makes no claim of compliance with any of them. The reasoning is in DR-004.

Model card and data statement

Decision records

Each record states the decision first, then the context, the options, why, what happened and what I’d change. Records are never edited once accepted. A later record supersedes them instead.

  1. DR-001 · decided Semester 1, 2024 (COMP90042 group project) · recorded 6 October 2026, in hindsight · Accepted

    TF-IDF keyword retrieval instead of dense retrieval

    Retrieve evidence by matching a claim's stemmed words against TF-IDF keyword tags on every passage, scored by cosine similarity plus tag overlap, and do not attempt dense (embedding) retrieval.

    Read the record
  2. DR-002 · decided Semester 1, 2024 (COMP90042 group project) · recorded 6 October 2026, in hindsight · Accepted, with known flaws

    A from-scratch Transformer encoder instead of an LSTM

    Classify each claim with a Transformer encoder trained from scratch on the claim and evidence stems, chosen over a four-layer LSTM by validation accuracy.

    Read the record
  3. DR-003 · decided October 2026 (the web revival) · recorded 6 October 2026 · Accepted

    A pruned, read-only evidence index for serverless hosting

    Ship a 39,666-passage subset of the evidence corpus in a 13 MB read-only SQLite file with the web app, prove that it reproduces the full-corpus retrieval for every dataset claim, and label every free-text result with whether it is verified.

    Read the record
  4. DR-004 · decided 6 October 2026 · Accepted

    Bring-your-own-key LLM features, called from the browser, with a local audit log

    Offer LLM features (a second opinion on a claim and a paired LLM-versus-classifier evaluation) only through the visitor's own API key, called directly from their browser, with every call written to an audit log kept in that browser.

    Read the record