Skip to content
Climate Claim Checker

Tour

The site in three short walkthroughs

Each recording is a scripted run of the live site with captions on screen, so you can see the main workflows without clicking through them. Everything shown is the real app, except the LLM results in the third walkthrough: they come from a labelled mock run, because no model was called and no API key was entered.

Walkthrough 1 of 3

Explore the dev set

Browse the 154 dev claims, narrow them to the ones where retrieval found a gold passage, then open one to compare the annotators' evidence with what the 2024 system retrieved and how it voted.

Transcript: what happens on screen
  1. Step 1 (0:00). All 154 dev claims: the annotators' label, the model's verdict and the gold evidence it found

    The Explore page opens on four headline figures with 95% intervals: mean evidence F of 0.043, 13 of 154 claims with any gold passage found, verdict accuracy of 38.3% and 47% of claims on the fallback path.

  2. Step 2 (0:07). Filter to the claims where the 2024 retrieval found at least one gold passage

    Choosing “Found gold” in the Evidence filter narrows the list to the 13 claims where retrieval found a gold passage.

  3. Step 3 (0:12). Open a claim to compare gold and retrieved evidence side by side

    Dev claim 752, “[South Australia] has the most expensive electricity in the world”, opens on its own page with its claim tags and the annotators' label, Supports.

  4. Step 4 (0:17). The retrained Transformer's verdict three ways, with its class probabilities

    Three cards show the verdict under the notebook protocol, one claim at a time, and with gold evidence. Each has bars for all four class probabilities and says whether it matches the label.

  5. Step 5 (0:24). Retrieved passages (left) against the annotators' gold evidence (right); hits are marked gold

    The 2024 submission retrieved four passages, two of them gold, so precision is 0.50, recall 1.00 and F 0.67. Each passage shows its tags, cosine similarity, tag overlap and score.

  6. Step 6 (0:30). Switch between the submitted 2024 retrieval and the faithful re-runs, with P, R and F per run

    The tabs switch between the saved 2024 output, the re-run of the submitted rule, the notebook's rule and the raw-claim variant, each with its own scores.

Walkthrough 2 of 3

Check a claim

Type a new claim and run the original pipeline on the server: TF-IDF retrieval over a pruned Wikipedia index, then the from-scratch Transformer, with every score and word contribution shown.

Transcript: what happens on screen
  1. Step 1 (0:00). Check a claim: type any sentence about climate science

    On the Try page the claim “Sea levels are rising faster than ever because glaciers are melting.” is typed into the claim box. The submitted 2024 scoring rule is selected.

  2. Step 2 (0:08). Run the original 2024 pipeline: TF-IDF evidence retrieval, then the from-scratch Transformer

    Pressing “Check this claim” sends the sentence to the server, which runs the TypeScript ports of the original preprocessing, retrieval and classifier.

  3. Step 3 (0:14). The verdict with all four class probabilities, beside the model's dev accuracy and its 95% CI

    The model answers Supports with 46.1% probability, ahead of not enough info at 27.3%. A note says the model is right on 37.7% of dev claims (95% CI 30.4% to 45.5%), no better than always answering “supports”.

  4. Step 4 (0:22). The passages it retrieved, each with its cosine similarity and tag-overlap scores

    Four retrieved passages about glacier and ice-sheet melt and sea-level rise are listed with their shared tags, cosine, overlap and combined score. A banner explains that this sentence was searched against a 39,666-passage subset.

  5. Step 5 (0:28). Under the hood: the notebook's preprocessing, the TF-IDF tag vector and each word's exact contribution

    The preprocessing panel shows the tokens kept and dropped, the tag TF-IDF weights, and a chart of how much each word pushes the verdict towards each label.

  6. Step 6 (0:34). Or pick a one-click example, checked against the full 1.19M-passage search at build time

    Clicking the example “Arctic sea ice has been shrinking for decades.” runs it straight away. A green note confirms that it returns the same passages as the full 2024 search over all 1.19M passages.

Walkthrough 3 of 3

Original model vs LLM

The bring-your-own-key settings, then the paired evaluation harness filled with a seeded mock run (no model was called, no key was entered) to show how the comparison is scored and reported.

Mocked AI response for illustration. The LLM verdicts in this recording are a seeded mock run file loaded through “Load a saved run”. No model was called and no key was entered, so its accuracy figures describe the mock, not any real model. With your own key the LLM evaluation page runs the same comparison for real.

Transcript: what happens on screen
  1. Step 1 (0:00). Original model vs LLM: the same dev claims and the same evidence, scored as a paired comparison

    The LLM evaluation page sets out the protocol: the same seeded sample of dev claims, a structured JSON answer, retrieved and gold evidence conditions, and paired statistics.

  2. Step 2 (0:05). Bring your own key: AI is optional, and the key stays in this browser and goes only to the provider

    “Add key” opens the AI settings dialog: Anthropic (default) or OpenAI, the model, a key field, “Remember on this device” (off keeps the key in sessionStorage for this tab only) and a note that the key never reaches the site.

  3. Step 3 (0:15). No key is entered in this demo. Cancel, then load a saved run file instead

    The dialog is cancelled without entering a key.

  4. Step 4 (0:19). A seeded mock run (no model was called) loaded through “Load a saved run” [Mocked AI response for illustration]

    A run file generated by the tour script with a seeded random generator is loaded. Its model is named “mocked-llm-for-illustration” and every rationale says no model was called.

  5. Step 5 (0:23). Accuracy and macro-F1 with 95% intervals, the paired difference, McNemar's exact test [Mocked AI response for illustration]

    For each evidence condition the page shows accuracy (Wilson interval) and macro-F1 (bootstrap interval) for the mock and the classifier, the paired difference, McNemar's exact test, Cohen's h, citation validity and evidence F. The numbers describe the mock, not any real model.

  6. Step 6 (0:29). Claim by claim, each LLM verdict is labelled AI-generated beside the classifier's [Mocked AI response for illustration]

    The claim table lists each sampled claim with its gold label, the classifier's verdict and the mock's verdict for both conditions, marked right or wrong.

Key features at a glance

Desktop screenshots at 1440 by 900 in the light and dark themes, and two at phone width. Select one to enlarge it.

  • Landing, light theme. The question, the pipeline and a real dev claim with its evidence and both verdicts.
  • Landing, dark theme. The same page in the dark theme.
  • Check a claim. A typed claim, the verdict and all four class probabilities, beside the dev-accuracy CI.
  • Retrieved evidence. Each retrieved passage with its shared tags, cosine similarity and tag overlap.
  • Explore the dev set. Headline figures with 95% CIs, filters, and all 154 dev claims.
  • Gold vs retrieved evidence. One claim's retrieved passages against the annotators' gold evidence.
  • Results with intervals. Accuracy with Wilson CIs and bootstrap intervals for macro-F1 and retrieval.
  • Pipeline walkthrough. One claim followed through preprocessing, tagging, scoring, selection and classification.
  • Bring your own key. AI settings: your own key, kept in this browser and never sent to this site.
  • LLM evaluation (mocked run). The paired harness filled with a labelled mock run: the layout only, no model called.
  • Methods and decisions. Provenance, evaluation design, limitations, the AI use statement and decision records.
  • Model card. Intended use, training data, evaluation with intervals and known failure modes.
  • Phone: landing. The landing page at 390 px wide.
  • Phone: a verdict. A verdict and its probabilities at 390 px wide.

How these were made

  • A Playwright script, web/e2e/showcase.spec.ts, drives Google Chrome through each journey and records it. Run pnpm showcase to re-record everything; pnpm e2e runs the same journeys quickly as end-to-end tests.
  • The inputs are fixed: one typed claim, dev claim 752, and the harness’s default sample (20 claims, seed 42). The mock LLM run is drawn with its own seed (7), so every re-recording shows the same thing.
  • Captions, step lists and transcripts come from one file, so they always match. ffmpeg turns each recording into an H.264 MP4 for this page and a GIF for the README.
  • Everything else on the site is live: the recordings show the same pages you can visit, starting from Try a claim.