Skip to content
Climate Claim Checker

LLM evaluation · bring your own key

Would a modern LLM do better with the same evidence?

The 2024 classifier was trained from scratch because the course banned pretrained models. This page puts a large language model in its place: same claims, same retrieved passages, same scoring. You run it with your own API key, and the statistics say how much of any difference is more than noise.

The protocol

Same inputs

A seeded random sample of the 154 dev claims (default N = 20, seed 42). The LLM gets the claim and the very passages the 2024 classifier read, each tagged with its id.

Structured answer

It must return JSON: one of the four labels, the ids of the passages it relied on and a one-line rationale. The answer is validated with a schema; anything else counts as wrong.

Two conditions

Retrieved evidence is the like-for-like test. Gold evidence is an upper bound: what each system does when retrieval is perfect.

Paired statistics

Accuracy (Wilson interval), macro-F1 and the paired difference (bootstrap, 5,000 resamples, seed 2026), McNemar’s exact test, Cohen’s h, citation validity, the course’s harmonic mean, latency and tokens. Macro-F1 averages both systems over the same labels: those in the gold set or in either system’s verdicts.

The bar to beat, on all 154 dev claims

The classifier is right on 38.3% of claims [31.0%, 46.2%] with the retrieved evidence and 57.8% [49.9%, 65.3%] with the gold evidence. Always answering “supports” scores 44.2%. It never predicts refuted or disputed. Full intervals and paired tests are on the results page.

Your key, your bill

Calls go from your browser straight to Anthropic (default claude-haiku-4-5, or claude-sonnet-5-5) or OpenAI (default gpt-5-mini). The key is never sent to this site. Every call is written to the AI audit log in your browser, which you can export.

Run the comparison

A seeded random sample of dev claims. Each one is sent to the model once per evidence condition, straight from your browser.

Anthropic · Claude Haiku 4.5

No key yet

Evidence the LLM sees

Cost note. 40 calls, about 17,200 input and 4,800 output tokens, roughly US$0.04 with Claude Haiku 4.5. An estimate only: your provider bills your key directly, and this site never sees the bill or the key.

The sample: 20 claims

✓ right · ✗ wrong · LLM verdicts and rationales are AI-generated · every LLM call is in the AI audit log

Scroll the table sideways for the verdicts →

Each sampled claim with its gold label, the classifier verdict and the LLM verdict for each evidence condition
ClaimGoldRetrieved evidenceclassifier · LLM (AI-generated)Gold evidenceclassifier · LLM (AI-generated)
claim 1256‘Next year or the year after, the Arctic will be free of ice’DisputedSupports–Supports–
claim 1490the concentration of carbon dioxide in Earth’s atmosphere has climbed to a level last seen more than 3 million years ago — before humans even appeared on the rocky ball we call home.SupportsSupports–Supports–
claim 540The heatwave we now have in Europe is not something that was expected with just 1C of warmingSupportsSupports–Supports–
claim 1745The IPCC simply updated their temperature history graphs to show the best data available at the time.Not enough infoSupports–Not enough info–
claim 1515no one really knows if last year 2016 was a global temperature record.Not enough infoSupports–Supports–
claim 1087The particular signature of warming in 2016 was also revealing in another way, Overpeck said, noting that the stratosphere… saw record cold temperatures last yearNot enough infoNot enough info–Not enough info–
claim 1896Greg Hunt CSIRO research shows carbon emissions can be reduced by 20 per cent over 40 years using nature, soils and trees.Not enough infoNot enough info–Not enough info–
claim 342While members of the media may nod along to such claims [about changes in weather extremes], the evidence paints a different storySupportsSupports–Supports–
claim 609That’s because as Antarctica’s mass shrinks, the ice sheet’s gravitational pull on the ocean relaxes somewhat, and the seas travel back across the globe to pile up far away — with U.S. coasts being one prime destination.”SupportsSupports–Supports–
claim 173If there were [carbon emissions], we could not see because most carbon is black.RefutesNot enough info–Supports–
claim 1222The melting Greenland ice sheet is already a major contributor to rising sea level and if it was eventually lost entirely, the oceans would rise by six metres around the world, flooding many of the world’s largest cities.SupportsSupports–Not enough info–
claim 141Greenpeace didn’t save the whales, switching from whale oil to petroleum and palm oil didNot enough infoNot enough info–Not enough info–
claim 2093Through its impacts on the climate, CO2 presents a danger to public health and welfare, and thus qualifies as an air pollutantSupportsNot enough info–Supports–
claim 1638Worry about global warming impacts in the next 100 years, not an ice age in over 10,000 years.Not enough infoSupports–Not enough info–
claim 1259While the north-east, midwest and upper great plains have experienced a 30% increase in heavy rainfall episodes – considered once-in-every-five year downpours – parts of the west, particularly California, have been parched by drought.SupportsNot enough info–Not enough info–
claim 368The most recent IPCC report lays out a future if we limit global heating to 1.5°C instead of the Paris Agreement’s 2°C.DisputedSupports–Supports–
claim 1467The amount of carbon dioxide absorbed by the upper layer of the oceans is increasing by about 2 billion tons per year.Not enough infoSupports–Not enough info–
claim 1928NASA Finds Antarctica is Gaining Ice,DisputedSupports–Supports–
claim 988The unlikely scenarios are now, all of a sudden, becoming more probable than they once were thought to be,’ says Sweet.”Not enough infoNot enough info–Not enough info–
claim 2895The latest measurements involve the use of satellite gravimetry, estimating the mass of terrain beneath by detecting slight changes in gravity as a satellite passes overhead.SupportsNot enough info–Supports–

Reading the result fairly

  • Not the course’s rules. The 2024 system could not use pretrained models; an LLM is nothing but pretraining. This measures how much that constraint cost, not who did better coursework.
  • The labels are relative to the gold evidence. An LLM that follows its instructions will often answer “not enough info” on retrieved passages that miss the point, and be marked wrong. That is the retrieval bottleneck showing, which is why the gold condition is there.
  • The gold condition favours the classifier. Its gold-evidence verdicts come from the training epoch chosen by accuracy on these same dev claims with gold evidence, so its score there is optimistic and the LLM-minus-classifier gap is understated.
  • Citations are scored strictly. In evidence F, every distinct id the LLM cites counts as a prediction, so an id it was never shown is a wrong one, just as a wrong retrieved passage is for the 2024 system. Citation validity is recomputed from the cited ids, including for loaded files.
  • Possible contamination. The claims, labels and Wikipedia passages are public (the course published them on GitHub), so they may be in a model’s training data, which would flatter the LLM.
  • Small samples, one run. Twenty claims give intervals about 40 points wide. Model answers can also vary between runs; Haiku is called at temperature 0, which reduces but does not remove that.
  • What counts as failure. Calls that fail for infrastructure reasons (key, network, rate limit after retries) are excluded and reported. Answers the model gave but that are unusable (bad JSON, refusal, cut off) are scored as wrong.

Evaluation design and AI use statement