Dev claim · claim-2813 · #15 of 154
“In fact, the authors go on to estimate climate sensitivity from their findings, calculate a value between 2.3 to 4.1°C.”
Claim tags (sorted stems; highlighted when a passage shares them)
- author
- calcul
- climat
- estim
- fact
- find
- go
- sensit
- valu
The model’s verdict
The Transformer retrained from the notebook (the 2024 weights were never saved), shown three ways. The first is how the notebook evaluated it.
Notebook protocol
Dev claims predicted in batches of 16, in file order, from the 2024 retrieved evidence, as the notebook does.
- Supports
- 58.6%
- Refutes
- 31.2%
- Not enough info
- 3.8%
- Disputed
- 6.3%
One claim at a time
The same model run on a single claim, which is how the Try-it page runs it.
- Supports
- 53.6%
- Refutes
- 35.2%
- Not enough info
- 6.9%
- Disputed
- 4.3%
With gold evidence
Fed the human-annotated evidence instead of retrieved passages. This is the 'val accuracy' the training loop reports.
- Supports
- 26.1%
- Refutes
- 6.1%
- Not enough info
- 56.4%
- Disputed
- 11.4%
Retrieved vs gold evidence
Gold passages were picked by the dataset’s annotators. Retrieved passages are what the TF-IDF rule returned from all 1.19M passages. Switch between the submitted 2024 output and the re-runs.
2024 submission
The passages the team actually retrieved and submitted in 2024 (their saved output file). The file records only passage ids, so scores are recomputed with the submission rule.
- Path
- not recorded
- Found
- 0 of 5 gold
- P
- 0.00
- R
- 0.00
- F
- 0.00
#1evidence-1045588 Go and fuck yourself.
- go (shared with the claim)
- cos
- 0.455
- overlap
- 1.000
- score
- 1.455
- shared
- 1
#2evidence-553632 And we are going.
- go (shared with the claim)
- cos
- 0.455
- overlap
- 1.000
- score
- 1.455
- shared
- 1
#3evidence-58052 Imma albifasciella (Pagenstecher, 1900)
- go (shared with the claim)
- cos
- 0.455
- overlap
- 1.000
- score
- 1.455
- shared
- 1
#4evidence-324049 Go and fuck yourself".
- go (shared with the claim)
- cos
- 0.455
- overlap
- 1.000
- score
- 1.455
- shared
- 1
#5evidence-875993 The market value has been estimated at # 20m.
- market
- محمد
- valu (shared with the claim)
- estim (shared with the claim)
- cos
- 0.533
- overlap
- 0.500
- score
- 1.033
- shared
- 2
#6evidence-746532 The owner 's value was estimated at # 2.
- محمد
- owner
- valu (shared with the claim)
- estim (shared with the claim)
- cos
- 0.519
- overlap
- 0.500
- score
- 1.019
- shared
- 2
Gold evidence (5)
evidence-679700 The IPCC literature assessment estimates that TCR likely lies between 1 °C and 2.5 °C.
- like
- assess
- literatur
- lie
- ipcc
- estim (shared with the claim)
evidence-158204 For constant humidity they computed a climate sensitivity of 2.3 °C per doubling of CO2 (which they rounded to 2, the value most often quoted from their work, in the abstract of the paper).
- quot
- valu (shared with the claim)
- constant
- humid
- comput
- sensit (shared with the claim)
- paper
- abstract
- doubl
evidence-1049371 The 1990 IPCC First Assessment Report estimated that equilibrium climate sensitivity to a doubling of CO 2 lay between 1.5 and 4.5 °C (2.7 and 8.1 °F), with a "best guess in the light of current knowledge" of 2.5 °C (4.5 °F).
- assess
- knowledg
- equilibrium
- lay
- sensit (shared with the claim)
- ipcc
- doubl
- estim (shared with the claim)
evidence-358515 IPCC authors concluded ECS is very likely to be greater than 1.5 °C (2.7 °F) and likely to lie in the range 2 to 4.5 °C (4 to 8.1 °F), with a most likely value of about 3 °C (5 °F).
- like
- author (shared with the claim)
- lie
- محمد
- valu (shared with the claim)
- conclud
- rang
- ipcc
- greater
- ec
evidence-920160 The IPCC Fifth Assessment Report reverted to the earlier range of 1.5 to 4.5 °C (2.7 to 8.1 °F) (high confidence) because some estimates using industrial-age data came out low.
- assess
- fifth
- revert
- earlier
- estim (shared with the claim)
- ipcc
- data
- low
- confid