claim_id,source,label,span ts-claim-c1-scifact-no-global-truth,https://aclanthology.org/2020.emnlp-main.609.pdf,SUPPORTS,§2 Background and task definition ts-claim-c2-scifact-mixed-polarity,https://aclanthology.org/2020.emnlp-main.609.pdf,SUPPORTS,"§3.3 gold labels; Table 1 and §6.3 are system outputs / case study, not gold mixed labels" ts-claim-c3-ai-scientist-s2-novelty,https://arxiv.org/pdf/2408.06292,SUPPORTS,§3 Idea Generation (Semantic Scholar filter) ts-claim-cf1-contested-claim-level,https://arxiv.org/abs/2012.00614,SUPPORTS,we include claims for which both supporting and refuting evidence were found ts-claim-cf1-contested-claim-level,https://commons.diy/v0/spaces/team-science/repository/file?path=graph/tests/polarity_concordance.out.txt,SUPPORTS,Climate-FEVER: n(k>=2)=790 r=0.292 mixed=154 (19.5%) independence=473.7 (60.0%) ts-claim-mg1-noisy-tournament-selection,https://openalex.org/W157468466,NOINFO,"Genetic Algorithms, Tournament Selection, and the Effects of Noise. (title only; full read pending)" ts-claim-ps1-cramer-model-fails-at-two-scales,doi:10.1007/s00220-004-1222-4,SUPPORTS,"Contrary to what would be predicted on the basis of Cramér's model concerning the distribution of prime numbers, we develop evidence that the distribution of $ψ(x+H)- ψ(x)$, for $0\le x\le N$, is approximately normal with mean $\sim H$ and variance $\sim H\log N/H$, when $N^δ\le H \le N^{1-δ}$." ts-claim-rc1-contested-fraction-by-evidence-source,doi:10.1038/s41562-018-0399-z,SUPPORTS,We find a significant effect in the same direction as the original study for 13 (62%) studies ts-claim-rc1-contested-fraction-by-evidence-source,doi:10.1126/science.aac4716,SUPPORTS,Ninety-seven percent of original studies had statistically significant results. Thirty-six percent of replications had statistically significant results ts-claim-rc1-contested-fraction-by-evidence-source,doi:10.1126/science.aaf0918,SUPPORTS,We found a significant effect in the same direction as in the original study for 11 replications (61%) ts-claim-s1-novelty-not-significance,https://arxiv.org/pdf/2408.06292,SUPPORTS,"After idea generation, we filter ideas by connecting the language model with the Semantic Scholar API (Fricke, 2018) and web access as a tool (Schick et al., 2024). This allows The AI Scientist to discard any idea that is too similar to existing literature." ts-claim-s1-novelty-not-significance,scout-s1-dual-error-gloss,NOT_EVIDENCE,the dual error is keeping trivia that is merely unseen — Scout gloss; not in Lu §3. ts-claim-so1-contested-after-open-retrieval,https://arxiv.org/abs/2210.13777,SUPPORTS,"Of the 81 claims in SciFact-Open with at least 2 ECAPs, 16 of them (20%) have conflicting evidence." ts-claim-so1-contested-after-open-retrieval,https://commons.diy/v0/spaces/team-science/repository/file?path=graph/tests/polarity_concordance.out.txt,SUPPORTS,SciFact-Open: n(k>=2)=81 r=0.459 mixed=15 (18.5%) independence=57.6 (71.2%) ts-claim-th1-comparative-judgment-noise,https://doi.org/10.1037/h0070288,NOINFO,A law of comparative judgment. (title only; full read pending) ts-claim-z1-listwise-collapse-global-discrimination,https://arxiv.org/html/2601.05930,SUPPORTS,Table 3 reveals a scalability defect where Accuracy@1 drops from the pairwise baseline ts-claim-z1-listwise-collapse-global-discrimination,https://commons.diy/v0/spaces/team-science/repository/file?path=graph/tests/noisy_argmax.out.txt,REFUTES,"pairwise accuracy p=0.590: N=3 0.439, N=4 0.358, N=5 0.308, Spearman 0.219/0.223 vs reported 0.434/0.350/0.311 and 0.25/0.22 — independent per-comparison noise alone reproduces Table 3"