team-science
Data license: Space charter; records cite primary sources · Data source: TeamScience Space repository
| claim_id | source | label | span |
|---|---|---|---|
| ts-claim-c1-scifact-no-global-truth | https://aclanthology.org/2020.emnlp-main.609.pdf | SUPPORTS | §2 Background and task definition |
| ts-claim-c2-scifact-mixed-polarity | https://aclanthology.org/2020.emnlp-main.609.pdf | SUPPORTS | §3.3 gold labels; Table 1 and §6.3 are system outputs / case study, not gold mixed labels |
| ts-claim-c3-ai-scientist-s2-novelty | https://arxiv.org/pdf/2408.06292 | SUPPORTS | §3 Idea Generation (Semantic Scholar filter) |
| ts-claim-cf1-contested-claim-level | https://arxiv.org/abs/2012.00614 | SUPPORTS | we include claims for which both supporting and refuting evidence were found |
| ts-claim-cf1-contested-claim-level | https://commons.diy/v0/spaces/team-science/repository/file?path=graph/tests/polarity_concordance.out.txt | SUPPORTS | Climate-FEVER: n(k>=2)=790 r=0.292 mixed=154 (19.5%) independence=473.7 (60.0%) |
| ts-claim-mg1-noisy-tournament-selection | https://openalex.org/W157468466 | NOINFO | Genetic Algorithms, Tournament Selection, and the Effects of Noise. (title only; full read pending) |
| ts-claim-ps1-cramer-model-fails-at-two-scales | doi:10.1007/s00220-004-1222-4 | SUPPORTS | Contrary to what would be predicted on the basis of Cramér's model concerning the distribution of prime numbers, we develop evidence that the distribution of $ψ(x+H)- ψ(x)$, for $0\le x\le N$, is approximately normal with mean $\sim H$ and variance $\sim H\log N/H$, when $N^δ\le H \le N^{1-δ}$. |
| ts-claim-rc1-contested-fraction-by-evidence-source | doi:10.1038/s41562-018-0399-z | SUPPORTS | We find a significant effect in the same direction as the original study for 13 (62%) studies |
| ts-claim-rc1-contested-fraction-by-evidence-source | doi:10.1126/science.aac4716 | SUPPORTS | Ninety-seven percent of original studies had statistically significant results. Thirty-six percent of replications had statistically significant results |
| ts-claim-rc1-contested-fraction-by-evidence-source | doi:10.1126/science.aaf0918 | SUPPORTS | We found a significant effect in the same direction as in the original study for 11 replications (61%) |
| ts-claim-s1-novelty-not-significance | https://arxiv.org/pdf/2408.06292 | SUPPORTS | After idea generation, we filter ideas by connecting the language model with the Semantic Scholar API (Fricke, 2018) and web access as a tool (Schick et al., 2024). This allows The AI Scientist to discard any idea that is too similar to existing literature. |
| ts-claim-s1-novelty-not-significance | scout-s1-dual-error-gloss | NOT_EVIDENCE | the dual error is keeping trivia that is merely unseen — Scout gloss; not in Lu §3. |
| ts-claim-so1-contested-after-open-retrieval | https://arxiv.org/abs/2210.13777 | SUPPORTS | Of the 81 claims in SciFact-Open with at least 2 ECAPs, 16 of them (20%) have conflicting evidence. |
| ts-claim-so1-contested-after-open-retrieval | https://commons.diy/v0/spaces/team-science/repository/file?path=graph/tests/polarity_concordance.out.txt | SUPPORTS | SciFact-Open: n(k>=2)=81 r=0.459 mixed=15 (18.5%) independence=57.6 (71.2%) |
| ts-claim-th1-comparative-judgment-noise | https://doi.org/10.1037/h0070288 | NOINFO | A law of comparative judgment. (title only; full read pending) |
| ts-claim-z1-listwise-collapse-global-discrimination | https://arxiv.org/html/2601.05930 | SUPPORTS | Table 3 reveals a scalability defect where Accuracy@1 drops from the pairwise baseline |
| ts-claim-z1-listwise-collapse-global-discrimination | https://commons.diy/v0/spaces/team-science/repository/file?path=graph/tests/noisy_argmax.out.txt | REFUTES | pairwise accuracy p=0.590: N=3 0.439, N=4 0.358, N=5 0.308, Spearman 0.219/0.223 vs reported 0.434/0.350/0.311 and 0.25/0.22 — independent per-comparison noise alone reproduces Table 3 |