id,statement,domain,status,falsify,novelty_vs_graph,about_lom_id,quote,quote_locus,created_ts ts-claim-c1-scifact-no-global-truth,"Given a fixed scientific corpus, a claim is not assigned a global truth label; verification is a SUPPORTS / REFUTES / NOINFO relation on each claim–abstract pair, because a global label would require systematic review.",CS / NLP / metascience,proposed,"If a later primary methods paper shows a single corpus-level truth bit can be assigned without systematic review and still match a team-of-experts systematic-review verdict at high agreement on SciFact-style claims, then Wadden et al.’s §2 reason for refusing a global label does not hold.","Literature-grounded restatement of SciFact’s task definition vs ingested paper doi:10.18653/v1/2020.emnlp-main.609, not a model-invented title.",doi:10.18653/v1/2020.emnlp-main.609,"While SCIFACT claims are indeed verifiable assertions about scientific findings, accurately assigning a global truth label to a scientific claim (given a fixed scientific corpus) requires a systematic review by a team of experts. In this work we focus on the simpler task of assigning SUPPORTS or REFUTES relations to individual claim-abstract pairs.",Wadden et al. 2020 §2 PDF (Anthology 2020.emnlp-main.609.pdf),2026-09-01T22:28:39Z ts-claim-c2-scifact-mixed-polarity,"SciFact’s task definition allows one claim to be both supported and refuted by different abstracts; the authors report that mix on real COVID-19 system outputs (Table 1, §6.3) but it never occurs in the gold dataset, where each claim has a single label.",CS / NLP / metascience,proposed,"If a primary re-annotation of SciFact gold found frequent SUPPORTS+REFUTES abstract pairs per claim, drop the ‘never in the dataset’ clause and treat mixed gold labels as the default.",Tightens Scout’s Table-1 reading against ingested SciFact paper node; gold never mixed on train+dev (C2 count PASS).,doi:10.18653/v1/2020.emnlp-main.609,Although our task definition allows for a single claim to be both supported and refuted (by different abstracts) – an occurrence we observe on real-world COVID-19 claims (§6.3) – this never occurs in our dataset. Each claim has a single label.,Wadden et al. 2020 §3.3 PDF,2026-09-01T22:28:39Z ts-claim-c3-ai-scientist-s2-novelty,"The AI Scientist’s idea-generation filter discards ideas that are too similar to existing literature by querying the Semantic Scholar API (plus web access); novelty is therefore a retrieved-paper similarity judgment, not a check against a durable citation graph of ingested literature, and the pipeline can still emit a full conference-style manuscript.",CS / ML / automated science,proposed,"If a replication showed that Semantic Scholar similarity filtering plus the paper write-up step never published an idea already present in a citation graph of the ingested seed papers and their references, treat S2 similarity as a sufficient novelty-vs-graph proxy.","Grounded in Lu et al. §3 vs ingested node arxiv:2408.06292; S2 paperId is missing (ingest_error 429), not fabricated.",arxiv:2408.06292,"After idea generation, we filter ideas by connecting the language model with the Semantic Scholar API (Fricke, 2018) and web access as a tool (Schick et al., 2024). This allows The AI Scientist to discard any idea that is too similar to existing literature.",Lu et al. 2024 §3 Idea Generation,2026-09-01T22:28:39Z ts-claim-cf1-contested-claim-level,"In Climate-FEVER, 19.5% of claims with at least two polar evidence sentences are contested (both SUPPORTS and REFUTES), and that fraction does not rise with the number of polar sentences (k = 2 to 5: 20.6, 16.5, 22.2, 19.3%; trend z = 0.06), whereas independent draws would give 41 to 82%.",climate / NLP / claim verification,ready_to_test,A re-annotation or re-retrieval of Climate-FEVER in which the contested fraction among k>=2 claims rises with k (trend z > 1.6) or falls outside 12-28% withdraws this claim.,First Wikipedia/climate node; no citation path to any read paper; connects to C2 through the polarity-concordance concept.,arxiv:2012.00614,"While FEVER only contains undisputed claims, we include claims for which both supporting and refuting evidence were found.",Climate-FEVER arXiv:2012.00614 §3 (ar5iv; verbatim per Skeptic run),2026-09-02T15:10:00Z ts-claim-mg1-noisy-tournament-selection,"Under noisy fitness evaluation, tournament selection's probability of choosing the truly best individual falls as tournament size grows at fixed noise, and the effect is predicted by a Gaussian noise model.",CS / evolutionary computation,proposed,A full read of Miller & Goldberg 1995 showing no Gaussian-noise treatment of selection accuracy versus tournament size would withdraw this paraphrase.,First evolutionary-computation node in the graph; no citation path to any ingested paper.,openalex:W157468466,,"paraphrase from title and abstract-level knowledge; full-text read pending (CiteSeerX record, no OA PDF resolved)",2026-09-02T02:30:00Z ts-claim-ps1-cramer-model-fails-at-two-scales,"Cramér's model for primes fails in the same direction at two scales: Montgomery–Soundararajan (2004) give evidence that the variance of ψ(x+H)−ψ(x) is ~H log(N/H), not the Poisson ~H, for N^δ ≤ H ≤ N^(1−δ); and at the smallest scale H = ln x our test finds the probability of at least one prime in [x−ln x, x+ln x] exceeds the Poisson value 1−e^−2 by ~0.75/ln x across 10^6–10^18 (graph/tests/prime_short_interval.py). Less-than-Poisson variance and a higher-than-Poisson hit rate are the same fact: primes in short intervals are more evenly spread than independent coins.",mathematics / analytic number theory,proposed,"A decade in 10^18–10^21 where the excess times ln x leaves [0.5, 1.0]; or a proof/literature result that the leading correction to 1−e^−2 at H = ln x is of a different order than 1/ln x; or a reader finding that the MS2004 variance regime does not extend toward H ~ log N (their theorem needs H ≥ N^δ).",neighborhood by construction (cites the paper it is about); the bridge from the MS2004 variance statement to the H = ln x hit rate is not in any ingested paper. Pair ap-104bf56087.,doi:10.1007/s00220-004-1222-4,"Contrary to what would be predicted on the basis of Cramér's model concerning the distribution of prime numbers, we develop evidence that the distribution of $ψ(x+H)- ψ(x)$, for $0\le x\le N$, is approximately normal with mean $\sim H$ and variance $\sim H\log N/H$, when $N^δ\le H \le N^{1-δ}$.",abstract (arXiv math/0409258),2026-09-02T18:49:53Z ts-claim-rc1-contested-fraction-by-evidence-source,"The contested fraction among claims with two or more evidence documents depends on how the evidence was gathered: in direct-replication corpora (original = SUPPORTS, failed replication by the authors' primary criterion = REFUTES) it is 38–63% (Camerer 2016: 7/18; Camerer 2018: 8/21; OSC 2015: ~61/97), versus ~20% in annotator-built open-retrieval corpora (Climate-FEVER 19.5%, SciFact-Open 18.5%).",metascience / claim verification / replication,weakly_supported,"Per-study tables give a replication-corpus contested fraction below 25% in two of the three projects, or a retrieval corpus built by replication-style evidence gathering shows ~20%.",novel: no ingested paper compares retrieval-corpus contestedness with replication-project outcomes; pair ap-180fa20fea (op-004 × op-012).,doi:10.1126/science.aaf0918,We found a significant effect in the same direction as in the original study for 11 replications (61%),abstract (PubMed 26940865),2026-09-02T18:17:15Z ts-claim-s1-novelty-not-significance,"A claim can be graph-novel vs TeamScience JSONL and still be insignificant if it would not change a #177 rule, a cheapest test, or the next ingest walk.",metascience / TeamScience ops,proposed,"If objectives v0.1 are met by 25 graph-novel claims none of which changed a rule, a test, or an ingest decision, drop this criterion.","neighborhood (Skeptic #177 on Foster-Lu edge; statement ≠ C3). Ops criterion, not graph-novel science.",arxiv:2408.06292,"After idea generation, we filter ideas by connecting the language model with the Semantic Scholar API (Fricke, 2018) and web access as a tool (Schick et al., 2024). This allows The AI Scientist to discard any idea that is too similar to existing literature.",Lu et al. 2024 §3 Idea Generation (contiguous; C3 quote),2026-09-02T02:24:34Z ts-claim-so1-contested-after-open-retrieval,"SciFact-Open reuses the 279 SciFact test claims verbatim; none had two polar evidence abstracts in SciFact, and after retrieval over 500K abstracts 81 do, of which 15 (18.5%) are contested, against 71% expected under independence.",CS / NLP / claim verification,ready_to_test,"If SciFact-Open's claim ids/text do not match SciFact's, or a recount of data/claims.jsonl gives a contested fraction among k>=2 outside 12-28%, this claim is withdrawn.",SciFact-Open cites SciFact (neighborhood at paper level); the same-claims comparison is new to the graph.,arxiv:2210.13777,"Of the 81 claims in SciFact-Open with at least 2 ECAPs, 16 of them (20%) have conflicting evidence.",SciFact-Open §3.3 (ar5iv 2210.13777; verbatim per Skeptic run; public release recounts 15/81),2026-09-02T15:10:00Z ts-claim-th1-comparative-judgment-noise,"A comparative judgment between two stimuli is modeled as the sign of the difference of two normally distributed discriminal processes, so pairwise discrimination accuracy is a function of the ratio of true difference to noise.",psychometrics / statistics,proposed,A full read showing Thurstone 1927 does not model comparative judgment as normally distributed discriminal processes would withdraw this paraphrase.,First psychometrics node; no citation path to any ingested paper.,doi:10.1037/h0070288,,paraphrase; full-text read pending (paywalled; Crossref-verified DOI),2026-09-02T02:30:00Z ts-claim-z1-listwise-collapse-global-discrimination,The drop of Accuracy@1 from 61.3% (N=2) to 31.1% (N=5) in Table 3 indicates that the LLM judge lacks global discrimination capability beyond binary interactions.,CS / ML agents,contradicted,"If Table 3-style Accuracy@1 at N=8, 10, 15 falls more than 2 SE below an independent-noise comparator calibrated to the judge's pairwise accuracy (0.221, 0.191, 0.146 at p=0.59), then a listwise deficit beyond pairwise noise exists and this claim is restored.","Zheng is ingested (#164); this is Zheng's own Finding-2 interpretation as an atomic claim, distinct from Scout's collapse-vs-pairwise clause.",arxiv:2601.05930,"Extending the scope to global Listwise Ranking further magnifies this limitation, as Table 3 reveals a scalability defect where Accuracy@1 drops from the pairwise baseline (61.3% → 31.1%) while Spearman Correlation hovers at a notably low level (ρ ≈ 0.23), indicating that the model lacks global discrimination capability, failing to sustain consistency beyond binary interactions.","Zheng et al. arXiv:2601.05930 HTML, §5 Finding 2 and Table 3",2026-09-02T02:30:00Z