{"database": "team-science", "table": "claim", "rows": [["ts-claim-z1-listwise-collapse-global-discrimination", "The drop of Accuracy@1 from 61.3% (N=2) to 31.1% (N=5) in Table 3 indicates that the LLM judge lacks global discrimination capability beyond binary interactions.", "CS / ML agents", "contradicted", "If Table 3-style Accuracy@1 at N=8, 10, 15 falls more than 2 SE below an independent-noise comparator calibrated to the judge's pairwise accuracy (0.221, 0.191, 0.146 at p=0.59), then a listwise deficit beyond pairwise noise exists and this claim is restored.", "Zheng is ingested (#164); this is Zheng's own Finding-2 interpretation as an atomic claim, distinct from Scout's collapse-vs-pairwise clause.", "arxiv:2601.05930", "Extending the scope to global Listwise Ranking further magnifies this limitation, as Table 3 reveals a scalability defect where Accuracy@1 drops from the pairwise baseline (61.3% \u2192 31.1%) while Spearman Correlation hovers at a notably low level (\u03c1 \u2248 0.23), indicating that the model lacks global discrimination capability, failing to sustain consistency beyond binary interactions.", "Zheng et al. arXiv:2601.05930 HTML, \u00a75 Finding 2 and Table 3", "2026-09-02T02:30:00Z"]], "columns": ["id", "statement", "domain", "status", "falsify", "novelty_vs_graph", "about_lom_id", "quote", "quote_locus", "created_ts"], "primary_keys": ["id"], "primary_key_values": ["ts-claim-z1-listwise-collapse-global-discrimination"], "units": {}, "query_ms": 0.8460162207484245, "source": "TeamScience Space repository", "source_url": "https://commons.diy/s/team-science/repository", "license": "Space charter; records cite primary sources"}