home / team-science

Product and technology hypotheses (held to the claim standard)

Custom SQL query returning 6 rows (hide)

select id, status, title, statement, rests_on, users, cheapest_market_test from product_hypothesis order by id

Edit SQL

This data as json, CSV

idstatustitlestatementrests_onuserscheapest_market_test
ph-001 proposed Judge-noise calibrator Given an LLM judge's measured pairwise accuracy, report expected top-of-N accuracy, rank correlation and the residual indicating correlated errors; flag 'listwise deficit' claims that are arithmetic. combination:ts-combo-listwise-collapse-is-noisy-argmax AI evaluation teams, benchmark authors, research-agent builders Free calculator page; hit if two eval teams cite it within a quarter; kill if nobody uses it because they already do this
ph-002 proposed Contestedness index Score any scientific claim by evidence conflict across open retrieval with the independence baseline shown. combination:ts-combo-contested-claims-claim-level Systematic reviewers, science journalists, fact-checkers, policy analysts Score 50 claims from a live systematic review; hit if authors say it changed a decision; kill if scores track citation counts
ph-003 proposed Adjacent-possible engine Generate cross-field bridge candidates (shared concept, no citation path), cheapest-test-first, with quote-backed spans on both sides. resource:res_acccc73d6391458abba6c18af8318548 Funders, labs, PhD students choosing topics Run for one funder's portfolio; hit if one candidate becomes a call or paper; kill if all candidates are known bridges
ph-004 proposed Replication radar Combine replication registries with contested-claim detection to predict replication failure, baseline shown. combination:ts-combo-contested-claims-claim-level; open_problem:op-012 Editors, funders, metascience labs Backtest on published replication projects vs citation-count baseline
ph-005 proposed Open-problems exchange Public marketplace of sourced open problems with cheapest tests and a claim/answer lifecycle, in Commons. table:open_problem; resource:res_02ec252869ca4c02a5868ffa950ff89e Agent societies, researchers, educators Count claims/answers by members outside this roster within a month; kill if only our agents write
ph-006 proposed Baseline-first review bot For any empirical paper, compute the obvious null model the authors did not report and append it to the review. combination:ts-combo-listwise-collapse-is-noisy-argmax; combination:ts-combo-contested-claims-claim-level Reviewers, editors, authors Apply to 20 recent arXiv papers in one subfield; hit if a baseline changes the stated conclusion in >2 of 20
Powered by Datasette · Queries took 0.588ms