RPM 2608.13940 claim + citation neighborhood
| kind | id | label | extra |
|---|---|---|---|
| paper | arxiv:2608.13940 | AI Research Preference Models | |
| cites | arxiv:1711.09846 | Population based training of neural networks | arxiv:2608.13940 reference list (arXiv HTML bibliography) |
| cites | arxiv:2301.04104 | Mastering diverse domains through world models | arxiv:2608.13940 reference list (arXiv HTML bibliography) |
| cites | arxiv:2408.06292 | The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery | Foster et al. 2026 Related Work / References: Lu et al. (2024) The AI Scientist arXiv:2408.06292 |
| cites | arxiv:2502.13138 | Aide: Ai-driven exploration in the space of code | arxiv:2608.13940 reference list (arXiv HTML bibliography) |
| cites | arxiv:2510.02387 | Cwm: An open-weights llm for research on code generation with world models | arxiv:2608.13940 reference list (arXiv HTML bibliography) |
| cites | arxiv:2601.03267 | Openai gpt-5 system card | arxiv:2608.13940 reference list (arXiv HTML bibliography) |
| cites | arxiv:2601.05930 | Can We Predict Before Executing Machine Learning Agents? | Foster et al. 2026 References: Zheng et al. (2026) Can we predict before executing machine learning agents? arXiv:2601.05930 |
| cites | arxiv:2602.02660 | MARS: modular agent with reflective search for automated ai research | arxiv:2608.13940 reference list (arXiv HTML bibliography) |
| cites | arxiv:2602.06855 | AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents | arxiv:2608.13940 reference list (arXiv HTML bibliography) |
| cites | arxiv:2603.17863 | DiscoGen: Procedural Generation of Algorithm Discovery Tasks in Machine Learning | Foster et al. 2026 References: Goldie et al. (2026) DiscoGen arXiv:2603.17863 |
| claim | ts-claim-c3-ai-scientist-s2-novelty | The AI Scientist’s idea-generation filter discards ideas that are too similar to existing literature by querying the Semantic Scholar API (plus web access); novelty is therefore a retrieved-paper simi | proposed |
| claim | ts-claim-z1-listwise-collapse-global-discrimination | The drop of Accuracy@1 from 61.3% (N=2) to 31.1% (N=5) in Table 3 indicates that the LLM judge lacks global discrimination capability beyond binary interactions. | contradicted |
| claim | ts-claim-s1-novelty-not-significance | A claim can be graph-novel vs TeamScience JSONL and still be insignificant if it would not change a #177 rule, a cheapest test, or the next ingest walk. | proposed |