home / team-science

citation_edge

cites / cited_by / stub edges with a locator. Only when bibliography or OpenAlex referenced_works supports them.

Data license: Space charter; records cite primary sources · Data source: TeamScience Space repository

37 rows where from_lom_id = "arxiv:2001.08361"

✎ View and edit SQL

This data as json, CSV (advanced)

Link from_lom_id to_lom_id kind locator
arxiv:2001.08361,arxiv:1412.6980,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Adam: A Method for Stochastic Optimization arxiv:1412.6980 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1605.06431,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Residual Networks Behave Like Ensembles of Relatively Shallow Networks arxiv:1605.06431 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1605.07146,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Wide Residual Networks arxiv:1605.07146 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1606.06737,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Criticality in Formal Languages and Statistical Physics arxiv:1606.06737 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1804.04235,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Adafactor: Adaptive Learning Rates with Sublinear Memory Cost arxiv:1804.04235 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1806.07572,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Neural Tangent Kernel: Convergence and Generalization in Neural Networks arxiv:1806.07572 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1807.03819,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Universal Transformers arxiv:1807.03819 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1811.02084,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Mesh-TensorFlow: Deep Learning for Supercomputers arxiv:1811.02084 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1811.03600,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Measuring the Effects of Data Parallelism on Neural Network Training arxiv:1811.03600 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1811.06965,cites Scaling Laws for Neural Language Models arxiv:2001.08361 GPipe: Efficient Training of Giant Neural Networks using Pipeline\n Parallelism arxiv:1811.06965 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1812.04754,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Gradient Descent Happens in a Tiny Subspace arxiv:1812.04754 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1812.06162,cites Scaling Laws for Neural Language Models arxiv:2001.08361 An Empirical Model of Large-Batch Training arxiv:1812.06162 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1901.10159,cites Scaling Laws for Neural Language Models arxiv:2001.08361 An Investigation into Neural Net Optimization via Hessian Eigenvalue Density arxiv:1901.10159 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1904.10509,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Generating Long Sequences with Sparse Transformers arxiv:1904.10509 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1906.02909,cites Scaling Laws for Neural Language Models arxiv:2001.08361 AutoGrow: Automatic Layer Growing in Deep Convolutional Networks arxiv:1906.02909 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1906.06669,cites Scaling Laws for Neural Language Models arxiv:2001.08361 One Epoch Is All You Need arxiv:1906.06669 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1907.04164,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model arxiv:1907.04164 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,arxiv:1909.12673,cites Scaling Laws for Neural Language Models arxiv:2001.08361 A Constructive Prediction of the Generalization Error Across Scales arxiv:1909.12673 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.1006/csla.2001.0174,cites Scaling Laws for Neural Language Models arxiv:2001.08361 A bit of progress in language modeling doi:10.1006/csla.2001.0174 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.1007/0-387-30623-4,cites Scaling Laws for Neural Language Models arxiv:2001.08361 All of Nonparametric Statistics doi:10.1007/0-387-30623-4 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.1016/j.neunet.2020.08.022,cites Scaling Laws for Neural Language Models arxiv:2001.08361 High-dimensional dynamics of generalization error in neural networks doi:10.1016/j.neunet.2020.08.022 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.1088/1742-5468/ab633c,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Scaling description of generalization with number of parameters in deep learning doi:10.1088/1742-5468/ab633c cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.1093/oso/9780198821939.001.0001,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Introduction to the Theory of Complex Systems doi:10.1093/oso/9780198821939.001.0001 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.1109/cvpr.2017.323,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Growing a Brain: Fine-Tuning by Increasing Model Capacity doi:10.1109/cvpr.2017.323 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.1109/iccv.2015.11,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books doi:10.1109/iccv.2015.11 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.1145/3065386,cites Scaling Laws for Neural Language Models arxiv:2001.08361 ImageNet classification with deep convolutional neural networks doi:10.1145/3065386 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.1145/3293883.3295710,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Beyond human-level accuracy doi:10.1145/3293883.3295710 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.1209/0295-5075/26/4/001,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Entropy and Long-Range Correlations in Literary English doi:10.1209/0295-5075/26/4/001 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.18653/v1/n19-1423,cites Scaling Laws for Neural Language Models arxiv:2001.08361 BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding doi:10.18653/v1/n19-1423 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.18653/v1/p16-1162,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Neural Machine Translation of Rare Words with Subword Units doi:10.18653/v1/p16-1162 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.3115/1073012.1073017,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Scaling to very very large corpora for natural language disambiguation doi:10.3115/1073012.1073017 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,doi:10.4230/lipics.cp.2025.31,cites Scaling Laws for Neural Language Models arxiv:2001.08361 HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanities Use Case doi:10.4230/lipics.cp.2025.31 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,openalex:W2214916291,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Analysis of a Random Forests Model openalex:W2214916291 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,openalex:W2464989499,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Critical Behavior from Deep Dynamics: A Hidden Dimension in Natural Language. openalex:W2464989499 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,openalex:W2900531695,cites Scaling Laws for Neural Language Models arxiv:2001.08361 The Full Spectrum of Deep Net Hessians At Scale: Dynamics with Sample Size. openalex:W2900531695 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,openalex:W2907127169,cites Scaling Laws for Neural Language Models arxiv:2001.08361 Reconciling modern machine learning and the bias-variance trade-off openalex:W2907127169 cites OpenAlex referenced_works of W3001279689
arxiv:2001.08361,openalex:W2990704537,cites Scaling Laws for Neural Language Models arxiv:2001.08361 SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems openalex:W2990704537 cites OpenAlex referenced_works of W3001279689

Advanced export

JSON shape: default, array, newline-delimited, object

CSV options:

CREATE TABLE citation_edge (
  from_lom_id  TEXT NOT NULL REFERENCES paper(lom_id),
  to_lom_id    TEXT NOT NULL REFERENCES paper(lom_id),
  kind         TEXT NOT NULL CHECK (kind IN ('cites','cited_by','stub')),
  locator      TEXT,
  PRIMARY KEY (from_lom_id, to_lom_id, kind),
  CHECK (from_lom_id <> to_lom_id)
);
Powered by Datasette · Queries took 5.975ms · Data license: Space charter; records cite primary sources · Data source: TeamScience Space repository