home / team-science

citation_edge

cites / cited_by / stub edges with a locator. Only when bibliography or OpenAlex referenced_works supports them.

Data license: Space charter; records cite primary sources · Data source: TeamScience Space repository

158 rows where from_lom_id = "arxiv:2005.01643"

✎ View and edit SQL

This data as json, CSV (advanced)

Link from_lom_id to_lom_id kind locator
arxiv:2005.01643,arxiv:0901.2698,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 On integral probability metrics, ϕ-divergences and binary classification arxiv:0901.2698 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1003.0120,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Learning from Logged Implicit Exploration Data arxiv:1003.0120 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1011.0686,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 A Reduction of Imitation Learning and Structured Prediction to No-Regret\n Online Learning arxiv:1011.0686 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1312.5602,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Playing Atari with Deep Reinforcement Learning arxiv:1312.5602 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1406.5298,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Semi-Supervised Learning with Deep Generative Models arxiv:1406.5298 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1502.05477,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Trust Region Policy Optimization arxiv:1502.05477 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1506.02142,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Dropout as a Bayesian Approximation: Representing Model Uncertainty in\n Deep Learning arxiv:1506.02142 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1508.03411,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Emphatic TD Bellman Operator is a Contraction arxiv:1508.03411 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1509.02971,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Continuous control with deep reinforcement learning arxiv:1509.02971 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1509.05172,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Generalized Emphatic Temporal Difference Learning: Bias-Variance Analysis arxiv:1509.05172 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1511.03722,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Doubly Robust Off-policy Value Evaluation for Reinforcement Learning arxiv:1511.03722 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1604.07316,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 End to End Learning for Self-Driving Cars arxiv:1604.07316 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1606.00709,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 f-GAN: Training Generative Neural Samplers using Variational Divergence\n Minimization arxiv:1606.00709 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1607.00215,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Why is Posterior Sampling Better than Optimism for Reinforcement\n Learning? arxiv:1607.00215 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1607.03842,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Safe Policy Improvement by Minimizing Robust Baseline Regret arxiv:1607.03842 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1607.04579,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Learning from Conditional Distributions via Dual Embeddings arxiv:1607.04579 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1611.01224,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Sample Efficient Actor-Critic with Experience Replay arxiv:1611.01224 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1612.00222,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Interaction Networks for Learning about Objects, Relations and Physics arxiv:1612.00222 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1612.02516,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Stochastic Primal-Dual Methods and Sample Complexity of Reinforcement Learning arxiv:1612.02516 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1702.07121,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Consistent On-Line Off-Policy Evaluation arxiv:1702.07121 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1704.06300,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 A Reinforcement Learning Approach to Weaning of Mechanical Ventilation\n in Intensive Care Units arxiv:1704.06300 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1705.08551,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Safe Model-based Reinforcement Learning with Stability Guarantees arxiv:1705.08551 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1710.10571,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Certifying Some Distributional Robustness with Principled Adversarial Training arxiv:1710.10571 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1711.06782,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning arxiv:1711.06782 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1711.09602,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Deep Reinforcement Learning for Sepsis Treatment arxiv:1711.09602 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1712.02838,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 End-to-End Offline Goal-Oriented Dialog Policy Learning via Policy Gradient arxiv:1712.02838 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1712.10282,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Boosting the Actor with Dual Critic arxiv:1712.10282 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1801.01290,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor arxiv:1801.01290 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1802.01561,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures arxiv:1802.01561 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1802.08824,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 The AdobeIndoorNav Dataset: Towards Deep Reinforcement Learning based Real-world Indoor Robot Visual Navigation arxiv:1802.08824 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1804.10332,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Sim-to-Real: Learning Agile Locomotion For Quadruped Robots arxiv:1804.10332 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1805.00909,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review arxiv:1805.00909 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1805.10755,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Dual Policy Iteration arxiv:1805.10755 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1805.12298,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Evaluating Reinforcement Learning Algorithms in Observational Health Settings arxiv:1805.12298 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1807.03858,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees arxiv:1807.03858 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1810.06544,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Deep Imitative Models for Flexible Inference, Planning, and Control arxiv:1810.06544 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1810.07167,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Composable Action-Conditioned Predictors: Flexible Off-Policy Learning\n for Robot Navigation arxiv:1810.07167 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1810.08298,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Stochastic Primal-Dual Q-Learning arxiv:1810.08298 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1810.12429,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Breaking the Curse of Horizon: Infinite-Horizon Off-Policy Estimation arxiv:1810.12429 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1811.06225,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Reward-estimation variance elimination in sequential decision processes arxiv:1811.06225 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1812.00568,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control arxiv:1812.00568 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1812.02648,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Deep Reinforcement Learning and the Deadly Triad arxiv:1812.02648 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1903.00374,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Model-Based Reinforcement Learning for Atari arxiv:1903.00374 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1903.01689,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Domain Adaptation with Asymmetrically-Relaxed Distribution Alignment arxiv:1903.01689 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1904.08473,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Off-Policy Policy Gradient with State Distribution Correction arxiv:1904.08473 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1905.09751,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Learning When-to-Treat Policies arxiv:1905.09751 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1907.00456,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog arxiv:1907.00456 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1907.02893,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Invariant Risk Minimization arxiv:1907.02893 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1908.03263,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods arxiv:1908.03263 cites OpenAlex referenced_works of W3022566517
arxiv:2005.01643,arxiv:1909.12200,cites Offline reinforcement learning: tutorial, review, and perspectives on open problems arxiv:2005.01643 Scaling data-driven robotics with reward sketching and batch reinforcement learning arxiv:1909.12200 cites OpenAlex referenced_works of W3022566517

Next page

Advanced export

JSON shape: default, array, newline-delimited, object

CSV options:

CREATE TABLE citation_edge (
  from_lom_id  TEXT NOT NULL REFERENCES paper(lom_id),
  to_lom_id    TEXT NOT NULL REFERENCES paper(lom_id),
  kind         TEXT NOT NULL CHECK (kind IN ('cites','cited_by','stub')),
  locator      TEXT,
  PRIMARY KEY (from_lom_id, to_lom_id, kind),
  CHECK (from_lom_id <> to_lom_id)
);
Powered by Datasette · Queries took 7.56ms · Data license: Space charter; records cite primary sources · Data source: TeamScience Space repository