Preprints and Working papers

No entries yet.

Publications

[ICML-26]

Position: RL Researchers Need to Distinguish Between Solving Simulators and Using Simulators as a Proxy

International Conference on Machine Learning Position Paper Track (ICML) · 2026

PDF · Code ·

[JMLR-26]

Investigating the Histogram Loss in Regression

Journal of Machine Learning Research (JMLR) · 2026

PDF · Code ·

[ICLR-26]

Distributions as Actions: A Unified Framework for Diverse Action Spaces

International Conference on Learning Representations (ICLR) · 2026

PDF · Code · Video ·

[ICLR-26]

Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models

International Conference on Learning Representations (ICLR) · 2026

PDF ·

[AAMAS-26]

When is Offline Policy Selection Sample Efficient for Reinforcement Learning?

International Conference on Autonomous Agents and Multiagent Systems (AAMAS), · 2026

PDF ·

[CCE-26]

PC-Gym: Benchmark environments for process control problems

Computers & Chemical Engineering · 2026

PDF ·

[RLJ-26]

Forager: a lightweight testbed for continual learning with partial observability in RL

Reinforcement Learning Journal (RLJ) · 2026

PDF ·

[RLJ-26]

Forager: a lightweight testbed for continual learning with partial observability in RL

Reinforcement Learning Journal (RLJ) · 2026

PDF ·

[RLC-25]

An Analysis of Action-Value Temporal-Difference Methods That Learn State Values

Reinforcement Learning Conference (RLC) · 2025

PDF ·

[RLC-25]

Deep Reinforcement Learning with Gradient Eligibility Traces

Reinforcement Learning Conference (RLC) · 2025

PDF · Code · Video ·

[RLC-25]

Rethinking the Foundations for Continual Reinforcement Learning

Reinforcement Learning Conference (RLC) · 2025

PDF ·

[RLC-25]

Investigating the Utility of Mirror Descent in Off-policy Actor-Critic.

Reinforcement Learning Conference (RLC) · 2025

PDF ·

[RLC-25]

Value Bonuses using Ensemble Errors for Exploration in Reinforcement Learning

Reinforcement Learning Conference (RLC) · 2025

PDF ·

[ICML-25]

Position: Lifetime tuning is incompatible with continual reinforcement learning

International Conference on Machine Learning (ICML) · 2025

PDF ·

[ICLR-25]

q-exponential family for policy optimization

International Conference on Learning Representations (ICLR) · 2025

PDF ·

[PNAS-24]

Is Okham’s razor losing its edge? New perspectives on the principle of model parsimony.

Proceedings of the National Academy of Sciences (PNAS) · 2024

PDF ·

[JMLR-24]

Data-Efficient Policy Evaluation Through Behavior Policy Search

Journal of Machine Learning Research (JMLR) · 2024

PDF ·

[JMLR-24]

Goal-Space Planning with Subgoal Models

Journal of Machine Learning Research (JMLR) · 2024

PDF ·

[JMLR-24]

Empirical Design in Reinforcement Learning

Journal of Machine Learning Research (JMLR) · 2024

PDF ·

[RLC-24]

The Cross-environment Hyperparameter Setting Benchmark for Reinforcement Learning

Reinforcement Learning Conference (RLC) · 2024

PDF ·

[RLC-24]

Investigating the Interplay of Prioritized Replay and Generalization

Reinforcement Learning Conference (RLC) · 2024

PDF ·

[RLC-24]

Demystifying the Recency Heuristic in Temporal-Difference Learning

Reinforcement Learning Conference (RLC) · 2024

PDF ·

[JAIR-24]

Multistep Predecessor Models and Mitigating Errors due to Hallucinated Value in Dyna-Style Planning

Journal of AI Research (JAIR) · 2024

PDF ·

[ICML-24]

Compound Returns Reduce Variance in Reinforcement Learning

International Conference on Machine Learning (ICML) · 2024

PDF ·

[ICML-24]

Position Paper: Limitations of and Alternatives to Benchmarking in Reinforcement Learning Research

International Conference on Machine Learning (ICML) · 2024

PDF ·

[TMLR-24]

Offline Reinforcement Learning via Tsallis Regularization

Transactions on Machine Learning Research (TMLR) · 2024

PDF ·

[Neurips-24]

Real-time recurrent learning using trace units in reinforcement learning

Advances in Neural Information Processing Systems (Neurips) · 2024

PDF · Code ·

[AIJ-23]

Investigating the Properties of Neural Network Representations in Reinforcement Learning

Artificial Intelligence Journal (AIJ) · 2023

PDF ·

[MLJ-23]

GVFs in the Real World: Making Predictions Online for Water Treatment

Machine Learning (MLJ) · 2023

PDF ·

[NeurIPS-23]

General Munchausen Reinforcement Learning with Tsallis Kullback-Leibler Divergence

Advances in Neural Information Processing Systems (NeurIPS) · 2023

PDF ·

[TMLR-23]

Resmax: An Alternative Soft-Greedy Operator for Reinforcement Learning

Transactions on Machine Learning Research (TMLR) · 2023

PDF ·

[JMLR-23]

Scalable Real-Time Recurrent Learning Using Columnar-Constructive Networks

Journal of Machine Learning Research (JMLR) · 2023

PDF ·

[ICML-23]

Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement Learning

International Conference on Machine Learning (ICML) · 2023

PDF ·

[CoLLAs-23]

Measuring and Mitigating Interference in Reinforcement Learning

Conference on Lifelong Learning Agents (CoLLAs) · 2023

PDF ·

[JAIR-23]

Exploiting Action Impact Regularity and Exogenous State Variables for Offline Reinforcement Learning

Journal of Artificial Intelligence Research · 2023

PDF ·

[JMLR-23]

Off-Policy Actor-Critic with Emphatic Weightings

Journal of Machine Learning Research (JMLR) · 2023

PDF ·

[ICLR-23]

Greedy Actor-Critic: A New Conditional Cross-Entropy Method for Policy Improvement

International Conference on Representation Learning (ICLR) · 2023

PDF ·

[ICLR-23]

The In-Sample Softmax for Offline Reinforcement Learning

International Conference on Representation Learning (ICLR) · 2023

PDF ·

[AISTATS-23]

Asymptotically Unbiased Off-Policy Policy Evaluation when Reusing Old Data in Nonstationary Environments

International Conference on AI and Statistics (AISTATS) · 2023

[TMLR-22]

Representation Alignment in Neural Networks

Transactions on Machine Learning Research (TMLR) · 2022

PDF ·

[TPAMI-22]

Robust Losses for Learning Value Functions

Transactions on Pattern Analysis and Machine Learning (TPAMI) · 2022

PDF ·

[JMLR-22]

Greedification Operators for Policy Optimization: Investigating Forward and Reverse KL Divergences

Journal of Machine Learning Research (JMLR) · 2022

PDF ·

[TMLR-22]

No More Pesky Hyperparameters: Offline Hyperparameter Tuning for RL

Transactions on Machine Learning Research (TMLR) · 2022

[UAI-22]

Understanding and Mitigating the Limitations of Prioritized Replay

Uncertainty in AI (UAI) · 2022

PDF ·

[ICML-22]

A Temporal-Difference Approach to Policy Gradient Estimation

International Conference on Machine Learning (ICML) · 2022

PDF ·

[JMLR-22]

A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning

Journal of Machine Learning Research (JMLR) · 2022

PDF ·

[AISTATS-22]

An Alternate Policy Gradient Estimator for Softmax Policies

International Conference on on Artificial Intelligence and Statistics (AISTATS) · 2022

PDF ·

[ICLR-22]

Resonance in Weight Space: Covariate Shift Can Drive Divergence of SGD with Momentum

International Conference on Learning Representations (ICLR) · 2022

PDF ·

[IEEE-21]

Sim2Real in Robotics and Automation: Applications and Challenges

IEEE Transactions on Automation Science and Engineering · 2021

PDF ·

[NeurIPS-21]

Continual Auxiliary Task Learning

Advances in Neural Information Processing Systems (NeurIPS) · 2021

PDF ·

[NeurIPS-21]

Structural Credit Assignment in Neural Networks using Reinforcement Learning

Advances in Neural Information Processing Systems (NeurIPS) · 2021

PDF ·

[ICLR-21]

Fuzzy Tiling Activations: A Simple Approach to Learning Sparse Representations Online

International Conference on Learning Representations (ICLR) · 2021

PDF ·

[JAIR-21]

General Value Function Networks

Journal of AI Research (JAIR) · 2021

PDF ·

[NeurIPS-20]

An implicit function learning approach for parametric modal regression

Advances in Neural Information Processing Systems (NeurIPS) · 2020

PDF ·

[NeurIPS-20]

Towards Safe Policy Improvement for Non-Stationary MDPs

Advances in Neural Information Processing Systems (NeurIPS) · 2020

PDF ·

[JAIR-20]

Adapting Behaviour via Intrinsic Reward: A Survey and Empirical Study

Journal of AI Research (JAIR) · 2020

PDF ·

[ICML-20]

Gradient Temporal-Difference Learning with Regularized Corrections

International Conference on Machine Learning (ICML) · 2020

PDF ·

[ICML-20]

Selective Dyna-style Planning Under Limited Model Capacity

International Conference on Machine Learning (ICML) · 2020

PDF ·

[ICML-20]

Optimizing for the Future in Non-Stationary MDPs

International Conference on Machine Learning (ICML) · 2020

PDF ·

[ICLR-20]

Training Recurrent Neural Networks Online by Learning Explicit State Variables

International Conference on Learning Representations (ICLR) · 2020

PDF · Code ·

[ICLR-20]

Maxmin Q-learning: Controlling the Estimation Bias of Q-learning

International Conference on Learning Representations (ICLR) · 2020

PDF · Code ·

[AAMAS-20]

Maximizing Information Gain in Partially Observable Environments via Prediction Rewards

International Conference on Autonomous Agents and Multi-agent Systems (AAMAS) · 2020

PDF ·

[NeurIPS-19]

Meta-Learning Representations for Continual Learning

Advances in Neural Information Processing Systems (NeurIPS) · 2019

PDF · Code ·

[NeurIPS-19]

Importance Resampling for Off-policy Prediction

Advances in Neural Information Processing Systems (NeurIPS) · 2019

PDF · Code ·

[NeurIPS-19]

Learning Macroscopic Brain Connectomes via Group-Sparse Factorization

Advances in Neural Information Processing Systems (NeurIPS) · 2019

[IJCAI-19]

Planning with Expectation Models

International Joint Conference on Artificial Intelligence (IJCAI) · 2019

PDF ·

[IJCAI-19]

Hill Climbing on Value Estimates for Search-control in Dyna

International Joint Conference on Artificial Intelligence (IJCAI) · 2019

PDF ·

[ICLR-19]

Two-Timescale Networks for Nonlinear Value Function Approximation

International Conference on Learning Representations (ICLR) · 2019

PDF ·

[AAAI-19]

The Utility of Sparse Representations for Control in Reinforcement Learning

AAAI Conference on Artificial Intelligence · 2019

PDF ·

[AAAI-19]

Meta-descent for online, continual prediction

AAAI Conference on Artificial Intelligence · 2019

PDF ·

[NIPS-18]

An Off-policy Policy Gradient Theorem Using Emphatic Weightings

Advances in Neural Information Processing Systems (NIPS) · 2018

PDF ·

[NIPS-18]

Supervised autoencoders: Improving generalization performance with unsupervised regularizers

Advances in Neural Information Processing Systems (NIPS) · 2018

PDF ·

[NIPS-18]

Context-dependent upper-confidence bounds for directed exploration

Advances in Neural Information Processing Systems (NIPS) · 2018

PDF ·

[ICML-18]

Improving Regression Performance with Distributional Losses

International Conference on Machine Learning (ICML) · 2018

PDF ·

[ICML-18]

Reinforcement Learning with Function-Valued Action Spaces for Partial Differential Equation Control

International Conference on Machine Learning (ICML) · 2018

PDF ·

[IJCAI-18]

Organizing experience: a deeper look at replay mechanisms for sample-based planning in continuous state domains

International Joint Conference on Artificial Intelligence (IJCAI) · 2018

PDF ·

[UAI-18]

High-confidence error estimates for learned value functions

Uncertainty in Artificial Intelligence (UAI) · 2018

PDF ·

[UAI-18]

Comparing Direct and Indirect Temporal-Difference Methods for Estimating the Variance of the Return

Uncertainty in Artificial Intelligence (UAI) · 2018

PDF ·

[NIPS-17]

Multi-view Matrix Factorization for Linear Dynamical System Estimation

Advances in Neural Information Processing Systems (NIPS) · 2017

PDF ·

[ICML-17]

Unifying task specification in reinforcement learning

International Conference on Machine Learning (ICML) · 2017

PDF ·

[ICML-17]

Adapting kernel representations online using submodular maximization

International Conference on Machine Learning (ICML) · 2017

PDF ·

[UAI-17]

Effective sketching methods for value function approximation

Uncertainty in Artificial Intelligence (UAI) · 2017

PDF ·

[IJCAI-17]

Learning sparse representations in reinforcement learning with sparse coding

International Joint Conference on Artificial Intelligence (IJCAI) · 2017

PDF ·

[AAAI-17]

Accelerated Gradient Temporal Difference Learning

AAAI Conference on Artificial Intelligence · 2017

PDF · Code ·

[AAAI-17]

Recovering true classifier performance in positive-unlabeled learning

AAAI Conference on Artificial Intelligence · 2017

PDF ·

[NIPS-16]

Estimating the class prior and posterior from noisy positives and unlabeled data

Advances in Neural Information Processing Systems (NIPS) · 2016

PDF ·

[JMLR-16]

Identifying global optimality for dictionary learning

In submission to JMLR · 2016

PDF ·

[JMLR-16]

Nonparametric semi-supervised learning of class proportions

In submission to JMLR · 2016

PDF ·

[IJCAI-16]

Incremental Truncated LSTD

International Joint Conference on Artificial Intelligence (IJCAI) · 2016

PDF · Code ·

[AAMAS-16]

Investigating practical, linear temporal difference learning

Autonomous Agents and Multi-agent Systems (AAMAS) · 2016

PDF ·

[AAMAS-16]

A Greedy Approach to Adapting the Trace Parameter for Temporal Difference Learning

Autonomous Agents and Multi-agent Systems (AAMAS) · 2016

PDF ·

[JMLR-16]

An Emphatic Approach to the Problem of Off-policy Temporal-Difference Learning

Journal of Machine Learning Research (JMLR) · 2016

PDF ·

[ECML-15]

Scalable Metric Learning for Co-embedding

ECML PKDD · 2015

PDF ·

[AAAI-15]

Optimal Estimation of Multivariate ARMA Models

AAAI Conference on Artificial Intelligence · 2015

PDF · Code ·

[DCC-13]

Partition Tree Weighting

Data Compression Conference (DCC) · 2013

PDF ·

[NIPS-12]

Convex Multi-view Subspace Learning

Advances in Neural Information Processing Systems (NIPS) · 2012

PDF · Code ·

[ICML-12]

Off-Policy Actor-Critic

International Conference on Machine Learning (ICML) · 2012

PDF ·

[AISTATS-12]

Generalized Optimal Reverse Prediction

International Conference on Artificial Intelligence and Statistics (AISTATS) · 2012

PDF · Code ·

[AAAI-11]

Convex Sparse Coding, Subspace Learning, and Semi-Supervised Extensions

AAAI Conference on Artificial Intelligence (AAAI) · 2011

PDF · Code ·

[NIPS-10]

Interval Estimation for Reinforcement-Learning Algorithms in Continuous-State Domains.

Advances in Neural Information Processing Systems (NIPS) · 2010

PDF ·

[NIPS-10]

Relaxed Clipping: A Global Training Method for Robust Regression and Classification

Advances in Neural Information Processing Systems (NIPS) · 2010

PDF ·

[IJCAI-09]

Learning a Value Analysis Tool For Agent Evaluation

International Joint Conference on Artificial Intelligence (IJCAI) · 2009

PDF ·

[ICML-09]

Optimal Reverse Prediction: A Unified Perspective on Supervised, Unsupervised and Semi-supervised Learning

International Conference on Machine Learning (ICML) · 2009

PDF ·

Theses

[UA-09]

Regularized factor models

University of Alberta · 2009

PDF ·

[UA-09]

A General Framework for Reducing Variance in Agent Evaluation

University of Alberta · 2009

PDF ·