Research
Broadly, I am interested in improving the scalability of RL algorithms, via
paradigms like offline RL, unsupervised RL, and RL pre-training. My long-horizon goal is to develop
large RL models that can learn in a completely unsupervised manner from the vast
collections of unlabeled data on the Internet.
My current focus is on using offline goal-conditioned RL to accomplish the above.
|
|