Follow

Variance-Dependent Regret Bounds for Linear Bandits and Reinforcement Learning: Adaptivity and Computational Efficiency. (arXiv:2302.10371v1 [cs.LG]) arxiv.org/abs/2302.10371

· · feed2toot · 0 · 0 · 0
Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.