An algorithm is given, requiring no knowledge of the horizon, whose cumulative regret satisfies R_t, whose cumulative regret satisfies 1 + O(\sqrt{\ln \ln n / \ln n})\bigr)\sqrt{t \ln n / 2}$ simultaneously for every $t \ge 1$.
Yang Cai, Vineet Gupta, Yan-Chen Jiang et al.· 1 citation
We give a deterministic algorithm for online inverse linear optimization with regret $O(d)$, uniform in the horizon and $O(d^{2})$ time per round. A bound of this order was obtained recently by Dewasurendra, settling a question of Gollapudi et al.\ and of Oki and Sakaue, but by an improper rule that enumerates covers a...
Yang Cai, Anupam Gupta, Vineet Gupta et al.· 2 citations· ⚡1
The experiments demonstrate that this new deep learning framework can almost precisely replicate all known solutions from theory, expand to more complex settings, and be used to establish the optimality of new designs for data markets and make conjectures in regard to the structure of optimal designs.
S. Ravindranath, Yan-Chen Jiang, David C. Parkes· Neural Information Processin...· 17 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.