Fragility of Value under Imperfect Alignment
This paper presents a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world.
Winter Cross, Léo Cymbalista, A. Harwood et al.
· 0 citations