Skip to content

Author

Y. Shkolnikov

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Artificial Id: Drive and Persistent Alignment in Agentic AI

Results show that adaptive direction can emerge without being explicitly specified as a behavioral objective, and the same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries.

Y. Shkolnikov · 0 citations
#artificial intelligence Preprint Sep 2026

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retro...

Y. Shkolnikov · 1 citation
Preprint Jul 2026

Composable Trust for Language Models: A proven boundary and a measured defense

In a language model, instructions and data share one token stream, so nothing inside the model's generation can keep untrusted text from steering it. We develop a trust model that places the authority to act outside the model, in code: a source's standing, not its content, decides which operation runs and whether it ac...

Y. Shkolnikov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.