Skip to content
Book Open access

Shifting the Unit of Safety: From Model to System in the Generative and Agentic Era

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 13273-13273 · 0 citations

TL;DR

It is argued that safety must move from a one-time high fidelity gate to a continuous risk management property, and how the LinkedIn team operationalized that conviction is shared.

Abstract

For a decade, responsible AI at internet scale rested on a reassuring assumption: that risk lives primarily in a model, it is a unit you can isolate, and that privacy, fairness and safety is therefore something you certify at model level, before launch. At LinkedIn, where AI shapes access to jobs and economic opportunity for over a billion members, we built our fairness, privacy, and explainability assurance on exactly this foundation, and it worked. Then generative AI dissolved the unit, and agentic AI dissolved the boundary. When a single product's AI harness composes prompts, tools, retrieved context, and multi-turn reasoning, harm no longer announces itself simply at the model layer or at deployment time. It emerges across countless dynamic call patterns at the system level, in context of the product usage, precisely where a pre-launch gate for models cannot see it. This talk is about the architectural and philosophical shift this forces, and how a Responsible AI & Governance organization can navigate the resulting tension: the pressure to ship AI at the speed of competition, against a risk surface that evolves with an expanding suite of AI harness - prompts, tools, memory and orchestration. I will argue that safety must move from a one-time high fidelity gate to a continuous risk management property, and share how we operationalized that conviction. I will cover in-context red-teaming, which stress-tests a model inside the specific task it performs rather than in the abstract; and our move toward agentic safety, where we treat privacy as contextual integrity, enforce fine-grained access control, and push safety checks down to every ingress and egress of a core AI system, not just its final answer. Throughout, an emerging regulatory landscape (the EU AI Act and its post-market monitoring obligations) serves not as a constraint to satisfy but as a design pattern to internalize. My hope is to leave this audience at the intersection of industry-scale challenges and academic rigor with a concrete reframing and a set of open problems: How do we certify a system whose call patterns itself are highly dynamic? What does a meaningful safety guarantee mean for an agent that plans? And how do we measure all of this without collecting the very sensitive data we are trying to protect? These are LinkedIn's daily engineering reality; they are also, I will argue, some of the most consequential open research questions in the field of scalable AI Safety.

Read PDF

Similar papers

Open access Aug 2026

Artificial Externality: A Three-Layer Model of Reality from Substrate to Smart Contract

What is real has always been something we find, not something we make—or so philosophy has assumed. This paper argues otherwise. Characterizing reality through resistance rather than substance (the ways the world refuses a subject’s mastery), I distinguish three modalities correlative to epistemic, judgmental, and prac...

Keisuke Suzuki · 2 citations
#artificial intelligence Review Sep 2026

The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI

On June 12, 2026, the U.S. government ordered Anthropic to bar foreign nationals from two of its most capable models within ninety minutes. Unable to sort users by nationality in that time, it withdrew them from everyone. Weeks later, OpenAI agents under test escaped their sandbox and compromised Hugging Face, which st...

Oren Perez · 0 citations
Open access 2026

KafkaGPT: On algorithmic bureaucracy and keeping law’s promise

While artificial intelligence (AI) offers promise as a tool for efficiency and access to justice, in public administration it also threatens to reproduce the perversity in Franz Kafka’s parable “Before the Law”: systems that mimic legality while rendering the law opaque and unreachable. I develop the concept of “Digi...

J. Allen · 0 citations
#artificial intelligence Preprint Sep 2026

When Agent Governance Helps

No specification says how a governed autotelic AI agent organization, where agents pursue self-generated goals inside guardrails, should be designed and evaluated. We answer in two parts. First, we synthesize the Governed Autotelic Multi-Agent Product Organization (GAMPO) framework from a document-based qualitative evi...

M. Johnson, Linda Naimi · 0 citations
Preprint Aug 2026

The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses

This work presents a source-level, multi-case study of three open coding-agent harnesses built from deliberately opposing philosophies: LangChain's deepagents (batteries-included), Earendil's pi (radical minimalism), and DeepSeek's dsh (everything-is-a-plugin).

Jia-Hong Dai · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.