[Notes] Agent Security Is a Systems Problem

Last updated Agent Runtime Security Papers

I recently went over two articles from March-May 2026 that highlight how we must always assume that aligned models can make unexpected decisions. Perplexity’s Security Considerations for Artificial Intelligence Agents develops a layered defense strategy; Christodorescu et al.’s Agent Security Is a Systems Problem treats the model as untrusted and enforce security invariants outside it.

A boundary that holds only when the model correctly interprets an instruction or recognizes an attack is a behavioral expectation, not an enforceable restriction. Moving enforcement outside the model is necessary, but letting the same untrusted model freely define its own policy would recreate the original dependency. The articles recognize open research problems including verifiable translation of user’s intent from natural language into enforceable constraints.

Taking the untrusted model assumption seriously changes how we design agents. A model-generated tool call is a request for authority, not proof of authorization. These requirements become harder when the workflow itself emerges at runtime, rather than following a program whose resource needs are known in advance. There are excellent analogies in here for people working on code integrity, particularly dynamic/JITted code security.

Related designs explore how to enforce these boundaries: aflock places authorization and evidence generation in an external MCP server, while Grimlock uses eBPF-mediated sandbox boundaries and attested channels to enforce identity and constrain delegation outside agent code. Many of the projects I’m tracking in the bureado/awesome-agent-runtime-security collection are attempting to tackle this angle.