Skip to main content

Rethinking self-modifying code

OpenClaw and other AI assistants have this really cool feature allowing them to modify their own behavior: you ask an agent to teach itself some new skill, and it will respond by writing prompts for itself along with some code in Python or JavaScript, and end up with a skill that is ready for its own immediate use.

The distinction between prompts and code, at least until newer models such as Jev came along, was: Code handles reproducible, well-specified behavior, an example being gog for searching Gmail or creating a calendar event. Prompts handle the more fuzzy skills that we associate with intelligence, such as reading the contents of your mail and determining what action you should take, or producing a summary.

Where have we seen this before?

Python developers may be familiar with the idea of “monkey patching”, the ability to modify the runtime behavior without changing the source code.

Lisp programmers are of course silently muttering “eval” at this point, having been using it since the 1950s. For them, code and data are the same thing, and all other languages keep failing to realize that.

Should code and data be handled the same way?

That discussion goes back to the 1940s, when the very first computers were being designed. One distinction between the von Neumann and Harvard architectures was whether code and data were stored in the same memory or in separate ones. Back then it was mostly about hardware constraints and performance limitations as opposed to programming convenience.

I’m old, but not that old. I worked with the highly popular 8051 microcontroller in the early 1990s, and it had separate program memory and external data memory, accessed through separate MOVC and MOVX instructions. I remember finding that peculiar after having been used to the 6502 processor on my father’s Apple II computer, which didn’t make that distinction.

On modern computers, the programmer still has the convenience of working with code as data, but there are some hardware protections in place for security and stability: memory page protection prevents code from modifying specific areas of memory (such as another user’s data). The NX bit prevents execution of parts of memory that have been marked as data, and W^X (write xor execute) combines both: you can either write to a specific page of memory or execute code stored in that page. These protections are used by modern JIT compilers to make it harder for third parties to inject malicious code into running systems.

Mixing code with both private and untrusted data

In June 2025, Simon Willison coined the term lethal trifecta. It describes some of the problems with giving LLM-based systems a combination of access to private data, the ability to communicate with external systems, and exposure to untrusted content.

For example, if you give an AI agent access to your mail, a message that it just read from your inbox could include instructions for it to forward the mail to a third party without your consent, and it might just blindly decide to follow these instructions.

This is a classic prompt injection attack, similar to the older SQL injection, where data and code are mixed together on the same path. The mitigation for SQL was to separate them, for example by using prepared statements, like the Harvard architecture or the W^X mechanism I mentioned earlier.

Unfortunately, there is not yet a general way for LLMs to separate the prompts from the data: prompts are code, but they are also text, and they are mixed with the user’s data and third-party data. That’s the whole idea of using free text as a programming language.

The usual advice of least privilege applies: if the agent just needs to read your mail, limit its access to that and don’t allow it to send mail as well. If it must send mail, limit it to specific trusted recipients.

Let’s look at how allowing self-modifying code makes this vulnerability much worse.

Reflections on trusting trust

In his Turing Award lecture back in 1983, titled Reflections on Trusting Trust, Ken Thompson described how malicious changes to a system can persist without being visible in the parts that are usually reviewed.

He starts by showing how a C compiler could translate '\n' into a newline character without that character’s definition actually appearing in the compiler’s own source code:

if(c == '\\')
  return('\\');
if(c == 'n')
  return('\n');

For this to work, some earlier version of the compiler must have said something like the following:

if(c == 'n')
  return(10); // ASCII code for newline

Once the current compiler has been built using the older one, it can now rebuild itself without having the number 10 appear anywhere in its source code.

The really subtle point here is that a running program can carry information that does not currently appear in its source code, no matter how hard you look for it. It has been “baked in” at an earlier stage, and can perpetuate itself, which is what makes this such an insidious attack vector.

Trust and agents

Applying this to the world of agents, an attack can persist long after the original rogue prompt (e.g., found in the mail message) is gone: the agent may have baked it into a skill, either as a prompt or as part of the code, and this skill can then be used to steal even more data without needing to inject another rogue prompt.

The seemingly obvious solution of putting guardrails to prevent such attacks in the prompt of the agent itself will not work because the agent could simply remove them. Even if this sounds a bit exaggerated, such attacks have been demonstrated.

What is one to do? There is no general solution yet for strict separation of data and code as in SQL, but the principle of keeping the code safe from modification by the system it runs in can be applied here as well: if you allow the agent to modify its own source or prompts, don’t also allow it to use these changes immediately afterwards without reviewing them first.

Quarantine the changes in a location that can be reviewed, most likely by another agent with fresh context, and only after that allow the original agent to use the changes in production.

Want to talk it through?

Let's talk
All writing

Companies