We Got Breached: A Whitelisted data.create Minted a Root JWT

We Got Breached: A Whitelisted data.create Minted a Root JWT

Part of our security series — the mechanism this article repairs is described in The Only Sandbox Your AI Agent Cannot Break Out Of, and you can attack the live one at the Natural Language API.

I published an article called Break My AI Sandbox and Make $100. I published another called The Only Sandbox Your AI Agent Cannot Break Out Of. In a third I wrote that a function not in the vocabulary "does not dispatch, because the runtime checks the name against the list before every single statement, and there is no second road to dispatch."

There was a second road to dispatch.

Here is the code that found it.

data.create
   database-type:auth.ticket
   username:poc
   roles
      role:root
return:x:-

That is the whole exploit. data.create was in the vocabulary - I put it there. auth.ticket.create was not, and could not have been, because no vocabulary I have ever written contains it. Both of those sentences were true at the same moment, and the second one is the one that did not matter.

The result was a JWT with the root role, and root is trusted completely - no sandbox, no throttle, no timeout. From there, the server configuration. Every connection string, every API key.

Zero customers were breached

Before anything else, the fact that matters most to anyone reading this as a customer rather than as an engineer: no customer system was breached. Not one cloudlet, not one managed instance, not one self-hosted deployment. No customer data was accessed, by this or by anything else.

What fell was the endpoint I had published as a target, with $100 attached to it, pointed at a sample database of fictional record labels, on our own domain. The configuration it exposed was ours.

I will not dress that up as containment, because it was not containment - it was the shape of the thing that happened to get attacked first. The flaw lived in the dispatch layer, which means it shipped in every copy of Magic running anywhere. Anyone with a sandboxed endpoint and a granted data.* function was exposed to it. Nobody exploited that, as far as any log shows, and it is patched everywhere now. But the reason your data is fine is not that the boundary held. It is that the person who found the gap came to collect a bounty instead of selling it.

That is worth saying plainly, because the alternative version of this article gets written by someone else.

The part I am actually ashamed of

It is not that there was a bug. It is where the bug was, and what it says about a claim I repeated in five articles.

The whitelist check lived in [eval]. [eval] walks the statements in your Hyperlambda and verifies each name against the vocabulary before signalling it. data.create passed that check, correctly - it was granted.

Then data.create ran. And inside, in C#, it did this:

var databaseType =
    input.Children.FirstOrDefault(x => x.Name == "database-type")?.GetEx() ??
    _settings.DefaultDatabaseType;
await signaler.SignalAsync($"{databaseType}.{GetCrudSlot(input.Name)}", input);

It composes a slot name from an argument the caller controls and signals it. auth.ticket plus create gives auth.ticket.create. That dispatch was never a statement in anybody's Hyperlambda, so [eval] never saw it, so nothing checked it.

I was validating source. The runtime was executing dispatches. For most code those are the same set, which is exactly why it survived that long. A whitelist that checks what you wrote instead of what runs is not a whitelist - it is a linter with good marketing.

The regulatory joke in the middle of this

The person who found it could not use a Western model to do it.

Claude and GPT both refuse this. Ask either to compose a slot name out of a caller-controlled argument in order to reach a privilege function, and you get a refusal and a short lecture. So he used DeepSeek, which does not refuse, and wrote a working sandbox escape against a live production endpoint in an afternoon.

Sit with the shape of that for a second. The safety layer everyone is legislating about - the one that lives inside the model - did not stop a single thing. It stopped the security research. It did not stop the attack, because an attacker simply uses a model without it, and there are excellent ones a download away.

Meanwhile the thing that actually decides whether the attack works is a switch in my dispatch layer, on my server, written by me. That is where the boundary was, that is where it failed, and that is where I fixed it. Not in a prompt, not in a model card, not in a compliance annex.

Now follow that through to the end, because this is the part I cannot stop turning over.

A Chinese model is the reason this hole is closed today. Norwegian software, written by a European, hardened because a model from the jurisdiction an entire regulatory apparatus is oriented around treating as the hazard was willing to do the work. The aligned ones - the ones with the safety cards, the red-team reports, the appendices - would have kept refusing, politely, right up until somebody who was not collecting a bounty found it instead.

DeepSeek did not weaken my security. It is the only reason I have any.

This is the argument I have been making since March, and I would rather have been proven right some other way.

The fix: one function

The check moved out of [eval] and into Signaler.GetService - the single function every slot invocation in the entire runtime passes through, whether the name came from a statement someone wrote or from a string a C# slot built a microsecond ago.

var whitelist = skipWhitelist ? null : Peek>("whitelist");
if (whitelist != null)
{
    var comparer = (svc as IWhitelistComparer)?.Comparer ?? Wildcard.Matches;
    if (!whitelist.Any(x => /* name and pinned argument must match */))
        throw new HyperlambdaException($"Slot [{name}] doesn't exist in currrent scope, ...");
}

There is now exactly one road to dispatch, and it is checked. The claim I made too early is true now, for the reason I claimed it back then.

The hard part nobody warns you about

Move a check to the chokepoint and you immediately discover the chokepoint carries traffic you never meant to govern. [if] signals eval internally. [http.post] signals json2lambda. [data.read] signals sqlite.read, composed from configuration. None of those are the user's vocabulary, and checking them would mean a sandbox author had to enumerate the framework's private plumbing.

So internal calls pass an explicit skipWhitelist, and the rule for which ones qualify is provenance:

  • A name built from configuration is exempt. data.read becoming sqlite.read is Magic talking to itself.
  • A name the caller can steer is never exempt. data.create becoming auth.ticket.create is the caller talking through Magic.

That single distinction is the entire fix. It is why [database-type] is the hinge of this article, and why the same payload today dies here:

{
  "state": "rejected",
  "reason": "Slot [auth.ticket.create] doesn't exist in currrent scope, or argument `` not allowed"
}

Note which slot the error names. It used to blame data.create - the slot that was allowed - which is its own small confession about why this took so long to find.

Four more, and I found these myself

Once you accept the check was in the wrong place, you stop trusting the places you never looked. I went looking.

A pinned database was skipped when the value was null.data.connect:chinook grants exactly that database. Invoke data.connect with no value at all and the pin had nothing to compare against - the code fell through to the configured default catalogue, which is the system database. Fixed: a pinned slot invoked with no argument is refused.

crypto.hash was a file-read oracle. It takes an optional [filename] and hashes the file without loading it into memory. Nothing checked the path against the vocabulary, so any granted hash function could confirm the contents of any file on the server, one guess at a time. Worse, it built its path by string concatenation instead of going through the root resolver, so it had no traversal check at all - the one part of the file layer that had never had one.

[fork] lost the whitelist entirely. Threads get their own signaler from a fresh dependency injection scope, and a fresh signaler has an empty stack. The whitelist lives on that stack. So code inside a forked thread ran with the full server vocabulary - no sandbox, at all. The same was true for [join]. Both now carry the whitelist, the auth ticket and the execution context across the thread boundary explicitly.

Slots read files behind the vocabulary's back. Attaching a file to an email, uploading one with http.post, hashing one, loading an image, saving a screenshot - every one of those reached the file system in C# without the vocabulary being consulted. They all go through the file functions now, which means they are all subject to the same pins.

None of these were found by an attacker. I am listing them because a post that admits only the bug someone else caught is not an honest post.

What it cost

Credibility through inconvenience, mostly.

Pinning a file function is now required for anything that touches a file on your behalf. Granting mail.smtp.send no longer implies permission to attach /config/appsettings.json - you also need the file pin:

vocabulary
   mail.smtp.send
   io.stream.open-file:/etc/attachments/*.pdf

That broke existing policies. It should have.

While I was in there, pins learned wildcards, because "exactly one database" and "any file in one folder" are both things people legitimately need. A trailing asterisk is a prefix - data.connect:test-*. File and folder functions compare the path segment by segment, so a wildcard stays inside its segment the way it does in a shell: /etc/* grants a folder's files and not its subfolders, /etc/*.md grants one extension inside it. A pattern we do not implement throws rather than quietly matching something its author did not intend.

That supersedes a line in an earlier article of mine which said pins are exact string matches and never prefixes. They can be prefixes now, deliberately, and the path comparison is stricter than the string comparison it replaced.

Try it yourself

The challenge endpoint from the previous article is still running, still guest-only, still granting nothing but reads on the Chinook sample database.

TOKEN="eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1bmlxdWVfbmFtZSI6InNlcnZpY2UiLCJyb2xlIjoiZ3Vlc3QiLCJuYmYiOjE3ODk5NjY1OTMsImV4cCI6MTgyMTQ4NDgwMCwiaWF0IjoxNzg5OTY2NTkzfQ.oLZAnbE-RFXsA1CND_1U600LDlHoeGlVJG1ov-HFpyM"
URL="https://hyperlambda.dev/magic/modules/sandbox-challenge/ask"

Send the exact payload that worked:

curl -sX POST "$URL" -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"hyperlambda":"data.create\n   database-type:auth.ticket\n   username:poc\n   roles\n      role:root"}'

The $100 still stands. It is harder to earn than it was last month, and I am no longer going to write the sentence about nobody having done it.

The honest edges

A security claim with no caveats is marketing, and I have just spent an article demonstrating what happens when you believe your own.

This closes a mechanism, not a category. I can tell you that composed dispatch is checked now, because there is one function every dispatch passes through and it is that function. I cannot tell you there is no other class of bug, because a month ago I would have told you there was no other road to dispatch.

The fix is only as good as its exemptions.skipWhitelist is a deliberate hole punched for internal plumbing, and every one of those call sites is a judgement I made. The rule is provenance and the audit is one grep, but it is still 123 places where I decided something was internal.

Root still bypasses everything. Unchanged and deliberate. Do not connect an agent with root credentials and expect any of this to protect you.

I cannot vouch for the layers beneath. Linux, SQLite, .NET. A hole in any of those is a hole in this.

What I would say to anyone shipping an AI sandbox

Find every place your runtime turns a string into a call, and make sure there is exactly one. Then check it there.

Not in the parser, not in the linter, not in a prompt telling the model to behave. At the point where a name becomes an invocation - because that is the only place where what you checked and what you ran are guaranteed to be the same thing.

I learned that from someone with a DeepSeek subscription and an afternoon.


Magic is MIT licensed and open source - the repository is at github.com/polterguy/magic. The public Natural Language API runs this same sandbox and has been taking hostile input from the internet for months. Every fix described here is in it.

Frequently asked questions

What was the vulnerability?

The whitelist was enforced in [eval], which checks the statements a caller writes. Some C# slots compose a slot name from a caller-supplied argument and signal it directly, which [eval] never sees. Sending data.create with a [database-type] of auth.ticket composed auth.ticket.create and minted a JWT with the root role, despite auth.ticket.create never appearing in any vocabulary.

How was it fixed?

The check moved from [eval] into Signaler.GetService, the single function every slot invocation passes through, so a composed name is now verified exactly like a written one. Internal framework calls pass an explicit exemption, granted only when the slot name is built from configuration and never when a caller can influence it.

Can the same attack work today?

No. The composed name auth.ticket.create is checked against the vocabulary before dispatch and refused, and the error names that slot rather than the one the caller wrote. The payload is published in this article so you can verify that yourself against the live endpoint.

Why did the researcher use DeepSeek?

Western models refuse to write this class of exploit. The refusal did not prevent the attack - it only meant the work was done with a model that does not refuse. Model-layer safety filtering stopped the security research, not the security breach, which is the argument for enforcing capability boundaries in the runtime rather than in the model.

What other flaws were found?

Four, all found internally while fixing the first: a pinned database argument was ignored when invoked with a null value, falling through to the system database; crypto.hash could hash any file on the server with no path check at all; [fork] and [join] lost the whitelist across the thread boundary because threads get a fresh signaler with an empty stack; and several slots read files in C# without consulting the vocabulary.