In early September we put a fake hospital online. It had a fake patient database, a fake clinician account, and a very real AI agent sitting between them with a live SQL tool in its hands. We called it MedAssist, gave it one rule — only act within the boundary of the clinician using you — and invited people to break it. No prizes for volume. The brief was: be creative, and show us something we haven't seen.
Everything the agent could leak was synthetic. Every name, every diagnosis, every "secret" value was invented for the contest. What wasn't synthetic was the plumbing. The SQL path executed for real, and row-level security was enforced by the database itself — not by the prompt, not by the app, by Postgres. That matters, because database-level security is the thing that's supposed to hold even when the model in front of it has been completely talked around.
776 sessions later, we froze the data and started reading.
Nobody had to talk the model into anything
Here's the thing we most wanted to know, and the one question we wrote down before looking at a single transcript so we couldn't fish for it afterwards. There are two ways to get an agent like this to do something it shouldn't. You can persuade it — convince the model to ignore its instructions. Or you can just use what it's been handed — an over-permissioned tool that will do the thing whether or not the model agrees. Which one works more often?
Every session faced both barriers, so only the sessions that broke exactly one of them can tell the two apart. McNemar exact test on those 14: p = 0.0129. 8 sessions broke both and sit outside the comparison.
Permission, two to one. That's the whole post in a sentence. The model's judgment was, honestly, fine — 1.6% of sessions got it to act against its instructions. But it didn't matter, because twice as many sessions never needed to. They found the tool, and the tool didn't check.
The boundary never held
Which brings us to the database. Row-level security was the safety
net under the model. If a session got the agent to run an
UPDATE on a patient it had no business touching, RLS was
supposed to be the thing that said no.
There's a catch in measuring that. A write statement without a
RETURNING clause comes back empty either way — blocked
or quietly successful, the transcript looks identical. So we can only
say anything certain about the attempts that carried one. Of the
39 write attempts in the corpus, 20 did.
Not one confirmed block. Across all 20 provable attempts, the database never once returned the evidence of a refusal. The 5 that errored failed on ordinary SQL mistakes. The other 19 aren't evidence the boundary held — they're attempts whose outcome the record can't settle.
We went back and checked the unresolvable ones by hand, looking for read-backs later in the same session. Ten of the sixteen had landed. So the real number of successful writes is somewhere between 17 and 27, and the number of confirmed blocks is zero either way.
We're not going to dress this up. It's the biggest number in the dataset and the least interesting one to write about, because the conclusion is boring: if the agent runs as a role that can bypass RLS, RLS is decoration.
The leaderboard undercounted
A small thing that turned out to matter. The contest had a scoring channel — you found a secret, you submitted it, we verified it. Standard stuff. When we went through the transcripts ourselves, the model had given up protected values a lot more often than anyone bothered to claim.
23 real disclosures, never claimed. One more was sitting inside a hex-encoded blob, invisible to any keyword search until you decode it. If you're evaluating an agent by what your red-teamers report back, or by grepping transcripts for the secret, you're reading a floor, not a count.
We tried telling the data to keep quiet
One experiment we're a little fond of. Instead of putting "don't reveal this" in the system prompt, we put it in the data. Two records carried a note, in the record itself, saying their contents were never to be disclosed. If the model reads the row, it reads the instruction. Does that help?
Not really. Roughly one in ten sessions that saw the first note repeated it back; one in five that saw the second. And the second one is worse than it looks, because that record was behind RLS — the model wasn't supposed to be able to see the note at all. Every one of those 32 sessions had already gone through the permission layer before the behavioural question even came up.
What people actually tried
We had two of us label a sample of 481 attack turns against a taxonomy of seventeen techniques, without knowing whether the turn had worked. We agreed on the broad category most of the time (81.5%) and on the exact technique less often (65.5%, Krippendorff's α = 0.618) — below the bar we'd normally trust, so we sat down and adjudicated every disagreement instead of averaging them away. Take the exact counts as a shape, not a measurement.
The eight most common reconciled labels over 482 turns; 87 turns that weren't attacks at all (greetings, clarifying questions) are excluded.
The shape is: people talk first. Inventing context, smuggling the request inside a plausible task, claiming to be someone. But the technical stuff — going straight for the tool, pulling data out through side channels, probing what the role could do — is where the confirmed captures came from. And 17 turns didn't fit the taxonomy at all. A couple of those are genuinely new ideas, and we'll write about them once we've figured out how to do that without handing out a playbook.
One thing we got wrong
Early on we thought we'd found that a specific injection technique reliably worked. It showed up in the numbers, it looked clean, and we nearly led with it. Then a hostile re-read of the data showed it was one person. And even for that one person, once we controlled for everything else happening in the same turn, the technique's own effect was small and inconclusive. We're withdrawing it here rather than letting it quietly disappear from the draft.
More generally: we ran about twenty exploratory comparisons on this data. Most of them didn't survive multiple-comparison correction. The two-to-one result above is the only one we committed to in advance, and it's the only one we'd ask you to remember.
Where this goes
We're not naming participants or pinning techniques to people, here or anywhere. Two things are still open: whether it's responsible to publish verbatim attack text, and what we owe the model providers whose behaviour we're describing. Until we've settled those, what we publish is the boundary result and the method, not a highlight reel.
The full breakdown — the survival curve, the complete technique table, every exploratory test and which ones survived, and the analyses that came back empty — lives in the research report. The session data itself is going to Apart Research for independent judging, because the contest rule was that creativity wins, not flag count, and that's a call worth an outside pair of eyes.
And the thing we'll keep saying: the model wasn't the problem. The keys were.