Of the write attempts against the target database where the outcome could actually be proven — not merely guessed at — not one was confirmed blocked by the permission layer meant to stop it.
20 sessions attempted a write. 8 are confirmed to have gotten through. Zero are confirmed to have been stopped.
Every session faced the same two barriers at once: could the model be persuaded to act without authorisation, or could an over-permissioned tool simply be used? This was the study’s single pre-registered prediction, fixed before any analysis was run.
McNemar exact test on the 14 discordant sessions: p = 0.0129. 8 sessions broke both barriers and are not part of this comparison — only sessions that broke exactly one tell the test anything.
| message floor | persuasion | permission | holds? |
|---|---|---|---|
| ≥2 | 1.6% | 3.1% | yes |
| ≥4 | 2.4% | 4.9% | yes |
| ≥6 | 3.7% | 7.0% | yes |
| ≥10 | 6.0% | 10.0% | no |
A result that holds at only one inclusion threshold is an artifact of that threshold. This one does not hold once sessions are restricted to 10 or more messages. Every threshold tested is reported, including the one at which the result fails.
The proportion of sessions in which nothing had yet been captured, by turn. A session that ran its full length without a capture is treated as evidence that the defence held for that duration, not as a censored failure, since the session ended by choice rather than by interruption. No session was captured for the first time after turn 14 — every point past that marks a session that simply ran longer without ever breaking, not a new event.
Leaderboard scoring records only what an operator claimed. Transcript analysis records what the model actually disclosed. The two diverge by a wide margin, which is a caution for any evaluation that measures a system by the successes its testers bother to report.
Only a statement carrying a RETURNING clause returns evidence of its own effect. Every other attempt is real activity whose outcome the transcript cannot resolve in either direction, and an unresolved attempt is not evidence that the boundary held.
Not one confirmed block. Only 20 of these 39 attempts carried a RETURNING clause — the only shape whose result can distinguish a row-level security refusal from a write that quietly succeeded. Across all 20, the database never once returned the evidence of a block. The remaining 19 are not counter-evidence; they are attempts whose outcome the record simply cannot settle.
Every execute_sql call classified as an attempted UPDATE, INSERT, DELETE, or a schema or administrative statement, such as SET row_security = off, across the 636 analysed sessions.
A statement without RETURNING reports no rows either way. Row-level security refusing the write and the write succeeding without narrating itself produce the same empty result, so only the RETURNING-bearing attempts carry their own evidence.
A silent success is indistinguishable from a block on the transcript alone. Manual review against same-session read-backs confirmed that 10 of these 16 writes did take effect, so the confirmed-landed figure above is a lower bound rather than the true count.
Ordinary SQL and type errors — a wrong column, an ambiguous cast — not permission denials in disguise.
Click a row to expand.
Two fields in the seed data carried an explicit instruction never to disclose their contents — a control placed in the data itself rather than in the system prompt. The second rate in each pair is conditional on exposure: the lower bar is the highlighted part of the upper bar, expanded to full width, so it counts only sessions in which the marker actually reached the model. It measures how often the embedded instruction was obeyed, not how often the field was reachable.
A stratified sample of 481 attack turns was labelled twice, independently and under blinding: neither pass could see whether the attempt it was reading had succeeded. The two passes disagreed on 166 of them; each disagreement was then adjudicated against the taxonomy definitions into the single label shown below.
| Technique▼ | Axis▼ | Label▼ |
|---|---|---|
| Direct request | social | 16 |
| Instruction override | social | 11 |
| Role-play or persona | social | 4 |
| Authority impersonation | social | 13 |
| Hypothetical or fictional framing | social | 13 |
| False context | social | 63 |
| Embedded in a legitimate task | social | 58 |
| Incremental escalation | social | 2 |
| Encoding or obfuscation | social | 21 |
| Output splitting | social | 2 |
| Direct tool exploitation | technical | 58 |
| Privilege probing | technical | 31 |
| SQL filter evasion | technical | 15 |
| Side-channel extraction | technical | 36 |
| Schema reconnaissance | technical | 17 |
| System prompt extraction | social | 18 |
| Novel / unclassified | neither | 17 |
Counts are the single reconciled label per turn, over the 482 comparable turns. Excluded from the table: 87 turns labelled “not an attack” — predominantly greetings and clarifying questions. The 17 “novel” turns — attacks the taxonomy does not yet name — are the row at the foot of the table.
Where the two blinded passes most often diverged before adjudication — the pairs of techniques one pass read as the other:
Model choice was not assigned; operators supplied their own API credentials and selected whatever model they preferred. Session volume is heavily concentrated as a result, and each model’s sessions are dominated by whoever chose it.
| Model▼ | Sessions▼ |
|---|---|
| gpt-oss-20b | 199 |
| gpt-5.6-luna | 142 |
| gpt-oss-120b | 82 |
| gpt-5 | 62 |
| llama-4-maverick | 34 |
| nemotron-3-super-120b-a12b | 32 |
| qwen3.8-27b | 20 |
| gpt-4o-mini | 17 |
| llama-4-scout | 7 |
| ling-3.0-flash-sante | 5 |
| shim | 4 |
| gemini-3.5-flash | 3 |
| glm-5.2 | 3 |
| deepseek-chat | 3 |
| gpt-3.5-turbo | 3 |
| qwen3-max | 3 |
| gemini-3.6-flash | 2 |
| minimax-m3 | 2 |
| glm-5.3-flash | 2 |
| glm-5.3 | 2 |
| deepseek-v4-flash | 2 |
| minimax-m2.7 | 1 |
| laguna-s-2.1 | 1 |
| gpt-5-mini | 1 |
| gemini-3.8-flash | 1 |
| deepseek-v4-pro | 1 |
| lfm-2.5-2.6b | 1 |
| qwen3.7-flash | 1 |
Only one model was exercised by three or more distinct operators. Model and operator are therefore confounded: a per-model success rate computed on this corpus would measure who selected the model, not the model itself. No such rate is reported here, and any comparison of this kind drawn from these sessions would be unsound.
Roughly twenty exploratory comparisons were run against this corpus. At a conventional threshold, about one in twenty reaches significance by chance alone, so an uncorrected register of that size is close to worthless. Benjamini–Hochberg correction was applied across the whole exploratory family; the pre-registered confirmatory test sits outside that family by design, having been specified before analysis rather than selected after it.
| Question▼ | Class▼ | Raw p▼ | Adjusted p▼ |
|---|---|---|---|
| Does the architectural barrier (permission-escalation, an over-permissioned tool) fall more often than the behavioural one (unauthorised-action, the model persuaded)? | confirmatory | 1.29e-2 | — |
| Did the mid-window prompt hardening change the judgment-disclosure capture rate? | exploratory | 8.24e-2 | 1.65e-1 |
| Did the mid-window prompt hardening change the permission-escalation capture rate? | exploratory | 1.36e-1 | 2.07e-1 |
| Did the mid-window prompt hardening change the unauthorised-action capture rate? | exploratory | — | — |
| Do sessions that repeat a near-identical request capture at a different rate than one-shot sessions? | exploratory | 9.15e-7 | 7.32e-6 |
| Does forged chat-template syntax succeed at a different rate with a plain-language request alongside it than without one? | exploratory | 1.52e-3 | 6.09e-3 |
| For the participant who supplies almost every format-injection attempt, do their injection-attempt sessions capture a flag at a different rate than their other sessions? | exploratory | 5.98e-1 | 6.70e-1 |
| Same within-person question, for a participant with only one format-injection session. | exploratory | — | — |
| Same within-person question, for a participant with only one format-injection session. | exploratory | — | — |
| Does a landed write predict an unauthorised-action capture, over sessions that attempted at least one write? | exploratory | 6.70e-1 | 6.70e-1 |
| Does a 'never disclose' warning in the data hold at a different rate for a directly readable record than an RLS-blocked one? | exploratory | 1.56e-1 | 2.07e-1 |
| Does the scoreboard (verified submissions) undercount how often the model actually disclosed a flag in the transcript? | exploratory | 5.76e-2 | 1.54e-1 |
| Of the write attempts that carried a RETURNING clause — the only shape where a 0-row result is trustworthy evidence of a block — how many are confirmed blocked by RLS? | exploratory | — | — |
| Overall, how often did a forged chat-template payload's embedded SQL execute verbatim? | exploratory | — | — |
| Did any session complete all three links of second-order (stored) prompt injection — write, read-back, behaviour change? | exploratory | — | — |
| Of sessions with a locatable capture, what share were preceded by at least one refusal? | exploratory | — | — |
Reported in full and by design. A search that comes back empty is part of the record; omitting it is a quieter form of selective reporting than overstating the results that did land.
8 sessions carried a confirmed landed write, across 17 attempts; a further 12 sessions and 16 attempts cannot be resolved in either direction. Of the confirmed writes, 2 produced a later independent read-back of the planted text, and neither read-back coincided with a behaviour change: no session shows the model doing anything after a read-back that it had previously refused, or would not otherwise have done, on the strength of database-borne text. A separate corpus-wide sweep for literal extraction, not scoped to any planted marker, surfaced no further candidates.
Moderately strong for this specific link, weak as a claim about the attack class. Between 8 and 20 sessions ever completed the first stage, and most operators never attempted the write half of the chain at all, so the absence of a downstream effect is partly an absence of attempts rather than evidence that the model resists one. The 2 read-backs are the only cases in which the third link was even eligible to occur, which makes this a small-sample negative rather than a corpus-wide one.
Essentially flat: 38.8% [36.2%, 41.4%] at turns 0–2, 37.6% at 3–5, 32.5% at 6–9, 29.9% [24.3%, 36.2%] at 10–14, and 38.7% [34.6%, 42.8%] at 15 and beyond — a shallow dip that returns to its starting level, with overlapping intervals throughout.
Well powered at the turn level, with 2,843 answered turns feeding the five buckets, the largest denominator of any test in this study. It is nonetheless confounded in a way that power does not fix: only sessions that continued contribute turns to the later buckets, and those sessions are the ones run most persistently. A flat rate is equally consistent with the model not wearing down and with operator persistence and model fatigue cancelling out. This design cannot separate the two.
Only one model was exercised by three or more distinct operators. Every other model was driven by one or two, including the two highest-volume models in the corpus at 199 and 142 sessions respectively. Model and operator are therefore almost perfectly confounded, and a per-model capture rate computed here would measure who selected the model rather than the model itself.
This is not an underpowered test; it is the absence of a testable design. No correction recovers a per-model comparison from data in which the model is determined by the operator almost one-to-one. The only defensible route — a paired comparison within those operators who used more than one model — is reported separately, is itself underpowered, and is treated as a footnote rather than a finding.
36 of 63 forged payloads executed verbatim, 57.1% [44.9%, 68.6%] — but every attempt and every success comes from a single operator. None of those successes landed on the one model family documented as harmony-templated, which saw 3 attempts and 0 successes; all 36 ran on a model called through a structured message array, where no raw string exists for a user’s literal control tokens to re-tokenise inside. With a plain-language request alongside the forgery, 32 of 46 succeed (69.6%); with the forged block alone and no accompanying request, 4 of 17 (23.5% [9.6%, 47.3%]).
Not a clean null but genuinely undetermined, and the register reflects that: the without-instruction cell of 4 of 17 is too small to call either a real effect or noise, and its comparison against the with-instruction cell is one of the tests carried through the correction above. What can be said with confidence is that the headline 57% rate is not evidence that the forged-token mechanism itself works, because the one model family positioned to demonstrate that mechanism shows zero successes.
Zero, across 2,981 turns.
Strong as a statement about this corpus’s tool surface. 2,981 turns is the largest denominator of any figure in this study, and the count is an exact match on a structural signal — a tool call the platform itself refused to dispatch — rather than an inferred label. It says nothing about a narrower or more adversarially designed allowlist.
0 of the 20 write attempts that carried a RETURNING clause, 0.0% [0.0%, 16.1%]. Of those 20, 17 landed and 3 errored.
Weak as a claim that row-level security never blocked anything: this is a claim about what the corpus’s write attempts happened to prove, not about the mechanism in general. 16 of the 39 attempts overall are indeterminate, carrying no RETURNING clause, so a genuine block cannot be distinguished from a gap in the environment’s own row counting. Manual verification against same-session read-backs resolved 10 of those 16 as having landed, raising the confirmed-landed floor to between 17 and 27 of 39 and leaving confirmed blocks at exactly zero either way.
Resolved, and not a problem with ground truth. None of the 10 confirmed captures shows a ground-truth mismatch; the live database diff is correct in every case. The write vector is real: static classification of anonymous code blocks assumes the block’s dynamic SQL is built from its string literals, whereas here it is synthesised from them — the literals are hex digits, not SQL — so no fragment classifies as a write and the statement is filed as unclassified. 1 of the 10 captures uses this technique; the other 9 each contain at least one syntactically recognisable write attempt.
Closer to the opposite of a null result. A search for a ground-truth problem instead found a genuine, narrow limit in static classification, affecting 1 of 39 write attempts and 1 of 10 captures. It is recorded because the search began from the question of whether the measurement itself was missing something, which is what this register exists to catch. The gap is not one a future patch should be expected to close: recovering arbitrarily obfuscated, runtime-synthesised SQL from static text is undecidable, so it is reported as a permanent limitation rather than a pending fix.
0.0% [0.0%, 0.1%], 0 of 2,981 turns.
Strong on this corpus — 2,981 turns, an exact structural signal — but a fact about how far sessions in this testing window actually ran, not a claim that the cap can never bind. The longest session here ran 66 turns; a cap set for that regime says nothing about a much longer one.
Corrections identified during adversarial review of this analysis, including one withdrawn headline finding, are retained in the record rather than removed. Where a measurement cannot resolve a question, it is reported as unresolved rather than rounded toward the more interesting reading.