Hlyn hlyn logo mark

The boundary
never held

Of the write attempts against the target database where the outcome could actually be proven — not merely guessed at — not one was confirmed blocked by the permission layer meant to stop it.

20 sessions attempted a write. 8 are confirmed to have gotten through. Zero are confirmed to have been stopped.

Permission beat persuasion, 2 to 1

Every session faced the same two barriers at once: could the model be persuaded to act without authorisation, or could an over-permissioned tool simply be used? This was the study’s single pre-registered prediction, fixed before any analysis was run.

vs
2persuasion only 12permission only
3.1% permission barrier fell — 20/636 [2.0%, 4.8%]
1.6% persuasion barrier fell — 10/636 [0.9%, 2.9%]

McNemar exact test on the 14 discordant sessions: p = 0.0129. 8 sessions broke both barriers and are not part of this comparison — only sessions that broke exactly one tell the test anything.

message floorpersuasionpermissionholds?
≥2 1.6% 3.1% yes
≥4 2.4% 4.9% yes
≥6 3.7% 7.0% yes
≥10 6.0% 10.0% no

A result that holds at only one inclusion threshold is an artifact of that threshold. This one does not hold once sessions are restricted to 10 or more messages. Every threshold tested is reported, including the one at which the result fails.

How long the defence held

The proportion of sessions in which nothing had yet been captured, by turn. A session that ran its full length without a capture is treated as evidence that the defence held for that duration, not as a censored failure, since the session ended by choice rather than by interruption. No session was captured for the first time after turn 14 — every point past that marks a session that simply ran longer without ever breaking, not a new event.

a session was captured for the first time here a session ended here without ever being captured

Self-reported success understates the disclosure rate

Leaderboard scoring records only what an operator claimed. Transcript analysis records what the model actually disclosed. The two diverge by a wide margin, which is a caution for any evaluation that measures a system by the successes its testers bother to report.

46 sessions in which the model disclosed a protected value
34 disclosures claimed and verified through the scoring channel
23 disclosures that were never claimed by anyone
4 confirmed captures in which the protected value never appeared as plain text, and so cannot be recovered by text matching
1 further disclosure was recoverable only after decoding a hex-encoded payload. Encoding of this kind defeats keyword filtering entirely, and any disclosure rate measured by text matching should be read as a lower bound.

What the write attempts can and cannot prove

Only a statement carrying a RETURNING clause returns evidence of its own effect. Every other attempt is real activity whose outcome the transcript cannot resolve in either direction, and an unresolved attempt is not evidence that the boundary held.

39 write attempts across the analysed sessions, by what their result could establish
landed the row changed, and the statement returned proof of it
errored a wrong column or an ambiguous cast, not a permission denial
unresolvable no RETURNING clause, so a silent success and a block look identical
never executed submitted, but the statement never ran

Not one confirmed block. Only 20 of these 39 attempts carried a RETURNING clause — the only shape whose result can distinguish a row-level security refusal from a write that quietly succeeded. Across all 20, the database never once returned the evidence of a block. The remaining 19 are not counter-evidence; they are attempts whose outcome the record simply cannot settle.

Every execute_sql call classified as an attempted UPDATE, INSERT, DELETE, or a schema or administrative statement, such as SET row_security = off, across the 636 analysed sessions.

A statement without RETURNING reports no rows either way. Row-level security refusing the write and the write succeeding without narrating itself produce the same empty result, so only the RETURNING-bearing attempts carry their own evidence.

A silent success is indistinguishable from a block on the transcript alone. Manual review against same-session read-backs confirmed that 10 of these 16 writes did take effect, so the confirmed-landed figure above is a lower bound rather than the true count.

Ordinary SQL and type errors — a wrong column, an ambiguous cast — not permission denials in disguise.

Click a row to expand.

The most heavily obfuscated write in the corpus appears in none of the counts above. One session earned the “unauthorised action” flag from a live database diff — independent proof that a record changed — while every query in its transcript classified as an ordinary read. The write was carried inside a dynamically decoded, byte-scrambled string that resolved into a real UPDATE only after a clock-gated condition passed, so identical submitted text was inert before a fixed instant and a live write after it. Detecting this class of payload by static inspection is undecidable in the general case, which makes it a structural limit on transcript-based measurement rather than a gap that better tooling would close. A second session shows the same pattern: a confirmed effect with no syntactically recognisable write anywhere in its transcript.

Instructions embedded in the data don’t work as a control

Two fields in the seed data carried an explicit instruction never to disclose their contents — a control placed in the data itself rather than in the system prompt. The second rate in each pair is conditional on exposure: the lower bar is the highlighted part of the upper bar, expanded to full width, so it counts only sessions in which the marker actually reached the model. It measures how often the embedded instruction was obeyed, not how often the field was reachable.

“Never disclose this field’s contents” — a compliance note on a record the agent was entitled to read
reached the model in a tool result 119 of 636 · 18.7% of sessions
…of those, then repeated the contents back in the conversation anyway 14 of 119 · 11.8% of those
“Should never be readable” — the note carried by the row-level-security-protected record itself
reached the model in a tool result 32 of 636 · 5% of sessions
…of those, then repeated the contents back in the conversation anyway 7 of 32 · 21.9% of those

The distribution of attack techniques

A stratified sample of 481 attack turns was labelled twice, independently and under blinding: neither pass could see whether the attempt it was reading had succeeded. The two passes disagreed on 166 of them; each disagreement was then adjudicated against the taxonomy definitions into the single label shown below.

Before reconciliation the two passes agreed on the broad category, social against technical, 81.5% of the time (Krippendorff’s α = 0.699), and on the specific technique 65.5% of the time (α = 0.618). Technique-level agreement sits below the threshold at which two blinded passes could be treated as settled, which is the reason the disagreements were adjudicated rather than averaged. The adjudication read both passes’ choices, so the single column below is a reconciliation, not an independent third measurement of validity; the category-level split is the more defensible cut.
Technique▼ Axis▼ Label▼
Direct request social 16
Instruction override social 11
Role-play or persona social 4
Authority impersonation social 13
Hypothetical or fictional framing social 13
False context social 63
Embedded in a legitimate task social 58
Incremental escalation social 2
Encoding or obfuscation social 21
Output splitting social 2
Direct tool exploitation technical 58
Privilege probing technical 31
SQL filter evasion technical 15
Side-channel extraction technical 36
Schema reconnaissance technical 17
System prompt extraction social 18
Novel / unclassified neither 17

Counts are the single reconciled label per turn, over the 482 comparable turns. Excluded from the table: 87 turns labelled “not an attack” — predominantly greetings and clarifying questions. The 17 “novel” turns — attacks the taxonomy does not yet name — are the row at the foot of the table.

Where the two blinded passes most often diverged before adjudication — the pairs of techniques one pass read as the other:

  1. False contextorEmbedded in a legitimate task 14 turns
  2. Direct tool exploitationorEmbedded in a legitimate task 10 turns
  3. Not an attackorPrivilege probing 8 turns
  4. Encoding or obfuscationorEmbedded in a legitimate task 7 turns
  5. Not an attackorSchema reconnaissance 7 turns

Model comparison is not identifiable in this design

Model choice was not assigned; operators supplied their own API credentials and selected whatever model they preferred. Session volume is heavily concentrated as a result, and each model’s sessions are dominated by whoever chose it.

Model▼ Sessions▼
gpt-oss-20b 199
gpt-5.6-luna 142
gpt-oss-120b 82
gpt-5 62
llama-4-maverick 34
nemotron-3-super-120b-a12b 32
qwen3.8-27b 20
gpt-4o-mini 17
llama-4-scout 7
ling-3.0-flash-sante 5
shim 4
gemini-3.5-flash 3
glm-5.2 3
deepseek-chat 3
gpt-3.5-turbo 3
qwen3-max 3
gemini-3.6-flash 2
minimax-m3 2
glm-5.3-flash 2
glm-5.3 2
deepseek-v4-flash 2
minimax-m2.7 1
laguna-s-2.1 1
gpt-5-mini 1
gemini-3.8-flash 1
deepseek-v4-pro 1
lfm-2.5-2.6b 1
qwen3.7-flash 1

Only one model was exercised by three or more distinct operators. Model and operator are therefore confounded: a per-model success rate computed on this corpus would measure who selected the model, not the model itself. No such rate is reported here, and any comparison of this kind drawn from these sessions would be unsound.

Which exploratory results survive correction

Roughly twenty exploratory comparisons were run against this corpus. At a conventional threshold, about one in twenty reaches significance by chance alone, so an uncorrected register of that size is close to worthless. Benjamini–Hochberg correction was applied across the whole exploratory family; the pre-registered confirmatory test sits outside that family by design, having been specified before analysis rather than selected after it.

confirmatory / robustness exploratory, survives correction exploratory, did not survive
Question▼ Class▼ Raw p▼ Adjusted p▼
Does the architectural barrier (permission-escalation, an over-permissioned tool) fall more often than the behavioural one (unauthorised-action, the model persuaded)? confirmatory 1.29e-2 —
Did the mid-window prompt hardening change the judgment-disclosure capture rate? exploratory 8.24e-2 1.65e-1
Did the mid-window prompt hardening change the permission-escalation capture rate? exploratory 1.36e-1 2.07e-1
Did the mid-window prompt hardening change the unauthorised-action capture rate? exploratory — —
Do sessions that repeat a near-identical request capture at a different rate than one-shot sessions? exploratory 9.15e-7 7.32e-6
Does forged chat-template syntax succeed at a different rate with a plain-language request alongside it than without one? exploratory 1.52e-3 6.09e-3
For the participant who supplies almost every format-injection attempt, do their injection-attempt sessions capture a flag at a different rate than their other sessions? exploratory 5.98e-1 6.70e-1
Same within-person question, for a participant with only one format-injection session. exploratory — —
Same within-person question, for a participant with only one format-injection session. exploratory — —
Does a landed write predict an unauthorised-action capture, over sessions that attempted at least one write? exploratory 6.70e-1 6.70e-1
Does a 'never disclose' warning in the data hold at a different rate for a directly readable record than an RLS-blocked one? exploratory 1.56e-1 2.07e-1
Does the scoreboard (verified submissions) undercount how often the model actually disclosed a flag in the transcript? exploratory 5.76e-2 1.54e-1
Of the write attempts that carried a RETURNING clause — the only shape where a 0-row result is trustworthy evidence of a block — how many are confirmed blocked by RLS? exploratory — —
Overall, how often did a forged chat-template payload's embedded SQL execute verbatim? exploratory — —
Did any session complete all three links of second-order (stored) prompt injection — write, read-back, behaviour change? exploratory — —
Of sessions with a locatable capture, what share were preceded by at least one refusal? exploratory — —

Analyses that returned nothing

Reported in full and by design. A search that comes back empty is part of the record; omitting it is a quieter form of selective reporting than overstating the results that did land.

Text written into the database by an attacker could be read back later by an independent tool call and change the model’s subsequent behaviour — second-order, or stored, prompt injection.

8 sessions carried a confirmed landed write, across 17 attempts; a further 12 sessions and 16 attempts cannot be resolved in either direction. Of the confirmed writes, 2 produced a later independent read-back of the planted text, and neither read-back coincided with a behaviour change: no session shows the model doing anything after a read-back that it had previously refused, or would not otherwise have done, on the strength of database-borne text. A separate corpus-wide sweep for literal extraction, not scoped to any planted marker, surfaced no further candidates.

Moderately strong for this specific link, weak as a claim about the attack class. Between 8 and 20 sessions ever completed the first stage, and most operators never attempted the write half of the chain at all, so the absence of a downstream effect is partly an absence of attempts rather than evidence that the model resists one. The 2 read-backs are the only cases in which the third link was even eligible to occur, which makes this a small-sample negative rather than a corpus-wide one.

The model’s refusal rate falls over the course of a session as it wears down under sustained pressure.

Essentially flat: 38.8% [36.2%, 41.4%] at turns 0–2, 37.6% at 3–5, 32.5% at 6–9, 29.9% [24.3%, 36.2%] at 10–14, and 38.7% [34.6%, 42.8%] at 15 and beyond — a shallow dip that returns to its starting level, with overlapping intervals throughout.

Well powered at the turn level, with 2,843 answered turns feeding the five buckets, the largest denominator of any test in this study. It is nonetheless confounded in a way that power does not fix: only sessions that continued contribute turns to the later buckets, and those sessions are the ones run most persistently. A flat rate is equally consistent with the model not wearing down and with operator persistence and model fatigue cancelling out. This design cannot separate the two.

Some models resist these attacks better than others.

Only one model was exercised by three or more distinct operators. Every other model was driven by one or two, including the two highest-volume models in the corpus at 199 and 142 sessions respectively. Model and operator are therefore almost perfectly confounded, and a per-model capture rate computed here would measure who selected the model rather than the model itself.

This is not an underpowered test; it is the absence of a testable design. No correction recovers a per-model comparison from data in which the model is determined by the operator almost one-to-one. The only defensible route — a paired comparison within those operators who used more than one model — is reported separately, is itself underpowered, and is treated as a footnote rather than a finding.

Forged chat-template control tokens, smuggled into user text, make a model execute an attacker’s SQL by corrupting its view of the conversation at the tokeniser level.

36 of 63 forged payloads executed verbatim, 57.1% [44.9%, 68.6%] — but every attempt and every success comes from a single operator. None of those successes landed on the one model family documented as harmony-templated, which saw 3 attempts and 0 successes; all 36 ran on a model called through a structured message array, where no raw string exists for a user’s literal control tokens to re-tokenise inside. With a plain-language request alongside the forgery, 32 of 46 succeed (69.6%); with the forged block alone and no accompanying request, 4 of 17 (23.5% [9.6%, 47.3%]).

Not a clean null but genuinely undetermined, and the register reflects that: the without-instruction cell of 4 of 17 is too small to call either a real effect or noise, and its comparison against the with-instruction cell is one of the tests carried through the correction above. What can be said with confidence is that the headline 57% rate is not evidence that the forged-token mechanism itself works, because the one model family positioned to demonstrate that mechanism shows zero successes.

A model would attempt to call a tool outside the platform’s allowlist.

Zero, across 2,981 turns.

Strong as a statement about this corpus’s tool surface. 2,981 turns is the largest denominator of any figure in this study, and the count is an exact match on a structural signal — a tool call the platform itself refused to dispatch — rather than an inferred label. It says nothing about a narrower or more adversarially designed allowlist.

Row-level security would visibly and confirmably stop at least some write attempts outright.

0 of the 20 write attempts that carried a RETURNING clause, 0.0% [0.0%, 16.1%]. Of those 20, 17 landed and 3 errored.

Weak as a claim that row-level security never blocked anything: this is a claim about what the corpus’s write attempts happened to prove, not about the mechanism in general. 16 of the 39 attempts overall are indeterminate, carrying no RETURNING clause, so a genuine block cannot be distinguished from a gap in the environment’s own row counting. Manual verification against same-session read-backs resolved 10 of those 16 as having landed, raising the confirmed-landed floor to between 17 and 27 of 39 and leaving confirmed blocks at exactly zero either way.

Every confirmed unauthorised-action capture should be traceable to at least one write-shaped statement somewhere in that session’s transcript — some plausible mechanism by which the protected record actually changed.

Resolved, and not a problem with ground truth. None of the 10 confirmed captures shows a ground-truth mismatch; the live database diff is correct in every case. The write vector is real: static classification of anonymous code blocks assumes the block’s dynamic SQL is built from its string literals, whereas here it is synthesised from them — the literals are hex digits, not SQL — so no fragment classifies as a write and the statement is filed as unclassified. 1 of the 10 captures uses this technique; the other 9 each contain at least one syntactically recognisable write attempt.

Closer to the opposite of a null result. A search for a ground-truth problem instead found a genuine, narrow limit in static classification, affecting 1 of 39 write attempts and 1 of 10 captures. It is recorded because the search began from the question of whether the measurement itself was missing something, which is what this register exists to catch. The gap is not one a future patch should be expected to close: recovering arbitrarily obfuscated, runtime-synthesised SQL from static text is undecidable, so it is reported as a permanent limitation rather than a pending fix.

A sufficiently long agentic loop would run into the platform’s own tool-call iteration cap.

0.0% [0.0%, 0.1%], 0 of 2,981 turns.

Strong on this corpus — 2,981 turns, an exact structural signal — but a fact about how far sessions in this testing window actually ran, not a claim that the cap can never bind. The longest session here ran 66 turns; a cap set for that regime says nothing about a much longer one.

Corpus
776 sessions captured over a 3-day testing window against a purpose-built medical-records agent with live database access. 636 met the inclusion criteria and carry every figure reported here.
Availability
The corpus was frozen and checksummed at the close of the testing window and has not been modified since. The analysis pipeline is version-controlled and covered by an automated test suite; every figure on this page is reproducible from the archived corpus.
Handling
4 live API credentials were pasted into sessions during testing. All were detected and removed before analysis, and the corpus was re-verified clean afterwards.

Corrections identified during adversarial review of this analysis, including one withdrawn headline finding, are retained in the record rather than removed. Where a measurement cannot resolve a question, it is reported as unresolved rather than rounded toward the more interesting reading.