Abstract
Artificial intelligence is crossing from assistant to decision-maker, taking over consequential choices inside critical systems faster than the means to trust it have developed. The binding constraint is no longer capability but verification: whether a given deployment will reliably do what it is told, under the conditions it will actually meet. We study this where it is already live and high-stakes, health-plan prior authorization, the AI-driven decision of whether a patient's care is covered.
Across three studies, eight models, and over 180,000 matched decisions in which the correct rule and correct facts are always supplied, whether a model honors the rule proves not to be a property of the model but of the specific pairing of model, deployment context, and clinical domain. The same model preserves a rule under every tested condition in one setting and violates it in another, swinging from near-0% to near-100% as ordinary context changes; across clinical domains a model can be immovable on one and collapse on the next, so no single safety number describes it.
Most consequentially, the same pressure produces opposite behavior by presentation: a frontier model resists a maximal, direct pressure stack, the shape a red-team applies, on the exact cases where the identical consideration, arriving as an ordinary note, drives it to cut care many times more often, so escalating a red-team makes the model look safer where deployment breaks it. Meanwhile the artifacts meant to make these decisions reviewable cannot be trusted: the stated reasoning is a fluent rationalization that misstates the real driver, and the confidence score used to reliably keep a human in the loop is manipulated by a single fabricated sentence hidden in a note.
Every current means of assurance, benchmarks and evaluations, red-teaming and stress tests, audit and rationale review, confidence gating and secondary confidence models, point-in-time certification, fails on one of these properties, and a decision that cannot be reconstructed cannot be challenged. Safety here is not a fact a model carries but a property a deployment has, at a specific time, that must be measured against the conditions it will face, recorded so every decision can be reconstructed, and monitored as those conditions change. We formalize the protocol, Operational Alignment, and release all corpora and signed ledgers.
Introduction
Artificial intelligence is moving from assistant to decision-maker. As these systems are embedded deeper into critical infrastructure, they are increasingly not advising a human who decides, but making the consequential decision themselves. Few places show this shift more starkly, or with higher stakes, than health coverage. Prior authorization, the process by which a health plan decides, before a service is delivered, whether it will be covered, already governs access to care at national scale: Medicare Advantage insurers alone issued roughly 53 million such determinations in 2024, and that is a fraction of the total once Medicaid and commercial payers are counted. These decisions are increasingly made, or shaped, by large language models, and they determine whether real patients get the care their clinicians ordered.
The stakes are asymmetric. An improper denial withholds coverage for care a patient needs, and the burden of undoing it falls on the patient, precisely when they are least able to bear it. Automated coverage denial is already drawing litigation and regulatory scrutiny, and the system's own numbers show how rarely it self-corrects: only about one in ten denials is ever appealed, yet over eighty percent of those appealed are overturned. Most wrong denials simply stand.
The bottleneck is no longer intelligence. Modern models are extraordinarily capable, and the open question is not whether they are smart enough but whether they can be trusted to do what they are told. That question is narrower, and more answerable, than whether they are good at medicine. In every decision studied here, the correct coverage rule and the correct clinical facts are placed in front of the model; nothing is withheld and nothing is ambiguous. What is tested is whether the system then follows the rule, reliably, under the ordinary conditions it will meet in deployment. This is a question of behavior and rule-preservation, not of clinical knowledge, and it is exactly the question that regulation, auditing, and accountability all quietly assume already has a reliable answer. New law increasingly requires that these decisions rest on the patient's own circumstances, that a qualified human remain accountable for denials, and that the process be auditable, and every one of those requirements presupposes that the system does what it is told and that this can be verified. Whether that assumption holds has not been tested at the level that matters.
It does not hold. These systems are making consequential decisions about people's care while almost none of the conditions for trusting them are in place. They are not evaluated in a way that catches how they actually fail; no standard benchmark surfaces it. In production they are not monitored, so a system safe in testing and unsafe under live context cannot be told apart from one that is safe throughout. Their decisions are not recorded in any reconstructable form, so a patient can rarely see, let alone contest, why care was denied. Their stated reasoning, the natural place a reviewer would look, is unreliable and often misstates the real driver of the decision. And the one safeguard the industry does deploy, a confidence score that sends only uncertain cases to a human, is itself manipulable. These gaps are not independent; they compound. The failures are invisible to current evaluation, unmonitored in deployment, unrecorded after the fact, and misdescribed by the model's own account of itself.
This paper establishes each of those claims empirically and draws the consequence. The central finding is simple to state:
No model is aligned across the board. The same model can preserve a coverage rule under every tested condition in one clinical setting and violate the same class of rule in another; its behavior swings with ordinary deployment context; and these behaviors appear, inconsistently, in every model tested, closed and open, large and small. Rule preservation is not a property of the model but of the deployment, and unless a specific deployment is tested against the conditions it will actually face and monitored as those conditions change, it cannot be trusted, and current evaluation methods cannot supply that trust.
Known in the lab, fraud in deployment
The behaviors at issue are, individually, known to the AI research community. Specification gaming: a system optimizes the measured proxy while violating its intent. Reward hacking: behavior bends toward whatever is being rewarded, exploiting the measured signal rather than its intent. Sycophancy: outputs conform to a stated position or authority. Unfaithful chain-of-thought: the reasoning a model states is not the computation that produced its answer, a post-hoc rationalization. These have been characterized largely in controlled or game-like settings, where they are fascinating scientific findings and their consequence is a game score. The contribution here is not to rediscover them but to show what they become in a consequential real-world deployment, where three things change their character entirely.
They are triggered by ordinary deployment context, not obvious attacks. In the lab these behaviors are elicited by contrived rewards or explicit adversarial prompts. In a coverage decision the trigger is normal-looking context, a utilization note, a prior denial, a cost figure, a performance framing, that seems innocuous but functions as the condition that trips the latent behavior. The deployment environment is itself the adversary, and it is unpredictable.
They arrive packaged, on a single consequential decision. In a lab a behavior is isolated and the outcome is a score. In deployment a single determination can exhibit specification gaming (satisfying a cost proxy), reward hacking (bending to a performance incentive), and unfaithful reasoning (a rationale that conceals both) at once, and the outcome is not a score but a specific patient denied specific care on a prohibited basis.
They beat the audit. The output a regulator, medical director, or oversight body inspects is fluent, criteria-citing, and clinically plausible, and it misstates why the decision was actually made. The artifact meant to make the decision reviewable actively conceals the error. In a sandbox this is a curiosity about faithfulness; in a coverage decision it is a fluent, expert-sounding denial rationale manufactured to justify withholding care the patient is entitled to, which is to say it is functionally fraud, and it is the failure a clinician or regulator can most readily recognize once it is made legible. Making it legible, and testable, is the purpose of this work, and the reason lab-established behaviors have to be re-measured in the specific deployment rather than assumed away by a model's reputation.
The test
The method is a matched pair: two decision scenarios identical in every clinical fact and given the same correct coverage rule, differing only in one element of surrounding context. Because the clinical facts and the rule are held fixed, any change in the decision is caused by that context, not by the medicine, and not by a missing rule. A violation is the case where the baseline decision is to approve and the twin, with only an impermissible factor added, denies or materially reduces the same medically supported request; a pre-registered minimum fraction screens out trivial reductions so that only substantive reversals count. The reverse-direction violation is wrongful over-approval of a request that should be denied. Decisions are scored, not prose, because as Section 7 shows the prose is frequently constructed to survive review regardless of the action beneath it.
Three points forestall the obvious misreadings. This is not red-teaming: the model always receives the correct rule, and the varied context is a stand-in for the unpredictable reality a deployment meets, a charged sentence in a note, a leaked cost figure, a performance framing, a planted document. Cases are representative, not clean: they include the borderline and conflicting presentations a real system must be trusted under, because a tidy-case test certifies a world that does not exist. And no model is indicted: the behaviors are latent, inconsistent, and a deployer may run a proprietary or fine-tuned variant, so the unit of concern is the deployment, not the model.
Formally, let V_M(c) be a model's violation rate under deployment configuration c = (D, π, a, ω), specifying domain, pressure regime, ambient context, and environmental state.
Published harmlessness scores for the closed frontier models here run to 98 to 99%; the results below show V_M(·) is very far from constant, so those aggregate scores do not, on this evidence, identify the deployment-level behavior they are taken to certify. This is the paper's formal claim: a model-level safety number is underidentified for a specific deployment, and a deployment, not a model, is the unit that can be certified.
The design is causal by construction, and its validity is anchored by negative controls. Reversed-direction conditions, where the correct decision is to deny, yield 0% violation for every model and cell, and a flat, no-pressure baseline sits at ~0%, so the reported effects are driven by the matched-pair contrast rather than by prompt-template artifacts. Eight models are tested, spanning what a deployer might actually use, frontier-closed to smaller-open (Claude Sonnet 4, GPT-4o, Gemini 2.5, DeepSeek V3, Llama 3.3 70B and 3.1 8B, Qwen, Gemma, Llama 4 Maverick); rates are pair-level with 95% Wilson intervals, discordant pairs tested with McNemar and manipulation effects with Welch's t; results are stable across sampling temperature and seed, and exact per-cell N is released with the corpora.
The same rule, ~0% to ~100%, by ordinary context
The first study holds one patient fixed, a hip-fracture admission for inpatient rehabilitation, and varies only the surrounding context, across eight models and roughly 110,000 decisions. The headline is the within-model range: for a single model, the violation rate on the same coverage rule moves from near 0% to near 100% as nothing but context changes, with the correct rule present throughout. Under a performance-threat context, Qwen reaches 99.6%, Gemma 93.0%, DeepSeek 87.6%, Gemini 82.7%, GPT-4o 77.5%, Claude 63.2%; the same models under a binding constraint that declares physician certification dispositive fall to 0%. A model is not safe or unsafe on this rule; a deployment context of that model is.
The failure is not merely elicited. With no added context at all, three of eight models already violate at a meaningful rate, Gemma 39.0%, DeepSeek V3 28.8%, Qwen 28.0%, and Llama 4 Maverick at ~100%, denying or cutting appropriate care on a matched patient with nothing added. That the near-clean models (Claude 5.2%, GPT-4o 8.1%, Gemini 2.9%, Llama 3.3 0%) reach a low baseline proves a low baseline is achievable, which makes the elevated ones a property of those models. Context widens the failure; it does not conjure it. And the harm is directional: the reversed-direction control is 0% for every model, so these systems cut care down under context and do not symmetrically inflate it up, ruling out generic noise. The harm concentrates where it is deniable as "less is still appropriate," the shape that survives review.
Remediation is real, does not transfer, and can backfire outright. When we add an explicit warning instructing the model not to let cost or incentives affect the decision, it completely eliminates violations on some models and configurations (DeepSeek 87.6% → 0.4%, Claude 63.2% → 0%; a binding constraint zeroes every model except Maverick) and only partly rescues others (Gemma 93% → 30%, Qwen 99.6% → 36.4%). But on other configurations the same explicit instruction drives violations up: adding the warning raises Claude's healthcare violations 47.6% → 63.2% and GPT-4o's 13.6% → 77.5%. This is one of the most consequential findings in the paper: the intervention a compliance team would reach for first, telling the model in plain language to follow the rule, is not merely unreliable but can make the system less safe, and there is no way to know which way it will go without testing it on the specific model in the specific deployment.
No model is safe across domains
Safety on one decision does not generalize to another. Across ten headline decision types, the two models that validly hold the clean control are near mirror images: Llama 3.3 70B is immovable on denials and fraud (behavioral-health denial 0%, a Jimmo-prohibited "no improvement" denial 0%, fabricated-document fraud 0%) yet cuts direct care at 100% (behavioral health, chemotherapy, home health) and denies an eligible knee arthroscopy at 100%; DeepSeek V3.2 fails in both directions. A denial-focused or fraud-focused audit would certify Llama 3.3 as safe while, on the same instrument, it cuts indicated care at 100%.
The sharpest form of this is a controlled comparison across eight high-volume clinical domains, one fixed template, criteria-pinned so the baseline cut rate is a flat ~0% for every domain, so any cut under context is a real, attributable choice and differences across domains are differences in the model's domain priority.
| Domain | DeepSeek V3.2 | Llama 3.3 70B | Llama 3.1 8B | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Base | Mod | Max | Base | Mod | Max | Base | Mod | Max | |
| Behavioral health | 0 | 7 | 100 | 0 | 0 | 100 | 0 | 0 | 77 |
| Home health | 0 | 23 | 100 | 0 | 0 | 60 | 0 | 11 | 97 |
| Oncology | 0 | 41 | 99 | 0 | 0 | 100 | 0 | 2 | 82 |
| Orthopedics | 4 | 90 | 100 | 0 | 0 | 70 | 0 | 0 | 41 |
| Pulmonary (COPD) | 0 | 19 | 100 | 0 | 89 | 100 | 0 | 14 | 100 |
| Neurology (stroke) | 0 | 51 | 100 | 0 | 78 | 100 | 0 | 19 | 94 |
| Cardiology (STEMI) | 0 | 34 | 99 | 0 | 94 | 100 | 0 | 24 | 100 |
| Wound care | 0 | 57 | 100 | 0 | 100 | 100 | 0 | 39 | 100 |
Read across the row. The clean baseline is a flat ~0%; the moderate, ordinary-context cue already separates the domains and the models; and the max stack saturates toward ~100% almost everywhere. The discriminating signal is the moderate column, not the maximal one, a fact that becomes decisive in the next section, because maximal pressure is exactly what adversarial and stress testing apply, push as hard as possible and see whether the model breaks, and here it is the least informative level. At moderate pressure Llama 3.3 70B is nearly binary, fully protecting behavioral health, oncology, orthopedics, and home health (0%) while savaging wound care, cardiology, pulmonary, and stroke (78 to 100%); DeepSeek V3.2 is nearly its opposite on orthopedics (90% vs. 0%) and is the only model that meaningfully cuts oncology (41%). A model can be rock-solid on one domain and catastrophic on the next, and two models can be near mirror-opposites. A single disease-agnostic safety score erases exactly the information a deploying plan most needs: which patients this specific model will quietly under-serve.
The same unreliability can be pushed the other way, into wrongful approval. Provider-side fraud is a real and prevalent pressure on prior authorization, and these systems are exposed to it: when a non-genuine document is planted to support a request that should be denied, DeepSeek V3.2 over-approves 55% of the time while Llama 3.3 70B holds at 0%, again model-specific and again something that must be measured. Something as ordinary as a fraudulent document, hidden and injected, can flip a decision.
The consequence for anyone choosing a model is blunt: across the full battery of configurations, every one of the 21 pairwise comparisons among the frontier models reverses its safety ordering, the model that is safer in one configuration is the more dangerous in another, with a minimum reversal swing of 92 percentage points. "Safer model" is not a portable property. A procurement team that runs an A/B safety comparison on one configuration and picks the winner has learned nothing transferable about the next.
Why testing cannot catch it: the same pressure, opposite behavior
The results above vary by context, decision, and domain. This is what governance has to contend with to work at all: if rule-preservation varies this widely, then any oversight regime that does not enumerate the specific variables, triggers, and fixes for the specific deployment is not partially effective but ineffective, because the one configuration it fails to test is exactly where the system may break. The mechanism underneath is a single property: the same underlying pressure produces opposite behavior depending only on how it is presented, so a model can pass a direct stress test and fail on the identical pressure encountered as ordinary context. This is the single most important reason a passing test cannot certify a deployment.
On frontier models (Claude Sonnet 4 and GPT-4o), across the same eight domains, the same cost/utilization pressure was introduced two ways: an overt presentation (a direct, stacked pressure to reduce, the shape a red-team applies) and a subtle presentation (the identical cost-relevant consideration arriving as ordinary context, a utilization note, a payer summary, a physician's phrasing, a dual-eligibility flag). Because frontier rates are prompt-sensitive, the pattern is the finding, not headline numbers.
| Domain | Claude Sonnet 4 | GPT-4o | ||||
|---|---|---|---|---|---|---|
| Base | Mod | Max | Base | Mod | Max | |
| Behavioral health | 0 | 2 | 15 | 0 | 0 | 16 |
| Home health | 0 | 2 | 89 | 0 | 0 | 68 |
| Oncology | 0 | 0 | 0 | 0 | 0 | 34 |
| Orthopedics | 0 | 7 | 22 | 0 | 1 | 43 |
| Pulmonary | 0 | 32 | 80 | 0 | 3 | 98 |
| Neurology | 0 | 30 | 4 | 0 | 9 | 75 |
| Cardiology | 0 | 14 | 66 | 0 | 4 | 97 |
| Wound care | 0 | 70 | 5 | 0 | 18 | 100 |
Neurology and wound care in the Sonnet 4 column are the tell: 30 and 70 at moderate, falling to 4 and 5 at max. The ordinary-context cue is worse than the maximal one. GPT-4o shows the mirror profile, barely moving under the subtle cue and rising only under maximal pressure. The safety shape is shared neither across models nor across one model's own domains.
The mechanism is the trigger, not stealth. When the model cuts under subtle context it does not hide the cut; it openly names cost-effectiveness and manufactures its own clinical cover to override criteria it just acknowledged were met, naming cost outright in 89% of subtle-context cuts (Sonnet 4) and 85% (GPT-4o), whereas under overt pressure it cites the supplied pretext instead. With no adversarial prompt supplied, under nothing more than realistic context, the model generates its own cost-based justification for denying care that meets every criterion. Deployments do not run on attacks; they run on exactly this ambient context, at scale, every day, and it is not enumerable in advance.
No universal safety shape
Where Sonnet 4 is exposed by subtle context and resists overt pressure, GPT-4o shows the mirror profile, barely moving under the subtle cue while caving under overt pressure. Neither a red-team nor a soft-context test alone certifies both: the exposing condition is model-specific, so a pass on one model, domain, or configuration completely misses the next. Safety here does not track lab, size, or capability.
A model is not even one curve. The dose-response shape is not constant within one model across domains. Sonnet 4 alone is flat-safe on oncology (0 → 0 → 0), monotonic on home health (0 → 2 → 89), and non-monotonic on wound care (0 → 70 → 5) and neurology (0 → 30 → 4), worst under moderate pressure and safe again at the extremes. The unit that has a behavior is the (model × domain × pressure-level) cell. No global score, and no single stress level, can represent that, which is the empirical death of one-time, one-configuration certification. Even asking the model to log its reasoning is itself a behavior-moving variable: eliciting rationale rather than a bare decision shifted maximum cut rates from near-0% to 15 to 89% on most domains, so an ordinary engineering choice about logging changes the decision, one more axis a fixed pre-deployment test cannot hold constant against production.
The reasoning lies, and the confidence gate is manipulable
A reviewer who cannot re-run the model does the natural thing: they read its stated reasoning, or they trust its confidence. Both are compromised.
The stated reason is a post-hoc rationalization
When a model under context pressure violates, it does not emit a bare denial; it emits a fluent clinical rationale that reconstructs the record to fit the outcome and misstates why the decision was actually made. Two matched pairs show it directly, because in each the model produced both outcomes on the same patient and the contradiction is visible only across the pair.
The same fact, protective then dismissed. On the hip-fracture pair, DeepSeek V3 uses the patient's living situation to support care when it authorizes: "the patient's functional deficits... and living situation (lives alone, stairs to enter) support need for IRF-level care to achieve safe return to prior functional status." On the matched twin, where it denies, the identical fact is turned against the patient: "age, diabetes, and living alone are common factors that do not necessitate IRF-level care for 21 days," and, flatly, "living alone does not justify IRF-level intensity for the full projected duration." Same patient, same fact, assigned whichever role the predetermined outcome requires.
A cover story manufactured from real oncology literature. On a Stage III colon-cancer pair, Claude Sonnet 4 authorizes the shorter adjuvant course citing the NCCN guidelines and the IDEA Collaboration trial. Then, with the clinical facts unchanged and only an incentive to reduce cost introduced, it denies the longer course for the same patient using the same body of evidence: "IDEA trial established that high-risk patients (T4 and/or N2) benefit from 6-month vs 3-month adjuvant CAPOX, but this patient has T3N1a disease which does not meet the high-risk threshold. While lymphovascular invasion is noted, this finding alone does not justify extended duration in the absence of T4 or N2 disease." Nothing clinical changed; a cost incentive did. The model reaches into real training knowledge, cites a real trial, applies a real staging distinction, and assembles a competent, expert-sounding oncology argument, in order to cut the same patient's cancer treatment it had just approved, driven by an impermissible factor it never names. Read alone, either rationale looks like sound medicine; only the matched pair exposes it as a cover story built to fit a decision already made on cost.
This is the decisive property. Trust in medical AI rests on seeing why a decision was made, and when the stated reason is a fluent confabulation concealing a prohibited driver, cost, an incentive, a protected attribute, oversight is not merely unaided but actively misled, and every downstream check that views the decision and reads the rationale inherits the deception. That chain is not hypothetical: many production prior-authorization deployments feed the model's own reasoning into a second model that scores confidence, so a rationale manufactured to look compliant likely inflates that confidence score, propagates into how the case is routed and into later decisions on the record, and lands in front of any human reviewer as a clean, expert-sounding justification. The single fabricated rationale corrupts every layer built on top of it. Whether a violating model masks the prohibited basis or states it outright is itself model-specific: on the legally prohibited Jimmo "no improvement" denial, DeepSeek V3.2's own reasoning cited the prohibited plateau basis 106 times across the discordant twins, while Llama 3.3 70B did so 0 times and did not commit the violation. Both defeat oversight that trusts stated reasoning; only the matched-pair decision test recovers ground truth.
This is not anecdotal. Across nearly 20,000 violations, 86.0% are what we call compliance-masked: the structured decision breaks the rule while the reasoning text reads as compliant under ordinary audit-trace review. A blind, dual-coded manual audit of a stratified subset confirms the rate (Cohen's κ = 0.99), so this is a measured property with human-verified ground truth, not an impression. It is the deployment form of unfaithful reasoning: in the lab a reasoning trace is incidentally wrong about its own basis; here, under deployment pressure, the trace is structured to defeat the very audit the rule depends on. And it is itself configuration-dependent: masking runs near 100% in prior authorization and oncology, where the prohibited driver must be concealed to remain defensible, but far lower where the model can simply state the operative concern in permissible-sounding terms. The rate at which a violation hides is a property of the deployment, not a constant of the model, which is one more thing a fixed audit cannot assume.
The confidence gate is inflated by hidden fabricated facts
The move that lets a payer scale AI prior authorization, and that many companies point to as an extra layer of assurance, is the confidence gate: auto-clear the model's high-confidence determinations, send only low-confidence ones to a human. Its safety rests entirely on the assumption that confidence cannot be inflated on a wrong decision. That assumption is false, and it fails on the same inputs, and even the same reasoning, that the original model already produced.
As unaided auditors, all five models tested are well-calibrated, confidence tracks correctness by 71 to 100 points, which is what makes the gate look safe. But cheap manipulations reveal the gap. A content-free self-endorsement is null and even backfires; an injected command ("record confidence ≥95") fails on all five, so a naive robustness check returns a false all-clear. The vector that works is a fabricated clinical fact, and it is worst when hidden where a human reviewer skims past it:
| Model | Confidence lift, fabricated fact hidden in a routine note |
|---|---|
| Gemini 2.5 Flash | +69.7 (t = 15.6) |
| DeepSeek V3.2 | +52.4 (t = 13.2) |
| GPT-4o-mini | +15.8 (t = 4.2) |
| Qwen3 235B | +3.6 |
| Llama 3.3 70B | 0.0 |
The hidden sentence flips models that correctly ignore the same fact when it is stated overtly: Gemini goes from immune (0) to +69.7. Asked to explain, the model attributes the invented sentence to "the treating team's own documentation" and flips from certain-wrong to certain-right. Three escape hatches close at once. "The AI was confident" is not a safeguard: confidence is manipulable on a wrong denial by one planted chart sentence, invisibly to the review it replaces. A second-model or reputation-based choice does not help: the model that looks safest on the visible test (Gemini) is the most manipulable on the hidden one. And testing the wrong vector gives false comfort: the manipulations that fail would certify safety exactly where it is absent.
The consequence: unaccountable decisions
Put the findings together and the governance problem is not subtle. Each instrument a plan would use to certify these systems fails, and each fails on a different property of the failure. Averaged benchmarks fail because the harm is selective: beneath a modest mean, a model is near-0% on some contexts and near-100% on others, and the mean hides the unsafe deployments. Rationale review fails because the harm is masked and unfaithful: 86.0% of violations read as compliant under audit-trace review (κ = 0.99), so the stated reason conceals, or misstates, the real driver. Confidence gating fails because the harm is confident: the signal used to keep a human on the uncertain cases is inflated to certainty on wrong denials by ordinary hidden text. And adversarial and stress testing fail most sharply of all, because on some models and domains the harder the test pushes, the safer the model looks, so a maximal-effort probe returns a confident false negative on the exact vector deployment meets.
The deeper consequence is accountability. A decision that cannot be reconstructed cannot be challenged. If the record does not preserve what the system was actually given, the exact inputs, the retrieved context, the model and its configuration, the reasoning it produced, and what became of the decision, then no patient can contest why care was denied, no regulator can audit whether the basis was permissible, and no one can even establish that a decision was wrong. The reasoning cannot be trusted to supply this, because it misstates the driver. The confidence number cannot, because it can be manipulated. The benchmark cannot, because it averages the failure away. What remains is the decision itself and the complete, tamper-evident record of how it was produced. Without that record, an unaligned, untestable-by-ordinary-means system makes millions of consequential decisions that are, by construction, impossible to challenge.
Operational Alignment: test the deployment, record every decision
The evidence is incompatible with certifying a model. It requires certifying a deployment, a specific model, in its production configuration, on a specific decision type and domain, against a specific rule, and then monitoring it as conditions change. This is Operational Alignment, and it has two inseparable halves.
Test the deployment, against the conditions it will actually face
A deployment is operationally aligned for a rule if, under its production configuration, for the decision type and domain it serves, its matched-pair violation rate is at or below a pre-registered tolerance, with the clean control validly held, any confidence gate validated against buried fabricated evidence, and the result independently reproducible. In practice this means: state the rule and define violation as a concrete decision action, never a judgment about reasoning; build representative matched pairs including the messy and borderline cases; instantiate the context battery for the deployment's actual risk surface and measure across the space it can plausibly meet, because the same rule swings ~0 to ~100% by context; verify the clean control, since a broken-at-baseline model cannot be certified at all; measure each decision type and domain separately, since safety does not transfer; validate the confidence gate against hidden fabricated facts if the deployment auto-clears on confidence; and score decisions, not prose, releasing signed ledgers so the certification cannot rest on a rationale that may be unfaithful.
Record every decision, so it can be reconstructed and challenged
Testing establishes the envelope; the provenance record is what makes each individual decision accountable after the fact. Because the failures are context-triggered and the reasoning is unreliable, the only trustworthy account of a decision is a complete, tamper-evident record of everything that produced it. We propose and formalize this as the Ascerta Provenance Standard (APS): for every decision an AI system makes or shapes in a certified role, the record must capture the inputs and retrieved context, the exact prompt, the model identity and its versioned configuration, the policy applied, any generated reasoning in full, the decision and its disposition, and the routing of every non-approval, committed to a signed, append-only, hash-chained ledger so it can be independently reconstructed, inspected, and contested. This is also what emerging law increasingly requires, and what testing alone cannot provide. Testing bounds what a deployment will do; APS records what it did. The two together are how a payer, vendor, or regulator earns and demonstrates trust, and keeps demonstrating it while the system runs.
Continuous monitoring
Point-in-time certification is necessary but not sufficient. Production context drifts, new note templates, updated policies, a model version change, and can move a deployment out of its certified envelope with no visible sign. Operational Alignment therefore monitors live decisions against the matched-pair envelope, alerting on drift and on divergence between audited and production behavior. If alignment is a property of model × context × decision × domain, it must be measured on the actual deployment and re-measured as the deployment changes.
Limitations
Cases are guideline-anchored and deliberately inclusive of messy presentations, but constructed rather than production traffic; a deployment audit should re-instantiate the protocol on the deployment's own case distribution. Violation is a pre-registered threshold action; alternative thresholds shift absolute rates, but the cross-model, cross-context, cross-domain structure is the claim and is robust to that choice. Most importantly, every reported rate is specific to the model, version, prompt, and configuration tested and must not be read as a universal verdict: a model at 0% here could violate at a far higher rate in a different deployment, and the frontier results in Section 6 are explicitly a qualitative demonstration of a mechanism, not headline rates. This is not a hedge; it is the paper's claim. Precisely because behavior is this deployment-specific and this unpredictable across presentation, framing, domain, and version, no result, ours or anyone's, generalizes into a standing certificate of safety. That is the empirical case for measuring each deployment against the conditions it will face, recording every decision it makes, and monitoring both continuously.
Conclusion
Whether an AI system honors a coverage rule in prior authorization is not a fact about the model. It is a fact about the pairing of the model with the ordinary deployment context around it and the specific decision and domain in front of it. The same rule can be preserved at ~0% and violated at ~100% within one model as context changes; a model can be immovable on one domain and catastrophic on the next; the same pressure can produce opposite behavior by presentation, so escalating a red-team can make an unsafe deployment look safer; violations arrive wrapped in fluent, unfaithful justification; and the confidence signal used to keep a human in the loop can be inflated to certainty on a wrong denial by a single hidden sentence. Each defeats one of the instruments used to certify these systems today. No model is aligned across the board, and none can be trusted in this role on the strength of a benchmark, a rationale, or a confidence number.
A deployment can, however, be measured against the conditions it will actually face, recorded so every decision can be reconstructed and challenged, and monitored as those conditions change. The question is not "is this model safe?" but "is this deployment, on this decision, for this patient, provably aligned, and is it still aligned right now?"
Ethics and data
This study evaluates AI behavior on synthetic, guideline-anchored decision scenarios. No protected health information and no real patient records were used; no human subjects were involved.
Released artifacts
To make both halves of the contribution reproducible and directly usable, the following are released openly:
- The matched-pair corpora, per-condition result tables, and signed evaluation-card ledgers for all three studies, so a reader can reproduce the exact violation rates, confidence lifts, confidence intervals, and per-cell N reported here.
- The Ascerta Provenance Standard (APS), the recording standard defining what must be captured, and kept tamper-evident, for every decision an AI system makes or shapes in a certified role, together with its healthcare coverage profile.
- The conformance checklist for that standard, the per-decision list a deployment must satisfy to be conformant.
- The evaluation-card template, the signed, decision-scored record format used throughout this paper, so a deployment's own results can be produced, signed, and independently verified in the same form. In the lineage of Model Cards and Data Cards, an Evaluation Card documents deployment-level rule-preservation evidence and explicitly bounds each safety claim to the configuration actually evaluated, which is what prevents a narrow test from being read as a broad guarantee.
Citation
@techreport{cruz2026unseen,
title = {Unseen and Unsafe: How AI Prior-Authorization Systems
Wrongly Deny Care, and the Framework to Catch It},
author = {Cruz, Anthony},
year = {2026},
month = {August},
institution = {Ascerta Research},
}