A physics engine, and a verification tool enforced on top of it. PhysWall inverts a closed, published, non-linear physical law at a single measurement point — and refuses when the inverse is not unique. Seven laws, one engine, and the same refusal in all of them.
Every model answers. That is the problem, and it is not the one we are arguing about.
The loudest debate in AI safety is about goals: deception, power-seeking, a target specified badly enough to matter. Those are real questions and this essay is not about them.
It is about a quieter failure that is already in production, and that has a number attached to it.
ConstraintBench measured six frontier models across two hundred engineering tasks in February 2026. The finding:
“feasibility, not optimality, is the primary bottleneck. The best model achieves only 65.0% feasibility.” arXiv:2602.22465
One third of what a leading model returns cannot be built. Not “is suboptimal” — cannot exist. And nothing in the reply says so.
A second measurement, from a clinical workshop at ICML 2026, is sharper. Studying acute kidney injury prediction on MIMIC-IV, the authors asked whether the model knew when it was wrong:
“the LLM’s epistemic self-awareness is entirely absent.” EIML@ICML 2026
Not weak. Absent. And their framing of why it matters is exact: a confident, incorrect prediction for a critically ill patient, with no signal that the reasoning is unreliable, is qualitatively different from a model that flags its own uncertainty.
One third of what the best model returns cannot be built. Not “suboptimal” — cannot exist. And nothing in the reply says so.
This is not an oversight. It is what the training rewards.
“Most AI models are trained on maximum likelihood... This creates a test-taker mentality: the AI would rather guess and get it right 70% of the time than admit it doesn’t know and get a zero.”
A system optimised to produce the most likely continuation will produce one. Always. Calibration work is real and worth doing, and it is fighting the objective function rather than expressing it.
Which means the property has to come from somewhere the incentive is reversed.
“In 2026, AI safety stops being about intent and becomes an engineering problem. Systems must show their reasoning. Every output must trace back to real documents... Guardrails can’t rely on model judgment; they require deterministic enforcement. The rule becomes simple: if a system can’t show where its answer came from, it doesn’t belong near patient-impacting decisions.” Sorcero, January 2026
Three requirements, and none is about alignment:
Here is the whole argument in one exchange. The question is ordinary and the answer is the part that does not exist elsewhere.
The arithmetic doesn’t hold here.
At a 4% rate and 40 visitors, fewer than
two conversions are expected.
Every significance test rests on a normal approximation,
and that needs about five. Nothing can be concluded at
this size — not even “no change”.
Get to 4,783 per group and the
question becomes answerable.
power.prop.test.
Run it yourself →A model would have said “a 30% improvement is promising”. A free significance calculator would have said “not significant”, which is true and not actionable. The third answer is the only one that tells you what to do next — and the only one that refuses to compute at all, because at that sample size the maths behind the question does not apply.
A checker that returns true or false is not this. It answers a question the caller already knew to ask, in one domain, and gets rewritten for the next one.
It refuses in kinds, not in binary. A measurement that admits two values sends you to measure again. A target no configuration reaches sends you to change the target. A result outside what the material can be sends you upstream to the reading. Three different noes, three different afternoons.
Its vocabulary is domain-independent. If the same six refusals cover an antenna, a river gauge and a table of competing hypotheses, the layer is checking the structure of the question, not the physics. That is testable: build tools that hold no physical law at all and see whether they refuse in the same words.
It warns about itself. A tool reporting seven independent bounds invites you to read agreement between two of them as two confirmations. If two are the same relationship from different sides, it has to say so:
“Running both and getting agreement is not two confirmations, and anyone treating it that way is double-counting.” from the bounds page
No model will tell you that. It has no reason to.
A refusal layer does not address deceptive alignment. It does not address power-seeking. It does not address a goal specified badly. It checks whether a measurement settles a question, and none of those four is a question about measurement.
Claiming otherwise would be the exact failure mode under discussion: answering a question that was not asked, fluently.
It narrows the surface where an unverified output becomes an acted-on decision.
That is smaller than “prevents catastrophe” and it is not a small thing. The gap between what a model returns and what a person does with it is where the thirty-five percent lands. Today that gap is filled by a human who may or may not check. A deterministic layer makes checking the default and refusal a first-class result rather than an error state.
And there is a second-order argument that matters more.
Habits set now persist. If the ecosystem normalises “every consequential output passes a deterministic check with stated provenance”, that norm is in place before anything more capable arrives. If it does not, the default is what it is today: a confident number, no signal, and a person who trusts it because nothing told them not to.
You do not get to install that habit later, under pressure, on a system you no longer fully understand.
Twelve tools across RF engineering, hydrology, photonics, hydraulics, finance, sports statistics and competing-hypothesis analysis. Seven possible answers; six are refusals. Every law is somebody else’s — Bode 1945, Fano 1950, Airy 1845, Manning 1889, Landauer 1961, Heuer 1999, Miller & Sanjurjo 2018 — named on the page that uses it, together with where that law stops.
Four of the twelve hold no physical law at all, and refuse in exactly the same words as the eight that do. That is the part worth arguing with: it means the vocabulary is not borrowed from physics, and the same shape should transfer.
One tool is held back because it has never seen a real measurement, and its page says so rather than showing a number.
It is not a safety system. It is an existence proof that the three requirements above can be met at once, in one vocabulary, across unrelated domains — by one person, in weeks, using laws published decades ago.
If that is buildable at this scale, the argument that it is impractical at any other scale needs a better reason than it currently has.
See all twelve → How it was built →Every measured figure here is cited and checkable. None of them is our measurement.