PhysWall · The Physical-Logic Gateway · four-dimensional gap decomposition

Three hypotheses, and the evidence eliminates two

A physics engine, and a verification tool enforced on top of it. PhysWall inverts a closed, published, non-linear physical law at a single measurement point — and refuses when the inverse is not unique. Seven laws, one engine, and the same refusal in all of them.

List what happened and what might explain it. It reports which explanations the evidence rules out — and says so when two survive rather than picking one.

Heuer’s word for what matters is diagnosticity: evidence consistent with every hypothesis has zero diagnostic value, however strong it feels. ACH is a structured analytic technique against confirmation bias — you count the inconsistent entries, not the consistent ones.

Not which one sounds best, and not which one you started with. Which one is left after everything that contradicts it has been removed.

The method is not ours. It is the Analysis of Competing Hypotheses, written down by Richards J. Heuer Jr. for the US intelligence community: list the hypotheses, list the evidence, and score each piece against every hypothesis, looking to disprove rather than to confirm. Heuer's own warning is the part most tools drop — that the matrix should not dictate the conclusion. So when two explanations are still standing, this says two are standing. Picking between them is a judgement about the case, and it is yours.

What it does

You list what happened, and what might explain it. It works out which explanations survive, which are ruled out and by what, and how much of the surviving one is actually supported rather than merely unchallenged.

Before weighing explanations it is worth checking the number they are explaining: a total that does not match its parts needs arithmetic, not a hypothesis.

The verdict scale is the same one the physics apps use:

CERTAIN        the evidence forces it
THEORETICAL    consistent, and nothing supports it directly
INTERPRETIVE   depends on a reading of the evidence
UNDETERMINED   more than one survives, and they conflict
NO VERDICT     the question cannot be answered as asked

Why this one matters most

Every other tool here rests on a closed physical law — Bode-Fano, Landauer, the error function. This one has none. No formula, no constants, nothing to invert.

And the same machinery works. Which means what these tools share is not the physics: it is the structure. That is a claim we could only test by removing the physics entirely, and this is that test.

⚠ What it is not for

This is not a tool for deciding legal cases or predicting how a court will rule. An engineer reviewing it said so plainly, and he is right. Legal reasoning is probabilistic and human. This is neither, and using it that way would be a mistake we would rather prevent than disclaim.

What it does is check whether an argument holds together — which explanations the evidence rules out, and what is left. That is useful in a legal context the way a spellchecker is useful to a novelist, and no more than that.

What it will not do

It will not pick a winner when two explanations both survive. It reports that they conflict and stops, which is the answer when the evidence genuinely does not decide.

It also separates two things most tools merge: how well the surviving explanation is supported, and whether it is the only one left. Those are independent, and combining them into one score loses both.

Click to open it → Or the basketball one

Runs in your browser. Nothing you type is sent anywhere.

Where people use this shape of reasoning

Incident review. Fault diagnosis. Historical argument. Anything of the form "several things could explain this — which ones can I rule out?" The formal name is consistency-based diagnosis, and it goes back to Reiter in 1987. We did not invent it; we made a version you can run.

The same engine reads antenna bandwidth, conductor loss and heat limits — with the physics back in.

WHAT 26 OF 30 IS, AND IS NOT

The coding here was blind: each report was read and its cause coded from the report alone, before any comparison. That matters more than the number.

WHERE IT SITS

A review of 25 aviation inter-rater studies puts the accepted agreement range at 70% to 88%. 26 of 30 is 86.7% — inside it, near the top.

AND WHY MOST PUBLISHED NUMBERS ARE NOT COMPARABLE

Roughly half the published kappas in this field are post-discussion. One marine study reports it plainly: the raw figures were 0.45 and 0.39, the raters then talked, and the recalculated figures — 0.72 and 0.64 — are what got published. A helicopter study says outright that no reliability measure can be computed from a consensus coding at all.

0.39 is below the usual acceptability floor. 0.64 is not. Same coders, same reports, one conversation in between.

AND WHAT WE ARE NOT CLAIMING

Percent agreement is not kappa. Kappa corrects for chance and percent agreement does not, so the same 86.7% gives a kappa anywhere from 0.61 to 0.82 depending only on how the categories are distributed. Until that distribution is published, this number should not be compared to a published kappa — and comparing them anyway is the failure this whole site is about.

PhysWall was developed and architected by Gadi Zion.

Built on PhysWall — the same engine reads antenna bandwidth, conductor loss, bit erasure and heat limits. It answers what the measurement implies, and refuses when the measurement cannot say.

Check this instead of believing it. Every number here reproduces from a source that is named, and the claims that turned out wrong are still printed next to what replaced them. The same engine runs all of these — it asks how much a measurement allows you to conclude, and refuses the same way in every field. The same engine runs all of these — it asks how much a measurement allows you to conclude, and refuses the same way in every field. How to check each one →