Sutrix vs. Claude API · Reproducibility · July 2026
Same study. Same primary result.
Every Sutrix run.
We ran Sutrix and an unassisted Claude API model five times on each of three synthetic ground-truth studies. Both systems found the planted effects and correctly returned no finding on the null study. The measured difference was numerical reproducibility.
Primary-result reproducibility
“Identical” means the computed primary statistic matched exactly across repeated runs. On the predictor-pain study, Claude’s p-value ranged from approximately 1e-20 to 3e-15, and its selected test and effect-size sign varied between runs. Sutrix’s primary statistics were byte-identical in all 15 runs.
Both systems recovered the signal. Sutrix kept the number fixed.
On these small, well-behaved synthetic studies, neither arm produced a wrong-but-unflagged analytical decision. The benchmark does not claim a silent-error advantage; it isolates whether the same study returns the same primary numerical result.
Disclosed limitation
Sutrix’s primary results were stable, but its planner did not always select the same number of secondary tests for the RCT study. That changed the multiplicity-adjusted family across runs. Pinning the statistical analysis plan removes that remaining source of drift.