Discuss a study
← All benchmarks

Sutrix vs. Claude API · Reproducibility · July 2026

Same study. Same primary result.
Every Sutrix run.

We ran Sutrix and an unassisted Claude API model five times on each of three synthetic ground-truth studies. Both systems found the planted effects and correctly returned no finding on the null study. The measured difference was numerical reproducibility.

Primary-result reproducibility

Identical study repeated 5 times
Claude API
Sutrix
Null-effect study
3 / 5 identical
5 / 5 identical
RCT PHQ-9 study
2 / 5 identical
5 / 5 identical
Predictor-pain study
1 / 5 identical
5 / 5 identical

“Identical” means the computed primary statistic matched exactly across repeated runs. On the predictor-pain study, Claude’s p-value ranged from approximately 1e-20 to 3e-15, and its selected test and effect-size sign varied between runs. Sutrix’s primary statistics were byte-identical in all 15 runs.

Both systems recovered the signal. Sutrix kept the number fixed.

On these small, well-behaved synthetic studies, neither arm produced a wrong-but-unflagged analytical decision. The benchmark does not claim a silent-error advantage; it isolates whether the same study returns the same primary numerical result.

Disclosed limitation

Sutrix’s primary results were stable, but its planner did not always select the same number of secondary tests for the RCT study. That changed the multiplicity-adjusted family across runs. Pinning the statistical analysis plan removes that remaining source of drift.

Method. Three synthetic ground-truth studies (null effect, RCT PHQ-9, predictor pain), five runs per arm per study. The competitor was unassisted Claude through the Anthropic API, not the Claude for Life Sciences product. Both arms used the same significance threshold and BH-FDR procedure. Run-level outputs and the scoring ledger are retained with the benchmark record.