Discuss a study
← All benchmarks

Sutrix vs. Julius AI · Concussion survey

A finished-looking workbook,
wrong about almost everyone.

We scored Julius AI’s best guided output and a validated Sutrix product run on the same real, de-identified concussion-knowledge survey—348 respondents and 152 columns—against the same biostatistician ground truth and 31-point rubric. Julius produced a clean, professional output. It was also silently wrong about most respondents.

The headline

31-point rubric
Julius AI
Sutrix
Structure recognition (of 12)
6
12
Domain operations (of 10)
8
10
Output quality (of 9)
4
9
Total (of 31)
18
31

Julius's best result across the levels we tested (guided, with the data dictionary provided). Tested 2026.

Julius silently reversed the gender of 305 of 309 respondents.

The answer key codes Male as 1 and Female as 2. Julius coded them the other way for all but four respondents — and nothing in the output flagged it. A researcher who trusted that workbook would run every analysis with the Male and Female groups swapped and never know. It made the same class of error on ethnicity, wrong for 302 of 313 respondents.

What else the rubric caught

  • It mis-scored the knowledge test — even with the answer key.

    One True/False item was scored backwards at every level, a systematic error carried across all 291 scored respondents. Giving Julius the scoring key did not fix it.

  • More documentation made it worse, not better.

    Given the full data dictionary, Julius scored lower than it did when merely guided — it could not integrate multiple documents, and read one as empty.

  • No reproducible code, at any level.

    The result cannot be re-run or traced back to the raw data. The biostatistician's reference for the same job is a 500-line, re-runnable script.

Why Sutrix scored a perfect 31

Sutrix codes against the researcher's validated answer key rather than a guessed one, so the gender direction and the knowledge score match the ground truth. Every result is re-derived by an independent second program before it ships — on this survey that pass agreed on every number — and the cleaned dataset regenerates bit-for-bit from saved code. The output that looks finished and the output that is correct are not the same thing; the rubric measures the second, and Sutrix cleared all 31 checks.

Method. One real de-identified concussion-knowledge survey (348 respondents × 152 columns). Ground truth is a biostatistician's cleaned dataset, answer key, and analysis. Both tools received the same evaluation rubric. Julius received guided instructions and the data dictionary; the Sutrix figure comes from a later validated product run. This compares the completed artifacts, not identical product interfaces or configurations. Every failure above was verified cell-by-cell against Julius's output. Raw outputs are available on request.