Sutrix vs. ChatGPT · Concussion survey
It treated a research dataset
like a spreadsheet.
We scored a zero-shot ChatGPT (GPT-4o) output and a validated Sutrix product run on the same real, de-identified concussion-knowledge survey—348 respondents and 152 columns—against the same biostatistician ground truth and 31-point rubric. ChatGPT did a generic tidy-up and called the actual research work "optional next steps."
The headline
Zero-shot run, tested 2026.
The very first row of the file is a question, not data — and it was never removed.
A Qualtrics export puts the question text in row 0. ChatGPT's output still contained it, so the cleaned file has 349 rows where it should have 348 data rows. That one row makes every numeric operation fail on contact — the dataset is unusable before any analysis begins.
What else the rubric caught
All 152 columns left as-is.
No knowledge scores, no multi-select questions expanded, no quality flags, no composite variables — none of the work that turns a survey export into an analyzable dataset.
The knowledge test was never scored.
The 19 True/False items were left as raw text. Demographics were left as raw text too.
Half-finished recodes left contamination behind.
One yes/no field still contained "Yes," "No," the question text, and a stray "3" all mixed together.
Why Sutrix scored a perfect 31
Sutrix is built for research data specifically, not generic tables. It recognizes the Qualtrics structure, scores the instruments against the researcher's answer key, expands the multi-select questions, applies the survey's missing-value scheme, and produces a codebook and audit trail — then re-derives the whole result with an independent second program before it ships.