Open data, followed all the way down to one person inside it.
Every piece here starts with a public dataset and ends with a story. The numbers can be checked: any figure set in mono opens its own evidence in a column to the right, so you never have to leave the sentence you are reading to find out whether it is true.
open to work
02pieces
02public datasets
271occupations
4,810teenagers
work
files
evidence · where the line falls
The distribution of questionnaire totals across all 4,810 teenagers. It runs smoothly from 0 to 51 with no gap anywhere — there is no natural place in it where “well” becomes “unwell”. The orange bar is the line the source file draws, at 14.
total 0median 5 · mean 7.2751
cutoff used by the file
14
teenagers at or above it
787 of 4,810
share flagged
16.4%
questionnaire range
0 – 51
median total
5
mean total
7.27
the cutoff was verified, not assumed
The source ships a binary depressed column without stating what produced it. So the pipeline tries every cutoff in range and reports how well each reproduces that column across all 4,810 rows. Exactly one is perfect, and the build asserts it — if a future version of the dataset changed the line, the pipeline would fail rather than quietly publish a wrong claim.
cutoff
agreement with the file’s own column
bdi_total ≥ 11
93.16%
bdi_total ≥ 12
95.72%
bdi_total ≥ 13
97.98%
bdi_total ≥ 14
100.00%
bdi_total ≥ 15
97.78%
bdi_total ≥ 16
96.49%
bdi_total ≥ 17
94.80%
bdi_total ≥ 18
93.45%
bdi_total ≥ 19
92.77%
bdi_total ≥ 20
91.83%
bdi_total ≥ 21
90.52%
bdi_total ≥ 14 reproduces the column exactly (100%). The next best, 13, gets 97.98% — close, and wrong.