Open data, followed all the way down to one person inside it.
Every piece here starts with a public dataset and ends with a story. The numbers can be checked: any figure set in mono opens its own evidence in a column to the right, so you never have to leave the sentence you are reading to find out whether it is true.
open to work
02pieces
02public datasets
271occupations
4,810teenagers
work
files
evidence · two raters, same jobs
Both columns are per-occupation exposure ratings of the same 271 jobs — one by a human rater, one by GPT-4. They correlate +0.834895% CI [+0.7947, +0.8676], with a mean absolute gap of 0.0767 and a maximum of 0.643.
GPT-4 rates these far more exposed than the human did
occupation
human
gpt-4
gap
Medical transcriptionists
0.232
0.875
+0.643
Bookkeeping, accounting, and auditing clerks
0.314
0.802
+0.488
Court reporters and simultaneous captioners
0.521
0.958
+0.437
Music directors and composers
0.281
0.711
+0.430
Computer hardware engineers
0.318
0.727
+0.409
Musicians and singers
0.133
0.408
+0.275
and these far less
occupation
human
gpt-4
gap
Concierges
0.700
0.467
−0.233
Fitness trainers and instructors
0.235
0.015
−0.220
Survey researchers
0.844
0.625
−0.219
Agricultural and food scientists
0.658
0.456
−0.202
Childcare workers
0.340
0.141
−0.199
Public relations specialists
0.788
0.591
−0.197
Read the two lists as pairs. Transcription, bookkeeping, court reporting, composition — work that arrives as text or symbols. Childcare, fitness instruction, concierge work — work that requires being in the room. Writers and authors sits at 0.774 human against 0.877 GPT-4, a gap of +0.103, ranked 47 of 271: GPT-4 thinks writing is even more exposed than the human rater does.
advancewhat this can and cannot be used for
not independent
These raters are not independent of each other in any strong sense — both are scoring the same occupation descriptions, and the human rater may have had model output available. Treat the agreement as a consistency check, not as replication.
the pattern is the point
A correlation of +0.8348 between two raters means the disagreements are a small minority of cases. What makes them worth showing is that they are not scattered — they sort cleanly by medium of output, the same split the ability layer produces.
direction unknown
Nothing here says which rater is right. It is entirely possible that GPT-4 is correct about transcription and the human rater is correct about concierges. The claim is only that the two disagree along one axis, and that the axis is not cognitive difficulty.