kdkodatajuanmaulana29@gmail.com

evidence · two raters, same jobs — kodata

Column 2: evidence · two raters, same jobs

index

Open data, followed all the way down to one person inside it.

Every piece here starts with a public dataset and ends with a story. The numbers can be checked: any figure set in mono opens its own evidence in a column to the right, so you never have to leave the sentence you are reading to find out whether it is true.

open to work

02pieces
02public datasets
271occupations
4,810teenagers

work

files

evidence · two raters, same jobs

Both columns are per-occupation exposure ratings of the same 271 jobs — one by a human rater, one by GPT-4. They correlate +0.8348 95% CI [+0.7947, +0.8676], with a mean absolute gap of 0.0767 and a maximum of 0.643.

GPT-4 rates these far more exposed than the human did

occupationhumangpt-4gap
Medical transcriptionists0.2320.875+0.643
Bookkeeping, accounting, and auditing clerks0.3140.802+0.488
Court reporters and simultaneous captioners0.5210.958+0.437
Music directors and composers0.2810.711+0.430
Computer hardware engineers0.3180.727+0.409
Musicians and singers0.1330.408+0.275

and these far less

occupationhumangpt-4gap
Concierges0.7000.467−0.233
Fitness trainers and instructors0.2350.015−0.220
Survey researchers0.8440.625−0.219
Agricultural and food scientists0.6580.456−0.202
Childcare workers0.3400.141−0.199
Public relations specialists0.7880.591−0.197

Read the two lists as pairs. Transcription, bookkeeping, court reporting, composition — work that arrives as text or symbols. Childcare, fitness instruction, concierge work — work that requires being in the room. Writers and authors sits at 0.774 human against 0.877 GPT-4, a gap of +0.103, ranked 47 of 271: GPT-4 thinks writing is even more exposed than the human rater does.

advancewhat this can and cannot be used for
not independent
These raters are not independent of each other in any strong sense — both are scoring the same occupation descriptions, and the human rater may have had model output available. Treat the agreement as a consistency check, not as replication.
the pattern is the point
A correlation of +0.8348 between two raters means the disagreements are a small minority of cases. What makes them worth showing is that they are not scattered — they sort cleanly by medium of output, the same split the ability layer produces.
direction unknown
Nothing here says which rater is right. It is entirely possible that GPT-4 is correct about transcription and the human rater is correct about concierges. The claim is only that the two disagree along one axis, and that the axis is not cognitive difficulty.