kdkodatajuanmaulana29@gmail.com

What Makes Writing Writing — kodata

Column 2: What Makes Writing Writing

index

Open data, followed all the way down to one person inside it.

Every piece here starts with a public dataset and ends with a story. The numbers can be checked: any figure set in mono opens its own evidence in a column to the right, so you never have to leave the sentence you are reading to find out whether it is true.

open to work

02pieces
02public datasets
271occupations
4,810teenagers

work

files

What Makes Writing Writing

introduction

dataset
Will AI Take My Job? Exposure, Skills and Wages
source
kylefengkfeng209/will-ai-take-my-job-exposure-skills-and-wages
retrieved
2026-07-29

What this data is

271 occupations in the United States, covering 81,559,900 workers. For each one the file carries what the Bureau of Labor Statistics publishes — median pay, how many people do it, its projected growth to 2034, the education it asks for — joined to how strongly the job draws on each of 21 cognitive abilities, and to scores for how exposed it is to artificial intelligence.

The interesting part is that exposure is scored twice, at two different levels of description: once for each cognitive ability, and once for each occupation taken whole. One row is one occupation.

Abstract

A job is made of abilities, so a job’s exposure ought to be predictable from the exposure of its parts. It is not. The two layers of this dataset run in opposite directions (−0.367), sharing only 13.5% of their variance. Exposure also rises with pay rather than falling with it.

Exactly one of the 21 abilities scores negative — Originality, the capacity for ideas nobody has had — and it is the ability writers rank highest on. Yet writing is rated the 4th most exposed occupation in the country, and moves further between the two orderings than any other job in the file. The variable the ability model is missing is not an ability. It is the medium the work comes out as: words, or objects.

What this piece asks

  1. How exposed is each occupation, and does that track pay the way we assume?
  2. If a job is built from abilities, does predicting its exposure from those abilities agree with the score the dataset publishes for the job itself?
  3. Which occupation disagrees most between the two — and is that claim robust to how the comparison is scaled?
  4. What kind of variable would explain the direction of the errors?

How the data is structured

filerowscolsone row iscolumns used here
ai_job_exposure.csv27123an occupationtitle, soc_code, wage, employment, growth, education, ai_exposure_llm_human, ai_exposure_llm_gpt4, ai_exposure_level
occupation_cognitive_profile.csv27123an occupation, as 21 ability levelssoc_code plus all 21 ability columns, in the file's own column order
cognitive_ability_ai_exposure.csv213a cognitive abilitycognitive_ability, cognitive_family, ai_exposure_score
hold this in mind
The per-ability scores and the per-occupation scores come from different methodologies and were not built to be compared with each other. Comparing them is my analytical decision, not a claim the dataset makes. The negative correlation shows that the two measure different constructs — not that either one is wrong.

Prose is written for anyone. Every figure set in mono opens its own evidence in a column to the right. Blocks marked advance hold the method, the confidence intervals, and the objections — open them if you want the working, skip them and the piece still reads.

piece 01 · the premise

Of the 271 occupations measured in this dataset, writing ranks most exposed to artificial intelligence.

Only three sit above it: survey researchers, translators, and public relations specialists. Below it sit 267 others — including almost every job we have spent a decade imagining would go first. Janitors are down there, scoring : twenty-eight times lower than writing.

That is not the finding. That is only the premise — and it is already enough to turn one old assumption over. In this data, exposure to AI rises with pay rather than falling with it . The jobs the dataset calls most exposed are, on the whole, the better-paid ones.

finding 01 · the only negative one

Of 21 cognitive abilities, exactly one scores negative.

The dataset measures exposure twice, at two different levels of explanation. Once for each cognitive ability a job might demand, and once for each job as a whole. Start with the abilities.

The most exposed ability is — arranging things according to a rule. Then memorisation, then perceptual speed. All of them tidy, finite operations, the kind where an answer can be marked right or wrong.

And at the bottom of the list, alone below zero: . The ability to come up with ideas nobody has had. It is the only one of 21 that scores negative — meaning that the more a job depends on it, the less this dataset thinks AI can do it. On that ability, writers sit in the th percentile.

The whole instrument is worth looking at once, because the fable at the end of this page is about it: all , in the order the source file lists them, with how much writing demands each one.

finding 02 · the dataset against itself

If both measures are measuring the same thing, they have to agree.

A job is made of abilities. So it should be possible to predict a job’s exposure from the exposure of the abilities it is built from, weighted by how much the job demands each one. Do that for all 271 jobs, then compare the prediction against the score the dataset actually publishes.

They do not agree. They are not even weakly related — they run in opposite directions . The more a job is built out of abilities this dataset calls exposed, the less likely that same dataset is to call the job exposed. The two layers share only of their variance; the other 86.5% is unaccounted for.

And one occupation moves further between those two orderings than any other. Built from its parts, writing should rank of 271 — almost the safest job in the country. Rated whole, it ranks . That is a displacement of , the largest of all 271.

advancewhy this is stated as displacement rather than as a residual

An earlier version of this piece led with a residual instead: writing sat +0.831 above its fitted value, 2.64 SD out, rank 1 of 271. I tested whether that survives a different scaling choice. It does not, so the claim was demoted and this block exists to say so.

what is robust
The correlation between the two layers is −0.3670 95% CI [−0.4658, −0.2591], Spearman −0.3990. Pearson is unchanged by any linear rescaling of the prediction, so this is robust by construction, and the interval is nowhere near zero. Rank displacement is also scale-free: it compares two orderings and never touches a unit.
what is not
Residual size and residual rank depend on how the prediction is mapped onto the rating’s scale. Shipped method — standardise the prediction, match the rating’s mean and SD — puts writing at rank 1. Ordinary least squares puts it at rank 17 with a residual of +0.273 (1.55 SD). Still notable. Not “the largest departure in the dataset”.
the objection
You might say OLS is simply the right answer and I should use it. The OLS slope here is -5.1569 — negative. Its best-fitting line therefore predicts that jobs built from more-exposed abilities are rated less exposed, which bakes in the very inversion this piece is trying to show. A residual against that line cannot answer the question “if parts implied the whole, where would this job sit?”. My rescaling answers that question, but it is an assumption, not a regression — and calling its output a “residual” invited exactly the wrong reading.
what would break it
If the ability-exposure weights were rescaled non-linearly, or if the ability demand levels were centred before weighting, the displacement could change. That has not been tested. The correlation would survive both.

You can look up any occupation in the file and see where the two measures put it.

predicted is where the occupation ranks when its exposure is built up from the 21 abilities it demands; rated is where the dataset puts it when judging the occupation whole; moves is the distance between those two positions, out of 271. A large positive number means the dataset treats the job as far more exposed than its own parts imply. Writing moves +263 places, the largest displacement in the file. Bars are scaled within each measure so two columns can be read against each other, not against zero.

Look at which way the errors go. Everything predicted too safe works through words. Everything predicted too exposed works through objects. A mechanic’s work really does demand high perceptual speed — but that speed is spent through hands, in three dimensions, on a thing that can be dropped. The ability model cannot see bodies. It only sees operations.

There is a third rating in this file, and it takes the same side. Alongside the human rater, every occupation was also scored by GPT-4. The two agree at , which sounds like agreement until you look at where they part. GPT-4 rates medical transcription higher than the human rater does — 0.232 against 0.875, from fairly safe to almost entirely exposed, on the same job. Bookkeeping clerks, court reporters, composers: same direction.

And the occupations the human rater marks higher than GPT-4 does are childcare workers, fitness instructors, concierges. Work whose content is being present with a person. So a model asked to rate exposure over-weights work that looks like text and under-weights work that requires a body in a room — which is the same failure the ability layer makes, found again by a measure that shares none of its machinery.

This is not proof that one of the two raters is wrong. Most likely both are right and measuring different things: one the cognitive operation, the other the practical substitutability of a whole occupation. But they were published side by side, in one dataset, under one word: exposure.

the bridge · findings that force decisions

This is the layer that separates data as a foundation from data as decoration. Change a number in the left column and the story on the right has to be rewritten.

finding

Writers and authors is the largest disagreement in the dataset: predicted rank 267, rated rank 4, out of 271.

decision

The story cannot be about being replaced. It has to be about being misread — the distance between what a job is made of and what a job is judged to be.

finding

Originality is the only one of 21 abilities scoring negative (−0.144), and it is the ability writers rank highest on.

decision

The turn has to happen on that single quality, and it has to be the quality that cannot be entered into a column. Not a metaphor; the mechanism.

finding

Writing's projected growth is +4%, “As fast as average”. Of 90 highly exposed occupations, 72 are projected to grow and 13 to shrink.

decision

No apocalypse. Nothing is allowed to end inside the story, because not one number supports an ending. What is left is ambiguity.

the story

The Weigher of Trades

A fable. It contains no figures at all — the numbers have had their say already, and what is left of them here is only their shape.

Once there was a kingdom that heard a rumour. An engine was coming, the rumour said, that could do what people do.

The king was not a cruel man, only a careful one. He did not want to wake one morning and find half his subjects idle and the other half unaware. So he sent out a clerk with a brass instrument and a ledger, and told him to weigh every trade in the kingdom, so that when the engine arrived he would know whom to move and whom to leave alone.

The instrument was a good one. Along its edge ran a row of small notches, and each notch stood for something a pair of hands might be asked to do. To remember. To sort. To notice quickly. To speak, and to be understood. To picture a thing that was not in the room. To think of something nobody had thought of yet.

The clerk would come to a workshop, watch the work for an afternoon, and ask: how much of this does your trade require? And of this? And of this? Then he turned the notches to match, and the instrument settled, and gave a number, and the number went into the ledger.

He weighed the roofers. The instrument said: safe.
He weighed the boilermakers. Safe.
He weighed the butchers, and the men who mend the great wheels, and the women who lay the cables along the top of the world. Safe, safe, safe.

It was pleasant work, and the ledger filled, and by autumn the clerk came to the last house on the last road, where a woman sat at a table, writing.

He watched her for an afternoon, as he had watched all the others, and then he turned his notches.

Remembering: some. Sorting: a great deal, though not in any order he could see. Noticing quickly: hardly any. Picturing a thing not in the room: a very great deal. Speaking so as to be understood: as much as any trade in the kingdom, and then more.

The instrument settled. It gave its number. And the number said: safe. Safer than the roofers. Safer than the butchers. Safer, in fact, than almost every trade the clerk had weighed all summer.

He wrote it down. He wished her a good winter. The ledger went up to the king, and in the last house on the last road everybody slept well.

In the spring the engine arrived.

It was smaller than anyone expected, and it did not want anything, and this was somehow worse. The king, being careful, did not consult his ledger. He put the question to the engine directly: of all the trades in my kingdom, which could you do?

The engine answered without hesitating, the way a river answers a slope. It named a trade, and then another, and then another. And fourth on its list — fourth of all the trades in the kingdom, ahead of the clerks and the counters and every other trade the ledger had marked in red — it named the woman in the last house on the last road.

The king had the two lists laid side by side on a long table, and stood looking at them for some time. They were the same trades. They were the same kingdom. And they were in very nearly the opposite order.

Here is the thing about the notches on the brass instrument.

Of all of them, every one but a single notch worked the same way: turn it up, and the engine grew stronger. The more remembering a trade required, the better the engine could do it. The more sorting, the better. The more noticing quickly, the better.

One notch, and only one, ran the other way. The notch for thinking of something nobody had thought of yet. Turn that one up and the engine faltered, and turned back, and had nothing to offer.

It was the one notch the woman had turned nearly all the way.

So both readings were true, and that is the part the king could not get past. She was made, more than almost anyone in his kingdom, out of the single thing the engine could not do. And her trade had been named fourth.

The clerk, when he was finally asked, took a long time to answer, and when he did he did not defend his instrument.

It weighs what the work asks of you, he said. It has no notch for what the work comes out as. A mechanic and a writer can turn the very same notches — both must notice quickly, both must hold a shape in mind, both must judge when a thing is finished. But one of them puts their hands on an object that can fall. My instrument cannot see the falling. It can only see the noticing.

And there was something else, which he said more quietly. The instrument asks how much. It cannot ask which one. Which sentence. In which paragraph. After how many were thrown away. There is no notch for that and there never will be — because a notch has to compare one person against another, and what makes her sentences hers is precisely that they cannot be compared.

Nothing happened as fast as anyone expected. That is usually how it goes.

The engine stayed. Some trades thinned and some thickened, and not always the ones the lists had promised. The woman in the last house on the last road went on writing, and was sometimes paid more than before and sometimes less, and no year was quite like the last.

Both lists went into a drawer in the palace, where they are still, one on top of the other, each perfectly correct. Nobody has ever taken them out and set them side by side again, and so nobody has noticed that they disagree.

Source. Will AI Take My Job? Exposure, Skills and Wages, Kaggle — kylefengkfeng209/will-ai-take-my-job-exposure-skills-and-wages. Retrieved 2026-07-29. 271 occupations, 81,559,900 workers covered, 21 cognitive abilities.

Limits. The per-ability scores and the per-occupation scores come from different methodologies and were not designed to be compared directly. The negative correlation above shows that they measure different constructs — not that either one is wrong. Comparing them is my analytical decision, not a claim the dataset makes. Every figure set in mono is traceable to a source file; the fable is fiction, and deliberately contains no figures at all.