evidence · which ratios survive an interval — kodata
Column 3: evidence · which ratios survive an interval
index
Open data, followed all the way down to one person inside it.
Every piece here starts with a public dataset and ends with a story. The numbers can be checked: any figure set in mono opens its own evidence in a column to the right, so you never have to leave the sentence you are reading to find out whether it is true.
4,810 teenagers — 2,446 boys and 2,364 girls. For each one: how much leisure screen time they report, four separate measures of how they sleep, and a 21-question depression questionnaire kept both as a total and as its 21 individual answers.
The questionnaire total runs from 0 to 51, with a median of 5. One row is one teenager. The file also ships a yes/no depressed column, without saying what produced it.
Abstract
Screen time is associated with depression scores here, and the association is tiny: +0.1344, accounting for 1.8% of why these teenagers differ from one another. Sleep quality accounts for 14.5% — 8 times as much — and holding hours slept constant absorbs 46% of the screen association.
The larger finding is about how such a claim gets made. Saying anyone is depressed requires drawing a line across a smooth distribution, and saying anyone is a heavy user requires drawing a second one. Neither line is usually reported. Moving both across their plausible range yields 130 risk ratios from 1.15× to 2.97×, all arithmetically correct, of which 38 cannot be distinguished from no difference at all. And breaking the questionnaire into its 21 items reveals no symptom to point at: the association is spread almost evenly across all of them.
What this piece asks
How strong is the screen-time association really, next to everything else measured on the same children?
What has to be decided before “N times more likely to be depressed” can be written at all?
Across every defensible pair of those decisions, what range of answers does this one dataset support?
Is the association attached to any particular symptom, or spread across all of them?
screen_normal_day_1to6 and bdi_item_01 … bdi_item_21
hold this in mind
This is cross-sectional. Nothing here can tell you which way any arrow points — a teenager who is already unhappy may reach for a screen rather than the reverse. The partial correlations show how much of a weak association travels with sleep; they are a decomposition, not a causal adjustment. And the questionnaire items ship unlabelled, so they are reported by number and never read as specific symptoms.
Prose is written for anyone. Every figure set in mono opens its own evidence in a column to the right. Blocks marked advance hold the method, the confidence intervals, and the objections — open them if you want the working, skip them and the piece still reads.
piece 02 · the premise
You have read the headline. Screen time is linked to teenage depression. In this data, that is true — and it is also almost nothing.
4,810 teenagers, each with a screen-time measure and a 21-question depression questionnaire. The correlation between them is . Real, positive, and in a sample this size not a fluke.
But a correlation of that size accounts for of why these teenagers differ from one another in how they answered. The other 98.2% is something else. Screen time is not the story of teenage unhappiness. It is a rounding error that got a press cycle.
Here is what does better, in the same file, measured on the same children: how well they sleep, at — the explanatory power of screens. And nearly half of the screen-time association, of it, disappears the moment you account for how many hours they actually slept.
If percentages of variance mean nothing to you, here is the same gap in plain units. Both columns are average depression score, climbing as you go down. One staircase is shallow and one is steep.
more screen time →level 16.03level 25.9level 37.24level 48.16level 58.76level 610.16swing of 4.26 points
Same 4,810 teenagers, same questionnaire, same scale. Going from the least to the most screen time moves the average by points. Going from the best to the worst sleep moves it by . Both climb steadily, so neither is noise — they are just very different sizes.
finding 01 · two lines nobody reports
“More likely to be depressed” requires drawing two lines first.
A questionnaire total is a number from 0 to 51. To say somebody is depressed you have to decide where the number stops being ordinary — and to say somebody is a heavy user you have to decide that too. The source file draws the first line at , which is why it reports 16.4% of these teenagers as depressed. Nothing forced that number.
Both sliders below move real lines. Every ratio they produce was computed in Python and committed to the data file — the page is not calculating anything, it is showing you 130 results that already exist.
Teenagers with more than about 4.6 hours of leisure screen time a day are 1.43× more likely to be depressed.
…where “more than about 4.6 hours” means a screen-time index of 3 or more (2,804 of 4,810 teenagers), and “depressed” means a questionnaire total of 14 or more — 524 heavy users and 263 others. 18.7% versus 13.1%. 95% CI [1.24×, 1.63×]
measured on
These are the two lines the source file itself draws. Every other position below is equally available to anyone reporting this dataset.
≥2
≥3
≥4
≥5
≥6
cutoff 5130 true headlines, 1.15× to 2.97× · hatched = 38 whose interval includes 130
That is the finding. Not that the number is wrong — every one of those 130 ratios is arithmetically correct. The finding is that the sentence is a choice, and the two decisions that determine it are the two things a published claim almost never states.
And there is a third lever, which the switch above the sliders exposes. Hold both thresholds exactly where the source file puts them and change nothing but the subgroup: the pooled answer is 1.43×, boys alone 1.75×, girls alone 1.31×. The correlation itself barely differs by sex — against — so this is not a finding about boys. It is the same arithmetic showing that a threshold amplifies whatever it is pointed at, and that a claim can be made a third larger by a choice nobody reports.
There is a harder version of this. Of those 130 ratios, cannot be told apart from no difference at all — their confidence intervals include 1. They are still real numbers, still correct, still quotable. Those are the hatched cells in the grid above.
advanceintervals, and why the grid is hatched
method
Each cell is a risk ratio with a 95% interval from the Katz log method: se = √(1/a − 1/n₁ + 1/c − 1/n₀), exponentiated. Analytic rather than bootstrapped, so the pipeline stays deterministic and needs no seed.
result
92 of 130 cells (70.8%) have an interval clear of 1. The remaining 38 do not, and are almost all at the extremes where the flagged group collapses to a few dozen teenagers.
the point estimate
The headline correlation itself is solid: +0.134495% CI [+0.1065, +0.1620] at n = 4,810. Small is not the same as uncertain — this association is precisely estimated and genuinely tiny, which is a different and more interesting problem than being noisy.
objection
You could argue that reporting the widest ratio is a straw man, since no careful analyst would cut at 30. Fair. The point is not that anyone does it deliberately, but that the sentence carries no trace of which lines were used, so a reader cannot tell a careful choice from a convenient one.
finding 02 · nothing to point at
There is no symptom screen time attaches to. It attaches to all of them, faintly.
If screens were doing something specific to these teenagers, it should show up somewhere specific — worse sleep, or more self-blame, or less appetite. So take the questionnaire apart and correlate screen time against each of its 21 questions on its own.
010203040506070809101112131415161718192021
Correlation between screen time and each of the 21 questionnaire items separately. Mean 0.0799, standard deviation only 0.0174 — 17 of 21 sit between 0.0619 and 0.0979. The source file ships the items unlabelled, so they are shown by number.
Rebuilt from the remaining 21 items, the score still correlates with screen time at +0.1344
You cannot take the finding apart by removing symptoms, because it does not sit in any of them. Note what this does not establish: sleep quality, which explains eight times as much and has an obvious mechanism, loads just as evenly across these 21 items. So an even spread says something about the questionnaire, not about screens. What it does establish is narrower and still worth having — there is no symptom here for anyone to point at.
I expected a spike and went looking for one. There isn’t a spike. My first reading of that was: no single symptom means no specific harm, so what this is really measuring is a general mood in how someone fills out a long form late at night.
Then I ran the same breakdown for sleep — which explains as much and has an obvious mechanism — and its profile is just as flat. Screen time’s spread relative to its size is ; sleep quality’s is — slightly more uniform, not less. So flatness cannot tell a mechanism apart from the absence of one. It is a property of this questionnaire, whose items all move together against whatever you correlate them with.
Which leaves a smaller, sturdier claim: there is no symptom here to point at. If you want to argue that screens damage teenagers in some particular way, this dataset cannot show you where — and it cannot show you where for sleep either. That is a limit of the instrument, and anyone claiming a mechanism from data like this runs into it.
advancethe claim I had to withdraw here
This is the second claim in this project that an audit demoted, and I would rather leave the trail visible than quietly restate it.
what I wrote
That the flat loading was “the signature of a general response tendency rather than of a mechanism”. That reads flatness as informative about screen time specifically.
the test
Run the identical per-item breakdown for all five drivers in the file. Dispersion relative to size: screen time 0.218, sleep quality 0.187, hours slept 0.215, weekend midsleep 0.366, social jetlag 0.581. Share of total loading carried by the five strongest items: screen time 30%, sleep quality 28.5%.
so
Sleep is at least as flat as screens on both measures. The original inference does not follow, and the piece now claims only that no symptom can be located — which is true, checkable, and about the questionnaire rather than about phones.
what survived
One genuine pattern did come out of it: item 19 is the strongest correlate of every sleep measure in the file — quality, hours, weekend midsleep, and social jetlag. The source ships the items unlabelled, so I will not guess what it asks.
advanceordinal data, overlapping intervals, and a subgroup check
pearson?
The screen index is ordinal and the items run 0–3, so Pearson is defensible but not obviously right. Spearman agrees throughout: screen × total is +0.1455 on ranks against +0.1344 on values. Nothing in this piece turns on the choice.
is it really flat?
Per-item correlations run +0.0374 to +0.1133, mean +0.0799, sd 0.0174. Every item carries a 95% interval of roughly ±0.028 at this sample size, so almost all 21 overlap one another. The honest statement is not “all items are equal” but “these data cannot distinguish them” — which is the same obstacle anyone claiming a specific mechanism would face.
21 tests
No multiple-comparison correction is applied, and none is needed for the argument: I am not claiming any single item is significant. I am reporting a failure to find a standout, and corrections make that failure more likely, not less.
the drop series
Removing items shortens the questionnaire, so the rebuilt totals are not on one common scale and the correlations are not strictly comparable. The direction is still informative; the exact decline is not. Stated in the pipeline as drop_series_caveat.
subgroups
Split by sex, the association is essentially identical — boys +0.1376 (n = 2,446), girls +0.1409 (n = 2,364). A null result, reported because it was run.
the bridge · findings that force decisions
Change a number in the left column and the fable on the right has to be rewritten.
finding
The same 4,810 teenagers yield risk ratios from 1.15× to 2.97× across 130 combinations of two thresholds that are never reported.
decision
The antagonist has to be the line, not the device. A story about a villainous object would contradict the data; a story about a decision is what the data actually shows.
finding
Per-item loading is flat: mean +0.0799, sd 0.0174, with 17 of 21 items inside a band of 0.0360.
decision
No single symptom may be named. Whoever measures in the story has to hope for one clear answer and fail to find it — because that is what happened here.
finding
Sleep quality explains 8× the variance screens do, and controlling hours slept absorbs 46% of the screen effect.
decision
The real correlate must be present in the story and ignored by everyone in it — said out loud once, late, and not written down.
the story
The Line on the Wall
A fable. It contains no figures at all — the numbers have had their say already, and what is left of them here is only their shape.
There was a town at the foot of a hill where, one spring, the children began to look tired.
Not ill. Tired. They came down to breakfast slowly and went up to bed late and answered questions a beat after they were asked. It was the kind of thing that is obvious to everybody and provable by nobody, and so the town argued about it for a season without getting anywhere.
What the council wanted was a number. How many of our children are unwell? Because you cannot put they seem tired in a ledger, and you cannot ask the province for help with a feeling.
So they sent for a physician, and the physician came with a list of twenty-one questions.
⁂
She asked every child all twenty-one, and wrote the answers down, and added them up, and each child came away with a tally. A small tally meant a child who answered the way children usually answer. A large one meant a child who did not.
Then the council asked her for the number, and she said she could not give them one yet, because she had tallies and they had asked for a count.
They did not follow, so she explained. There is no place inside a child where well becomes unwell. There is only the tally, which runs smoothly from the smallest to the largest with no gap anywhere along it. If you want a count, someone has to draw a line across the tally and say: above this, unwell. And that someone will be me, and I will be guessing.
Draw it, said the council.
So she went to the wall of the meeting hall, where the tallies had been chalked up from smallest to largest, and she drew a line across them, and above the line there were some children and below it there were many, and the council wrote down the number of the ones above and sent it to the province, and that was the count.
⁂
By then the town had already decided what the cause was.
A peddler had come through two years before selling small glass panes that lit up, and now every child in the town had one, and every parent could tell you exactly how late their child had been holding it. The panes were new. The tiredness was new. The council spent four meetings on the panes.
Prove it, they told the physician. Show us the children with the panes are the ones above the line.
And she said: I can show you that, and I would like you to watch how I do it.
She went back to the wall. She rubbed out her line and drew a new one, lower, and said: here, nearly every child in the town is above the line, and the ones with the panes are only slightly more likely to be among them, and you would conclude that the panes hardly matter. Then she drew it higher, near the top, and said: here, very few children are above the line, but those few hold the panes far more often than the rest, and you would conclude that the panes are a catastrophe.
She kept moving it. Each time she moved it she read out a sentence, and each sentence was true, and no two sentences agreed about how much the panes mattered.
I can give you any answer you would like, she said, and I will not have lied to you once. Tell me which sentence you want and I will tell you where to put the line. What I cannot do is tell you where the line belongs, because it does not belong anywhere. It only has to be somewhere.
⁂
One of the councillors — the youngest, who had been quiet — asked a better question. She asked: of your twenty-one questions, which one do the pane-holders answer worst?
The physician said she had hoped very much that there would be an answer to that.
Because one question would have been a mechanism. One question answered badly and the rest answered normally would have meant the panes were doing something in particular, and something in particular can be understood, and sometimes undone. She had gone through the twenty-one looking for it.
What she found instead was that the pane-holders answered every question a little worse. Not one of them badly. All of them faintly, and by almost exactly the same faint amount, as though the panes did not cause any particular unhappiness but only made a child slightly readier to say yes to a long list of questions late in the evening.
You could take away her worst question and ask the other twenty. The pattern held. You could take away ten. It held. There was nothing to remove, because there was nothing in any one place to begin with.
⁂
It was late, and the meeting had gone on, and the physician said one more thing before she packed up her list.
She said: of everything I measured on your children, the panes moved the tally least. What moved it most, by a long way, was how they slept. And they are not sleeping. Some of that is the panes. A good deal of it is the new road, and the mill that runs at night, and the fact that half of them are up before it is light to work the terraces.
Nobody wrote that down. There was no line for it, and no peddler to blame for it, and the road had been very expensive.
The count went to the province, with the line where she had first drawn it, because a line has to be somewhere. The panes were forbidden by the following spring, and then permitted again the spring after, on the grounds that nothing much had changed.
The children went on looking tired. The mill went on running. And in all of it, from the first meeting to the last, nobody thought to ask them what time they had gone to bed.
Source.Screen Time vs Mental Health (ML-ready), Kaggle — kylefengkfeng209/screen-time-vs-mental-health-ml-ready. Retrieved 2026-07-30. 4,810 teenagers (2,446 boys, 2,364 girls), 21 questionnaire items. Pipeline: , committed data: .
Limits. This is cross-sectional: nothing here can tell you which way any arrow points, and a teenager who is already unhappy may reach for a screen rather than the reverse. The partial correlations are not a causal adjustment; they show how much of a weak association travels with sleep, not that sleep is the cause. The questionnaire items ship unlabelled, so they are reported by number and never interpreted as specific symptoms. Every figure set in mono is traceable to the committed data file; the fable is fiction, and deliberately contains no figures at all.
evidence · which ratios survive an interval
Every cell of the grid, at the source file’s own heavy-user threshold of 3, with a 95% interval from the Katz log method. Rows in orange are the ones whose interval includes 1 — arithmetically correct, and indistinguishable from no difference at all.
cutoff
ratio
95% interval
heavy flagged
others flagged
≥ 5
1.21×
1.14× – 1.28×
1,579
936
≥ 6
1.24×
1.16× – 1.32×
1,390
801
≥ 7
1.27×
1.18× – 1.37×
1,227
691
≥ 8
1.29×
1.19× – 1.40×
1,069
592
≥ 9
1.31×
1.20× – 1.44×
944
514
≥ 10
1.33×
1.21× – 1.47×
827
444
≥ 11
1.33×
1.19× – 1.49×
726
390
≥ 12
1.38×
1.23× – 1.55×
654
339
≥ 13
1.40×
1.23× – 1.59×
585
299
≥ 14
1.43×
1.24× – 1.63×
524
263
≥ 15
1.50×
1.29× – 1.74×
460
220
≥ 16
1.45×
1.24× – 1.70×
414
204
≥ 17
1.50×
1.27× – 1.79×
364
173
≥ 18
1.54×
1.28× – 1.85×
322
150
≥ 19
1.58×
1.30× – 1.91×
302
137
≥ 20
1.59×
1.30× – 1.96×
272
122
≥ 21
1.70×
1.35× – 2.14×
233
98
≥ 22
1.93×
1.50× – 2.48×
213
79
≥ 23
1.95×
1.49× – 2.55×
188
69
≥ 24
2.00×
1.50× – 2.67×
168
60
≥ 25
2.16×
1.58× – 2.95×
154
51
≥ 26
2.45×
1.73× – 3.47×
137
40
≥ 27
2.32×
1.61× – 3.34×
120
37
≥ 28
2.35×
1.59× – 3.47×
105
32
≥ 29
2.60×
1.70× – 3.96×
98
27
≥ 30
2.68×
1.72× – 4.19×
90
24
cells in the whole grid
130
interval clear of 1
92 (70.8%)
interval includes 1
38 (29.2%)
ratio range
1.147× – 2.967×
The weak cells cluster at the extremes, where the flagged group shrinks to a few dozen teenagers. That is exactly where the most quotable ratios live.