---
title: "Claim Ledger: The Average Is Nobody's Result"
description: "Of 255 studies on AI-assisted colonoscopy, 21 split the result by who held the scope. They disagree."
author: "Harry Floyd"
publication: "The Durability Curve"
canonical: "https://durabilitycurve.com/claims/the-average-is-nobodys-result/"
essay: "https://durabilitycurve.com/blog/the-average-is-nobodys-result/"
substack: "https://harryfloyd.substack.com/p/the-average-is-nobodys-result"
published: "2026-08-04"
last_verified: "2026-08-02"
law: "Law IV"
claims: 35
struck: 0
---

# Claim Ledger: The Average Is Nobody's Result

*Of 255 studies on AI-assisted colonoscopy, 21 split the result by who held the scope. They disagree.*

Essay: https://durabilitycurve.com/blog/the-average-is-nobodys-result/  
Ledger (canonical, cite this): https://durabilitycurve.com/claims/the-average-is-nobodys-result/  
Published 2026-08-04 · last verified 2026-08-02

## What the essay claims

Walking in: "The trial says the assistant improves detection by eight points. That is the effect. My people will get roughly that."

Walking out: "That is an average over operators the tool may be affecting in opposite directions, and almost nobody checks. When they do check they disagree about who benefits — in randomised trials, on the same endpoint. The average is a property of their operator mix, not of the tool, so it was never going to transfer to mine. I can check my own in an afternoon."

The load-bearing move: The average is not a weak estimate of a single effect. It is a mixture of effects that can have opposite signs, weighted by a staffing decision. Povyakalo 2013 is the existence proof; the count is the evidence that the field mostly does not look; the eight that print numbers are the evidence that even looking is not enough unless the numbers are published.

## The claim ladder

Strongest first. Each rung names whose behaviour it measures and what it does not reach.

| # | Claim | Whose behaviour it measures | What it does not reach |
|---|---|---|---|
| 1 | Of 255 abstracts on AI-assisted colonoscopy, 21 report the effect split by an operator property and 234 report an average and nothing else | The reporting choices of one literature, at the abstract layer, as of 2026-08-02 | These are deposited abstracts. A study can split in its full text and never say so here. The count measures what is visible at the layer the field summarises itself in. It is also a floor: three patterns each found what the others missed |
| 2 | The 21 contradict each other about who benefits. 7 find the larger effect in weaker operators, 4 in stronger, 3 no interaction, 3 nothing significant after stratifying, 1 divergence over time. Both directions appear in randomised trials on the same endpoint | Published subgroup results in one procedure | Does not adjudicate. Several are underpowered after stratification, several are post-hoc, and the piece must not pick a winner. The spread is the finding |
| 3 | Only 8 of the 21 print per-stratum estimates. None reports a test of the interaction. The rest report which subgroup reached significance, which is not a test of a difference | The reporting and inference choices of the studies that did look | Does not establish the interaction is real in any of them. It establishes that the published evidence cannot tell you |
| 4 | The standard has no field for it. CONSORT-AI 5(iv) asks what expertise users needed; the guidance asks for differences across population subgroups; no item asks for the result across operator subgroups | The reporting guideline, verified item by item at PMC7598943 | Scoped to CONSORT-AI. DECIDE-AI's items were not read and no claim is made about guidelines in general |
| 5 | Povyakalo 2013 is the existence proof. Reanalysing a CAD mammography study whose average effect was null: +0.016 sensitivity (95% CI 0.003–0.028) for the 44 least-discriminating readers on 45 easier cancers, −0.145 (95% CI 0.034–0.257) for the 6 most-discriminating on 15 harder ones | 50 readers, 180 mammograms, one reanalysis | Exploratory and post-hoc by the authors' own description, with thresholds derived from the same regression and six readers in the key cell. It proves an average can hide opposite signs. It is not a law, and colonoscopy does not replicate it |
| 6 | The average is a property of your operator mix, not of the tool. Change who is on shift and the measured effect changes with no change to the tool | Nothing. An arithmetic consequence of rung 2 | Does not quantify how much any published estimate would move. A reason the number does not transfer, not a correction to it |

Scope limits travel in the same sentence as the number, in the body, never in a footnote.

## What would make it wrong

Any of these lands and the piece is wrong, not adjustable.

Full-text checking of a random sample of the 234 finds that most do split in the full text. Highest-probability failure by a distance.

A fourth search pattern finds a large new tranche (>8) all three of mine missed. The floor is then too low to support "rarely".

An independent adjudication of the same 255 abstracts returns materially more than 21 under the stated definition. The count is then an artefact of my reading, not the literature.

A reporting-guideline item for operator-stratified results turns out to exist in CONSORT-AI. Rung 4 fails outright. (Verified against the item list; the residual risk is a later update.)

Someone has already published this count. The prior-art sweep found the ecological version (39216648, 38272274) and the mechanism (Povyakalo) but not the count. If the count exists, the object is not an object.

## The evidence, row by row

Status: VERIFIED = primary source opened and the quoted words read off it by a checker who did not write the essay · EXECUTED = a first-party run, the claim is what it printed · REPORTED = carried from a source not opened in full.

### 1. Of 255 abstracts with text in the corpus, 21 report the assistant's effect split by a property of the operator and 234 report an average and nothing else

- Status: VERIFIED (read 2026-08-02)
- Quote: "HEADLINE: 21/255 studies report the assistant's effect split by a property of the OPERATOR. = 8.2% 234 report an average over operators and nothing else."
- Source: first-party run by The Durability Curve, unassisted_baseline_audit.py stdout, operator-split headline block

### 2. Among primary records by PubMed's own publication type the figure is 21 of 157

- Status: VERIFIED (read 2026-08-02)
- Quote: "Among primary records only: 21/157 = 13.4%"
- Source: first-party run by The Durability Curve, same run, headline block

### 3. Of the 21, only 8 print per-stratum effect estimates for every stratum; 3 print one stratum and a bare null; 10 report only a direction or a significance verdict

- Status: VERIFIED (read 2026-08-02)
- Quote: "SPLIT_FULL 8 per-stratum estimates for every stratum · SPLIT_PARTIAL 3 one stratum estimated, the other a bare null · SPLIT_DIRECTION 10 a direction or a p-verdict, nothing poolable"
- Source: first-party run by The Durability Curve, same run, verdict block

### 4. The splits disagree: 7 find the larger effect in weaker operators, 4 in stronger, 3 no interaction, 3 nothing significant after stratifying, 1 divergence

- Status: VERIFIED (read 2026-08-02)
- Quote: "LOWER 7 ... HIGHER 4 ... UNIFORM 3 ... NONE 3 ... DIVERGES 1"
- Source: first-party run by The Durability Curve, same run, direction block

### 5. Three deliberately dissimilar detectors were run; the obvious one reached 15 records of which 7 are real splits, missing 14

- Status: VERIFIED (read 2026-08-02)
- Quote: "pattern A: reached 15 of which real splits 7 precision 0.47 coverage 0.33 of the 21 found" · "Pattern A alone, unread, would have reported 15 records of which 7 are real, missing 14."
- Source: first-party run by The Durability Curve, same run, pattern-performance block

### 6. Pattern B reached 39 records at precision 0.36 and coverage 0.67 of the found set; pattern C reached 13 at precision 0.69 and coverage 0.43

- Status: VERIFIED (read 2026-08-02)
- Quote: "pattern B: reached 39 of which real splits 14 precision 0.36 coverage 0.67 of the 21 found" · "pattern C: reached 13 of which real splits 9 precision 0.69 coverage 0.43 of the 21 found" · "'coverage' is the share of the found positives a pattern reached, NOT recall"
- Source: first-party run by The Durability Curve, same run, same block

### 7. Every record any pattern reached was hand-read, and the script refuses to print the headline while one is unread

- Status: VERIFIED (read 2026-08-02)
- Quote: "HEADLINE: UNAVAILABLE — 20 record(s) reached by a pattern are unread." (the refusal branch, observed firing before adjudication was complete)
- Source: first-party run by The Durability Curve, same script, headline guard

### 8. The PubMed query returns 259 records, 255 with abstracts

- Status: VERIFIED (read 2026-08-02)
- Quote: "PubMed hits: 259 (retmax 400, fetched 259)" · "classified : 255 (no abstract: 4)"
- Source: first-party run by The Durability Curve, same run, header

### 9. In a seeded 20-record sample of the primary average-only class, 9 had open full text and none reported an operator split in its results

- Status: VERIFIED (read 2026-08-02)
- Quote: Sample drawn with random.Random(20260802) from 136 eligible records; PMC full text retrieved for 9; operator-split patterns returned 0 result-section hits in 8, and in the ninth the only hit was future-tense text in a trial protocol
- Source: first-party full-text check by The Durability Curve, seeded sample, scratchpad sample20.json; PMCIDs 10532435, 10381252, 11528738, 11512038, 9796278, 13054503, 12627440, 10011797, 12618418

### 10. Eleven of the twenty sampled records are not in PMC and their full texts were not read

- Status: VERIFIED (read 2026-08-02)
- Quote: NCBI ID converter returned "Identifier not found in PMC" for 11 of the 20 sampled PMIDs
- Source: first-party full-text check by The Durability Curve, same check, ID-converter response

### 11. Reanalysing a CAD mammography study, use of CAD was associated with a 0.016 increase in sensitivity (95% CI 0.003-0.028) for the 44 least discriminating radiologists on 45 relatively easy cancers

- Status: VERIFIED (read 2026-08-02)
- Quote: "Use of CAD was associated with a 0.016 increase in sensitivity (95% confidence interval [CI], 0.003-0.028) for the 44 least discriminating radiologists for 45 relatively easy, mostly CAD-detected cancers."
- Source: Povyakalo, Alberdi, Strigini & Ayton, Med Decis Making 2013, deposited abstract, Results
- URL: https://pubmed.ncbi.nlm.nih.gov/23300205/

### 12. For the 6 most discriminating radiologists, sensitivity decreased by 0.145 (95% CI 0.034-0.257) on the 15 relatively difficult cancers

- Status: VERIFIED (read 2026-08-02)
- Quote: "However, for the 6 most discriminating radiologists, with CAD, sensitivity decreased by 0.145 (95% CI, 0.034-0.257) for the 15 relatively difficult cancers."
- Source: Povyakalo et al., Med Decis Making 2013, deposited abstract, Results
- URL: https://pubmed.ncbi.nlm.nih.gov/23300205/

### 13. The original study detected no significant average effect, and the reanalysis found CAD helped the less discriminating readers and hindered the more discriminating ones

- Status: VERIFIED (read 2026-08-02)
- Quote: "It indicates that, despite the original study detecting no significant average effect, CAD helped the less discriminating readers but hindered the more discriminating readers."
- Source: Povyakalo et al., Med Decis Making 2013, deposited abstract, Conclusions
- URL: https://pubmed.ncbi.nlm.nih.gov/23300205/

### 14. That reanalysis covered 50 readers interpreting 180 mammograms both with and without computer support

- Status: VERIFIED (read 2026-08-02)
- Quote: "reanalyzing data from a published study where 50 professionals (\"readers\") interpreted 180 mammograms, both with and without computer support"
- Source: Povyakalo et al., Med Decis Making 2013, deposited abstract, Background
- URL: https://pubmed.ncbi.nlm.nih.gov/23300205/

### 15. The reader and case strata were derived after the fact from the same regression, and the authors call the method exploratory

- Status: VERIFIED (read 2026-08-02)
- Quote: "Using regression estimates, we obtained thresholds for classifying a posteriori the cases (by difficulty) and the readers (by discriminating ability)." · "Our exploratory analysis method reveals unexpected effects."
- Source: Povyakalo et al., Med Decis Making 2013, deposited abstract, Method and Conclusions
- URL: https://pubmed.ncbi.nlm.nih.gov/23300205/

### 16. The authors said such differential effects should be assessed when evaluating CAD and similar warning systems

- Status: VERIFIED (read 2026-08-02)
- Quote: "Such differential effects, although subtle, may be clinically significant and important for improving both computer algorithms and protocols for their use. They should be assessed when evaluating CAD and similar warning systems."
- Source: Povyakalo et al., Med Decis Making 2013, deposited abstract, Conclusions
- URL: https://pubmed.ncbi.nlm.nih.gov/23300205/

### 17. In a population-based randomised trial of 4,824 surveillance colonoscopies, CADe increased ADR among low-performing endoscopists (45.5% vs 52.1%; aRR 1.15, 95% CI 1.01-1.30) but not among high-performing ones (65.9% vs 63.5%; aRR 0.96, 95% CI 0.88-1.05)

- Status: VERIFIED (read 2026-08-02)
- Quote: "Of 5,309 randomized surveillance colonoscopies, 4,824 surveillance colonoscopies were included in the final analysis" · "CADe increased ADR among low-performing (ADR<54.5%) endoscopists (45.5% vs 52.1%; aRR 1.15 [95% CI 1.01-1.30]), but not among high-performing endoscopists (65.9% vs 63.5%; aRR 0.96 [95% CI 0.88-1.05])."
- Source: Endoscopy 2026, Galician screening programme, PMID 42365851, deposited abstract, Results
- URL: https://pubmed.ncbi.nlm.nih.gov/42365851/

### 18. That subgroup analysis was prespecified

- Status: VERIFIED (read 2026-08-02)
- Quote: "Prespecified subgroup analyses evaluated endoscopist baseline performance (median ADR in the standard arm)."
- Source: Endoscopy 2026, PMID 42365851, deposited abstract, Methods
- URL: https://pubmed.ncbi.nlm.nih.gov/42365851/

### 19. In a 3,059-patient multicentre randomised trial the gain was larger in experts: ADR of experts 42.3% vs 32.8% (P < .001) and of nonexperts 37.5% vs 32.1% (P = .023)

- Status: VERIFIED (read 2026-08-02)
- Quote: "From November 2019 to August 2021, 3059 subjects were randomized to AI-assisted colonoscopy (n = 1519) and conventional colonoscopy (n = 1540)." · "ADR of expert (42.3% vs 32.8%; P < .001) and nonexpert endoscopists (37.5% vs 32.1%; P = .023)... were all significantly higher in the AI-assisted colonoscopy."
- Source: Clin Gastroenterol Hepatol 2023, PMID 35863686, deposited abstract, Results
- URL: https://pubmed.ncbi.nlm.nih.gov/35863686/

### 20. A 2026 study expected the greatest increase among low-volume and junior endoscopists

- Status: VERIFIED (read 2026-08-02)
- Quote: "The greatest increase was expected among low-volume and junior endoscopists."
- Source: Dis Colon Rectum 2026, PMID 41919624, deposited abstract, Objective
- URL: https://pubmed.ncbi.nlm.nih.gov/41919624/

### 21. It reported that seniors rose from 51.7% to 59.6% (p = 0.02) and juniors from 56.2% to 62.9% (p = 0.12), and concluded the tool correlated with increased detection for experienced endoscopists

- Status: VERIFIED (read 2026-08-02)
- Quote: "Of 12 senior and 12 junior endoscopists, the seniors had a statistically significant increase ( p = 0.02) from 51.7% to 59.6%, whereas juniors did not (56.2% to 62.9%, p = 0.12)." · "Computer-aided detection in colonoscopy correlated with an increased adenoma detection rate for experienced endoscopists."
- Source: Dis Colon Rectum 2026, PMID 41919624, deposited abstract, Results and Conclusions
- URL: https://pubmed.ncbi.nlm.nih.gov/41919624/

### 22. Those two reported increases are 7.9 points and 6.7 points, a gap of 1.2 points, and only the larger crossed p<0.05

- Status: VERIFIED (read 2026-08-02)
- Quote: 59.6 − 51.7 = 7.9 · 62.9 − 56.2 = 6.7 · 7.9 − 6.7 = 1.2. Subtraction only; no model, no reanalysis, and no claim that either estimate is wrong
- Source: first-party arithmetic on row 21's quoted rates by The Durability Curve, derived from row 21, recomputed by hand 2026-08-02

### 23. The same study reported the same pattern on two further operator axes: high-volume 51.3% to 59.4% (p = 0.01) against low-volume 56.5% to 63.1% (p = 0.13), and surgeons 49% to 62% (p = 0.02) against gastroenterologists 56.9% to 60.8% (p = 0.15)

- Status: VERIFIED (read 2026-08-02)
- Quote: "Colorectal surgeons had a statistically significant increase in adenoma detection rate from 49% to 62% ( p = 0.02), but gastroenterologists did not (56.9% to 60.8%, p = 0.15)." · "When comparing 12 high-volume and 12 low-volume endoscopists, the high-volume group had a statistically significant increase (51.3% to 59.4%, p = 0.01), whereas the low-volume group did not (56.5% to 63.1%, p = 0.13)."
- Source: Dis Colon Rectum 2026, PMID 41919624, deposited abstract, Results
- URL: https://pubmed.ncbi.nlm.nih.gov/41919624/

### 24. Pooling two randomised trials, CADe raised ADR (RR 1.29, 95% CI 1.16 to 1.42) while examiner experience did not (RR 1.02, 95% CI 0.89 to 1.16), and the authors concluded experience plays a minor role

- Status: VERIFIED (read 2026-08-02)
- Quote: "use of CADe (RR 1.29; 95% CI: 1.16 to 1.42) and colonoscopy indication, but not the level of examiner experience (RR 1.02; 95% CI: 0.89 to 1.16) were associated with ADR differences in a multivariate analysis." · "Experience appears to play a minor role as determining factor for ADR."
- Source: Gut 2022, AID-1 + AID-2, PMID 34187845, deposited abstract, Results and Conclusion
- URL: https://pubmed.ncbi.nlm.nih.gov/34187845/

### 25. The title of the paper naming the inference error is "The Difference Between 'Significant' and 'Not Significant' is not Itself Statistically Significant"

- Status: VERIFIED (read 2026-08-02)
- Quote: "The Difference Between \"Significant\" and \"Not Significant\" is not Itself Statistically Significant"
- Source: Gelman & Stern, The American Statistician 60(4):328-331, 2006, author's deposited PDF, title, page 1
- URL: http://www.stat.columbia.edu/~gelman/research/published/signif4.pdf

### 26. Their point is not that thresholds are arbitrary but that large changes in significance can correspond to small, non-significant changes in the underlying quantities

- Status: VERIFIED (read 2026-08-02)
- Quote: "Rather, we are pointing out that even large changes in significance levels can correspond to small, nonsignificant changes in the underlying quantities."
- Source: Gelman & Stern, 2006, deposited PDF, abstract, page 1
- URL: http://www.stat.columbia.edu/~gelman/research/published/signif4.pdf

### 27. The reporting standard for AI trials asks investigators to state what level of expertise was required of users

- Status: VERIFIED (read 2026-08-02)
- Quote: "CONSORT-AI 5 (iv) Extension Specify whether there was human–AI interaction in the handling of the input data, and what level of expertise was required of users."
- Source: CONSORT-AI extension, Nat Med 26:1364-1374, 2020, checklist table, item 5 (iv), PMC full text
- URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC7598943/

### 28. Its guidance encourages exploring differences in performance across population subgroups, and no item asks for the result across operator subgroups

- Status: VERIFIED (read 2026-08-02)
- Quote: "Beyond this, investigators should also be encouraged to explore differences in performance and error rates across population subgroups."
- Source: CONSORT-AI extension, 2020, discussion of item 19 extension, PMC full text; the full 14-item extension list was read and contains no operator-stratified reporting item
- URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC7598943/

### 29. The COLO-DETECT protocol pre-registered a subgroup analysis by colonoscopist type

- Status: VERIFIED (read 2026-08-02)
- Quote: "Subgroup analyses will be conducted on colonoscopist type (i.e., non-BCSP accredited vs. BCSP accredited) and indication for colonoscopy (screening vs. symptomatic)."
- Source: COLO-DETECT trial protocol, Colorectal Dis 2022, PMID 35680613, protocol full text, statistical analysis section, PMC9796278
- URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC9796278/

### 30. The same protocol said a range of colonoscopist experience was anticipated and desirable, and specified how it would be measured

- Status: VERIFIED (read 2026-08-02)
- Quote: "A range of experience amongst participating colonoscopists is anticipated and desirable; to facilitate subgroup analysis by colonoscopist experience, it will be assessed by BCSP accreditation status, lifetime procedure numbers and yearly procedure numbers for the last 3 years"
- Source: COLO-DETECT protocol, 2022, protocol full text, PMC9796278
- URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC9796278/

### 31. The COLO-DETECT results reported one average across 2032 participants at 12 NHS hospitals: adenomas detected in 56·6% versus 48·4%, adjusted odds ratio 1·47 (95% CI 1·21-1·78)

- Status: VERIFIED (read 2026-08-02)
- Quote: "We did a multicentre, open-label, parallel-arm, pragmatic randomised controlled trial in 12 National Health Service (NHS) hospitals (ten NHS Trusts) in England" · "Between March 29, 2021, and April 6, 2023, 2032 participants (1132 [55·7%] male, 900 [44·3%] female; mean age 62·4 years [SD 10·8]) were recruited and randomly assigned" · "Adenomas were detected in 555 (56·6%) of 980 participants in the CADe-assisted colonoscopy group versus 477 (48·4%) of 986 in the standard colonoscopy group, representing a proportion difference of 8·3% (95% CI 3·9-12·7; adjusted odds ratio 1·47 [95% CI 1·21-1·78], p<0·0001)."
- Source: COLO-DETECT, Lancet Gastroenterol Hepatol 2024, PMID 39153491, deposited abstract, Methods and Findings
- URL: https://pubmed.ncbi.nlm.nih.gov/39153491/

### 32. Its randomisation was stratified by age group, sex, indication and NHS Trust, and its abstract names no colonoscopist subgroup

- Status: REPORTED (read 2026-08-02)
- Quote: Stratification verbatim: "with stratification by age group, sex, colonoscopy indication (screening or symptomatic), and NHS Trust". The absence of a colonoscopist subgroup is an absence in the deposited abstract only — the paper is not in PMC and the full text was not read.
- Source: COLO-DETECT, Lancet Gastro Hep 2024, PMID 39153491, deposited abstract, Methods
- URL: https://pubmed.ncbi.nlm.nih.gov/39153491/

### 33. The field's current answer is a study-level one: pooling twenty-eight randomised trials and 23,861 participants, an expert-only subgroup showed a similar effect size (RR 1.19, 95% CI 1.11-1.27) and the review concluded the benefit holds irrespective of endoscopist experience

- Status: VERIFIED (read 2026-08-02)
- Source: Gastrointest Endosc 2025 systematic review, PMID 39216648, deposited abstract, Results and Conclusions
- URL: https://pubmed.ncbi.nlm.nih.gov/39216648/

### 34. A second meta-analysis reached the same study-level conclusion across twenty-four randomised trials and 17,413 colonoscopies

- Status: VERIFIED (read 2026-08-02)
- Quote: "Twenty-four RCTs involving 17,413 colonoscopies (AI assisted: 8680; non-AI assisted: 8733) were included." · "Type of AI system used or endoscopist experience did not affect overall improvement in ADR."
- Source: Gastrointest Endosc 2024, PMID 38272274, deposited abstract, Results and Conclusions
- URL: https://pubmed.ncbi.nlm.nih.gov/38272274/

### 35. Bainbridge stated in 1983 that unused physical skills decay and a monitoring operator becomes an inexperienced one

- Status: VERIFIED (read 2026-08-01)
- Quote: "Unfortunately, physical skills deteriorate when they are not used, particularly the refinements of gain and timing. This means that a formerly experienced operator who has been monitoring an automated process may now be an inexperienced one."
- Source: Bainbridge, Automatica 19(6):775-779, 1983, §1.1.1 Manual control skills
- URL: https://ckrybus.com/static/papers/Bainbridge_1983_Automatica.pdf

## Cite

- A claim: "[claim text]" (Floyd, Harry, 2026, https://durabilitycurve.com/claims/the-average-is-nobodys-result/)
- The essay: Floyd, Harry (2026). The Average Is Nobody's Result. The Durability Curve. https://durabilitycurve.com/blog/the-average-is-nobodys-result/
- This ledger: Floyd, Harry (2026). Claim Ledger: The Average Is Nobody's Result [structured claims with sources]. The Durability Curve. https://durabilitycurve.com/claims/the-average-is-nobodys-result/

Quote with attribution and a link. Say if you changed the wording. Not licensed for model training.
