# Galleri screening policy comparison: methods

Prepared September 15, 2026. This is an illustrative model, not a validated clinical forecast or a head-to-head trial result.

## Current article baseline (September 23 revision)

The main article and figure 3 now use **colonoscopy every ten years**, with the explicitly hypothetical **20% additional colorectal-incidence reduction** described under “Central colonoscopy policy.” The headline ages 65–74 compare population-weighted outcomes: 2.5× detections / 5.5% more false-positive events for women, and 5.8× / 28% for men. Prevention is counted separately.

Current files: `galleri-colonoscopy-policy-model.json` / `.csv` (five-year ages 50–74 plus nine scenarios) and `galleri-policy-age-bands.json` / `.csv` (30–49 illustrative, 50–64, 65–74). The core FIT files and their FDA scenarios remain available as explicitly labeled historical alternatives; the following original FIT sections document those files, not the main chart's current baseline.

## Original FIT model: question and units

Compare selected USPSTF grade A/B cancer-screening options with those same screens plus annual Galleri and with annual Galleri alone in 1,000 people, separately for registry female and male populations and five-year age bands from 50–54 through 70–74.

Outputs describe a typical year in an ongoing, fully adhered-to program, with screening schedules staggered across people. They are not the yield of a first visit at which every test is due. Cancer counts are annual incident cancers caught or not caught by the modeled screens. False positives count test results without the target cancer, not unique people, recalls resolved without imaging, biopsies, or harms. Tests with precancer targets can yield useful findings even when there is no cancer.

The baseline selects annual FIT, biennial mammography when applicable, five-year cervical cotesting when applicable, and annual low-dose lung CT for eligible people. It does not apply every alternative screening test to each person. Optional grade C PSA screening is excluded. The comparison stops before age 75, when the set of recommended tests starts to change further. It does not extrapolate Galleri evidence below age 50.

The Galleri-alone scenario includes only Galleri detections and false-positive tests, using the same performance assumptions as the combined policy. All “extra” outcome fields are differences from the USPSTF baseline and can be negative for Galleri alone. This comparison does not model the prevention benefits of standard screening.

## Inputs: published estimates versus assumptions

| Input | Central value | Status / source |
|---|---:|---|
| All-cancer and site incidence | By five-year age band and sex | SEER observed 2019–2023 rates, all races |
| Mammography sensitivity / false-positive rate | 86.9% / 11.1% | Historical BCSC 2007–2013 benchmark |
| FIT sensitivity / false-positive rate | 74% / 6% | USPSTF evidence review; specified assay threshold |
| Lung CT sensitivity / false-positive rate | 78.6% / 5.3% | Retrospective Lung-RADS analysis of NLST, after baseline |
| Cervical cotest cancer sensitivity | 94.1% | Kaufman et al., 2020; cotests within 12 months before cancer |
| Cervical non-cancer positive rate | 7% | Assumption; varied to 5% and 10% |
| Fraction of female population with a cervix | 75% | Assumption used for test counts; not an adjustment to observed all-women cancer incidence |
| Lung population eligibility by age | BRFSS eligible count / Census population | 2022 data; same proportion applied to both sexes because joint age-by-sex eligibility is unavailable here |
| Fraction of lung cancers arising in eligible people | 60% | Assumption; varied to 40% and 80% |
| Galleri cancer-type sensitivities | CCGA values × 39.3/51.5 | Unvalidated transport assumption combining case-control and prospective overall estimates |
| Galleri false-positive rate | 0.4% | Full PATHFINDER 2 reported May 2026 |
| Detectable period before clinical presentation | One year for every cancer and modality | Assumption; varied to six months and two years |
| Adherence | 100% | Assumption, including follow-through sufficient to identify screen-detected cancers |
| Detection overlap within cancer type | Independent | Central assumption; also calculate mathematical bounds |

These are not matched-population test performance estimates. In particular, the cotesting study is retrospective laboratory evidence rather than a prospective, representative screening cohort. Its cancer sensitivity is not sensitivity for cervical precancer. The 7% non-cancer positive rate is deliberately labeled an assumption, not specificity measured in that paper. Lung performance comes from a retrospective application of Lung-RADS and need not describe every contemporary screening program. Age-specific test performance is held constant where the model lacks directly comparable estimates.

### Recommendations

- [Breast: biennial mammography at 40–74](https://www.uspreventiveservicestaskforce.org/uspstf/recommendation/breast-cancer-screening).
- [Colorectal: screening at 45–75](https://www.uspreventiveservicestaskforce.org/uspstf/recommendation/colorectal-cancer-screening); this model selects annual FIT.
- [Cervical: screening through 65](https://www.uspreventiveservicestaskforce.org/uspstf/recommendation/cervical-cancer-screening); this model selects five-year cotesting and assumes adequate prior screening permits stopping after 65. One fifth of the 65–69 band remains age-eligible, using equal single-year weights.
- [Lung: annual CT at 50–80 with at least 20 pack-years, currently smoking or quit within 15 years](https://www.uspreventiveservicestaskforce.org/uspstf/recommendation/lung-cancer-screening).
- [Prostate: grade C individual decision at 55–69](https://www.uspreventiveservicestaskforce.org/uspstf/recommendation/prostate-cancer-screening), excluded from the A/B baseline.

### Performance sources

- [BCSC screening mammography benchmarks, 2007–2013](https://www.bcsc-research.org/index.php/statistics/screening-performance-benchmarks-archive/screening-sens-spec-false-negative-archive).
- [USPSTF colorectal evidence and recommendations](https://www.uspreventiveservicestaskforce.org/uspstf/recommendation/colorectal-cancer-screening).
- [Pinsky et al., Performance of Lung-RADS in the National Lung Screening Trial](https://pmc.ncbi.nlm.nih.gov/articles/PMC4705835/): after-baseline estimates, not first-round estimates.
- [Kaufman et al., Contributions of Liquid-Based Cytology and HPV Testing in Cotesting, 2020](https://doi.org/10.1093/ajcp/aqaa074), Table 3: 1,123 of 1,193 cotests within 12 months of a cancer diagnosis positive by either component, 94.1%.
- [Klein et al., CCGA validation, 2021](https://doi.org/10.1016/j.annonc.2021.05.806); [complete cancer-class sensitivity table](https://www.galleri.com/hcp/galleri-test-performance).
- [Full PATHFINDER 2 results reported May 2026](https://grail.com/press-releases/grail-presents-pathfinder-2-results-of-more-than-35000-participants-showing-the-galleri-test-substantially-increased-cancer-detection-with-robust-performance-and-favorable-safety-at-2026-a/): all-cancer episode sensitivity 39.3%, specificity 99.6%. Episode sensitivity includes cancers diagnosed within 12 months and is not a directly measured probability of detecting a cancer throughout a fixed preclinical window.

### Incidence and eligibility

[SEER Explorer](https://seer.cancer.gov/statistics-network/explorer/application.html?site=1&data_type=1&graph_type=3&compareBy=sex) provides observed incidence per 100,000 person-years for 2019–2023. Divide by 100 to obtain cases per 1,000 person-years. These are not 2026 projected national case counts or age-adjusted rates. The all-site series includes reportable cancers such as in situ bladder cancer; it is not restricted to invasive malignancy in every site.

The repository pins raw JSON for all sites used in `inputs/seer-site-{site}.json`. Each was retrieved from this endpoint with its corresponding numeric site ID:

```text
https://seer.cancer.gov/statistics-network/explorer/source/content_writers/render_region_5.php?site={site}&data_type=1&graph_type=3&compareBy=sex&race=1&rate_type=1&stage=101&advopt_precision=1&advopt_display=2
```

`inputs/seer-formats.json` records variable labels. The response's `year_range=5` denotes 2019–2023 in that snapshot; age codes 211–215 denote 50–54 through 70–74; sex codes 3 and 2 denote female and male. Site 1 supplies the all-cancer total. The code rejects missing series and negative residual incidence. SEER and Census use different sex codes, handled explicitly.

[Prevalence of Lung Cancer Screening in the US, 2022](https://pmc.ncbi.nlm.nih.gov/articles/PMC10958241/), JAMA Network Open 2024;7:e243190, provides estimated eligible populations by age under the 2021 criteria:

| Age | Eligible population |
|---|---:|
| 50–54 | 2,063,840 |
| 55–59 | 2,639,255 |
| 60–64 | 3,338,111 |
| 65–69 | 2,440,591 |
| 70–74 | 1,946,129 |

The denominator sums 2022 single-year age populations from the [Census vintage 2023 national age/sex estimates](https://www2.census.gov/programs-surveys/popest/datasets/2020-2023/national/asrh/nc-est2023-agesex-res.csv), pinned as `inputs/census-2023-vintage-agesex.csv`. Eligibility is not actual uptake: the model assumes all eligible people complete screening. Survey-based counts and Census totals are approximate, so eligibility fractions are also varied by ±25%. A separate assumption is required for the proportion of lung cancer cases within the eligible population; the population eligibility fraction alone cannot supply it.

## Cancer-type mapping

The model uses nonoverlapping SEER sites, then assigns the remaining all-cancer incidence to a residual category. Mappings are approximate; for example, kidney/renal pelvis and soft-tissue categories do not exactly match the CCGA classes. It does not add broad lymphoma or leukemia totals on top of their component sites.

| SEER site ID and category | Assigned CCGA sensitivity before scaling |
|---|---:|
| 55 Breast | 30.5% |
| 47 Lung and bronchus | 74.8% |
| 66 Prostate | 11.2% |
| 20 Colon and rectum, including appendix | 82.0% |
| 53 Melanoma | 46.2% |
| 71 Bladder, invasive and in situ | 34.8% |
| 72 Kidney and renal pelvis | 18.2% |
| 86 Non-Hodgkin; 83 Hodgkin lymphoma | 56.3% each |
| 58 Corpus and uterus, NOS | 28.0% |
| 92 ALL; 93 CLL | 41.2% each |
| 96 AML; 97 CML | 20.0% each |
| 40 Pancreas | 83.7% |
| 3 Oral cavity/pharynx; 46 Larynx | 85.7% each |
| 80 Thyroid | 0% |
| 35 Liver and intrahepatic bile duct | 93.5% |
| 89 Myeloma | 72.3% |
| 18 Stomach | 66.7% |
| 17 Esophagus | 85.0% |
| 61 Ovary | 83.1% |
| 57 Cervix | 80.0% |
| 34 Anus, anal canal and anorectum | 81.8% |
| 51 Soft tissue including heart | 60.0% |
| 76 Brain and other nervous system | 0%, conservative assumption without a separate CCGA estimate |
| 38 Gallbladder | 70.6% |
| Residual: all sites minus listed sites | 50.8%, assigned CCGA “other” sensitivity |

The central scenario multiplies these sensitivities by 39.3/51.5. This does not calibrate the weighted result in every population to 39.3%, and it is not evidence that every cancer type suffers the same reduction in screening sensitivity. It is one transparent scenario. Alternatives retain unscaled CCGA sensitivities or assign a flat 39.3% to every cancer type. No zero is a claim of biological impossibility.

## Equations

Let `lambda[k]` be annual incidence of cancer type `k` per 1,000 people. Let `D` be the assumed preclinical detectable duration, `I` the screening interval, and `s` the test's sensitivity conditional on an opportunity to detect that cancer.

For a random phase of the screening schedule, set `q = floor(D/I)` and `r = D/I - q`. The modeled probability of catching the cancer before clinical presentation is:

```text
capture(s, I, D) = 1 - (1-s)^q × (1-r×s)
```

This explicitly assumes independent errors if there are repeated opportunities. With a one-year window and a two-year mammography interval, half of cancers have a screening opportunity, producing `0.5 × sensitivity` in this scenario. That is not a general rule that biennial programs have half the sensitivity of a test. A two-year window changes that result. Cancer-type and modality-specific detectable windows would be needed for a more realistic natural-history model.

Multiply lung capture by the assumed share of cancers in eligible people. Multiply cervical capture by the age-eligible share. Other cancers receive zero baseline capture unless targeted by a selected recommended screen.

For each type let `S[k]` and `G[k]` denote the resulting standard and Galleri capture probabilities. Central combined detection assumes independence within type:

```text
caught_standard = sum(lambda[k] × S[k])
caught_hybrid = sum(lambda[k] × (S[k] + G[k] - S[k]×G[k]))
missed = sum(lambda[k]) - caught
detected_percent = 100 × caught / sum(lambda[k])
```

The lower and upper possible hybrid detection probabilities for each cancer type, conditional on its marginal probabilities, are `max(S[k], G[k])` and `min(1, S[k]+G[k])`. Sum these weighted by incidence to obtain the overlap bounds in the downloadable results. These are mathematical overlap bounds, not confidence intervals or a full uncertainty analysis.

For each test, false-positive results are approximated by:

```text
false_positive_tests = (eligible_population_per_1000
                        - eligible_target_cancer_incidence × D)
                       × false_positive_rate / interval
```

This uses incidence times the explicitly assumed duration as a simplified preclinical pool. It does not measure actual screening prevalence or account for depletion after earlier detection. For Galleri, the target pool is all cancers. For other tests it is the targeted cancer. Summing expected test counts requires no assumption that different tests' false positives are independent. Those sums cannot be converted into the probability that a person has at least one false positive without additional information about within-person overlap.

## Sensitivity analyses and limitations

The downloadable JSON contains central results and ten original one-at-a-time alternatives, plus eight separate FDA-assay scenarios described below: six-month and two-year detectable windows; unscaled CCGA and flat 39.3% Galleri profiles; 40% and 80% lung-case eligibility; 5% and 10% cervical non-cancer positivity; and ±25% lung population eligibility. These are chosen scenarios, not empirically established lower and upper bounds. They do not cover every uncertain parameter or combinations of changes.

At age 65–69, changing the detectable window from six months to two years changes modeled hybrid detection from 28.7% to 75.6% in women and 20.0% to 51.2% in men. This uncertainty is much larger than the overlap bounds for the central scenario.

Additional limitations:

- Current SEER incidence reflects existing screening and prevention, not an unscreened counterfactual population. Holding it fixed across strategies misses prevention and changes in diagnosis timing.
- Incidence is a flow of new diagnoses. The fixed detectable window is a modeling assumption needed to relate it to screening opportunities; the data do not establish that window.
- No stage distribution, lead-time, overdiagnosis, competing mortality, survival, cancer deaths prevented, precancer treatment, cost, workup completion, or procedure harms are modeled.
- Galleri's reported episode sensitivity is not identical to the per-opportunity sensitivity used in this simplified model. Transporting case-control sensitivities into prospective screening, even after scaling, is unvalidated.
- Repeated-test errors and errors between modalities may be correlated. Cancer detection varies with stage and other characteristics even within the same type.
- Lung eligibility uses a common age-specific fraction for both sexes and an assumed share of eligible cancer cases. Individual smoking histories and health eligibility are not simulated.
- Cervix presence is simplified to a common fraction. Prior hysterectomy indications, screening history, mastectomy, family history, and other individualized eligibility are not simulated.
- Registry sex categories are proxies for screening populations, not a substitute for anatomy-based care.
- Age-band incidence and eligibility are held fixed; the model does not follow people aging into or out of screening. The age-65 cervical cutoff is approximated with equal weights.
- Counts refer to cancers, not necessarily distinct cancer patients. The model does not explicitly simulate multiple primary cancers in one person.
- Positive results have different meanings and consequences across modalities. A cancer-only false-positive count understates the benefit of some precancer findings and is not a count of unnecessary procedures.

## Reproduction

From the repository root:

```bash
python3 scripts/blog/galleri/policy_model.py
python3 -m unittest discover -s scripts/blog/galleri -p 'test_policy_model.py'
python3 scripts/blog/galleri/render-policy-chart.py
```

The renderer needs Matplotlib. The model and tests use only the Python standard library and pinned input files; no live network calls are required. `policy_model.py` also copies this document to the public downloads directory.

Outputs:

- `public/blog/galleri-screening-policy-model.csv`: all 30 central policy/group rows.
- `public/blog/galleri-screening-policy-model.json`: central parameters, modality counts, overlap bounds, and all alternative scenarios.
- `public/blog/galleri-uspstf-screening-by-age-sex.png`: pre-rendered comparison chart.
- `public/blog/galleri-screening-policy-methods.md`: this methods document.

Automated checks cover cancer-count conservation, overlap bounds, screening-opportunity calculations, zero/perfect added-test limits, false-positive increments, population denominators, and the cervical age cutoff. They verify implementation arithmetic, not clinical validity.

## External trial benchmarks (September 18, 2026)

Run `python3 scripts/blog/galleri/benchmark-trials.py` to regenerate `public/blog/galleri-trial-benchmarks.json`. This is a descriptive comparison; no model parameters were fitted to these results.

- **PATHFINDER 2:** Use the 32,007 performance-evaluable participants, not the 35,878 enrolled denominator. The 173 Galleri detections plus 31 A/B screening detections give 204/31 = 6.58 (described as approximately 6.5-fold in the source). Galleri yield is 5.41 per 1,000 participants; excluding 22 recurrent cancers gives 151 new primaries, or 4.72 per 1,000. The 287 positive results minus 173 confirmed cancers give 114 false-positive results, or 3.56 per 1,000 participants. This last denominator includes people with cancer and is therefore not the formal false-positive rate. Sources: [full fact sheet](https://grail.com/wp-content/uploads/2026/05/Pathfinder2_FactSheet_FINAL.pdf), [ASCO abstract](https://www.asco.org/abstracts-presentations/260912).
- **NHS-Galleri:** The reported intervention counts are 937 Galleri-detected cancers plus 236 from usual screening, compared with 290 screen-detected cancers in control. The count ratio is 1,173/290 = 4.04, and the difference is 883 across the three-round trial. We do not convert this into an annual rate without matched person-time. [Source, page 3](https://grail.com/wp-content/uploads/2026/05/NHS-GalleriASCO_Factsheet_FINAL.pdf).
- **Model:** Central age/sex scenarios yield ratios of 1.88–3.57 and incremental annual detection of 1.32–6.41 per 1,000. A comparable numerical order of magnitude does not validate the model. PATHFINDER's detection routes are not a randomized counterfactual; NHS screening practice, attendance, age distribution, and the initial versus repeat screening rounds differ from the modeled full-adherence annual program.

A proper calibration needs trial age/sex-specific incidence, screening eligibility and attendance, prior screening, detection routes, and follow-up matched to the model. The available aggregate summaries cannot determine whether the model's age-specific estimates are accurate or conservative. Its headline and chart are explicitly presented as scenarios.


## Screening-specific appendix

`galleri-uspstf-by-screening.csv` and `.json` break the central USPSTF policy into FIT, mammography, cervical cotesting, and eligible lung CT across all age/sex groups. Each denominator remains 1,000 people of the stated age and sex, not 1,000 tests or eligible participants. True positives equal target-site incidence times the modeled screening capture probability. Missed target cancers equal target-site incidence minus those true positives, including people outside eligibility and cancers arising between scheduled tests. They are not solely negative test results in tested people. False-positive tests use the same modality-specific calculation as the main model.

Caught cancers and false-positive tests sum to the USPSTF policy totals. Missed target cancers do not sum to all missed cancers: the main model also includes cancers outside these four sites. The article displays ages 65–69; downloads include all five age bands. Cervical results in ages 65–69 reflect eligibility only at age 65, with the model's equal single-year age weights and cervix-present assumption.


## FDA briefing sensitivity scenarios (September 21, 2026)

Source: [FDA executive summary for the September 23 meeting](https://www.fda.gov/media/194909/download). The proposed assay differs from MCED-V2 (pp. 8–9). The main chart and its central parameters remain unchanged. `fda_scenarios.py` adds eight named, reproducible scenarios to the main downloadable JSON.

For each study, one scenario replaces the CCGA-derived profile with new-primary cancer episode sensitivities from Table A5 (pp. 48–51) and the proposed assay's overall false-positive rate (PATHFINDER 2 0.15%; NHS 0.26%). New-primary counts avoid including recurrences in incident cancer sensitivity: e.g. PATHFINDER breast is 4/33, not Table A4's 14/53. Each profile maps FDA categories to the existing nonoverlapping SEER sites; grouped lymphoid, pancreatic/biliary, and other categories remain approximate matches. Where fewer than five new cancers were observed, or for unmatched/residual sites, use that study's overall new-primary sensitivity (90/268 or 259/820). This fallback is an assumption and does not imply an observed sensitivity for missing types. The JSON preserves counts and fallback flags. Unmatched sites and the residual default to the same fallback. No additional CCGA scaling is applied.

The other three scenarios for each study retain its subtype profile and use age-specific false-positive rates: point estimate, lower endpoint, and upper endpoint, obtained by complementing specificity in Tables A6-1/A7-1 (pp. 52/54). Subtract the upper specificity limit for the lower false-positive limit and vice versa. Apply the 50–59 estimate to ages 50–54 and 55–59, the 60–69 estimate to 60–64 and 65–69, and the 70–79 estimate to 70–74. Apply the same rate to both sexes; these tables do not supply joint age/sex rates. All marginal endpoints move together in the endpoint scenarios: these are not simultaneous confidence bounds or model confidence intervals.

These remain exploratory first-round episode-sensitivity inputs transported into a steady-state model with a one-year detectable window, independent detection within sites, and full adherence. They are not measured repeated-screening outcomes, do not establish that changing assays caused changes in sensitivity, and are not calibrated age-specific forecasts. The FDA explicitly notes uncertainty about episode sensitivity as true sensitivity (pp. 11, 26). The NHS registry does not capture recurrent cancers and categorizes test-positive recurrences as non-cancer (p. 27); its false-positive rates consequently use a different outcome definition. The FDA's first-round results are distinct from the three-round NHS trial endpoints in the main article.


## Introductory chart: eligible screening populations

The introductory chart is a separate model from the age/sex policy comparison. Run `python3 scripts/blog/galleri/colonoscopy_model.py`, then `python3 scripts/blog/galleri/screening_population_model.py`, then render with `python3 scripts/blog/galleri/render-screening-chart.py` (requires matplotlib). Outputs are `galleri-uspstf-eligible-populations.json`, `.csv`, and `galleri-uspstf-by-screening.png`. The annual panel is per 1,000 people in each estimated eligible population: breast (women 40–74), cervical (people with a cervix 21–65), and lung (adults 50–80 meeting smoking eligibility). A separate panel shows colonoscopy per 1,000 examinations using a study of ages 50–84; it labels the recommended 45–75 range separately. Endpoints are inclusive. Do not add these rows or compare their absolute counts as if they describe the same population.

We use the pinned Census 2022 single-year age/sex population weights. Each single-year age takes its SEER five-year-band incidence, including age 21 from 20–24, age 75 from 75–79, and age 80 from 80–84. Annual target cases, caught cases, and false-positive tests are summed across eligible ages and sexes, then divided by the estimated eligible population and multiplied by 1,000. The same one-year detectable window and full-adherence capture formula apply. Missed cases are target cancers within the eligible population not caught by the selected schedule; they do not include cancers in the excluded population. All rates remain illustrative, not guideline-issued forecasts.

Breast uses biennial mammography with the existing performance inputs. The introductory colorectal panel uses per-examination outcomes, as detailed below; its annual outcome fields are null to prevent denominator confusion. Cervical uses Pap every three years at 21–29 and cotesting every five years at 30–65, both permitted by the [current USPSTF guideline](https://www.uspreventiveservicestaskforce.org/uspstf/recommendation/cervical-cancer-screening). Pap cancer sensitivity is 1,015/1,193 (85.1%) from [Kaufman et al., 2020](https://doi.org/10.1093/ajcp/aqaa074), a retrospective analysis of the cytology component of cotesting within 12 months of cancer diagnosis; its transfer to standalone screening at 21–29 is an assumption. Pap non-cancer positivity is assumed to be 5%; it is not a measured cancer-specific false-positive rate. Cotest inputs remain 94.1% and 7%. We estimate cervix-present population as 75% of the female registry population at each age, dividing uncorrected cervical cancer incidence by that eligible denominator. This constant fraction is a simplifying anatomy assumption, not an age-specific hysterectomy estimate; the data use registry sex rather than gender identity.

Lung uses annual CT, 78.6% sensitivity and 5.3% false-positive rate. Eligible populations for 50–79 use the [2022 BRFSS study](https://pmc.ncbi.nlm.nih.gov/articles/PMC10958241/) counts already used by the main model plus 1,098,421 at 75–79. The 75–79 eligible fraction is assumed to apply at age 80 because the source omits that age. We estimate 60% of lung cancers at each age occur in smoking-eligible people, as in the main model; this is an assumption, not a directly measured eligible-cohort incidence rate. Eligible incidence is those estimated cases divided by eligible population, rather than general-population incidence applied to eligible smokers. Both sexes share the age-specific eligibility fraction.

The estimates do not quantitatively remove people unable to undergo curative treatment or with other health exclusions, nor model detailed high-risk histories. They omit prevention through treating precancer and use current observed incidence rather than a natural-history simulation. These limitations, especially cervical positivity/anatomy and lung-risk assumptions, limit interpretation. The later age/sex policy chart and its appendix breakdown remain unchanged.


### Colonoscopy per examination

`colonoscopy_model.py` produces `galleri-colonoscopy-per-exam.json` and `.csv`. Source counts and eligibility are pinned in `inputs/deep-c-colonoscopy-2014.json` from [Imperiale et al. 2014, Figure 1 and Table 2](https://doi.org/10.1056/NEJMoa1311194). These are screening yields, not incidence. The older-weighted cohort is not standardized to USPSTF ages 45–75. The study used colonoscopy as the reference test, so its misses cannot be directly observed.

For an illustrative, mutually exclusive partition per 1,000 examinations:

- `cancer_caught = observed cancer count / sample size × 1000`.
- `cancer_missed = cancer_caught × (1 / cancer_sensitivity − 1)`.
- `precancer_detected = (advanced precancer + nonadvanced adenomas) / sample size × 1000`.
- `precancer_missed = precancer_detected × (1 / precancer_sensitivity − 1)`.
- `target_negative = 1000 − cancer_caught − cancer_missed − precancer_detected − precancer_missed`.
- `false_positive_interventions = target_negative × (1 − specificity)`; `true_negatives = target_negative × specificity`.

Use assumed cancer sensitivity 0.95 and specificity 0.86 from [USPSTF Table 7](https://www.ncbi.nlm.nih.gov/books/NBK570821/table/ch2.tab7/). Precursor sensitivity 0.85 is an additional assumption: the source gives size-dependent lesion-level sensitivity, not a pooled person-level estimate for this study. Scenarios use 0.75 and 0.95, not confidence bounds. Precancer categories are the study's categories; small serrated lesions remain unresolved. True negatives and false positives are modeled, not observed. The 14% conditional false-positive rate means non-precancerous polyp interventions, not cancer misdiagnoses. All six outcome categories sum to 1,000; detected precancers are not cancers caught or false positives.

This supersedes the introductory chart's prior 88-per-1,000-exam proxy. The revised FP estimate also allows for inferred missed precancer and uses one cohort's cancer/precancer yields. The prior GIQuIC age-weighted approximation is not silently presented as the same population.

### Central colonoscopy policy

The main article now uses the colonoscopy model. Its detailed appendix chart uses `galleri-colonoscopy-policy-model.json`/`.csv`, generated by `colonoscopy_model.py` and rendered with `render-policy-chart.py --colonoscopy`. It shows all five age bands, both sexes, and three policies; downloads include nine combinations of prevention (0%, 20%, 40%) and detectable window (0.5, 1, 2 years).

This is a mature-program approximation, not a cohort simulation. It retains the original FIT model's non-colorectal assumptions. Colonoscopy replaces FIT: interval ten years, cancer sensitivity 0.95, random-phase capture formula. No observed initial-screen yield is divided by ten. At the central one-year detectable window, capture probability for remaining CRC is 0.095. That is an assumed opportunity-to-detect probability, not test sensitivity or the total effectiveness of a colonoscopy program.

Let `I` be observed age/sex-specific SEER colorectal incidence. For guideline and hybrid policies, assume additional prevention `P = p × I`, with central `p = 0.20`. Remaining incidence is `I − P`. Apply detection to that remaining pool. Both policies have identical prevention; Galleri alone retains baseline incidence and has zero modeled prevention. Hybrid capture uses `s + g − s*g` within each site; overlap bounds are updated for the remaining CRC pool. For each policy, `caught + missed + prevented = baseline all-cancer incidence`. Detection percentages use cancers remaining after prevention as denominator. Prevented cancers are shown separately in the appendix table and downloads, not added to the green detection segment.

The prevention parameter is explicitly a hypothetical **additional** reduction relative to already-screened US incidence, not total efficacy versus no screening and not a trial calibration. It avoids claiming a known counterfactual unscreened incidence. Applying it immediately assumes a mature program: there is no lag, aging, adenoma progression, repeat-round change, surveillance, competing mortality, stage shift, or adherence dynamics. No survival or life-years estimate follows.

Routine colonoscopy FP interventions use `max(0, 1000 × (1 − observed_adenoma_fraction) − remaining_CRC × detectable_years) × 0.14 / 10`. Age/sex-specific observed adenoma fractions come from [GIQuIC Table 3](https://pmc.ncbi.nlm.nih.gov/articles/PMC9398947/#T3), pinned locally. This remains a proxy: missed adenomas and serrated-only precancers are unresolved; using observed detection as prevalence can overstate the negative pool. The older screening sample also includes elevated-risk participants. Surveillance is excluded, understating procedure burden. These annual age-specific inputs differ from the pooled examination illustration above.

The Galleri false-positive denominator increases slightly when prevention removes cancers. True negatives are not aggregated across modalities: counts are test events, not unique people. The colonoscopy FP endpoint excludes precancer, while other tests retain the main model's cancer-only endpoints; differences in FP counts are not equivalent differences in harm.

Reproduction:

```sh
python3 scripts/blog/galleri/colonoscopy_model.py
python3 scripts/blog/galleri/screening_population_model.py
python3 scripts/blog/galleri/render-screening-chart.py
python3 scripts/blog/galleri/render-policy-chart.py --colonoscopy
python3 -m unittest discover -s scripts/blog/galleri -p 'test*model.py'
```

## Simplified chart age bands

`age_band_model.py` generates `galleri-policy-age-bands.json` and `.csv` for figure 3. It averages annual outcome rates using 2022 Census sex-specific population weights across 30–49, 50–64, and 65–74. Detection percentages are ratios of aggregated rates. The oldest group is deliberately labeled 65–74: the model does not support a general 65+ claim. The broad bands now aggregate the colonoscopy model, including prevention counts. The original FIT outputs remain separately available.

The younger group is an illustrative extrapolation of Galleri performance, not a validated estimate or recommendation. [Manufacturer information](https://www.galleri.com/safety-information) describes elevated-risk adults such as ages 50+. In ages 30–49, the model permits cervical cotesting throughout, mammography starting at 40, colonoscopy starting at 45, and no lung CT. Below 45, the model assigns neither colonoscopy detection nor additional prevention. Suppressed younger site rates are unallocated to individual cancer classes: total incidence remains intact, and the additional residual uses the existing residual Galleri sensitivity assumption. This limits subtype-based interpretation in the younger group.

Generate broad-band inputs before rendering: `python3 scripts/blog/galleri/age_band_model.py`, then `python3 scripts/blog/galleri/render-policy-chart.py`. Chart outcome-number columns are omitted; numerical tables and detailed captions are in the article's chart appendix. Shared axes and units remain in the images. Colonoscopy per-exam outcomes keep a separate denominator from annual outcomes.


## Paired assay tradeoff cited in the article

[GRAIL sponsor briefing, p. 104, Table 34](https://www.fda.gov/media/194910/download#page=104) compares the two assays among participants in both analyzable sets. Among 303 cancer cases: MCED-V2 detected 102 + 16 = 118 (38.94%); revised Galleri detected 102 + 4 = 106 (34.98%). Among 21,116 participants without cancer: MCED-V2 positives were 22 + 55 = 77 (0.3647%); revised Galleri positives were 22 + 10 = 32 (0.1515%). The sponsor explicitly attributes the lower false-positive rate and lower sensitivity to more stringent classifier training. These paired estimates differ in population from the 39.3% full-study result used in the main model; this paragraph does not change its assay inputs.
