Blue Camel Consulting

Subgroup replication monitor

What a delivered subgroup scores once the search is over

The sponsor, Quillmere Therapeutics, is invented and every cohort is synthetic, drawn from a fixed seed and computed in your browser.

Exec summary, preset program, seed 417

Running

Scoring the preset program.

Why. A subgroup kept as the best of candidates on patients is a winner's curse estimate: part of its apparent AUC is the search. The confirmatory arm is the first measurement the search never saw. This layer watches that arm, every refit and the placebo arm after delivery, and states in advance what pages a person.

Every figure above is computed at load by the same code that fills the tables below.

The checks behind the summary

Running

Scoring the synthetic programs.

Adjust the bounds or draw a new program
Page rules, stated before the data, and their measured false-page rates
  • Page when the drop from discovery AUC to confirmatory AUC has a one-sided 95% bootstrap lower bound above the replication bound, or when the confirmatory arm's 95% interval tops out under 0.60 AUC, so that arm shows no useful separation. Either needs the confirmatory cohort to meet the minimum size.
  • Watch when the drop passes the bound and neither page condition holds, because a discovery estimate on 50 patients is too wide to size the drop. A confirmatory cohort under the minimum is a watch whatever it shows, and each row says which reason applies.
  • Page on a refit when the delivered pair wins under 10% of bootstrap refits of the refreshed data. Watch when it wins under 70%, or when the refit picks a different pair but the delivered one still wins at least 10%.
  • Page when the drug minus placebo AUC gap estimate is under the floor in two successive cohorts. One cohort under the floor is a watch.

This is a monitoring alarm with a stated false-page rate, not a confirmatory test. Selection and its error control belong before delivery; this layer runs after it, on refits, new sites and successive cohorts. The rates here are fixed text, measured offline by running this page's own code at the default settings on simulated programs in which nothing changed, one Monte Carlo draw each, on seed blocks from 31,000,000 to 39,000,999. Replication rule: it paged 6 of 2,000 pre-specified signatures (0.3%), and 145 of 2,000 (7%) when the true drop sits exactly at the bound, against a nominal 5%. Refit rule: a selection-stability alarm, not a change detector. Where nothing changed it paged 147 of 800 refits (18%, from 400 programs with two refreshes each), 129 of them where the 80-patient search had delivered a weaker pair than the population's best. It pages more often when several pairs are about equally good: 137 of 400 refits (34%) with three interchangeable markers. Placebo rule: it reads estimates, not intervals, so its rate depends on how close the true gap sits to the floor. With the true gap held at 0.35 it paged 0 of 1,000 six-cohort runs; held at 0.25, 28 (3%); at 0.20, 226 (23%); at the 0.15 floor itself, 676 (68%).

Runbook for this check

    Method, and what this cannot see

    How the numbers are made. Each synthetic patient has 30 baseline variables and a responder label. Responders sit higher on a few variables; on most of the rest there is no difference at all. A signature is two variables, its score is their average, and its separation is the AUC of that score between responders and non-responders, counted from the drawn patients. Each delivered signature in the first preset comes from its own synthetic program and is the best of all 435 two-variable pairs searched on its own discovery cohort of 50, and its confirmatory arm is drawn from the same population, so the drop is the optimism of choosing the best of 435 on 50 patients, not a change in the patients. The delivered AUC here is the searched cohort's own apparent AUC, the number a naive report would carry. A method that reports a nested cross-validated or held-out estimate is built to remove this optimism, and the monitor checks what is left after delivery. Measured offline with this page's code, a five-fold nested estimate of the same search on the same 50 patients sat within 0.02 AUC of the population value on average, where the apparent AUC sat 0.07 to 0.22 above it, over 150 draws per signature. The population AUC under each signature is its true value in the generating population, known only because the data are synthetic. Intervals are 95% percentile bootstrap intervals over 400 resamples. The drop resamples both cohorts but does not repeat the search inside each resample, so the discovery interval is optimistic too. Two changes on the page are planted: the second data refresh adds two new sites where a non-clinical variable carries signal, and in the placebo preset placebo responders drift toward the signature across cohorts. The placebo cohorts are fresh patients each time, not cumulative accrual, and the gap between drug arm AUC and placebo arm AUC is a screen for a prognostic rather than predictive signature, not a test of the treatment-by-marker interaction.

    What this cannot see. Any real sponsor's subgroups, cohorts or validation. Here a signature scores every patient and nothing abstains. A Model-Derived Subgroup with a No Call covers a called subset (the study below reports 27.7% to 40.4% of each cohort called), so the monitor would also bound the called share in the confirmatory arm. The study's datasets (CATIE, CAN-BIND, COMPASS) are not used or reproduced here. No real trial data and no patient information were used anywhere.

    Context. Geraci et al. (AI, 2026) describe NetraAI's Model-Derived Subgroups as compact, inspectable hypotheses that may inform enrichment "after external validation". This page sketches a monitor for the step after delivery, on an invented sponsor. Built by Jeff Pinto. His CAMH-era work trained decoupled classifiers per patient subgroup on psychiatric clinical records, closing an accuracy parity gap from 35% to 1% across 140 permutations.

    Source: Geraci J, Qorri B, Cumbaa C, et al. Interpretable Subgroup Discovery with Abstention in Small, Heterogeneous Clinical Trials: A Retrospective Multi-Dataset Study. AI 2026, 7(9), 349. doi:10.3390/ai7090349.