Statistical validation reports
The measurement record behind the package defaults. Each report reads the canonical summary tables produced by the benchmark harness in this repository; the package itself ships only the user-facing vignettes.
- Simulation study: calibration, power and FDR
The simulator, then every default measured on it: the random modes and reference df, the CellType:condition term, the 0.99.17 switches; power, recall, FDR control, gene categories, spatial layouts, and timing. - The real cohort: the null tail, its cause, and the fix
A 55-patient CosMx cohort against its own niche-shuffle null: why real and null were indistinguishable, the per-gene tail, the composition confound, the nested intercept, the reversed verdict and its stress tests. - The two-stage estimator
The patient-level estimator, paired on the same simulated datasets and on the real cohort: calibration, power, the cells appetite, the negative control, the adversarial scenarios, and when to use it. - Combining correlated niche p-values: Brown vs Cauchy
Calibration and power of the two combiners, and why Cauchy needs two-sided input. - Gene-set inference: calibrating spiGSEA
The correlation calibration, the self-contained test's failure modes, and the competitive default. - What was tried and rejected
The alternatives that were built, measured and removed, each with the number that decided it: the between-sample stratum, cell subsampling, random slopes as the fix, the QL dispersion, the gene filter, and more.