> ## Preface to the published copy > > This is the pre-specified statistical analysis plan for the study reported at > [creativetouchrotherham.co.uk/lab/insights/food-addiction-criteria-bmi](https://creativetouchrotherham.co.uk/lab/insights/food-addiction-criteria-bmi). > It is published under CC BY 4.0 alongside the three scripts that implement it: > [`extract.py`](extract.py), [`analyse.py`](analyse.py) and [`make_figures.py`](make_figures.py). > > **Everything below this preface is verbatim.** Nothing has been rewritten, tidied, updated to > match what the analysis found, or clarified for a general reader. That is deliberate. A > pre-registration whose text is improved after the results are known is no longer a > pre-registration, and the whole value of this document rests on it being the same text that was > committed on 26 July 2026, before any model was fitted. Four consequences a reader should expect: > > 1. **It is written in the future tense**, because at the time of writing the analysis had not > been run. "Will be reported" means it had not yet been. > 2. **Some numbers in the body are superseded.** §4 plans around discordant-phenotype cells of > n=182 and n=407; after exclusions the real cells were 155 and 394, and §9 Deviation 1 records > that, along with the recomputed power. Where the body and the deviations log disagree, the > deviations log is later and correct. Both are left standing so the correction is visible > rather than absorbed. > 3. **§9, the deviations log, is the only section written after results were seen**, which is > what a deviations log is for. Each entry is dated. Two entries exist. Neither is a tidying-up > edit. > 4. **It is a working document, not a paper.** It argues with an earlier version of its own > design, names the specific ways it could be wrong, and in one place (Deviation 1) records a > structural mistake in its own reasoning that was caught only when the model failed to > converge. It is published as it was used. > > **A word the article avoids and this document does not.** The body calls itself a > pre-registration throughout, starting with its first line. The article deliberately does not. It > says "pre-specified", because nothing here was deposited with a public registry, and > "pre-registration" is a term most readers reasonably hear as meaning exactly that. The plan's > own usage is the looser methodological one, a plan fixed and committed before the analysis > rather than a plan lodged with an independent third party. Both §8 below and the article itself > are explicit that no third party vouches for the timing. > > The word is left standing anyway. Correcting a plan's vocabulary after its results are known is > a smaller edit than correcting its models, but it is the same kind of edit, and a document that > insists on leaving superseded numbers visible cannot then quietly improve its own terminology. > Read "pre-registration" throughout the body as "pre-specified analysis plan". > > **Cross-references to documents that are not published.** The body cites two internal working > files. `01-literature-positioning.md` is the literature review and validity-threat assessment > the plan was built on; its substance appears in the article's discussion and limitations > sections. `00-STATE.md` is the project's running technical log; the item-storage behaviour it is > cited for is explained in full in the article's section on the criterion 7 defect. Neither is > needed to read this plan, and neither contains anything that would change how it should be read. > > **What is not published, and will not be.** The individual quiz responses. Creative Touch > publishes aggregated findings only, and running these scripts against our database is something > only we can do. The plan and the scripts are published so that every decision between the raw > responses and the published figures can be inspected, which is the part that ordinarily has to > be taken on trust. > > **Roles, and who did what.** The body refers to a "coordinator" and an "implementing agent". > Every decision recorded in this plan was made by the project owner, including both entries in > the deviations log and, specifically, the Deviation 1 ruling that was made with a result already > visible. The analysis scripts were drafted and run with LLM assistance under that direction, > which is what "implementing agent" refers to. We name it rather than tidying the wording away, > for two reasons. A document that discloses a ruling made with the answer in view, and then > quietly conceals how the code was written, is disclosing selectively. And the concern the > disclosure raises is answerable: the plan was committed before the analysis ran, the scripts are > published, and every number in the article can be traced from the first to the second without > taking anyone's word for who typed it. > > **On timestamps.** This plan was committed to version control before the analysis ran, and the > analysis script was committed 88 minutes later. Both timestamps are ours, not a registry's. The > article says so plainly and does not describe the study as pre-registered in the clinical-trials > sense. Publishing the plan does not upgrade that claim; it just means the plan itself can now be > read rather than merely asserted. --- # Statistical Analysis Plan (SAP) ## YFAS 2.0 criterion-level analysis — Creative Touch original dataset ## Version 1.0 — committed 2026-07-26, prior to any inferential query This is a pre-registration. Everything in this document is decided in advance of seeing any inferential output. No model has been fitted, no p-value has been computed, no coefficient has been inspected. Descriptive counts used to power this plan (endorsement rates, exclusion counts, discordant-cell sizes) are permitted under the pre-registration convention that non-inferential sample characterisation may precede a locked analysis plan — see §2 for the exact list of what was and was not looked at before this document was written. British English throughout. Any change to this plan after commitment must be recorded in §9 (Deviations Log), with the date, the change, and the reason. A deviations log with no entries at publication time is the intended, honest outcome — it means the plan was followed exactly. --- ## 1. Primary research question and hypotheses (carried forward from Step 1, unchanged) **Primary question:** Is the YFAS 2.0 → adiposity association uniform across the eleven diagnostic criteria, or is it concentrated in criteria whose content is entangled with body size and its social consequences, **after conditioning on overall food-addiction severity**? **H1 (primary, confirmatory):** Among respondents of equal overall food-addiction severity, BMI predicts endorsement of the content-confounded criteria (C4 activities given up, C5 use despite harm, C8 interpersonal problems, C9 role failure, C10 hazardous use) more strongly than it predicts endorsement of the pharmacological-core criteria (C6 tolerance, C7 withdrawal, C11 craving). **H2 (secondary, confirmatory — promoted from exploratory per coordinator decision):** The criterion-endorsement profile of the discordant phenotype "high symptom count at normal BMI" differs from "high symptom count at obese BMI" — specifically, the normal-weight-high-symptom group is hypothesised to show relatively higher endorsement of the pharmacological-core criteria and relatively lower endorsement of the content-confounded criteria than the obese-high-symptom group. **H3 (secondary, exploratory):** Criterion profile varies by age band. No directional hypothesis is pre-registered — the literature audit found no prior age-by-criterion finding to anchor a direction. **H4 (secondary, confirmatory replication attempt):** Criterion 9 (role obligation failure) shows a sex-differential pattern, replicating Saffari et al. 2022's gender-DIF finding on this specific criterion. No other criterion is pre-registered to show a sex effect; a sex effect appearing on a different criterion instead of C9 is a failure to replicate, not a new discovery to chase. Full rationale, falsification conditions, and literature grounding for H1–H4 are in `01-literature-positioning.md`. This SAP does not repeat that reasoning; it operationalises it. --- ## 2. Why the primary model is rest-score conditioned (not the raw 11-model design proposed in Step 1) Step 1 proposed 11 independent logistic regressions of `criterion_met ~ BMI + age + sex`, ranking the 11 raw BMI coefficients against each other. **This is superseded.** Two problems make the unconditioned design incapable of testing H1 as stated: 1. **Circularity by construction.** Every one of the 11 criteria is a component of the same underlying construct that BMI is already known to track in aggregate (the whole-sample monotonic symptom-count gradient across BMI classes). An unconditioned model will find all 11 BMI coefficients positive, because BMI predicts overall severity and overall severity mechanically predicts each component criterion. Ranking 11 coefficients that are all positive for the same underlying reason does not distinguish content-confounded criteria from pharmacological-core criteria — it just re-derives the aggregate gradient eleven times. 2. **Difficulty-scale confound.** C9 sits at 27.9% whole-sample endorsement; C7 sits at 58.9%. These are different points on the item-difficulty continuum (in the IRT sense used by Chapron 2023 and Saffari 2022). Raw odds ratios at different baseline endorsement rates are not on a comparable scale for the purpose of ranking "which criterion is more BMI-sensitive" — a criterion with low base-rate endorsement can show an inflated or deflated OR for reasons unrelated to content. The content-confounding hypothesis, properly stated, is conditional: **does BMI predict endorsement of a specific criterion among people who are equally food-addiction-symptomatic overall?** This is the only framing that isolates a content effect from a severity effect. It is also, per the Step 1 literature audit (§1.5 of `01-literature-positioning.md`), exactly the distinction between *differential validity against an external criterion* and *DIF against the latent trait* — and conditioning on a severity measure is what makes this design commensurable with both Chapron 2023 (IRT-conditioned) and Saffari 2022 (Rasch-conditioned), neither of which could have been reconciled against an unconditioned raw-OR design. ### 2.1 Rest-score construction For each criterion *k* (k = 1..11), define: ``` rest_score_k = (total criteria met, 0–11) − (criterion k met, 0/1) ``` Range: 0–10. This is the respondent's symptom burden **excluding the criterion currently being modelled**, avoiding the circularity of regressing a criterion on a total score that already contains it (which would mechanically bias the coefficient). Rest-score is computed fresh per criterion-model — eleven different rest-score variables, one per criterion, each built by subtracting that criterion's own indicator from the 0–11 total. ### 2.2 Primary confirmatory specification **Design:** stacked long-format dataset. Each respondent contributes 11 rows (one per criterion, C1–C11), so n_rows = 11 × n_respondents. Each row carries: `criterion_id` (categorical, 11 levels), `criterion_met` (0/1, the outcome), `BMI` (continuous, respondent-level, repeated across their 11 rows), `rest_score_k` (varies by row — the rest-score appropriate to that row's criterion), `age`, `sex`, `respondent_id` (cluster variable), and `criterion_set` (categorical: `content_confounded` = {C4,C5,C8,C9,C10}, `pharmacological_core` = {C6,C7,C11}, `neither` = {C1,C2,C3}). **Model:** ``` GEE, logit link, binomial family, exchangeable working correlation, clustered on respondent_id criterion_met ~ C(criterion_id) + BMI:C(criterion_set) + rest_score + age + sex ``` restricted to rows where `criterion_set ∈ {content_confounded, pharmacological_core}` — C1, C2, C3 rows are excluded from this fitted model entirely, not merely excluded from interpretation. This is the single confirmatory test for H1. **Criterion-specific intercepts are mandatory, and the specification above is not interchangeable with `BMI * criterion_set`.** The eight criteria in the confirmatory model differ substantially in base endorsement rate (whole-sample: C4 33.2%, C5 52.4%, C8 29.8%, C9 27.9%, C10 33.2% versus C6 36.6%, C7 58.9%, C11 44.8%). A model carrying only a two-level `criterion_set` term would force a single common intercept across all five content-confounded criteria and another across all three core criteria, misattributing baseline difficulty differences to the set factor — **reintroducing, in the confirmatory model itself, the very difficulty-scale confound that §2 point 2 gives as the reason for abandoning the unconditioned design.** `C(criterion_id)` therefore supplies a free intercept per criterion (8 levels), and the BMI slope is estimated at set level via the `BMI:C(criterion_set)` interaction. Note that `criterion_set` is a deterministic function of `criterion_id`, so a `criterion_set` **main effect** is collinear with the criterion-level intercepts and must be omitted; only the `BMI:C(criterion_set)` product term is estimated. Implement in `statsmodels` via the formula interface exactly as written, and assert the resulting design-matrix rank in the script rather than trusting the formula parser to drop the right term. **The confirmatory test statistic:** the Wald test (or robust-SE equivalent under GEE) on the `BMI × criterion_set` interaction contrast, specifically the difference in the BMI coefficient between `content_confounded` and `pharmacological_core` levels of `criterion_set` (set `pharmacological_core` as the reference level so the interaction term is read directly as "content-confounded slope minus core slope"). **One p-value. One confidence interval. This is the test H1 lives or dies on.** **Why GEE over a mixed-effects (GLMM) model:** GEE with exchangeable working correlation gives population-averaged (marginal) odds ratios, which is the correct target for the research question as stated ("does BMI predict endorsement, on average, across respondents") rather than the subject-specific odds ratios a GLMM random-intercept model would give. GEE is also more robust to misspecification of the within-respondent correlation structure — exchangeable is a reasonable default for 11 criteria of the same underlying construct with no a priori temporal or hierarchical ordering, and GEE's sandwich variance estimator is robust even if the true correlation structure isn't exactly exchangeable. statsmodels' `GEE` class supports this directly. ### 2.3 Secondary/descriptive output (per-criterion, not confirmatory for H1) The confirmatory model in §2.2 estimates BMI slopes at **set** level only (two slopes), so it cannot by construction yield eleven per-criterion coefficients. These are obtained from a **separate, explicitly secondary model**, pre-specified here as: ``` GEE, logit link, binomial family, exchangeable working correlation, clustered on respondent_id criterion_met ~ C(criterion_id) + BMI:C(criterion_id) + rest_score + age + sex ``` fitted across all eleven criteria (C1–C3 included here, unlike the confirmatory model), giving one BMI slope per criterion. These eleven coefficients are reported as **secondary, descriptive output**, under Benjamini-Hochberg FDR correction at q=0.05 across the eleven tests. These answer "which specific criteria show the effect" as a follow-up to the single confirmatory contrast in §2.2 — they do not themselves constitute the test of H1, and must not be reported as though 11 significant p-values confirm H1 in the absence of the confirmatory interaction contrast being significant. ### 2.4 Retained: unconditioned models, explicitly relabelled The 11 raw (unconditioned) `criterion_met ~ BMI + age + sex` logistic regressions proposed in Step 1 are retained in the analysis script and reported, but **explicitly labelled as descriptive, not as a test of H1**. Their purpose is transparency (showing the reader what the naive/uncontrolled picture looks like) and as a didactic contrast against the conditioned model — the gap between the raw picture and the rest-score-conditioned picture is itself a finding worth reporting, since it demonstrates the circularity problem in §2 point 1 empirically rather than just asserting it. --- ## 3. Secondary arm: IRT/DIF convergent-methods check **Status:** retained, and strengthened by the primary model's conditioning. Under an unconditioned primary design, IRT-DIF would have tested a different question (item-latent-trait relationship) from the primary analysis (raw criterion-BMI association). Under the rest-score-conditioned primary design, both approaches are now trait-conditioned — the GEE model conditions on rest-score (a sum-score proxy for the trait), and IRT-DIF conditions on the latent trait estimate directly. **This makes the IRT-DIF arm a genuine convergent-methods check on the primary finding, not a separate question.** **Method (modelled on Chapron et al. 2023, PMID 37666092):** fit a unidimensional 2-parameter logistic (2PL) IRT model to the 11 binary criteria (item difficulty and discrimination parameters per criterion), then test uniform and non-uniform DIF for BMI as the grouping variable (BMI split at the median of the analytic cohort, plus a continuous-BMI DIF sensitivity check if the software path allows it). Compare DIF-flagged criteria against the H1 content-confounded/pharmacological-core grouping. **statsmodels capability note:** statsmodels does not have a dedicated IRT/DIF module. This arm will be implemented using `statsmodels.discrete.discrete_model.Logit` per item within an iterative 2PL estimation routine, or — if within-scope package availability allows within the WSL venv — the Python `girth` or `py-irt` packages, which run in the same statsmodels-compatible numpy/scipy environment. **This choice is deferred to the implementing script and must be recorded in that script's header comment**, but under no circumstances will this arm be implemented in R. If no adequate Python IRT package is available in the WSL venv at implementation time, this arm will be simplified to a Mantel-Haenszel DIF test (stratified by rest-score decile, comparable to a coarse ordinal-conditioning DIF test), which is implementable directly with `scipy.stats` and pandas cross-tabulation — **this substitution, if needed, must be logged in §9 as a deviation with the reason (package unavailability), decided at implementation time, not after results are seen.** **Reporting:** as a table of per-criterion DIF flags (significant/non-significant, direction) set alongside the primary GEE per-criterion coefficients from §2.3, for the reader to see convergence or divergence directly. --- ## 4. Discordant phenotypes (H2) — pre-specified confirmatory secondary analysis **Cell sizes (already computed, non-inferential, as supplied):** - Normal BMI (18.5–<25) with ≥6 symptoms: n=182 - Obese (≥30) with ≥6 symptoms: n=407 - Obese with ≤1 symptom: n=99 These are large enough to support a confirmatory (not merely descriptive) comparison. H2 is tested as follows. **Outcome:** for each respondent in the two comparison groups (normal-BMI-high-symptom, n=182; obese-high-symptom, n=407), compute the proportion of their endorsed criteria that fall in the content-confounded set (C4,C5,C8,C9,C10) versus the pharmacological-core set (C6,C7,C11) versus neither (C1,C2,C3) — a compositional (proportion) outcome per respondent, restricted to their endorsed criteria only. **Test:** two-sample comparison of the content-confounded-share proportion between the two groups. Given the outcome is a bounded proportion (0–1) with a moderate-sized but not huge n per group, use a logistic regression on the pooled criterion-level binary outcome (content_confounded vs pharmacological_core endorsed-criterion membership, restricted to endorsed criteria only, one row per endorsed criterion per respondent) with `group` (normal-BMI-high-symptom vs obese-high-symptom) as the predictor and a cluster-robust (GEE, clustered on respondent) standard error, mirroring the primary model's approach for methodological consistency. This is implementable in statsmodels via `GEE` with a binomial family exactly as in §2.2, just on a different subsample and a different grouping predictor. **Total symptom count must be adjusted for, and this is not optional.** Both groups are defined by ≥6 symptoms, but that is a floor, not a match: the obese group is very likely to sit higher within the 6–11 range (the whole-sample gradient runs 3.38 at normal BMI to 6.85 at obese III). Because the compositional outcome is computed over *endorsed criteria only*, a respondent endorsing 10 criteria mechanically has a different achievable content-confounded share from one endorsing 6 — with 5 content-confounded, 3 core and 3 neither criteria available, the share is bounded differently at different totals. An unadjusted comparison would therefore partly measure the severity difference between the groups rather than the profile difference H2 is about, which is the same error the primary model's rest-score conditioning exists to avoid. The H2 model is therefore: ``` endorsed_criterion_is_content_confounded ~ group + total_symptom_count + age + sex ``` as a GEE clustered on respondent, restricted to endorsed criteria in the content-confounded and pharmacological-core sets. The unadjusted comparison is additionally reported as descriptive, labelled as such, for the same didactic reason given in §2.4. The group means of `total_symptom_count` are reported alongside, so the reader can see directly how much severity separation the adjustment is absorbing. **A priori power statement:** with group sizes of 182 and 407 and an outcome proportion expected to sit in a plausible mid-range (0.3–0.6, based on the five content-confounded criteria out of eight relevant criteria, i.e. base rate ~0.5–0.6 under the null of no group difference), a two-proportion comparison at these sample sizes has power ≥80% to detect an absolute difference in content-confounded-share of approximately 12–15 percentage points at alpha=0.05 (two-sided), using a standard normal-approximation power calculation for two independent proportions; this will be re-verified exactly (not just approximately) using `statsmodels.stats.power` (`NormalIndPower` or equivalent) as a checked, reported number in the analysis output — the approximate figure above is stated here only to establish that the design is adequately powered for a pre-registration decision, not as the final reported power figure. **Obese-low-symptom group (n=99):** retained as a third descriptive comparison group (profile of the few obese respondents who report minimal symptoms) but not part of the confirmatory H2 test — the confirmatory test is the two-group comparison above. The n=99 group's criterion profile is reported descriptively alongside, since it is the "protective" or resilient discordant phenotype and is of independent interest, but no confirmatory hypothesis is pre-registered for it given its smaller size and the absence of a directional literature-grounded prediction for this specific cell. --- ## 5. A priori power analysis by criterion, using the analytic cohort's own observed endorsement rates **Permitted pre-analysis step:** criterion endorsement rates within the n=1,799 analytic cohort (after exclusions, see §7) are a non-inferential descriptive statistic needed to power the study, and are computed as the first script step, before any regression, GEE, or hypothesis test is run. This is standard pre-registration practice (a priori power calculations require the actual sample's characteristics) and is explicitly distinguished here from an inferential "peek." **Procedure:** for each of the 11 criteria, using the n=1,799 cohort: 1. Compute observed endorsement rate. 2. Using `statsmodels.stats.power.NormalIndPower` (or a logistic-regression-appropriate power approximation via simulated power under the planned GEE specification, if time allows — the simulation-based approach is preferred where feasible since it directly reflects the rest-score-conditioned model rather than a simple two-proportion approximation), compute the minimum detectable odds ratio (per 5-unit BMI increase) at 80% power, alpha=0.05 (two-sided), for that criterion's observed endorsement rate, adjusting for the two covariates (rest_score, age, sex) already in the model. 3. Report all 11 values in a table, explicitly flagging C8 (29.8% whole-sample rate) and C9 (27.9% whole-sample rate) as the two criteria with the lowest expected power — **both are central to H1's content-confounded set**, so this is a genuine limitation of the design that must be stated in the results discussion regardless of outcome, not discovered post hoc if these two criteria fail to reach significance. **Recompute note:** the 27.9%/29.8% figures above are whole-sample (n=6,105) rates carried forward from Step 1's literature-positioning document for illustrative purposes in this SAP. The actual power table in §5 step 2 will use the n=1,799 analytic cohort's own observed rates, which may differ from the whole-sample rates — this recomputation is itself the first script step and is not an inferential test. --- ## 6. Environment and implementation **Language: Python.** Not PHP, not R. `StatsCalculator.php` (site's existing PHP statistics helper) does not implement regression, GEE, or IRT — it provides `calculateTwoSamplePValue`, `calculateCorrelation`, `calculateCohenD`, `calculateChiSquareTest`, `calculateTrendTest`, and generic p-value helpers only. None of these cover the model classes this SAP specifies. Reusing it is not viable for this analysis and the parent execution plan's "reuse, do not rebuild" assumption does not apply here. **Runtime:** a committed Python script, executed in a WSL virtual environment with `statsmodels` installed (confirmed: neither Windows nor the base WSL environment has `scipy`/`pandas`/`statsmodels` pre-installed; these must be present in the project's dedicated analysis venv before the script is run — venv setup is an implementation-time task, not a SAP decision, but is noted here so the analysis is reproducible from a documented starting point). **Key statsmodels components to be used:** - `statsmodels.genmod.generalized_estimating_equations.GEE` with `family=sm.families.Binomial()`, `cov_struct=sm.cov_struct.Exchangeable()` — primary H1 model and H2 model. - `statsmodels.discrete.discrete_model.Logit` — descriptive per-criterion unconditioned models (§2.4) and rest-score-conditioned single-criterion sensitivity models. - `statsmodels.stats.multitest.multipletests` with `method='fdr_bh'` — Benjamini-Hochberg correction for the 11 secondary per-criterion p-values (§2.3). - `statsmodels.stats.power.NormalIndPower` (and/or a custom simulation loop built on `numpy`/`statsmodels.Logit` for the GEE-adjusted power estimate) — a priori power analysis (§5). - IRT/DIF arm (§3): package choice deferred to implementation, logged as a deviation if it departs from a pure-statsmodels implementation. **Data pull:** the analysis script connects to the `creativetouch` MariaDB `calculators` table (`module='food_addiction'`) directly via a read-only database connection, applies the exclusions in §7 in code (not by hand), and writes the resulting analytic dataframe to a versioned, timestamped CSV snapshot alongside the script's output, so the exact analytic sample is reproducible independent of any later database changes (see §8 on the data freeze). --- ## 7. Pre-specified exclusions (verbatim, to be implemented exactly as follows) Applied in this order, with the running sample size logged at each step in the script output: 1. **`age_range = 'under_18'` exclusion.** Research-ethics justification, not statistical: self-reported data from minors collected via an unsupervised online quiz without documented parental consent or ethics board sign-off on that collection pathway is excluded as a bright-line rule, not a sensitivity choice. Verified count: **99 of 1,898** in the research cohort (BMI + age_range + sex all present) fall in this band. **Analytic n after this exclusion alone: 1,799.** *Provenance note on these two figures:* 1,898 and 99 were counted on 2026-07-26 across the whole table, i.e. **without** the §8 freeze filter applied. Once `date_create <= '2026-07-25 23:59:59'` is enforced, both may fall by a small number of rows. The freeze-filtered counts are recomputed by the script as its first descriptive step and are the figures that will be reported; the numbers above are the pre-registration's best estimate at commit time, not a prediction the analysis is bound to reproduce exactly. 2. **BMI recomputation from raw height/weight.** The stored `bmi.category` field is not used for classification at any stage — its documented boundary overlaps (normal 18.5–25.0 overlapping overweight 25.0–30.0, and equivalent overlaps at 30/35/40) make it untrustworthy at every threshold. BMI is recomputed in the analysis script directly from `demographics.height` and `demographics.weight` using the standard formula (kg/m², converting units as needed per the stored unit convention in the `demographics` JSON), for every respondent in the cohort, before any other BMI-dependent step. 3. **BMI plausibility screen**, applied to the *recomputed* BMI value, not the stored one: exclude BMI <12 or BMI >80 kg/m². Verified: 5 known implausible rows (max recomputed BMI ≈106) in the full 6,105-row dataset; the exact count within the n=1,799 post-exclusion-1 cohort is confirmed in the script's first descriptive step and reported before any modelling begins. 4. **BMI category recomputation** (for descriptive/visual reporting only, not for the primary continuous-BMI model): non-overlapping WHO cut-points applied to the recomputed BMI — underweight <18.5; normal 18.5 to <25; overweight 25 to <30; obese I 30 to <35; obese II 35 to <40; obese III ≥40. 5. **Cohort restriction.** Analysis is restricted to the n=1,898 (pre-under_18-exclusion) / n=1,799 (post-exclusion) respondents with BMI, age_range, and sex all present — no imputation for missing covariates. This is a complete-case design, appropriate given the study's hypothesis-testing (not population-prevalence-estimation) purpose. This choice is itself a further layer of the self-selection threat (§8 of `01-literature-positioning.md`): completing all three demographic fields is itself an act of engagement that may correlate with symptom severity, and this is stated as an explicit limitation, not silently absorbed into the "cohort" framing. 6. **Sex-category exclusion for the primary model.** `sex = 'other'` and `sex = 'prefer_not_to_say'` (combined n≈24 in the full dataset; exact count within the n=1,799 cohort confirmed in the script's descriptive step) are excluded from the primary confirmatory GEE model on power grounds, but retained in a separate descriptive table reporting their criterion endorsement pattern without inferential claims attached. 7. **No deduplication.** `code` is the only unique identifier per row; a repeat tester is indistinguishable from a new respondent and no heuristic deduplication (by demographic/timestamp proximity or otherwise) will be attempted, since this would introduce arbitrary, unauditable exclusions that are worse than the known-imperfect status quo. This is stated as an acknowledged, unresolved limitation, not an exclusion rule. **Final analytic n:** 1,799 minus the BMI-implausibility count (step 3) minus the sex-exclusion count (step 6, primary model only — step 6 does not reduce the descriptive-table n). The exact final n is reported as the first line of script output before any modelling, and is not a number this SAP predicts in advance, since steps 3 and 6's exact within-cohort counts were not part of the verified figures supplied for this plan. --- ## 8. Stopping rule / no-peeking / data freeze **The dataset is live and growing** (6,105 rows and rising as of 2026-07-26, the date this plan and Step 1 were produced). To keep the analytic sample fixed, reproducible, and free of any incentive to keep collecting data until a desired result appears: - **Data freeze date: 2026-07-25.** The analytic sample is defined as all `calculators` rows with `module='food_addiction'` and `date_create <= '2026-07-25 23:59:59'` (server time). Any row created after this timestamp is excluded from this analysis regardless of when the script is actually run. - **The freeze boundary is deliberately in the past at the moment this plan is committed.** The draft of this SAP specified `<= 2026-07-26 23:59:59`, which was the commit day itself and therefore still open — submissions could have continued arriving inside the frozen window after the plan was committed, leaving the analytic sample undetermined at pre-registration time and the freeze unverifiable after the fact. Moving the boundary back one day makes the sample fully determined and auditable the moment the commit lands, at a cost of roughly one day of accrual against n≈1,799. Reproducibility is worth more than a handful of rows. - **Maximum `date_create` filter is hard-coded in the analysis script** as a WHERE clause, not left as an implicit "whatever is in the table when the query runs" — this makes the freeze enforceable and auditable independent of when the script executes. - **No inferential output has been inspected before this plan was committed.** The only quantities inspected prior to writing this SAP are the non-inferential descriptive counts explicitly listed in the coordinator's brief (under_18 count, discordant cell sizes) and the whole-sample criterion endorsement rates from the original verified-facts brief — none of these are outputs of a fitted model, a hypothesis test, or a p-value. - **No interim analysis.** The script runs once, end to end, against the frozen sample, and produces the full pre-specified output set (§2–§5) in a single pass. No partial results are inspected and no analytic choices are revisited based on an early look at any single hypothesis's result before the others are run. --- ## 9. Deviations Log *(Empty at commit time. Any departure from this plan — including implementation-time substitutions explicitly flagged as possible above, such as the IRT package choice in §3 — must be recorded here with: date, section affected, what changed, and why. An empty log at publication time means the plan was followed exactly; a populated log is not a failure, it is what makes this pre-registration meaningful rather than decorative.)* | Date | Section | Deviation | Reason | |---|---|---|---| | 2026-07-26 | §4 | H2 working correlation changed from `Exchangeable()` to `Independence()`. Nothing else changed — same GEE, same logit/binomial family, same formula including the mandatory `total_symptom_count` adjustment, same cluster-robust sandwich SEs on `respondent_id`, same population-averaged estimand. | **Forced by non-convergence, not chosen.** See the note below. | | 2026-07-26 | §§2.2, 2.3 | **Addition of an unplanned sensitivity analysis excluding criterion 7** from the pharmacological-core set, re-fitting the primary confirmatory contrast on the remaining C6/C11 core. The pre-registered primary model is **not** replaced. | **Measurement defect discovered post-analysis** in the instrument's own scoring code. See Deviation 2 below. | ### Deviation 1 — full record (§4 H2 working correlation) **What failed.** The model as pre-registered in §4 does not converge on the real analytic sample. Fitted with `Exchangeable()` at `maxiter=1000`, the dependence parameter goes negative and the fit diverges to `NaN` (`converged = False`, ρ = `NaN`). This was independently reproduced by the coordinator, not taken on the implementing agent's report. **Why it failed — a structural error in the plan, not a data problem.** The H2 outcome is *compositional within respondent*: each respondent's endorsed criteria are partitioned into at most 5 content-confounded and at most 3 pharmacological-core. Endorsing one content-confounded criterion therefore makes another *less* likely within the same cluster, so within-cluster dependence is **negative by construction**. The empirical residual correlation is **ρ = −0.1233** against the exchangeable structure's hard lower bound of **−1/(m−1) = −0.1429** at the observed maximum cluster size of 8 — close enough to the boundary that estimation breaks down. §4 carried the exchangeable structure over from the primary model **by analogy**. In the primary model that analogy is sound, because the eleven criteria are indicators of one construct and dependence is strongly positive. For a compositional outcome the analogy inverts. This is a specification error in the SAP that could have been caught at writing time by reasoning about the sign of the dependence, and was not. **Why `Independence()` is the minimal correct repair.** GEE point estimates are consistent under *any* working correlation; the working correlation affects efficiency only. Inference rests on the cluster-robust sandwich estimator, which is retained unchanged and remains valid under arbitrary within-cluster dependence, including negative dependence. The estimand H2 is stated in terms of — a population-averaged log-odds — is therefore identical to the pre-registered one. No alternative considered (respondent-level quasi-binomial or beta regression on the share) would have preserved §4's stated test as closely. **Disclosure — the ruling was made with the result visible.** The `Independence()` fit was computed *before* this deviation was authorised, as part of diagnosing the convergence failure, and was held under an explicit QUARANTINE label in the script output pending a ruling. The implementing agent did not substitute a model on its own initiative and reported no H2 confirmatory result. The ruling was then made by the project owner on 2026-07-26. This ordering is disclosed rather than concealed because it is the kind of thing a pre-registration exists to make visible. Two facts bound the risk it creates: the repair is the standard, most conservative option and is dictated by the failure mode rather than by the result; and **the amended model yields a null** (adjusted OR 0.9296, 95% CI 0.8563–1.0092, p = 0.0816), so the amendment cannot have been made to manufacture significance. Had the amended fit been significant, the honest course would have been to report it with this same disclosure attached and appropriately weakened confidence. **Consequence for H2.** H2 is **not supported**. The pre-specified direction (normal-BMI-high- symptom respondents showing a *lower* content-confounded share than obese-high-symptom respondents) is the direction observed, but the interval includes 1. Per §10 this is reported with the same prominence a positive result would have received. The unadjusted comparison (OR 0.8693, 95% CI 0.8014–0.9429, p = 0.00073) *is* significant, and the gap between adjusted and unadjusted is itself the finding §4's adjustment argument predicted — reported as such, and explicitly not substituted for the confirmatory result. **Power, recomputed against the real cells** (§4 committed to exact re-verification). The cells shrank after exclusions, from the 182/407 quoted in §4 to **155/394**: minimum detectable difference at 80% power is **12.68 pp** (was 11.96 pp), and power at a 12 pp difference is **0.7533** (was 0.8030). **H2 was therefore slightly under-powered relative to its pre-registered design**, which must be stated alongside the null rather than allowing the null to be read as a confident absence of effect. --- ### Deviation 2 — full record (criterion 7 measurement defect) **What was found, and when.** On 2026-07-26, *after* the pre-registered analysis had been run and reported, a defect was found in the scoring logic of the instrument that generated this dataset — `modules/calculators/food_addiction/controllers/Food_addiction.php`, the `$dsm` criterion-to-item map. Criterion 7 ("characteristic withdrawal symptoms") is defined as items `[1, 12, 13, 14, 15]`. It should be `[11, 12, 13, 14, 15]`. - **Item 1** — *"When I started to eat certain foods, I ate much more than planned"* — is criterion 1 content and is already correctly mapped to criterion 1. It is not a withdrawal item. It is the only item in the instrument mapped to two criteria. - **Item 11** — *"When I cut down on or stopped eating certain foods, I felt irritable, nervous or sad"* — is a withdrawal item, the affective counterpart to item 14's physical symptoms, and is mapped to **no criterion at all**. It is scored against a threshold and then discarded. - The two are not interchangeable even mechanically: item 1's endorsement threshold is 6, item 11's is 4. A single-digit transcription error (`11` written as `1`) accounts for both anomalies at once. **Magnitude.** A criterion is met if *any* one of its items crosses threshold, so item 1 alone sets criterion 7. Across the 6,105 frozen rows, 3,597 are criterion-7-positive and **686 of those (19.1%) are positive solely on item 1** — that is, solely on a non-withdrawal item. The rate of false *negatives* (respondents who would have met criterion 7 on item 11 alone) is **unknowable**, for the reason below. **Why the frozen data cannot be repaired.** Item responses are persisted only where their criterion was met (00-STATE §2). Item 11 belongs to no criterion, so its responses were **never written** for any respondent. The data required to rescore criterion 7 does not exist and never did. This is irreparable for the frozen cohort, not merely inconvenient. It also retrospectively explains the "item 11 never appears" observation recorded in 00-STATE §2, which was read at the time as a storage quirk rather than as the fingerprint of a scoring bug. **Consequences for the analysis.** Criterion 7 as measured here is a **contaminated indicator**: it is missing its affective withdrawal item and is contaminated with criterion-1 content. This matters disproportionately because C7 is one of the two counter-and-BH-significant findings (OR 1.085, q = 0.048, running against its own pharmacological core) and therefore load-bearing for the study's central caveat. `total_symptom_count`, and hence every `rest_score`, is also affected in both directions, so the defect touches the primary model and not only C7. **What is being done, and what is deliberately not.** The pre-registered primary model is **retained unchanged as the primary result**. A sensitivity analysis excluding C7 from the pharmacological core (leaving C6 and C11) is added and reported alongside it. **Declared in advance of running it:** C6 and C11 are both strongly negative (OR 0.912 and 0.870), so excluding C7 is **expected to strengthen H1**. That expectation is recorded here *before* the sensitivity result is seen, precisely so that a favourable outcome cannot later be presented as a discovery. It remains a sensitivity analysis whatever it returns. It will not be promoted to primary, and the pre-registered contrast of 1.0864 remains the headline figure. **Two variants, both specified before running, both reported regardless of what they show.** The defect contaminates C7 twice over: as an *indicator* (its own rows), and as a *contributor* to `total_symptom_count` and therefore to every other criterion's `rest_score`. The two variants separate those: - **S1 — set-membership only (the primary sensitivity).** Drop C7's rows from the stacked model so the pharmacological core becomes {C6, C11}. `rest_score` and `total_symptom_count` are left **exactly as in the primary model**, still computed over all eleven criteria. This changes one thing only, so it is directly comparable to the pre-registered contrast. - **S2 — full removal (secondary).** Drop C7 entirely and recompute `total_symptom_count` and every `rest_score` over the remaining ten criteria. This removes the contaminated variance from the conditioning variable as well, which is what the measurement concern strictly implies, at the cost of no longer being the same model conditioned differently. S1 is designated the primary sensitivity **because it is the more conservative of the two**, not because of anything either returns. If S1 and S2 disagree materially, that disagreement is itself reported; neither is selected on the basis of which is more favourable. **Unchanged by this deviation:** the extraction, the frozen snapshots, the freeze date, the primary model, all eleven per-criterion secondaries, and every figure already reported. The sensitivity is additive. Re-running the analysis script must reproduce every pre-existing number byte-for-byte, and that reproduction is treated as a required check on the amendment. **Also fixed, separately and forward-only.** The scoring bug itself is being corrected in the live calculator. No historical row is migrated, backfilled or rewritten; the frozen study input is untouched. Submissions after the fix date carry the corrected mapping and are therefore not directly comparable to those before it, which is recorded at the fix site in the code. --- ## 10. What gets reported regardless of outcome This analysis will be written up and published with the same prominence whether H1 is confirmed, disconfirmed, or produces a null/ambiguous result. The specific commitment: - If the confirmatory `BMI × criterion_set` interaction contrast (§2.2) is **not statistically significant** (or is significant in the opposite direction to H1), the resulting article's headline finding will be stated as: *"After accounting for overall food-addiction severity, BMI does not predict endorsement of body-size-adjacent criteria more strongly than it predicts endorsement of the pharmacological-core criteria in this sample — the well-documented YFAS–BMI relationship appears to be diffuse across the construct rather than concentrated in criteria whose content overlaps with the lived experience of a larger body."* This sentence, or a substantively equivalent one reflecting the actual direction and magnitude of a null/contrary result, will appear as the lead finding, not buried in a limitations section, and the article will not be reframed post hoc as being "really" about something else. - If H1 is confirmed, the article will state the finding with the calibrated hedging established in `01-literature-positioning.md` §3 — an association-pattern claim, not a mechanism claim, explicitly naming self-selection collider bias and reverse causation as alternative explanations the design cannot rule out. Under rest-score conditioning, a confirmed H1 is somewhat better evidence for the content-mechanism account than the unconditioned design would have supported, since severity has been explicitly partialled out — but "somewhat better evidence for" is not "established," and the article will not claim more than that. - The unconditioned models (§2.4), the IRT-DIF convergence check (§3), and the H2 discordant-phenotype result (§4) are all reported in full regardless of whether they align with or contradict the primary GEE finding. A convergence failure between the GEE and IRT-DIF arms is itself reported as a finding, not resolved by selectively emphasising whichever arm gives the more publishable answer. - H3 (age) and H4 (sex/C9 replication) are reported as stated in §1 regardless of outcome, including an explicit statement of whether the Saffari 2022 C9-gender finding replicates or fails to replicate in this sample. --- ## 11. Validity threats carried forward (calibrated, not deleted) Full detail in `01-literature-positioning.md` §3; summarised here with the calibration update from rest-score conditioning: - **Self-selection / help-seeking collider bias** remains the primary unresolved threat. Conditioning on rest-score does not address this — if both symptom severity and adiposity independently drive the probability of taking an online food-addiction quiz, the collider structure operates on the *joint* propensity to appear in this dataset at all, upstream of anything the within-sample regression can adjust for. This threat applies with equal force to the conditioned and unconditioned models and is not mitigated by the design changes in this SAP. - **Reverse causation** remains live for C4, C8, C9 specifically (role failure, interpersonal problems, activities given up could plausibly precede rather than follow weight gain). Rest-score conditioning does not distinguish forward from reverse causal direction — it isolates a content-specific association from a severity-driven one, but says nothing about which way any resulting content-specific association runs. - **Mechanism claim calibration, updated:** under the unconditioned design (Step 1's original proposal), the mechanism claim ("these criteria are confounded by lived experience of body size, not addictive process") was explicitly stated as untestable with this data — only an association-pattern claim was defensible. Under rest-score conditioning, a confirmed H1 result is **better evidence for** the content-mechanism account than an unconditioned finding would have been, because the conditioning rules out the simplest alternative explanation (that content-confounded criteria are just picking up general severity). It does **not** rule out self-selection collider bias or reverse causation, both of which remain fully live under the conditioned design. The calibrated statement for any resulting publication is: *this design can distinguish a content-specific BMI association from a severity-driven one; it cannot establish that the content-specific association, if found, reflects the proposed body-size-burden mechanism rather than reverse causation or selection effects.* This is a meaningfully stronger position than Step 1's original "cannot establish mechanism at all," but still short of a causal claim. - **under_18 respondents** are excluded per §7 step 1, not analysed under any circumstance, including sensitivity analyses. --- ## 12. Reporting standard **STROBE (cross-sectional)** is the primary reporting checklist for the study as a whole, as recommended in `01-literature-positioning.md` §5. **COSMIN** items relevant to measurement-properties reporting (sample size justification, DIF analysis method reporting) apply specifically to the IRT/DIF secondary arm (§3) and should be used to structure that subsection's write-up, without treating the whole study as a COSMIN measurement-validation paper — it remains a STROBE cross-sectional study with an embedded psychometric convergence check. --- ## Summary of what is locked by this document 1. Primary confirmatory test: one GEE interaction contrast (`BMI:C(criterion_set)`, content-confounded vs pharmacological-core), rest-score conditioned, with criterion-specific intercepts via `C(criterion_id)`. 2. Secondary confirmatory test: H2 discordant-phenotype criterion-composition comparison (n=182 vs n=407), GEE on pooled endorsed-criterion binary outcome. 3. Convergent-methods check: IRT-DIF for BMI, Chapron-method-modelled, Python/statsmodels-compatible implementation only. 4. Exploratory: H3 (age, no direction), H4 (sex, C9-specific replication attempt only). 5. Exclusions: under_18 (n=99 of 1,898 → analytic n=1,799 before further screens), BMI implausibility (<12 or >80, recomputed BMI), non-imputed complete-case cohort, sex `other`/`prefer_not_to_say` excluded from primary model only, no deduplication. 6. Data freeze: `date_create <= '2026-07-25 23:59:59'`, hard-coded in script, boundary deliberately in the past at commit time. 7. Multiplicity: BH-FDR q=0.05 on the 11 secondary per-criterion coefficients; the primary confirmatory test is a single contrast, not subject to multiplicity correction. 8. Environment: Python, statsmodels, WSL venv. No R. No PHP `StatsCalculator`. 9. Reporting commitment: null/contrary H1 result published with equal prominence, specific sentence pre-written in §10. 10. Deviations log open and empty; any departure from this plan recorded there with date and reason.