Skip to the main content

We Tested Our Own Food Addiction Quiz Against BMI, And One Of Our Predictions Ran Backwards

This is about our own data, not someone else's study. For our free food addiction quiz on this site, we set out to test a specific, written-down prediction about how it relates to body weight. One of the specific criteria we'd predicted would carry the link ran significantly the opposite way instead. If you've taken this quiz yourself, this is written with you in the room.

Does BMI predict all 11 food addiction criteria? A stylised forest plot shows most criteria moving together, with one highlighted result running in the opposite direction.

If you’ve ever taken the Yale Scale questionnaire, [1] a 35-item questionnaire whose answers are scored into eleven symptom criteria, asking things like whether you’ve tried and failed to cut down on or stop eating certain foods, or whether you’ve had problems with family or friends because of how much you overate, there’s a reasonable assumption sitting underneath it. It is a research and screening instrument rather than a diagnostic one, and “food addiction” is not itself a standalone diagnosis in the DSM framework the scale’s eleven criteria were adapted from, so a symptom count is a description of what you endorsed, not a verdict on what you have. The assumption is that the test measures one thing: an addictive relationship with food, the same way a substance-use questionnaire measures an addictive relationship with alcohol or nicotine. Score higher, the logic goes, and you have more of that thing. And because people in larger bodies tend to score higher on average, it’s tempting to read the whole scale as a single, uniform signal of “how addicted are you to food,” with body weight simply riding along as the visible consequence.

We wanted to know whether that assumption holds up criterion by criterion, not just as a total score. Specifically: does body mass index (BMI) predict all eleven parts of this questionnaire equally, or does it load more heavily onto the questions whose content is already entangled with what it’s like to live in a larger body: your family or friends being worried about how much you eat, giving up other important activities because of your eating, or continuing to eat foods you knew were physically dangerous given a health condition you already had? We pre-specified which five criteria would carry more of that signal and which three wouldn’t, before we ran a single test.

The short answer is: the link is real, but it isn’t where we expected it to be, not entirely. Our headline test passed. But the criterion that was supposed to be one of the clearest examples of the pattern we predicted, the one about giving up activities because of eating, ran significantly in the opposite direction. We think that’s more interesting, and more useful to report, than a clean confirmation would have been, and this piece walks through why.


How we tested this

This was run as a pre-specified analysis, not an exploratory dig, though we want to be precise about what that means here. Pre-registration, in the strict sense used in clinical trials, means depositing an analysis plan with an independent third party, such as a public registry, before the analysis is run, so that someone other than the researchers can vouch for the timing. We didn’t do that, and we’re not going to describe this study as pre-registered. What we did instead: we wrote the full statistical analysis plan, including the hypothesis, the exact statistical model and the rule for handling multiple comparisons, and committed it to version control before any model was fitted or any p-value inspected. That plan is published in full, as it was written, with its deviations log and without later tidying.

The analysis itself was run afterwards, in a separate committed script, against a frozen sample: every response submitted to Creative Touch’s free online food addiction quiz up to 25 July 2026, with no later submissions able to affect the result and no possibility of the freeze date being adjusted after the fact. The ordering is genuine. But the repository was private, with no external remote, and a commit timestamp is something the person making the commit controls, not an independent record. We’d use a public registry next time. Until then, take this as the discipline it is, not the credential it isn’t.

After excluding respondents under 18 (a bright-line research-ethics decision, not a statistical one; an unsupervised online quiz has no route to parental consent), screening out a handful of biologically implausible height/weight combinations, and requiring BMI, age band and sex to all be recorded, the descriptive analytic sample was 1,796 respondents. The primary confirmatory model, which additionally excludes the small number of respondents recording sex as “other” or “prefer not to say” for power reasons, ran on 1,773 respondents (14,184 stacked rows, since each respondent contributes one row per criterion in the primary design). Roughly 77% of respondents were female, and the sample skewed toward the US and UK. This is a self-selected group of people who chose to take a food addiction quiz online, not a population sample, and that matters for how far the findings generalise (see limitations, below).

How the analytic sample was arrived atExclusion flow diagram. 6,105 submissions up to the freeze date. 4,207 excluded for not recording BMI, age band and sex, leaving 1,898; none were removed by the integrity guard. 99 under-18s excluded, leaving 1,799. 3 implausible height and weight combinations excluded, leaving the descriptive analytic sample of 1,796. A further 23 respondents recording sex as other or prefer not to say were excluded from the primary model only, leaving 1,773.n = 6,105Quiz submissions up to the freeze date−0unreadable stored result−4,207one or more of the three missing−0failed the integrity guardn = 1,898BMI, age band and sex all recorded−99under 18, on ethics groundsn = 1,799Aged 18 or over−3implausible recomputed BMIn = 1,796Height and weight biologically plausibledescriptive analytic sample−23other / prefer not to sayn = 1,773Sex recorded as female or maleprimary confirmatory modelFreeze clause: module = ‘food_addiction’ AND date_create <= ‘2026-07-25 23:59:59’
The freeze was a genuine no-op rather than a convenient cut: the latest submission in the table was 2026-07-25, before the boundary the analysis plan had already committed to. Two of the checks removed nobody, and both are shown rather than quietly dropped, because a stated check that turns out to bind on no one is still evidence the check was applied. The under-18 exclusion is a bright line taken on research-ethics grounds, not a sensitivity choice: an unsupervised online quiz has no route to parental consent. The final step applies to the primary confirmatory model only; every descriptive and secondary figure in this article uses the full 1,796.

Why we couldn’t just compare the eleven questions directly

Before the result, it’s worth explaining why the obvious approach, checking whether BMI predicts each of the eleven criteria and then seeing which ones show the strongest link, would have been close to meaningless. This is the single most useful thing this study has to offer anyone reading a health questionnaire critically, food-related or otherwise, so it’s worth sitting with.

Every one of the eleven criteria is a component of the same total score, and that total score is already known to track body weight. Creative Touch’s own live analytics show mean symptom count rising from 3.38 at a normal BMI to 6.85 in the highest band. One caveat on that figure, which we explain in full further down: that total score includes criterion 7, and criterion 7 turns out to have been measured incorrectly for reasons unconnected to this specific analysis. Roughly a fifth of its positives come from a non-withdrawal item that’s itself strongly linked to BMI, so this gradient is very likely slightly overstated, in that specific direction, by an amount we haven’t attempted to quantify. It is not in doubt as a gradient. One criterion out of eleven, partly affected, doesn’t produce or reverse a rise from 3.38 to 6.85.

Run all eleven criteria against BMI on their own, without any adjustment, and the picture looks dramatic: all eleven show a positive association with BMI, and all eleven are statistically significant at p < 0.05, several with p-values so small they carry more than twenty zeroes after the decimal point. Criterion 2 (persistent desire to cut down, or repeated failed attempts to do so) shows an unconditioned odds ratio of 1.614 per 5 BMI units; criterion 10 (use in physically hazardous situations, such as continuing to eat foods known to be dangerous given an existing health condition, or being so distracted by food that you could have been hurt, for example whilst driving) shows 1.442; criterion 1 (eating larger amounts or for longer than intended) shows 1.399, at p = 3.4 × 10⁻²², about as small as p-values get in social science data.

Mean food addiction symptom count by BMI category, across the whole responding populationBar chart. Mean symptom count out of a possible eleven rises from 3.38 in the normal BMI band to 6.85 in the highest obesity band. The underweight band sits at 3.89, above normal, on only 92 respondents. Overweight 4.60, obese class one 5.73, obese class two 6.39. The proportion meeting the clinical significance criterion rises alongside it, from 36.3 per cent at normal BMI to 65.8 per cent at the highest band.mean number of criteria endorsed, out of a possible 110246810clinicallysignificantUnderweightBMI under 18.5 · n = 923.8940.2%NormalBMI 18.5 to 25 · n = 1,1143.3836.3%OverweightBMI 25 to 30 · n = 8244.6043.8%Obese IBMI 30 to 35 · n = 5005.7357.8%Obese IIBMI 35 to 40 · n = 2726.3965.1%Obese IIIBMI 40 and over · n = 3366.8565.8%Creative Touch live analytics, 3,138 respondents with a BMI recorded, as at 2026-07-26.Whole responding population, not the frozen study cohort.
This is the well-established gradient the rest of the article is built against, and it is not in dispute: more symptoms are endorsed, on average, further up the BMI range. Two qualifications. The total on the horizontal axis includes criterion 7, whose measurement was defective throughout, and roughly a fifth of its positives came from a non-withdrawal item that is itself linked to BMI, so this gradient is very likely slightly overstated by an amount we have not tried to quantify. One criterion out of eleven, partly affected, does not produce or reverse a rise from 3.38 to 6.85. And the underweight band sits above normal on 92 respondents, which is a reminder that the relationship is not a clean straight line, a point the article returns to in the secondary checks. These are the site’s live analytics across everyone who has taken the quiz, deliberately not the frozen study cohort, because this figure describes the background the study set out to interrogate.

That looks like overwhelming, uniform evidence. It isn’t evidence of anything specific at all. Every criterion shows a positive BMI association for the trivial reason that BMI predicts overall symptom severity, and overall severity mechanically drags every one of its own components upward. Comparing eleven numbers that are all inflated by the same underlying mechanism doesn’t tell you which criteria are distinctively linked to body weight. It just re-proves, eleven times over, something we already knew from the total score.

So we conditioned. For each criterion, we built a model that asks a sharper question: among people who are equally symptomatic overall, measured as their total score on the other ten criteria, so the outcome is never inside its own predictor, does BMI still predict whether they endorse this particular one? That single adjustment changes the picture completely. Once overall severity is held constant, criterion 1’s odds ratio collapses from 1.399 (p = 3.4 × 10⁻²²) to 1.014 (p = 0.71). An association with a p-value carrying twenty-one zeroes after the decimal point evaporates entirely, not because the earlier finding was wrong or fraudulent, but because it was answering a different question than the one that matters. Comparable shifts, in the region of +0.25 to +0.42 in odds-ratio terms, appear across all eleven criteria once conditioning is applied. This is the circularity problem our statistical plan was designed around, demonstrated with real numbers rather than just argued for on paper.

Every criterion’s BMI odds ratio before and after holding overall severity constantSlope chart. On the left, all eleven criteria show unconditioned odds ratios between 1.24 and 1.61, every one above 1 and every one statistically significant. On the right, after conditioning on the respondent’s score on the other ten criteria, the same eleven spread from 0.87 to 1.21 and straddle 1. Criterion 1 falls furthest, from 1.399 to 1.014. Full values are in the accompanying data table.0.911.11.21.31.41.51.6no assoc.UnconditionedBMI vs each criterion aloneConditionedamong people equally symptomatic overallC2 failed cut-down 1.20C10 hazardous use 1.19C5 use despite harm 1.11C7 withdrawal 1.09C8 interpersonal 1.08C3 time spent 1.08C1 larger amounts 1.01C9 role failure 1.00C6 tolerance 0.91C4 activities given up 0.88C11 craving 0.87Content-confoundedPharmacological coreNeither set
Show these figures as a table
CriterionUnconditioned ORConditioned ORShift
C2 failed cut-down1.6141.205−0.410
C5 use despite harm1.5231.112−0.412
C7 withdrawal1.5011.085−0.416
C10 hazardous use1.4421.187−0.255
C3 time spent1.4331.075−0.358
C8 interpersonal1.4051.084−0.321
C1 larger amounts1.3991.014−0.385
C9 role failure1.3120.998−0.314
C6 tolerance1.2860.912−0.374
C11 craving1.2690.870−0.399
C4 activities given up1.2440.875−0.368
Left: every one of the eleven criteria shows a positive, statistically significant association with BMI when tested on its own. That looks like overwhelming evidence and is close to meaningless, because BMI predicts overall severity and overall severity mechanically drags each of its own components upward. Right: the same eleven criteria after asking the sharper question, whether BMI still predicts endorsing this criterion among people equally symptomatic overall. The uniform picture disappears and the criteria spread out on both sides of 1. Criterion 1 (bold) falls furthest, from 1.399 at p = 3.4 × 10⁻²² to 1.014 at p = 0.71. n = 1,796.

The headline result, and the criterion that broke rank

With conditioning in place, our primary, pre-specified test compared two groups of criteria. The first, the content-confounded set, covers questions whose wording plausibly overlaps with the lived experience of being in a larger body, independent of any addictive process: giving up activities because of eating (criterion 4), continuing to eat despite knowing it was causing harm (criterion 5), interpersonal problems caused by how much you eat (criterion 8), failing to fulfil role obligations because of overeating (criterion 9), and hazardous use (criterion 10). The second, the pharmacological core, covers tolerance (criterion 6), withdrawal (criterion 7) and craving (criterion 11): the criteria that map most directly onto the biological mechanics of addiction, with the least obvious overlap with body-size stigma or its social consequences.

The eleven criteria, their hypothesised set, and how each behaved once overall severity was held constant
CriterionEndorsed byOdds ratio per 5 BMI units (95% CI)q
Content-confounded (hypothesised stronger BMI association)
4Important activities given up because of eating33.2%0.875 (0.812–0.944)0.0014
5Continued eating despite knowing it was causing harm50.8%1.112 (1.031–1.199)0.0132
8Interpersonal problems caused by how much you eat28.6%1.084 (0.997–1.178)0.0794
9Failure to fulfil role obligations because of overeating28.9%0.998 (0.922–1.079)0.9519
10Hazardous use (eating when it was physically dangerous)33.9%1.187 (1.087–1.296)0.0005
Pharmacological core (hypothesised weaker)
6Tolerance (needing more for the same effect)34.5%0.912 (0.846–0.983)0.0288
Withdrawal56.4%1.085 (1.008–1.169)0.0481
11Craving42.8%0.870 (0.812–0.933)0.0005
Neither set (reported for transparency, not adjudicating the hypothesis)
1Ate larger amounts, or for longer, than intended53.3%1.014 (0.940–1.095)0.7852
2Persistent desire, or repeated failed attempts, to cut down56.9%1.205 (1.112–1.306)< 0.0001
3A great deal of time spent obtaining, eating or recovering44.3%1.075 (0.995–1.161)0.0814
Odds ratios above 1 mean higher BMI predicted endorsing that criterion more often among respondents of equal overall severity; below 1, less often. This is the first table to use q, so: a q value is a p value adjusted to control the share of false positives expected across a whole family of tests, here the eleven criteria tested at once (Benjamini–Hochberg, q = 0.05). It is the figure to read rather than the raw p. Criterion 4 is shown in bold because it ran significantly counter to its own hypothesised set. † Criterion 7’s measurement was defective throughout the data collection period, so its row is not clean evidence. n = 1,796.

Our prediction was that BMI would predict the content-confounded criteria more strongly than the pharmacological-core criteria, even after conditioning on overall severity. That prediction held, under the primary model exactly as it was written down in advance. The content-confounded slope came out at OR 1.0608 per 5 BMI units; the pharmacological-core slope at OR 0.9764. The confirmatory test, a single, pre-specified contrast between the two, with no multiplicity correction because it was the one test the whole design was built around, gave a ratio of odds ratios of 1.0864 (95% CI 1.0253–1.1510, z = 2.8086, p = 0.0050), on n = 1,773 respondents across 14,184 stacked rows. One test, one p-value, exactly as specified in advance. H1 passed.

That is the headline. Here is the part that has to sit in the same breath as it, not several paragraphs later: the pattern underneath that headline number is considerably messier than that single number suggests, and one of the five criteria we predicted would run one way ran significantly the other way instead. Criterion 4, activities given up because of eating, was one of our clearest hypothesised examples of a content-confounded criterion. Looked at on its own fitted slope, rather than folded into the set-level average, it showed an odds ratio of 0.875 (95% CI 0.812–0.944, q = 0.0014 after correcting for multiple comparisons), meaning higher BMI was significantly less likely to predict endorsing this criterion at equal overall severity, not more. Criterion 7 (withdrawal), on the pharmacological-core side, also ran counter to its own set’s direction: OR 1.085 (95% CI 1.008–1.169, q = 0.048). We have to flag that number immediately rather than let it sit unqualified: after this analysis was run, we found that criterion 7 had been measured incorrectly for the calculator’s entire seventeen-month life. That result can no longer stand as clean evidence, for reasons the next section explains in full, and criterion 4 is now this study’s central counter-finding. Criterion 9 (role failure) also pointed the “wrong” way but was statistically indistinguishable from no effect at all (OR 0.998, q = 0.952).

Add it up and the set-level averages hide real internal disagreement: only 3 of the 5 content-confounded criteria sit above an odds ratio of 1 (spread across the set: 0.875 to 1.187), and only 1 of the 3 pharmacological-core criteria does (spread: 0.870 to 1.085, though that upper figure is criterion 7’s compromised result, not a clean one, for reasons the next section sets out). Much of the overall contrast is being pulled by criteria 6 and 11 (tolerance and craving) dragging the “core” side down, rather than by the content-confounded criteria uniformly rising. A summary that reports 1.0864 and stops there would leave a reader believing every content-confounded criterion behaved as predicted. Only three of the five did, and criterion 4 broke rank at a level that survives correction for multiple testing. That nuance is the actual finding here, not an asterisk on it.

Odds ratio for endorsing each criterion per 5 BMI units, holding overall severity constantForest plot of eleven criteria ordered by effect size. The five content-confounded criteria and three pharmacological-core criteria are shown together at the top, and they interleave rather than separating cleanly. Criterion 10 hazardous use is highest at 1.187 and criterion 11 craving lowest at 0.870. Criterion 4, activities given up, was hypothesised to sit high but sits at 0.875, significantly below 1. The three criteria in neither set are shown separately below. Full values are in the accompanying data table.0.80.911.11.21.3odds ratio per 5 BMI units (log scale)no associationC10 hazardous use1.187(1.087–1.296)C5 use despite harm1.112(1.031–1.199)C7 withdrawal †1.085(1.008–1.169)C8 interpersonal1.084(0.997–1.178)C9 role failure0.998(0.922–1.079)C6 tolerance0.912(0.846–0.983)C4 activities given up0.875(0.812–0.944)C11 craving0.870(0.812–0.933)C2 failed cut-down1.205(1.112–1.306)C3 time spent1.075(0.995–1.161)C1 larger amounts1.014(0.940–1.095)reported for transparency, not part of either hypothesised setContent-confounded set slope1.0608Pharmacological-core set slope0.9764Content-confoundedPharmacological coreNeither set
Show these figures as a table
CriterionSetOR per 5 BMI units95% CIq (BH)
C2 failed cut-downNeither set1.2051.112–1.306< 0.0001
C10 hazardous useContent-confounded1.1871.087–1.2960.0005
C5 use despite harmContent-confounded1.1121.031–1.1990.0132
C7 withdrawalPharmacological core1.0851.008–1.1690.0481
C8 interpersonalContent-confounded1.0840.997–1.1780.0794
C3 time spentNeither set1.0750.995–1.1610.0814
C1 larger amountsNeither set1.0140.940–1.0950.7852
C9 role failureContent-confounded0.9980.922–1.0790.9519
C6 tolerancePharmacological core0.9120.846–0.9830.0288
C4 activities given upContent-confounded0.8750.812–0.9440.0014
C11 cravingPharmacological core0.8700.812–0.9330.0005
Each criterion’s own fitted slope, ordered by effect size rather than by criterion number, so the sets can be seen to interleave. Neither set is homogeneous in the direction of its own average. Only 3 of the 5 content-confounded criteria sit above 1, and only 1 of the 3 pharmacological-core criteria does. Criterion 4 (ringed) was hypothesised to sit near the top and instead sits significantly below 1. The confirmatory test compared the two set-level slopes shown at the foot: ratio of odds ratios 1.0864 (95% CI 1.0253–1.1510, p = 0.0050). That ratio is on a different scale from the axis above and is deliberately not plotted on it. † Criterion 7’s result is compromised by the measurement defect described below. n = 1,773 respondents, 14,184 stacked rows.

One more result worth flagging on its own terms: criterion 2, persistent desire to cut down, or repeated unsuccessful attempts, turned out to be the single strongest BMI association in the entire study (OR 1.205, 95% CI 1.112–1.306, q = 0.0001), and it wasn’t part of either hypothesised set. It sat in a “neither” category, reported for transparency but never intended to adjudicate the main hypothesis. It’s a genuinely interesting number, and it’s worth being precise about what kind of evidence it is rather than waving it away. The test itself was pre-specified and multiplicity-corrected: criterion 2’s model is one of the eleven in the secondary per-criterion analysis, fitted exactly as planned and carried through the same Benjamini–Hochberg correction as the rest. What is post-hoc is singling it out, which we are only doing because it turned out to have the largest estimate. That makes its prominence here exploratory, not a separately hypothesised finding, and it is not a result we predicted or designed the study to detect.

We found a measurement defect in our own tool, and here’s exactly what it does to this result

Before you read any further, you need to know something that came to light after every number above had already been calculated and reported: the calculator that produced this dataset had been mis-scoring criterion 7, withdrawal, for its entire life.

Here’s what went wrong. The calculator’s scoring code assigns each of the eleven DSM-5-adapted criteria to a specific set of the underlying questionnaire’s items, and a criterion counts as met if any one of its assigned items crosses its own threshold. Criterion 7 should have been built from five items, one of which is the questionnaire’s genuine withdrawal item: “When I cut down on or stopped eating certain foods, I felt irritable, nervous or sad.” Instead, a single mistyped digit substituted a different item entirely: one that already belongs to criterion 1 and reads “When I started to eat certain foods, I ate much more than planned.” That’s not a withdrawal question. It’s a question about amount and duration, and it was already being counted correctly under criterion 1. Meanwhile, the real withdrawal item was scored against its own threshold and then, because of the same slipped digit, filed under no criterion at all and quietly discarded.

How criterion 7 was mis-wired, and how it is wired nowTwo panels. In the faulty version, item 1 (ate much more than planned) feeds both criterion 1 and criterion 7, and item 11 (felt irritable, nervous or sad when cutting down), the genuine withdrawal item, feeds no criterion at all and is discarded. In the corrected version, item 1 feeds only criterion 1 and item 11 feeds criterion 7 alongside items 12 to 15.How it was scored23 February 2025 to 27 July 2026Item 1“ate much more than planned”Item 11“felt irritable, nervous or sad”Items 12–15other withdrawal itemsCriterion 1larger amounts / longerCriterion 7withdrawalalso counted as criterion 7in no criterion at all, so never storedHow it is scored nowcorrected 27 July 2026Item 1“ate much more than planned”Item 11“felt irritable, nervous or sad”Items 12–15other withdrawal itemsCriterion 1larger amounts / longerCriterion 7withdrawalnow counted as criterion 7A single mistyped digit, 11 written as 1, put a criterion-1 item into criterion 7and left the real withdrawal item in no criterion at all.
The scoring map assigns each criterion a set of questionnaire items, and a criterion counts as met if any one of its items crosses its own threshold. Criterion 7 should have drawn on the genuine withdrawal item (item 11). Instead it drew on item 1, which asks about amount and duration and was already counted correctly under criterion 1, whilst item 11 was scored against its threshold and then filed under no criterion and discarded. Because the calculator only stores a respondent’s item answers when the criterion they belong to was met, item 11’s answers were never written to the database for anyone. That is why criterion 7 cannot be rescored retrospectively: the data needed to do it was never captured. 191 of the 1,012 respondents counted as endorsing criterion 7 (18.9%) qualified solely through the misplaced item.

This has been happening since the calculator first went live on 23 February 2025, so it touches every one of the 6,105 responses in the full dataset, and, by extension, every one of the 1,796 respondents in this study. Seventeen months of results, all of them.

Here’s the size of it. In our analytic cohort, 1,012 of 1,796 respondents (56.3%) endorsed criterion 7. Of those, 191 (18.9%) were positive solely because of the misplaced item, with no genuine withdrawal content contributing anything at all (this closely matches the 686 of 3,597, 19.1%, seen across the whole frozen dataset before our study-specific exclusions). Put another way: roughly one in five of the people this study counts as showing withdrawal symptoms were flagged on a question that has nothing to do with withdrawal.

That 18.9% is not a false-positive rate, and we should not let it be read as one. It measures how many recorded criterion-7 positives depended on the misplaced item among the answers that were kept. Because the genuine withdrawal item was never stored for anyone, we cannot know how many of those 191 people would have met criterion 7 anyway, correctly, through the question that was discarded. Some of them almost certainly would have. The right description is that 18.9% of criterion-7 positives were dependent on the faulty mapping, which is a statement about the scoring, not a count of people wrongly classified.

That’s the contamination we can measure. The false-negative rate, respondents who genuinely felt irritable, nervous or sad when cutting down but weren’t captured because the real withdrawal item was never assigned to any criterion in the first place, is something we can’t measure, and never will be able to. The calculator only stores a respondent’s answers to a criterion’s items when that criterion is met. Because the genuine withdrawal item was never mapped to any criterion, its answers were never written to the database for anyone, at any point in the tool’s history. There is no archived data to go back and rescore. This isn’t a gap better analysis can close later. It’s gone.

What this means for the finding you just read: criterion 7 was one of only two criteria running significantly against its own hypothesised set’s direction, alongside criterion 4. We can no longer treat that as two clean counter-findings. Criterion 4 was scored correctly throughout, untouched by this defect, and its result stands exactly as reported: it remains this study’s central counter-finding. Criterion 7’s result is now, at best, ambiguous, and quite possibly telling us something about a mislabelled question rather than about withdrawal itself. From here on, when we say this study found a criterion running the “wrong” way, we mean criterion 4.

We logged this the moment we found it, in the study’s deviations record (statistical analysis plan, §9, Deviation 2), and we ran a sensitivity check to see what removing the contaminated criterion does to the study’s one confirmatory result. That check was added after the fact, and we need to be exact about what was and wasn’t fixed in advance, because this is precisely the distinction the rest of the article asks you to hold us to. The sensitivity analysis was not in the original plan; it could not have been, since the defect had not been discovered when the plan was written. It is recorded as post-hoc and non-confirmatory in the results file itself. What was settled before either version was fitted: which of the two would count as the primary sensitivity result, and the direction we expected the estimate to move. That ordering is worth something, but it is not pre-registration, and we are not going to call it that.

Two versions: one that simply drops criterion 7 from the pharmacological-core comparison group (n = 1,773 respondents, 12,411 rows), and one that also recalculates every respondent’s severity adjustment to exclude criterion 7 entirely (same n). Both come out higher than our pre-specified figure: 1.1440 (95% CI 1.0752–1.2171, p = 2.09 × 10⁻⁵) and 1.1455 (95% CI 1.0750–1.2206, p = 2.79 × 10⁻⁵), against the pre-specified 1.0864 (95% CI 1.0253–1.1510, p = 0.0050).

Two things about those numbers, and both need care. First, this is not evidence for our hypothesis, and we want to be direct about that because it would be easy to misread it that way. We declared, in writing, before running either version, that removing criterion 7 should push the contrast higher: criterion 7 was the only member of the pharmacological core whose own slope sat above 1, so dropping it from that side of the comparison mechanically removes the one thing that was pulling the “core” average down. Getting a higher number after doing that isn’t a discovery. It’s exactly what we said would happen, for a reason that has nothing to do with whether the underlying hypothesis is true. The pre-specified figure of 1.0864 remains this study’s headline number. We are not promoting the sensitivity analysis above it, and neither should you.

Second, the fact that both versions land close together, 1.1440 and 1.1455, a difference of 0.0015 despite more than half the sample having a recalculated adjustment term in the second version, might look like proof the defect didn’t really matter. It isn’t. Both versions simply remove criterion 7 from the analysis. Neither one measures a correctly-scored criterion 7 in its place, because that data doesn’t exist and never will. Agreement between two ways of excluding something tells you the exclusion is stable. It doesn’t tell you what a properly measured criterion 7 would have shown.

The live calculator has been fixed. It now maps criterion 7 to the correct items, including the genuine withdrawal question, and the same correction has been made everywhere the old mapping was duplicated. No historical response has been changed or re-scored: this dataset, and this study, stand exactly as they were collected, defect included, because that’s the only honest way to report on data that’s already been analysed and frozen. Anyone comparing criterion 7’s prevalence before and after the fix date should expect a step change in the numbers that has nothing to do with anything real happening in the population, and everything to do with the tool finally asking the right question.

We’re not raising any of this to soften what we found. Criterion 4’s result doesn’t need criterion 7 to stand on. It stands on its own.

A pattern that’s tempting, and that we’re not going to pretend fits

There’s an alternative way to group these eleven criteria that looks neater than the one we specified in advance: split them into things other people can see, or that involve a history of trying and failing to control eating (failed cut-down, hazardous use, use despite harm, interpersonal problems) versus private internal states nobody else witnesses (craving, tolerance). It’s a satisfying story, and if we’d built the analysis around it after seeing the results, it would look like it explained the data rather well.

It doesn’t hold up, and criterion 4 is enough to show why: activities given up is about as socially visible a criterion as this scale contains, and it sits at the bottom of the association rather than near the top. We’d ordinarily also point to criterion 7, withdrawal, as a second example: a private internal state that sits near the top of the association, the opposite pairing, which would break this alternative story a second way. We can’t use it here. As the previous section explains, criterion 7’s own measurement is compromised, so its position in this ranking isn’t clean evidence for or against anything. One clear counter-example is still enough. Any grouping constructed after the fact to fit eleven data points can always be found. That is precisely why we committed to a specific grouping in advance, rather than sorting the criteria once we already knew the answer. We’re raising this alternative pattern here, explicitly, as a question worth testing properly in a future study with its own plan fixed in advance, not as a finding of this one, and not as a replacement for the hypothesis that only partly held up.

Does a normal-weight person with a high symptom count look different from an obese person with the same count?

Our second pre-specified question was narrower and more exploratory: among people who score highly on the questionnaire overall, does the specific pattern of which criteria they endorse differ depending on whether they’re at a normal BMI or classed as obese? The idea was that a normal-weight person with a high symptom count might lean more towards the pharmacological-core criteria (craving, tolerance, withdrawal), while an obese person with an equally high count might lean more towards the content-confounded criteria.

We could not detect a difference, and the honest reason is that the study wasn’t quite large enough to rule one out. After adjusting for each group’s overall symptom load (necessary, since both groups were defined by a floor of six or more symptoms, but the obese group sits higher within that range on average, 8.69 total symptoms versus 8.36), the adjusted comparison came out at OR 0.9296 (95% CI 0.8563–1.0092, p = 0.0816). The direction matched what we predicted, but the confidence interval crosses 1, so it does not meet our pre-specified threshold for support. This is an inconclusive null, not evidence of no difference. The unadjusted, purely descriptive version of the same comparison was significant (OR 0.8693, p = 0.00073). The gap between the adjusted and unadjusted numbers is itself informative, and is exactly what the adjustment was built to catch: some of that raw difference was simply the two groups’ differing overall severity, not a true difference in criterion profile.

Two things sit behind that null and deserve disclosure rather than a footnote. First, the comparison groups shrank between when we wrote the plan and when we ran the analysis, from the planned 182 normal-weight/407 obese respondents to the real cohort’s 155 and 394, which means our achieved statistical power against a 12-percentage-point difference was 0.7533, below the 0.80 the original design assumed. A study built to detect an effect at 80% power and running at 75% is meaningfully less able to rule a real effect out, and that should sit alongside the null result, not be discovered later by a sceptical reader doing their own maths.

What adjusting for overall symptom severity did to the phenotype comparisonTwo odds ratios on one axis with a reference line at 1. The descriptive, unadjusted contrast is 0.8693, 95% confidence interval 0.8014 to 0.9429, which does not include 1. The pre-specified adjusted contrast is 0.9296, 95% confidence interval 0.8563 to 1.0092, which does include 1. Adjustment moved the estimate towards 1 and widened the interval across it, so the result is an inconclusive null. The comparison groups also shrank from a planned 182 and 407 to an actual 155 and 394, taking power against a 12 percentage point difference from 0.8030 to 0.7533.0.800.850.900.951.001.05no differenceodds ratio, normal-BMI versus obese, for endorsing a content-confounded criterion (log scale)Descriptive only  not the confirmatory result0.869395% CI 0.8014–0.9429, p = 0.00073Pre-specified and adjusted0.929695% CI 0.8563–1.0092, p = 0.0816adjustment moved the estimateand widened the interval across 1Comparison groupsplanned 182 normal / 407 obeseactual 155 normal / 394 obesePower at a 12-point differencedesigned 0.8030achieved 0.7533Smallest difference detectableplanned 11.96 ppactual 12.68 ppModel departure, disclosed: the pre-registered exchangeable working correlation did not converge on thiscompositional outcome; it was changed to independence, robust standard errors unchanged. See below.
Show the underlying values
ContrastOdds ratio95% CIpWorking correlation
Unadjusted, descriptive only0.86930.8014–0.94290.00073Independence
Adjusted for total symptom count, age and sex (pre-specified)0.92960.8563–1.00920.0816Independence
Both intervals answer the same question; only the adjusted one answers it fairly. Because both groups were defined by a floor of six or more symptoms but the obese group sits higher within that range on average (8.69 symptoms against 8.36), part of the raw difference was simply that severity gap rather than any difference in criterion profile. Removing it moved the estimate from 0.8693 to 0.9296 and carried the interval across 1. That is an inconclusive null, not evidence the two groups are alike. The band beneath says why the distinction matters: the comparison groups shrank between writing the plan and running it, so the study reached 0.7533 power against a 12-point difference rather than the 0.80 it was designed for. A study running under its intended power is less able to rule an effect out, and that belongs beside the null rather than in a reader’s own arithmetic. n = 549 respondents, 3,231 endorsed criteria.

Second, we have to be transparent about a change we made to the statistical model itself, after seeing a result. The plan we committed in advance specified a particular statistical structure (an “exchangeable” working correlation, in the technical language of the method) for this comparison, mirrored from our primary model. That structure simply would not fit the data. The model failed to converge. The reason turned out to be structural: the outcome here is a proportion of a respondent’s own endorsed criteria, so it’s mathematically compositional, and dependence within a person’s own answers runs negative rather than positive, which broke the assumption the original plan had carried over from the primary analysis by analogy.

The fix, switching to a simpler, more conservative statistical structure (“independence,” in the same language, whilst keeping every other part of the model, including the robust standard errors, identical), is standard practice for this kind of failure, and it does not change what the estimate means. But in the interests of full disclosure: the fix was decided after the failed model, and the fitted result of the fix, had already been seen internally, under an explicit quarantine label pending a ruling from the project owner. We are naming that sequence because committing to a plan in advance is only meaningful if departures from it are disclosed exactly as they happened, not smoothed over. Two facts limit the risk this creates: the fix is the standard, most conservative repair for this specific failure mode, dictated by the mathematics rather than chosen to flatter a result, and the fixed model produced a null, which is not the kind of outcome anyone would engineer a deviation to reach.

What else we checked, and didn’t chase

A handful of secondary checks, all pre-specified, came back with results that don’t tidily support or undermine the headline finding, and per our own pre-specified reporting commitment, we’re printing all of them rather than only the ones that flatter the story.

Every pre-specified secondary check, what it returned, and what we did about it
CheckWhat came backNumbersStatusHow we’ve treated it
Is BMI’s relationship to these criteria a straight line?No. All three curvature tests came back significant.quadratic χ² = 8.00, p = 0.0047
set-specific p = 0.019
4-knot spline p = 0.0079
Reported, not promotedThe primary model stays linear because that specification was locked before any of this was tested. Both halves belong in the same sentence.
Does a different statistical method agree?Only partly. An item-response-theory check agreed on 6 of the 11 criteria.6/11 agreement
2PL, Swaminathan–Rogers DIF
Reported, not promotedThe two methods ask overlapping but distinct questions. Where they disagree, the disagreement is the finding; we have not adopted whichever told the tidier story.
Does the prior criterion-9 sex difference replicate?No. The specific result reported by Saffari and colleagues (2022) did not reproduce here.criterion 9: uniform p = 0.242
non-uniform p = 0.909
flagged instead: 2, 3, 5, 8, 11
(BH-corrected: 3, 5, 8 survive)
Reported, not promotedRecorded as a named replication failure rather than quietly substituting our own criteria for the one prior work flagged. The plan specified uncorrected p-values here, so all five flagged criteria are reported, with the BH split shown.
Does criterion endorsement vary by age?Yes, strongly — but with no direction pre-specified, because no prior literature gave us one to test.χ² = 149.29, df = 60
p = 1.4 × 10⁻⁹
ExploratoryDescriptive rather than confirmatory. A genuine pattern worth a dedicated study with its own plan fixed in advance, not a claim made here.
Each of these was written into the analysis plan before the data were touched, and all four are printed here whether or not they flatter the headline result — which is what the reporting commitment in that plan requires. “Reported, not promoted” means the finding is real and is not being used to rewrite the primary result. The temptation in a study like this is to notice that the curvature check came back significant, or that a second method disagreed, and reorganise the headline around whichever looked most interesting after the fact. Locking the primary specification in advance is what removes that option, and it only means anything if the checks that did fire are shown rather than filed away.

BMI’s relationship to criterion endorsement isn’t perfectly linear. A quadratic term added to the model was significant (p = 0.0047), as was a version allowing the curve to differ by criterion set (p = 0.019) and a more flexible spline-based check (p = 0.0079). We didn’t switch the primary model to account for this, because the simpler, linear specification was locked in before any of this was tested. But both facts belong in the same sentence: all three checks found evidence that a straight line is an incomplete description of BMI’s relationship to these criteria in this sample, and we deliberately didn’t chase that with the primary result. Note what those tests do and don’t establish: they say the linear specification is incomplete here, not what the true shape is.

A separate statistical approach agreed with our criterion-level results on only 6 of the 11 criteria. That approach is a differential-item-functioning analysis built on an item-response-theory model: it estimates each respondent’s underlying trait level, then tests whether a criterion behaves differently for people at the same trait level but different BMI. It is a related but distinct question from the direct regression contrast our primary model uses, and the comparison here is against the eleven criterion-level secondary results rather than against the primary set-level contrast. When two reasonable methods looking at overlapping questions agree on half the picture and disagree on the rest, the right response is to report the disagreement as a finding in itself, not to quietly prefer whichever method tells the tidier story.

We tried, and failed, to replicate a specific prior finding from the published literature. A 2022 study by Saffari and colleagues, using a Taiwanese university sample, found a gender-related difference specifically on criterion 9 (failure to fulfil role obligations). [2] In our sample, that same test came back null (uniform p = 0.242, non-uniform p = 0.909). Interestingly, sex-related differences did show up, just on different criteria than the one prior work had flagged. The plan specified uncorrected p-values for this arm, and on that basis five criteria were flagged: 2, 3, 5, 8 and 11. That list thins out if the same multiplicity correction used elsewhere in the study is applied to it. Criteria 3, 5 and 8 stay comfortably below q = 0.05; criterion 11 lands at q = 0.0503 and criterion 2 at q = 0.0647, both just the wrong side of the line. Reporting all five is what the plan called for, but a reader deserves to see which of them would survive the stricter test and which sit on the boundary. We’re reporting this as a failure to replicate a specific, named prior result, which is a more useful statement than either ignoring the prior study or quietly substituting our own finding for it.

Criterion endorsement also varies meaningfully by age band, tested with no pre-specified direction since no prior literature gave us one to test (joint test across the age-by-criterion interaction: χ² = 149.29, df = 60, p = 1.4 × 10⁻⁹). This is exploratory and descriptive rather than confirmatory, a genuine pattern worth a dedicated future study, not a claim we’re making strongly here.

What this study can, and can’t, tell you

None of the above should be read with more confidence than the design allows, and it’s worth being explicit about the limits rather than leaving them to a small-print section at the very end.

Height and weight were self-reported, not measured by a clinician, which introduces the well-documented tendency for both to be reported with some inaccuracy. [3] The sample is self-selected: people who go looking for an online food addiction quiz are not a random sample of the population, and are likely to differ systematically from people who don’t, in ways that could affect these results in either direction. The data are cross-sectional, collected at a single point in time, so nothing here can establish that one thing causes another, in either direction. And BMI itself is a genuinely crude measure of body composition. It cannot distinguish muscle from fat, doesn’t account for where fat is distributed, and Creative Touch’s own body composition explainer already makes this case in detail. We’re reinforcing that existing position here, not revisiting it, and a study like this one is exactly the kind of place where BMI’s bluntness matters most, because it’s the exposure variable the entire analysis rests on.

Most of the submissions we started with are not in this study, and that deserves saying plainly rather than leaving it to the flow diagram. Of 6,105 frozen submissions, 4,207 (68.9%) were dropped for missing at least one of BMI, age band or sex. This is a complete-case analysis, with no imputation. The reason is mostly structural rather than sinister: the demographics section of the quiz is explicitly optional, sits after the questionnaire itself, and carries a prominent “Skip to Results” button, so a large share of people simply never filled it in. Of the 6,105, 4,373 reached the section at all and 1,898 completed enough of it to be usable, which means the attrition splits roughly into people who skipped it outright and people who started it and left something blank. Neither group is a random sample of the whole. We have no way to rule out that people willing to state their height, weight, age and sex differ systematically from those who weren’t, in ways that could bear on exactly the relationships this study is about, and we did not model the missingness. Every number in this article describes the 1,796 who answered, not the 6,105 who took the quiz.

The model itself makes simplifying assumptions we can name but did not test. The primary model allows one shared age effect and one shared sex effect across all the criteria at once, and a single common coefficient for the severity adjustment. Our own secondary results give reason to doubt all three: the age-by-criterion interaction was strongly significant, sex-related differences appeared on several individual criteria, and there is no particular reason a single severity coefficient should fit criteria that differ as much as craving and role failure do. Because age and BMI are themselves related, incomplete adjustment for criterion-specific age effects could in principle leak into the very contrast the study rests on.

We have not run those alternative specifications, and that is a deliberate choice rather than an oversight. The primary model was fixed before any result was seen, and fitting new versions of it now, knowing what the answer looks like, is exactly the freedom that committing to a plan in advance is meant to remove. Naming the assumption is honest; quietly refitting until something fits better would not be. A study designed from the start to test these specifications is the right way to settle it. On the same principle, country was recorded but not adjusted for, because the plan did not specify it; given that family concern, hazardous-use interpretations and willingness to report weight all plausibly vary by culture, that is a real gap and not a trivial one.

Five further limitations are specific to this design, and one is serious enough that we’ve already given it a full section above rather than a footnote here. Criterion 7’s own measurement was compromised for the entire life of the tool that generated this data, contaminating both that criterion’s individual result and, via the severity adjustment every other criterion is conditioned on, the primary model itself, in a way that can’t be corrected retrospectively because the data needed to correct it was never stored. Deduplication isn’t possible: each quiz submission carries a unique code, so we cannot tell a genuine repeat-tester from a first-time respondent, and no attempt was made to guess.

The design has a validity ceiling that conditioning cannot remove. Content confounding, reverse causation, and a form of selection bias called collider bias (where two things that each independently make someone more likely to take the quiz, higher symptom severity and higher BMI, can create an association between them that has nothing to do with either causing the other) all predict exactly the same observable pattern we found. Conditioning on overall severity strengthens the case that any content-specific signal isn’t simply severity in disguise, but it cannot rule out reverse causation or selection bias, and this design tests an association pattern, not a mechanism. That distinction matters and we’re not going to blur it for a tidier conclusion.

Two of the specific criteria in our hypothesised content-confounded set, interpersonal problems (criterion 8) and role failure (criterion 9), were also the two lowest-endorsed, and therefore the two least statistically powered, criteria in the whole study, a limitation of the design that holds regardless of what the results happened to show. And finally, no analysis of the full 35-item questionnaire was possible. Individual item-level answers are only stored in Creative Touch’s database when the criterion they belong to was actually met, which means the missing data pattern is tied directly to the outcome itself. Treating a missing answer as equivalent to “answered zero” would manufacture a false correlation, so we didn’t attempt any analysis at that finer level.

Where this sits in the wider research

Two published studies have already asked versions of this question, and they disagree with each other, which is itself part of the framing here. Chapron and colleagues (2023), working with 508 people already in treatment for obesity or addiction, used item-response theory to test whether BMI changed how these same eleven criteria behaved, and found significant differences, but the published abstract doesn’t say which specific criteria carried that signal. [4] Saffari and colleagues (2022), working with 974 Taiwanese university students, ran a similar test and found none by weight status at all, only a single gender-related difference on criterion 9. [2] Two well-conducted studies, in different populations, reaching opposite conclusions.

The existing research does not tell one tidy story
StudySampleMethodWhat they foundPopulation
Saffari and colleagues, 2022974 Taiwanese university studentsItem-response theory (DIF)No differences by weight status at all. One sex-related difference, on criterion 9.A general student population, not help-seeking.
Chapron and colleagues, 2023508 people already in treatment for obesity or addictionItem-response theory (DIF)BMI-related differences found — but the published abstract does not say which criteria carried the signal.A clinical population at the severe end.
This study, 20261,796 self-selected online quiz respondentsSeverity-adjusted logistic GEE, hypothesis fixed in advanceNamed the specific criteria and tested a stated mechanism. Partly supported; criterion 4 ran counter to its own set.Self-selected and help-seeking, unlike either of the above.
Two well-conducted studies reached opposite conclusions, and the right-hand column is the most likely reason why: a general student population, a clinical population already in treatment, and a self-selected help-seeking one are not three samples of the same thing. We are not presenting this study as the tie-breaker. It sits in the gap between the other two — naming which criteria carry a BMI signal, and testing a mechanism stated in advance rather than scanned for afterwards — in a third population again, with its own limits set out above.

Our contribution sits in the gap between them: naming which specific criteria carry a BMI signal, testing a stated content-confounding hypothesis decided in advance rather than scanning the data for whatever turns up, and doing so in a large, self-selected, help-seeking sample unlike either prior study’s population. We are not claiming to be first to notice that BMI and this questionnaire interact. That question has already been asked twice, with conflicting answers. What hasn’t previously been done, as far as we could establish, is naming the specific criteria and testing a specific mechanism against them directly.

It’s also worth being clear about a gap in the wider literature: we did not identify a citable, pooled correlation figure for continuous food-addiction score against continuous BMI. The closest comparable published figures concern how many people meet the questionnaire’s own classification threshold (around 20%, 95% CI 18–21%, across a large pooled review) rather than the strength of the underlying dose-response relationship. [5] We’re stating that gap plainly rather than inventing a benchmark number to compare ourselves against. Separately, three independent validation studies in different populations have supported a single-factor structure for this questionnaire rather than several sub-dimensions. [6][7][8] Our finding doesn’t challenge that. It’s a finding about differential validity against an external measure (body weight), not a claim that the test itself needs restructuring.

What this means if you’ve taken this test

If you’ve taken this questionnaire and scored highly on some of these specific criteria, this research has something to say to you directly, and we want to say it carefully. Read plainly, part of what we found is that, at an equal level of overall symptom severity, being in a larger body predicted a slightly lower chance of endorsing the criterion built around stopping other important things, like work or time with family and friends, because of how much or how often you were eating certain foods, and a somewhat higher chance of endorsing things like your family or friends being worried about how much you overate, or continuing to eat foods you knew were physically dangerous given a health condition you already had. Read plainly, some of that second group of findings may be less about your own relationship with food than about how other people respond to your body.

We have to be honest about that word “may”. Our data cannot separate that reading from at least two others that predict exactly the same pattern, one of which is simply that people who already felt judged about their weight were more likely to go looking for a food addiction quiz in the first place. We’re not saying that to moralise, and we’re not saying it to undermine anyone’s own sense of whether food feels like a problem for them. We’re saying it because a test that mixes “how your body responds to food” with “how the people around you respond to your body” is measuring more than one thing, and knowing that can change how you read your own score.

This does not mean the questionnaire is broken, and it doesn’t mean the concept of food addiction is simply social stigma wearing a clinical label. Three separate validation studies, in different populations, have supported a single coherent underlying construct for this scale, [6][7][8] and our own headline test, that BMI predicts the content-adjacent criteria more than the pharmacological ones even after accounting for overall severity, passed. What we’d encourage, rather than discarding the score, is treating a high total less as a single verdict and more as an invitation to look at which specific things you endorsed, and asking honestly which of those feel like they’re about your relationship with food itself and which feel more like they’re about how your life has been shaped by other people’s reactions to your body. Those are different things to work on, and they call for different kinds of support.

How to read your own score once you have itA three-step sequence followed by two side-by-side fields. The steps run from your total score, to which particular criteria you endorsed, to what each of those actually asks about. The two fields are your own experience with food — craving, tolerance, eating more than intended, repeated failed attempts to cut down — and the context around food and your body — other people’s reactions, eating despite a health condition, effects on work and relationships. The two overlap: most scores draw on both, and a score is a starting point for interpretation rather than a single verdict.Your total scorehow many of the eleven you endorsedWhich particular ones you endorsednot just how manyWhat each of those actually asks aboutthe question belowWhich of these is this question really about?Your own experience with foodcravingtolerance — needing more for the same effecteating more, or for longer, than you meant totrying repeatedly to cut down, without successThe context around food and your bodyother people’s reactions and worryeating despite a health condition you hadeffects on work, or on time with other peopleactivities given upthese overlap;most scoresdraw on bothA score is a starting point for interpretation, not a single verdict.
This is a way of reading a score you already have, not a diagnostic tool and not a route to one. The split matters because the questionnaire mixes two different kinds of thing: how your body and mind respond to food, and how your life and the people in it have responded to your body. A high total can be built from either, or from both, and the number on its own does not tell you which. Nothing here suggests the second column is less real or less worth taking seriously — only that it points towards different kinds of support. If food, weight or eating patterns are causing you distress, this figure is not a substitute for a conversation with a GP or a mental health professional.

Where Creative Touch fits

We built the free YFAS 2.0 quiz that generated this dataset as an educational tool, not a diagnostic one, and this study doesn’t change that. If your own results have left you wanting to understand what the questionnaire is actually measuring before you interpret your score, our explainer on the Yale Food Addiction Scale covers what each part of it is designed to capture. If what you’re sitting with feels less like addiction and more like a preoccupation with food that’s followed a period of restriction or dieting, our piece on food restriction versus food addiction may be the more useful next read. And if this article has left you wondering how much weight BMI itself should carry in how you think about your body, our body composition guide walks through what BMI does and doesn’t capture, alongside our free body composition calculator. None of this replaces a conversation with a GP or a mental health professional if food, weight or eating patterns are causing you distress. These tools exist to help you understand your own patterns more clearly, not to stand in for that conversation.

Frequently asked questions

Does this mean the food addiction quiz is unfair to people in larger bodies?

Not unfair, but not uniform either. Our results suggest that at an equal level of overall symptom severity, the five criteria we grouped in advance as content-confounded showed a stronger average link to BMI than the three pharmacological-core criteria. That is a statement about the two group averages, not about each criterion individually: only three of the five sat above the no-association line, and one of them, activities given up, ran significantly the other way. It means your total score may partly reflect factors connected to body size alongside factors connected to eating behaviour itself, which is worth knowing when you interpret your own result, but doesn’t mean the tool is measuring something false.

Should I stop trusting my score on this quiz?

No. Three independent validation studies, in different populations, support a single coherent underlying construct for this questionnaire, [6][7][8] and nothing in our findings challenges that. What we’d suggest instead is looking at which specific criteria you endorsed rather than only the total, since some individual questions carry a demonstrably stronger link to body weight than others.

Does this study prove that body size causes people to endorse certain criteria?

No, and this is one of the most important limits of the design. This is a cross-sectional study, meaning everything was measured at one point in time, so it cannot establish that anything caused anything else, in either direction. Reverse causation (where the criteria came first and influenced weight over time) and selection effects (where people who are both heavier and more symptomatic were simply more likely to take the quiz in the first place) would produce an identical-looking pattern in this data. We can say there’s a specific, content-linked association pattern; we can’t say why it exists.

If part of my score reflects how other people react to my body, does that mean I don’t have a problem with food?

Not necessarily, and please don’t read it that way. The finding that some criteria (like family concern about your eating, or health problems from your GP) show a stronger link to body size doesn’t cancel out the criteria that are more clearly about your own relationship with food, like craving or difficulty cutting down. Both can be true for the same person. If food feels distressing or out of your control, that’s worth taking seriously and worth discussing with your GP, regardless of what this study found about the questionnaire’s structure.

Did the criterion 7 defect affect my own score if I’ve taken this quiz?

Possibly, in one specific way. If you were flagged as endorsing criterion 7 (withdrawal), some or all of that may have come from a question about eating more than planned rather than from anything about withdrawal itself. It’s also possible you experienced genuine withdrawal symptoms that weren’t credited towards that criterion at all, because the correct question was never wired into the scoring. The tool has now been fixed, so anyone taking it from here has their withdrawal criterion scored on the right questions. This affects one criterion out of eleven, not your whole result, and it was never the only route to a high total score. If it matters to you, retaking the corrected quiz will give you a more accurate read on withdrawal specifically.

Where can I read the underlying data and methods?

The analysis plan and all three scripts are published in full. The statistical analysis plan sets out the hypotheses, the models and the decision rules as they were committed before any result was seen, and carries its own log of the two departures from it. extract.py pulls the responses and applies every exclusion, analyse.py fits every model, and make_figures.py draws every chart. Between them they contain the frozen date range, each exclusion in the order it was applied, every model specification and every plotted coordinate, so the analysis can be read line by line rather than taken on trust. The individual responses themselves are not published, and will not be. See the data and methods note below for what that means in practice.

Data and methods availability

The pre-specified statistical analysis plan (SAP v1.0) was committed to our git history before any inferential analysis was run, and the analysis script was committed separately afterwards, 88 minutes later. Both are timestamped, and both are now published. The plan appears exactly as it was committed, future tense and all, including the places where its own planning numbers were later superseded and the deviations log that supersedes them; a pre-registration rewritten after the results are known is not one. Alongside it are the three scripts as they ran: extract.py (the data pull, the 25 July 2026 freeze written as a literal in the query, and every exclusion), analyse.py (every model, every pre-specified decision rule, both logged deviations) and make_figures.py (every figure). One redaction has been made and it is marked in the file: the database credentials in extract.py now read from environment variables. Nothing else differs. The ordering of those commits is still something we attest to rather than something a registry attests to, for the reasons given above, but the plan and the code are no longer in that category: anyone can read what was decided and what was run.

We publish aggregated findings, like the ones in this article, not individual responses. That’s a commitment in our privacy policy, and it’s one this article keeps: nobody gets access to identified, individual-level answers, on request or otherwise.

Two logged, dated deviations from the plan exist. One is the statistical model change for our second research question, described above in the discordant-phenotypes section. The other is the criterion 7 measurement defect and its sensitivity analysis, described in full in its own section above. Both are recorded in the SAP’s deviations log. One figure in this article, the proportion of criterion 7 endorsements arising solely from the misplaced item, sits outside the analysis script by necessity, since make_figures.py deliberately never touches the live database; it’s stated independently in two separate records, one of them the published plan’s own deviations log, and they agree.

Consent, and what respondents were told. Nobody could start the questionnaire without passing a consent step first. It required three separate confirmations: that the respondent consented to their anonymous responses and location data being collected and processed as described, that they had read the privacy policy, and that they were 18 or over. The same screen set out what is collected (questionnaire responses, approximate location derived from IP address at city, region and country level, and optional demographics), what it is used for, and what rights respondents have. Research use was stated up front rather than assumed after the fact: one of the three declared purposes was, in the quiz’s own words, to “build aggregate statistics to advance food addiction research”. This analysis is that purpose being carried out. Responses are anonymous by design: no full IP address and no directly identifying information is stored, and no account or email is required to take the quiz. The demographic questions this study depends on are labelled optional, sit after the questionnaire itself and can be skipped with a single button, which is why so many submissions lack them, and respondents can ask to see, correct or delete their data, or withdraw consent, at any time.

Competing interests. This is not independent research and we are not going to present it as such. Creative Touch built and operates the quiz that produced this dataset, discovered and fixed the scoring defect described above in its own code, designed the analysis, ran it, and is publishing the result on its own commercial website, which also links to related Creative Touch tools and content. That combination is worth stating plainly because it is exactly the arrangement a sceptical reader should want disclosed. It does not make the numbers wrong, and the pre-committed plan, the published scripts and the counter-findings reported above are all there so the work can be checked rather than trusted. Funding: none external; the study was carried out in-house at Creative Touch’s own cost. External involvement: none. The analysis has not been peer reviewed, and no independent statistician or clinician has reviewed it on our behalf.

References

  1. Gearhardt, A., Corbin, W., Brownell, K. (2016). Development of the Yale Food Addiction Scale Version 2.0. Psychology of Addictive Behaviors, 30(1), 113-121.

    doi: 10.1037/adb0000136
  2. Saffari, M., Fan, C.W., Chang, Y.L., Huang, P.C., Tung, S.E.H., Poon, W.C., Lin, C.C., Yang, W.C., Lin, C.Y., Potenza, M.N. (2022). Yale Food Addiction Scale 2.0 (YFAS 2.0) and modified YFAS 2.0 (mYFAS 2.0): Rasch analysis and differential item functioning. Journal of eating disorders, 10(1), 185.

    doi: 10.1186/s40337-022-00708-5
  3. Connor Gorber, S., Tremblay, M., Moher, D., Gorber, B. (2007). A comparison of direct vs. self-report measures for assessing height, weight and body mass index: a systematic review. Obesity reviews : an official journal of the International Association for the Study of Obesity, 8(4), 307-26.

    doi: 10.1111/j.1467-789X.2007.00347.x
  4. Chapron, S.A., Kervran, C., Da Rosa, M., Fournet, L., Shmulewitz, D., Hasin, D., Denis, C., Collombat, J., Monsaingeon, M., Fatseas, M., Gatta-Cherifi, B., Serre, F., Auriacombe, M. (2023). Does food use disorder exist? Item response theory analyses of a food use disorder adapted from the DSM-5 substance use disorder criteria in a treatment seeking clinical sample. Drug and alcohol dependence, 251, 110937.

    doi: 10.1016/j.drugalcdep.2023.110937
  5. Praxedes, D.R.S., Silva-Júnior, A.E., Macena, M.L., Oliveira, A.D., Cardoso, K.S., Nunes, L.O., Monteiro, M.B., Melo, I.S.V., Gearhardt, A.N., Bueno, N.B. (2022). Prevalence of food addiction determined by the Yale Food Addiction Scale and associated factors: A systematic review with meta-analysis. European eating disorders review : the journal of the Eating Disorders Association, 30(2), 85-95.

    doi: 10.1002/erv.2878
  6. Brunault, P., Berthoz, S., Gearhardt, A.N., Gierski, F., Kaladjian, A., Bertin, E., Tchernof, A., Biertho, L., de Luca, A., Hankard, R., Courtois, R., Ballon, N., Benzerouk, F., Bégin, C. (2020). The Modified Yale Food Addiction Scale 2.0: Validation Among Non-Clinical and Clinical French-Speaking Samples and Comparison With the Full Yale Food Addiction Scale 2.0. Frontiers in psychiatry, 11, 480671.

    doi: 10.3389/fpsyt.2020.480671
  7. Martini-Blanquel, H.A., Mendiola-Pastrana, I.R., Hernández-López, R.G., Guzmán-Covarrubias, D., Romero-Henríquez, L.F., Rivero-López, C.A., López-Ortiz, G. (2025). Cross-Cultural Adaptation and Psychometric Validation of the YFAS 2.0 for Assessing Food Addiction in the Mexican Adult Population. Behavioral sciences, 15(8), 1023.

    doi: 10.3390/bs15081023
  8. Linardon, J., Messer, M. (2019). Assessment of food addiction using the Yale Food Addiction Scale 2.0 in individuals with binge-eating disorder symptomatology: Factor structure, psychometric properties, and clinical significance. Psychiatry research, 279, 216-221.

    doi: 10.1016/j.psychres.2019.03.003