Dr. Manouchehr Hessabi
← All writing
7 min readpediatric health · autism research · measurement

Why GI symptom estimates in autism vary so widely

Estimates of gastrointestinal symptoms in autism run from about 4 percent to 97 percent across studies. What that spread reveals about measurement itself.

By Manouchehr Hessabi, MD, MPH

A reader who looks up how common gastrointestinal problems are among autistic children will find numbers that cannot all be right. Published estimates for having one or more gastrointestinal symptoms range from 4.2 percent to 96.8 percent across studies, with a median of 46.8 percent, as summarized by Holingue and colleagues in Autism Research in 2023.

That is very close to the entire span of possible values. Encountering two figures from opposite ends of it, a reasonable person concludes that science simply does not know.

There is a more useful reading. Most of that spread is a measurement result rather than a biological one. It reflects who was asked, what they were asked, whether the child could perceive and report an internal sensation, and whether the study used a validated instrument at all. The disagreement is worth studying in its own right, because it is an unusually clear lesson in how prevalence estimates are made.

Some definitions first. Prevalence is the proportion of a defined population that has a condition at a given time. A symptom is something a person experiences and reports, such as abdominal pain or nausea. A sign is something an observer can see, such as vomiting or the pattern of a child's bowel movements. That distinction between reported experience and observable event turns out to carry a great deal of weight here.

What the pooled estimates actually say

The largest synthesis of this literature is a systematic review and meta-analysis by Wang and colleagues, published in Frontiers in Psychiatry in 2022. It pooled 63 studies covering 131,416 participants with autism spectrum disorder and reported an overall gastrointestinal symptom prevalence of 48.67 percent, with a 95 percent confidence interval of 43.50 to 53.86.

The same analysis reported pooled estimates for individual symptoms:

  • Constipation: 26.17 percent
  • Bloating: 22.52 percent
  • Abdominal pain: 21.38 percent
  • Diarrhea: 19.92 percent
  • Vomiting: 5.99 percent

Those figures look precise, and the confidence interval on the overall estimate is narrow. Read on their own, they suggest a settled question.

They are not the most informative number the authors reported. That distinction belongs to a statistic most summaries omit: the heterogeneity, which for the overall estimate was an I-squared of 99.51 percent.

What I-squared of 99.51 percent means

I-squared describes what share of the variation observed across studies is attributable to real differences between those studies rather than to sampling chance alone. If several studies of the same population differ only because each drew a different random sample, I-squared is low. If they differ because they were, in effect, measuring different things, I-squared is high.

An I-squared near 100 percent means almost none of the disagreement between these studies can be explained by chance. Something systematic separates them.

This has a direct consequence for how the pooled figure should be read. Under heterogeneity that high, an average across studies describes a literature rather than estimating a stable parameter in a population. It tells you roughly where this body of work sits. It does not tell you the probability that a given autistic child has gastrointestinal symptoms.

The confidence interval deserves the same care. A 95 percent interval of 43.50 to 53.86 quantifies how precisely the average of these studies was located. It is not a statement that the true value in any population falls in that range, and it is not evidence that the studies agreed with one another. Precision about an average and agreement among inputs are different properties, and only one of them is present here.

None of this is a criticism of the authors, who reported the heterogeneity plainly so that readers could see it. The problem arises downstream, when a single percentage travels without the statistic that qualifies it.

Where the disagreement comes from

Several design choices can move a prevalence estimate substantially, and studies in this area differ on all of them.

Case definition. Which children counted as autistic, and how that was established, varies across studies. Whether children with co-occurring intellectual disability were included or excluded matters as well, for reasons the next section makes concrete.

Instrument. An open-ended question to a parent, a checklist written for one particular study, and a formally validated scale are three different measurements of what sounds like a single construct. They will not return the same number.

Recall window and threshold. Symptoms in the past week and symptoms ever experienced are different quantities. So is a threshold that counts occasional discomfort against one that requires a symptom to be persistent or clinically significant.

Comparison group. Many estimates are reported without one. A prevalence figure with no non-autistic comparison group tells a reader how common something is in the sample studied, but not whether it is more common than it would be otherwise.

The interoception and communication problem

Several of the most frequently reported gastrointestinal symptoms are internal experiences. Abdominal pain, nausea, and bloating are knowable to a researcher only if someone can perceive the sensation and communicate it. That makes their measurement dependent on the child's ability to report and on the caregiver's ability to infer.

The 2023 Holingue study makes this concrete. Among 308 autistic children, 36 percent of whom had co-occurring intellectual disability, about 49 percent of parents reported that their child experienced any gastrointestinal sign or symptom, with constipation at 32.25 percent, abdominal pain at 21.24 percent, and food refusal at 19.34 percent.

The finding that speaks to measurement is a different one. About 30 percent of parents reported uncertainty about at least one gastrointestinal sign or symptom. That uncertainty was not evenly distributed. For abdominal pain, 25.93 percent of parents of children with co-occurring intellectual disability were uncertain, compared with 6.06 percent of parents of children without. For bloating, the figures were 19.44 percent and 6.12 percent. Certainty about observable signs such as constipation and diarrhea was similar between the two groups.

The pattern is coherent. Uncertainty concentrates in the symptoms that require a report, and among the children for whom reporting is hardest. Observable signs, which a caregiver can witness directly, do not show the same gap.

This carries a methodological implication worth stating carefully. A study that treats parental uncertainty as absence, rather than recording it as its own category, would tend to undercount subjective symptoms, and would undercount them most in the children least able to describe an internal sensation. That is a direction of potential bias inherent in the measurement approach. It is not a claim about any particular study's findings.

Why better instruments matter more than another prevalence paper

Given the above, another study reporting another percentage adds relatively little. The more valuable contribution is a measure that different research groups can use in the same way.

An instrument earns that status by demonstrating several properties. Factor structure describes whether the items behave as though they measure one underlying thing or several. Internal consistency describes whether items intended to measure the same construct agree with one another. Test-retest reliability describes whether the measure returns similar results on separate occasions when the underlying condition has not changed. Convergent validity describes whether it agrees with other measures that theory says it should agree with. This is the same set of questions covered in the discussion of what a biomarker has to earn before it is used, applied to a questionnaire rather than a laboratory measure.

One recent effort in this direction is the Gastrointestinal Symptom Severity Scale, an instrument built on the Rome IV criteria, which are the standard consensus definitions used to classify functional gastrointestinal disorders. Martínez-González and colleagues evaluated it in a sample of 265 individuals with autism, mean age 9.44 years with a standard deviation of 4.99, and reported a confirmed unidimensional factor structure, good internal consistency, adequate test-retest reliability, and strong convergent validity (Digestive and Liver Disease, 2024, PubMed 38851976).

Those are the reported psychometric findings, not an endorsement, and a single evaluation is a starting point rather than a settled result. The relevant point is the direction of the work. Progress on this question depends more on measuring consistently than on measuring again.

How to read the next prevalence figure you see

A short set of questions makes most estimates easier to interpret:

  • Who was in the sample, and how was autism established?
  • What instrument produced the number, and had it been validated?
  • What recall window and what threshold were used?
  • Was there a comparison group?
  • Was heterogeneity reported, and if the figure is pooled, how large was it?

The honest summary is narrower than the headline numbers suggest. Gastrointestinal symptoms appear common enough among autistic children that careful inquiry is warranted, and existing evidence indicates they are frequently accompanied by caregiver uncertainty, particularly for symptoms that depend on self-report. The literature does not support any single precise prevalence figure, and a study reporting one without its measurement context should be read accordingly.

This piece deliberately sets aside the question of mechanism, which is a separate literature and is discussed in the explainer on what the gut-brain evidence supports. Related work in this area is collected under pediatric health and GI research, and the underlying peer-reviewed record is available in the list of publications.

About the author. Dr. Manouchehr Hessabi is a physician-epidemiologist and Senior Research Scientist at the BERD core of UTHealth Houston's Center for Clinical and Translational Sciences. See his peer-reviewed publications or research programs.