How researchers define functional GI disorders in children
How the Rome IV criteria define functional GI disorders in children, and why prevalence shifts with the criteria, the questionnaire and who answers.
By Manouchehr Hessabi, MD, MPH
Functional gastrointestinal disorders in children are defined by symptoms, using a consensus classification called the Rome criteria, rather than by a test or a visible structural problem. Because the definition is the measurement, how common these disorders appear to be depends heavily on which version of the criteria is used, which questionnaire collects the symptoms, and who answers it. A 2026 meta-analysis found that switching from one version to the next lowered the pooled prevalence of pediatric abdominal pain disorders from about 12% to about 8%.
What are functional gastrointestinal disorders in children?
Functional gastrointestinal disorders (FGIDs) are recurring digestive symptoms, such as abdominal pain, constipation or regurgitation, that are classified by their pattern rather than by an identifiable structural or biochemical cause. The 2016 revision of the classification, known as Rome IV, also introduced a second name for the group, disorders of gut-brain interaction, reflecting how the field now thinks about them.
A few terms recur in the research. Functional abdominal pain disorders (FAPDs) are a subgroup centered on pain. They include irritable bowel syndrome (IBS), in which pain is linked to bowel habits, and functional dyspepsia (FD), in which symptoms center on the upper abdomen. Functional constipation and infant regurgitation, repeated spitting up in infants, sit in other branches of the same classification.
The broader relationship between the gut and the nervous system is covered in the explainer on the gut-brain axis. This article stays with a narrower question: how researchers decide who counts.
How do the Rome IV criteria decide who counts?
The Rome criteria are consensus definitions developed over a two-decade process that culminated in Rome IV in 2016. The pediatric report for Rome IV, by Hyams and colleagues, made a change that matters for both clinicians and researchers. Earlier definitions required that there be "no evidence for organic disease." Rome IV replaced that phrase with "after appropriate medical evaluation the symptoms cannot be attributed to another medical condition".
The committee explained the reasoning: there is now evidence to support symptom-based diagnosis, so a functional disorder can be a positive diagnosis rather than whatever remains after everything else has been ruled out. In their words, the change allows "selective or no testing" to support the diagnosis.
For an epidemiologist, the important point is that a case definition, the rule that decides who is a case, works like a measuring instrument. Change the rule and you change the reading, even if the children being measured are exactly the same.
How common are these disorders in children?
Estimates are high, and they vary. In a US study, Robin and colleagues recruited 1,255 mothers of children aged 0 to 18 to complete an online survey about their child's symptoms. Using Rome IV, 24.7% of infants and toddlers aged 0 to 3 and 25.0% of children and adolescents aged 4 to 18 met symptom-based criteria for at least one functional disorder. The most common were infant regurgitation among infants (24.1%) and functional constipation among toddlers (18.5%) and among older children and adolescents (14.1%). Quality-of-life scores were lower in children who met criteria.
A systematic review by Vernon-Roberts and colleagues searched nine health databases and pooled 20 Rome IV prevalence studies covering 18,935 children. The median prevalence of any functional disorder was 22.2% in children under four and 21.8% in those aged four to eighteen. The ranges are the more telling figures: 5.8% to 40% in the younger group and 19% to 40% in the older group. Studies using the same version of the criteria still produced very different answers.
Why did changing from Rome III to Rome IV change the numbers?
The clearest evidence comes from studies that compare the two versions directly.
A systematic review and meta-analysis by Jeong and colleagues, published in Gut and Liver in 2026, pooled 67 studies of 1,712,737 children and adolescents in 34 countries. The overall prevalence of functional abdominal pain disorders was 10.89%. Under Rome III it was 11.84% (95% confidence interval, 10.29% to 13.61%); under Rome IV it was 8.42% (6.10% to 11.62%). The ranking also changed. Under Rome III, IBS was the most prevalent and functional dyspepsia the least. Under Rome IV, functional dyspepsia became the most prevalent, affecting about one in 23 individuals, compared with about one in 51 for IBS. The authors suggested that the stricter Rome IV framework likely explains the shift.
That review also reported heterogeneity of 95.90%. Heterogeneity, usually reported as I squared, describes how much the included studies disagree beyond what chance would produce. A value that high means the pooled figure is an average over very different studies, not a single shared truth.
A smaller study removed one source of that variation by testing the same children twice. Baaleman and colleagues gave adolescents aged 11 to 18 at a school in Cali, Colombia, the Rome III questionnaire and then the Rome IV version 48 hours later. Among the 96 who completed it correctly, 40.6% met criteria for a functional disorder under Rome III and 29.2% under Rome IV. Functional constipation fell from 31.3% to 13.5%, while functional dyspepsia rose from 0% to 11.5%.
Agreement between the two versions was summarized with kappa, a statistic that measures agreement beyond what chance alone would produce, where 1 is perfect agreement and 0 is chance level. The kappa was 0.34, which the authors described as minimal agreement. The same teenagers, two days apart, were classified quite differently depending on the rulebook.
Can the same criteria change move numbers in opposite directions?
Yes, and the reason is a useful lesson. Edwards and colleagues reviewed the records of 106 children aged 8 to 17 seen at a clinic for chronic abdominal pain, applying both sets of criteria through a single pediatric gastroenterologist. In that group, Rome IV produced more diagnoses, not fewer: functional dyspepsia 84.9% versus 52.8%, IBS 69.8% versus 34%, and overlap of the two 58.5% versus 17.9%.
So IBS was less common under Rome IV in the global community meta-analysis and roughly twice as common under Rome IV in this clinic sample. Both findings can be true. The clinic children had already been referred for weekly pain lasting at least eight weeks, and their classification came from a clinician reading structured histories rather than from a self-completed survey. Changing a definition interacts with who is being measured and how. The Edwards authors also noted that Rome IV seemed to produce greater variety within each diagnostic category, and that whether these labels predict treatment response is still an open question.
Who answers the questionnaire, and why does it matter?
For young children, symptoms are reported by a parent. Adolescents often report for themselves. Each approach has limits.
In the Robin survey, mothers reported on their children, and children were more likely to meet criteria if their parent also did (35.4% versus 23.0%). That is an association, a statistical link, and on its own it cannot separate shared genetics, shared environment, shared diet, or a parent who notices and reports symptoms more readily because they experience them too.
Self-report brings its own problems. In the Colombian study, 39 of 135 adolescents, 28.9%, were excluded because they did not follow the questionnaire's instructions. The authors cautioned that the limitations of using questionnaires to measure prevalence must be taken into account when reading their results.
Who gets recruited matters too. An online survey, a school sample and a specialty clinic each reach different children, a problem explained in more detail in the article on selection bias. Instruments built on the Rome criteria are also used in specialized populations, as discussed in how gastrointestinal symptoms are measured in autism research.
How should a careful reader read a prevalence figure for a symptom-defined condition?
When a headline says that a fixed share of children has a functional gut disorder, a few questions put the number in context.
- Which version of the criteria? Rome III and Rome IV figures are not interchangeable, as the direct comparisons show.
- Which instrument, and who answered it? A parent-completed online survey, an adolescent's questionnaire and a clinician's chart review measure different things.
- Community or clinic? A referred clinic sample will look very different from a school or population sample.
- Which age band and which country? The most common disorders change with age, and the pooled estimates span dozens of countries.
- How much do the studies disagree? A pooled average with very high heterogeneity is a summary of a varied literature, not a precise constant.
None of this means the conditions are not real or not burdensome. The Robin study found lower quality of life in affected children. It means that prevalence for a symptom-defined condition is always prevalence under a particular definition, measured in a particular way.
Closing thoughts
Functional gastrointestinal disorders show, in an unusually clean way, how much a case definition shapes what epidemiology reports. The same adolescents can be classified differently two days apart, and the same revision can push a diagnosis down in one population and up in another. Reading these studies well means reading the methods section as closely as the headline number.
More on related work is available through the pediatric health and GI research program page, and the full list of peer-reviewed work is on the publications page.