1st April, 2025

The Seine with Clothing on the Bank (ca. 1883–1884) by Georges Seurat
Childhood adversity does not affect everyone in the same way. Two children may experience similar events and follow very different developmental paths. One may develop persistent depression, anxiety, substance use, or difficulties in relationships. Another may experience temporary distress but recover well. A third may show few obvious symptoms until much later in life.
For clinicians, this raises an important question: can we predict who will suffer the most?
It is an appealing goal. If we could reliably identify children at highest risk after adversity, services could offer extra monitoring, earlier treatment, or more intensive support. Researchers have therefore developed prediction models using combinations of information such as the type and severity of adversity, family circumstances, socioeconomic conditions, previous mental health problems, biological markers, and characteristics of the child.
In principle, prediction models are well suited to this problem. They do not need to determine exactly why a child develops later problems. Their task is simply to distinguish, as accurately as possible, between people who will and will not experience a particular outcome.
A risk factor does not need to cause an outcome to be useful for prediction. If previous psychiatric treatment, school absence, family instability, or neighbourhood disadvantage improves our ability to predict later depression, it may be useful to include, even if the underlying causal pathways are complicated.
The difficulty comes when prediction models move from research papers into clinical screening. A model can perform reasonably well statistically while still being of limited practical value in a clinic. Consider a model designed to predict which maltreated children will develop depression within five years. Researchers may report a large association between predicted risk and later depression. The real world, however, needs answers to more concrete questions:
How many children identified as high risk will actually become depressed? How many children who later develop depression will the model miss? How many families will be offered additional intervention unnecessarily? Is there something useful we can do differently for the children classified as high risk?
Imagine that 10% of children exposed to adversity develop a particular severe mental health outcome. Even a reasonably accurate model may generate many false positives because most children will not develop the outcome. If clinicians screen hundreds of children, a substantial proportion classified as high risk may remain well.
That does not automatically make screening useless. Medicine routinely accepts some false positives when missing a serious condition would have major consequences. But the balance depends on what follows the screening result.
If being classified as high risk leads to a brief conversation, closer follow-up, or an intervention that is safe and acceptable, the threshold for screening may be relatively low. If it leads to stigma, intensive assessment, costly treatment, or involvement of additional services, false positives are much more consequential.
False negatives matter just as much. A model that identifies a small group at extremely high risk may appear attractive, but if most children who eventually develop problems fall outside that group, it is poorly suited to deciding who deserves support.
There is another challenge specific to adversity research. Maltreatment may be underreported or incompletely documented. Family circumstances change over time. Children disappear from longitudinal studies. Mental health outcomes may be missed when young people do not access services. For reasons such as these, a prediction algorithm built on incomplete or selective data typically performs much better in the original study than in the messy reality of clinical practice.
External validation is thus key. A model developed in one cohort should be tested in different populations, settings, ages, and healthcare systems, especially in the ones that screening is being implemented in. Performance should also be assessed across groups who may differ in access to care, socioeconomic position, ethnicity, or exposure to different forms of adversity.
Even a well-validated model does not answer the final question: does using it improve care?
A screening tool only becomes clinically valuable if acting on its predictions leads to better outcomes than existing practice. Ideally, this should be tested prospectively. Do clinicians identify vulnerable children earlier? Do children receive more appropriate support? Does functioning improve? Are unnecessary referrals avoided?
That is the standard prediction models, whatever their domain is, need to meet before they become useful tools in medical care.