Validation in urban dispersion is never easy. In such cases, there is evidence to consider: fluctuating meteorology, dynamic traffic emissions, urban complexity, uncertainty and spatial heterogeneity in both emissions and concentrations.
AI can aid this process but does not eliminate the need for validation; often, it actually requires additional validation.
ML can allow data-rich comparisons, identify model error patterns, identify where there is likely poor model performance, and assist with interpreting monitoring evidence. But it cannot determine for itself whether a model should be used for any particular application, for example for transport planning, exposure assessment or emergency response.
AI-assisted validation should thus be seen as an aid to, and not a substitute for, scientific judgement.
Why Validation Is Difficult in Urban Dispersion
Urban dispersion is a complex process. Urban dispersion model outputs will therefore be compared to real observations of a highly variable environment, with plumes travelling down a street, being deflected by buildings, recirculating into building streets, passing over or around obstacles, and being influenced by changes in emissions and wind direction.
As a consequence, machine learning in urban dispersion must be used with caution. The challenge is not just statistical but physical:
- A model may capture large-scale concentrations but not local peaks;
- A model may work well in open terrain and not well in streets;
- A model may have an acceptable long-term mean but fail to reproduce short time-scale events.
Validation is needed to determine where these models can be used reliably.
Where AI May Help Urban Dispersion Validation
AI may aid validation by comparing the results of the model against far more data than one can review manually.
An urban dispersion model will generally be compared to other observations, for example a single fixed monitor, a dense network of sensors, wind tunnel measurements, or a different urban dispersion model. If a large quantity of data exists for comparison in a given location or scenario, it may be useful for validation.
AI will not be able to assess directly the utility of the model, for example whether it provides adequate estimates for use in policy-making or operational decision-making, but may help identify where and why model results diverge from observations.
For example:
- Is the model good in certain conditions and not good in others, for example under moderate winds but not low wind speeds?
- Are model errors greater in certain locations, for example near junctions and building corners?
- Can the model capture short time-scale peaks but not long time-scales, or vice versa?
It is useful to see that validation is not just a numerical score but understanding of where and when a model may be reliable.
Comparison of Model and Urban Pollution Monitoring
A sensor network will provide more data on localised urban pollution than a single fixed monitor and will better capture any localised variations in concentrations compared to a fixed monitor.
Machine learning may be useful in comparing a model to this type of data. For example, machine learning may be useful for identifying where sensors have different results to the model, spatial patterns in model errors, or whether a model can adequately capture the timing of pollution movement across a sensor network.
This could be useful in helping urban pollution monitoring be a better validation tool for an urban dispersion model.
Of course, the monitoring data has to be reliable to be suitable. Flawed sensor placement, miscalibration, or a poorly designed network can skew a validation of your AI. A machine learning algorithm can digest bad evidence swiftly, but it cannot make poor measurements into robust validation data.
Assessing Short-Term Peaks
A key question for validating models is: can the model adequately assess short-term peaks?
A model can be well-performing relative to daily or hourly averages, while being unable to capture the short events relevant for exposure. This is likely to be during periods of congestion, periods of low wind speed, a change in wind direction, or poor street canyon ventilation.
AI could support a validation of your model, checking whether it fails around such short-term events. If you have high-frequency monitoring data in your network of interest, you could use AI to look at when the model is missing peaks in time, in size or at the location.
This is particularly relevant in situations where AI is used to predict short-term pollution peaks. A model should not just be evaluated in terms of whether it performs on average, it should be tested against whether it performs well around the short-term events it is intended to predict.
For policy applications this distinction is relevant; for a model of peak exposure, it needs to be tested against peak exposure.
Locating Hotspots and Spatial Errors
Similarly, AI could help identify where a dispersion model performs badly in space. A model could, for example, underestimate pollutant concentrations near a busy junction, overestimate them elsewhere on the same street, or fail to capture a hotspot caused by local recirculation.
These could be errors that are not easily identified if you rely on broader summary statistics. Machine learning algorithms could help identify these spatial patterns. By comparing model outputs with monitoring data across the street network of interest, the model errors could be identified.
This can support hotspot identification using AI, but it can also support a more robust validation of models. If a model cannot capture existing hotspots, it is less clear that it should be trusted to predict the location of new hotspots without further validation.
The important takeaway here is that spatial error is more than a simple technicality. It has an impact upon where monitoring is deployed, how exposure is understood and where attention is applied.
Why You Still Need Wind Tunnel Evidence
It is also important that your AI-assisted validation is not just dependent on routine monitoring data. Physical evidence is still useful in this context.
Wind tunnel studies can be used to evaluate model performance around key factors such as the position of the source, wind direction and local building geometry. This could enable a model to be evaluated in terms of whether it captures the physical processes, such as the recirculation zone, the plume, how separation occurs and the concentration gradients that occur.
When AI is used to evaluate a large body of data from modelling against wind tunnel data, this controlled physical evidence could help identify where a model fails physically rather than simply making a random error.
If the model is consistently missing the concentration patterns around the corner of a building, an AI algorithm might detect these errors but physical evidence can help explain it.
Using Physical and Digital Evidence Together
AI could be useful in bringing both physical and digital evidence together. A machine learning algorithm could be used to compare the output of a numerical model to wind tunnel data as well as field-based measurements of pollutant concentration and data from your network of sensors.
It could support you in understanding where each dataset agrees with one another and where they are in contrast. This would enable a more robust approach to physical and digital modelling.
It is one thing to use a digital model to explore a wide range of scenarios but another for you to be confident you can accurately predict a future scenario. Experiments offer the control, but lack the complexity of a real city. Field measurements offer realism, but are not easily repeated.
AI can aid this analysis but is incapable of judging the relative importance of these evidence sources: that will remain a scientific issue.
Validating Emergency Response
Validating models for emergency response presents a special problem: it may need to support advice in real time, under high uncertainty, with imperfect knowledge of the source, meteorology and urbanity.
An AI tool might help match model predictions against sensor measurements to deduce likely plume directions, or point out where it could be valuable to place additional sensors. However, it still needs to be validated for the circumstances of real-time, rapid urban dispersion.
If a model fails to identify a local pathway or underestimates a local peak during the process of urban airborne release planning, this will affect subsequent advice on what action to take and where. Thus, the validation needs to include how fast the model predicts transport through urban environments at street, intersection, building and canyon scale.
An automated, fast-acting system is of little use if its limitations are not understood beforehand.
What Is Being Validated
AI-assisted validation can and should go beyond testing whether model predictions match the measurements. It can test:
- whether it captures the timing of events;
- whether it predicts the correct peaks;
- whether it works for all wind directions;
- whether it accounts for variability in street networks and the spatial structure of the urban canopy;
- whether its performance changes under different configurations, such as junctions, canyons and open areas;
- whether its error is random, systematic or linked to underlying physical processes;
- whether it is fit for the purpose of the advice being generated.
All of this matters, because a model can appear good in one way and still be unsuitable for a policy use.
The Problem of AI Confidence
AI makes it appear as though validation has progressed when it has not. AI can produce sophisticated-looking results in the form of performance scores, performance maps, clusters or error categories.
It is still only as good as the underlying measurements used for the validation, and the tests themselves. There is a danger here of automated confidence: the impression that because validation is more data-intensive it must also be more reliable.
This is not always true. It does not necessarily provide a solution to the problem if models are validated against inadequate measurements, unrepresentative conditions, or a bad time scale. It may actually make the problem harder to see.
Validation therefore needs to remain transparent so that policy users can understand what the model was validated against, what it was not validated against, and where there is room for uncertainty.
An Appropriate Application for AI-Assisted Validation
AI is helpful in making model validation more informative if it is being used to aid rather than replace experts.
AI tools are useful when they facilitate comparisons of data sets, identify recurring errors, highlight model weaknesses, evaluate predictions of extreme conditions, and connect findings from field campaigns with experimental and measurement results. They can also guide where more measurement and experimentation should be prioritised.
However, urban dispersion remains a physical problem: models need to be judged against data that represent the urban streets, processes of the urban flow and the questions of policy.
The most appropriate use of AI is therefore a practical one, with clear boundaries; helping researchers understand model performance, and helping policy understand where to apply it with confidence and where to use it cautiously.
AI-assisted validation has a place if it is making uncertainty more explicit, rather than creating the illusion that model results are more certain than the science warrants.


