Urban dispersion models provide a way of calculating how pollutants travel through street and junction networks, building canyons, and across the wider urban area. They help complete missing measurement records, explore scenarios, and underpin decisions regarding monitoring network design, planning and emergency response.
Machine learning has a role to play here, but should not replace physical understanding, as urban air movement is strongly affected by the arrangement of the urban fabric, by vehicle and pedestrian activity, by turbulence, weather, and the location of the emission source. These are physical processes, not just data patterns; they need to be measured, tested and explained.
The right question therefore is not whether machine learning and artificial intelligence can replace urban dispersion models, but rather how they can help, and where they can introduce further uncertainty.
Why Urban Dispersion Is a Difficult Modelling Problem
Cities are complex fluid environments, so modelling urban dispersion is inherently difficult. For example, a pollutant released into the air at road level may travel along the length of a street, recirculate within a street canyon, flow around the corner of a building, enter a parallel street or rise above the roof tops to disperse across the wider urban environment.
A change in wind direction, traffic pattern, the height of buildings, the location of junctions, or the position of a roadside emission source will have an impact on how a plume will disperse, hence why street-level air pollution prediction is difficult even if we know reasonably well what the sources are.
Machine learning offers some possibilities because it is a powerful tool to understand the relationship between complex and high-dimensional data sets. However, the problem of modelling how a plume disperses in urban areas is not just one of data, but of physics too.
If the training data set does not sufficiently represent all the possible flow states a plume may encounter, the AI model may be biased by that limitation, rather than being able to solve it.
Where Machine Learning Can Help
There are circumstances where machine learning can contribute in an informed way to this work. It is particularly suited to the task of understanding a complex set of relationships that would be hard to interpret if relying on human judgement alone.
Air quality studies in urban areas are increasingly complex in this respect, often drawing on a wide range of information sources, including data from fixed monitoring stations, mobile monitors, vehicle traffic data, weather observations, building information, sensor networks and outputs from air quality models.
It can be a challenge to find any significant relationships between variables derived from all of these sources. In these circumstances, machine learning can assist by identifying recurring patterns, detecting unusual events, filling gaps in data and aiding the interpretation of time-evolving situations.
It can also provide a way of comparing model results to field observations that are too large or complex for a human analyst to process quickly and comprehensively.
This would be of particular value if we are assembling a body of evidence that is drawn from multiple sources, and if we have a good idea of the overall validity and limitations of that evidence. The process of turning field campaign findings into policy evidence typically involves the synthesis of field measurements, model results and human interpretation in the context of a wider evidence base.
Machine learning may contribute a valuable step in the process of connecting field and model evidence.
Using AI to Interpret Monitoring Data
Air quality sensor networks can provide large datasets on the spatial and temporal patterns of ambient pollution levels. One example where machine learning may contribute is in understanding these datasets.
A set of sensors may record consistent evidence of high concentrations in a given area during a specific wind direction. It may highlight evidence of a short-term peak in pollution levels that is not seen in the average daily data. It may differentiate between an anomaly at a given sensor location versus a wider pollution event.
In these circumstances, machine learning provides a useful tool for interpretation of the patterns that the sensor network is measuring.
Yet AI will not eliminate the need for sensor quality control. Bad data due to poor calibration, poor location, or unrepresentative data gathering still leads to bad evidence. While machine learning is effective at processing bad data, it cannot turn bad data into good data in any meaningful way.
AI and Short-Term Pollution Peaks
Short-term pollution events are hard to capture, because they may be very brief and depend heavily on local conditions. A single traffic queue, a change of signal timing, a poorly ventilated space, or a change in wind direction may cause short exposure peaks.
Machine learning can potentially support short-term event detection in high-frequency monitoring data. It may be able to be trained to spot events in high-frequency data that may signal short-term spikes, such as a change in traffic flow, a change in wind speed, or a change in readings across a monitoring network.
This could support an understanding of peak exposure in urban policy, where short peaks are relevant but averages may fail to capture these events.
The caveat though is that prediction is not equivalent to explanation. A machine learning model may be able to learn that peaks occur under certain conditions, without necessarily explaining the reasons why. This is important for policy. An AI model that can predict a peak but is unable to explain why the peak occurs may still not be useful in designing policy intervention.
Supporting, Not Replacing, Physical Models
Machine learning has many ways of supporting physical and numerical models, including by estimating uncertain model inputs, identifying sensitive model variables, emulating slower-running models, and comparing the results of many models.
For example, a machine learning model might be trained on the outputs of simulations in order to provide quick approximations of those models. This could potentially be useful for screening models or supporting fast analysis. It could also potentially support the identification of when a more detailed model and a less detailed model disagree, thereby drawing attention to those cases.
However, AI is not a substitute for physical modelling. Urban dispersion involves the physics of recirculation, flow separation, turbulence, and plume spread, which need to be understood through field measurements, experiments, and numerical modelling.
A better future for physical and digital modelling might involve more machine learning, but only as part of a wider evidence ecosystem.
The Importance of Wind Tunnel and Tracer Evidence
Machine learning needs data to train on, but that data should ideally contain more than just routinely gathered data.
Experiments remain valuable for separating out physical phenomena from spurious correlations. Wind tunnel modelling allows repeatable experiments on how source position, wind direction, and surrounding buildings shape dispersion.
Tracer data is similarly useful, because it helps us understand how things move from a known location through a real or a simulated environment. Evidence from urban tracer experiments helps us identify pathways and processes which are not captured by routine monitoring.
For AI, this type of evidence is valuable because it provides meaningful input. It connects patterns in data to the movement of material.
Where Machine Learning Can Mislead
Machine learning can mislead if its output looks more conclusive than it is.
A machine learning model can perform well on existing data but may fail when conditions are not present in the training data set. A model might find patterns at one location that are not transferable to other streets. That is the danger of producing detailed forecasts that appear plausible while relying on weak assumptions or limited training data.
It is particularly important in urban dispersion modelling because there may be rare events to consider, such as an emergency release, an unusual wind direction, a change in road geometry due to temporary roadworks, or unusual road conditions such as significant congestion that is not represented in the model’s training data.
Another potential problem is that AI-based models are susceptible to unrepresentative biases in the monitoring networks used to train the models. For example, if monitoring sensors are sited mainly near major roads, the AI may be less effective at predicting concentrations in smaller streets, courts and courtyards, or in enclosed spaces.
Similarly, if the training data reflect primarily typical conditions of meteorology and traffic, the AI may not produce useful predictions in less common meteorological conditions.
Validation Still Matters
Machine learning does not remove the need for model validation. It makes it even more important.
A model using machine learning for urban air quality modelling should be validated against appropriate empirical evidence to match the specific application. For example, if it is intended for use for the purposes of long-term exposure assessment, then it should be validated to match that time scale. If it is being used for short-term peak events, then it should be validated to match short time frames. If it is being used for emergency release modelling, it needs to be validated under conditions that represent the time frame needed for the rapid movement of a plume.
This is why model validation is a policy issue: AI-based results are likely to influence decision making for planning, monitoring or exposure assessment, or emergency responses, and therefore the uncertainties in their predictions must be understood before they can be used in the context of those decisions.
Model validation should assess accuracy, but also the usefulness, interpretability and suitability of the AI output for the specific decision-making purpose.
AI in Emergency Response Modelling
It is possible that there is a role for machine learning in emergency response to an accidental release, particularly where speed of interpretation is important. Machine learning could potentially be used to fuse together information from multiple data sources, including real-time sensor readings, meteorological forecasts and existing dispersion models, to provide an estimate of the direction and extent of the movement of a released material.
But in a dense urban setting, emergency response planning for urban airborne releases is dependent on understanding the movement of a plume on the fine scale through the streets, around the building perimeters, through junctions and down urban canyons.
So a fast answer from a machine learning model could provide an indication of where a plume may be, but it must be used only when the limitations of that indication are known.
There is also the added consideration of the cost of an error in an emergency response situation. A model may fail to predict a local pathway in the movement of a plume, or underestimate a local peak in concentration and mislead the monitoring or response strategy.
So, in an emergency response use case, AI should be used to assist emergency dispersion modelling, but never replace the validated dispersion models and the experience of the modellers interpreting their outputs.
A More Practical Role for Machine Learning
There are potential uses of AI in urban dispersion research in terms of enhancing the understanding of data, identifying underlying patterns, facilitating the design of monitoring networks, speeding up the assessment of potential release scenarios and assisting in the comparison of models.
AI should only be used in a practical way, with clearly identified roles and boundaries, and for which it is useful.
Machine learning can be a valuable tool for understanding data when it is linked with good-quality evidence from field campaigns, controlled experiments, validated dispersion models and well-defined policy-relevant questions. It will be a wasted investment when it is used to provide an answer to a physical problem with no connection to field evidence.
We need to develop a use of machine learning in urban dispersion modelling that is evidence-aware. The use of AI models in this context means that models must be built on data that is representative of the conditions to be simulated, validated against real data sets that were not used in the construction of the model, interpreted in a way that recognises that they are models that reflect physical realities, with communicated uncertainties and clear application requirements.
The future for AI use is not to replace dispersion models for understanding urban dispersion. It is to use machine learning within a stronger evidence and dispersion modelling framework.
When it is well used, machine learning can help researchers and policymakers better understand the complexities of data on emissions and concentrations in the urban environment. If used poorly, it has the potential to add a new level of spurious confidence to the understanding of urban dispersion problems.
The difference is whether AI is used as a tool to help understand urban dispersion, or to replace the understanding of urban dispersion.


