Update to Google’s AI weather model improves forecast accuracy
Update to Google’s AI weather model improves forecast accuracy
谷歌 AI 天气模型更新,预报准确度进一步提升
Google is one of the major players in AI (meaning machine learning) weather forecast model space. The models it and others generate have their strengths and weaknesses, but the main advantage is that they can have forecast performance similar to traditional models while requiring far less computing horsepower to run. That means they can be run more frequently. Google recently released version 3 of its WeatherNext model, with the biggest change being that it now ingests some satellite weather data, shortening the lag time between current weather conditions and generating a new forecast. The update is detailed in a white paper.
谷歌是人工智能(即机器学习)天气预报模型领域的主要参与者之一。谷歌及其他机构生成的模型各有优劣,但其主要优势在于,它们在实现与传统模型相当的预报性能的同时,所需的计算能力却远低于传统模型。这意味着它们可以更频繁地运行。谷歌最近发布了 WeatherNext 模型的第三版,其最大的变化在于现在能够摄入部分卫星天气数据,从而缩短了从当前天气状况到生成新预报之间的延迟时间。该更新的详细信息已在白皮书中公布。
Reanalysis
再分析
Many weather models make use of what’s called a “reanalysis,” which is a sort of model of its own. Reanalyses take in all kinds of weather data and combine them into a single, consistent global snapshot of the atmosphere. That requires that they provide estimates for conditions over locations without real-world measurements, because weather forecast models need to work with a global picture. Nearly all AI weather models have been relying entirely on reanalyses, with machine-learning algorithms training on global reanalyses and spitting out a weather map in the same format. There are some compromises there—the raw data sources themselves may contain some information that gets lost in the reanalysis blender, and these global snapshots are generally only produced every six hours.
许多天气模型利用所谓的“再分析”(reanalysis),这本身也是一种模型。再分析汇集了各种天气数据,并将它们整合成一个单一、连贯的全球大气快照。这要求它们必须为缺乏实测数据的地区提供状况估算,因为天气预报模型需要基于全球视角进行工作。几乎所有的 AI 天气模型都完全依赖于再分析,机器学习算法通过全球再分析数据进行训练,并输出相同格式的天气图。这其中存在一些折衷——原始数据源本身可能包含一些在再分析过程中丢失的信息,而且这些全球快照通常每六小时才生成一次。
Traditional weather forecast models often also take in other raw data, capturing as much information as possible to accurately represent the current state of the atmosphere so the model can use physics to simulate conditions forward. WeatherNext 3 now does some of this as well, adding in weather satellite data and upping the forecast frequency to hourly as a result. There are some other changes, too. The spatial resolution has been increased, and the machine-learning model is larger, prompting some process tweaks to limit the increased computational demands. They also added a separate machine-learning model trained on satellite-based precipitation estimates, meaning there are multiple precipitation forecasts available.
传统天气预报模型通常还会摄入其他原始数据,尽可能多地捕捉信息以准确呈现当前的大气状态,从而使模型能够利用物理学原理模拟未来的状况。WeatherNext 3 现在也实现了部分功能,通过加入气象卫星数据,将预报频率提高到了每小时一次。此外还有一些其他变化:空间分辨率得到了提高,机器学习模型规模也更大,这促使团队进行了一些流程调整,以限制计算需求的增加。他们还增加了一个基于卫星降水估算训练的独立机器学习模型,这意味着现在可以提供多种降水预报。
Unlike traditional models that use physical properties in a location to simulate physical processes, machine-learning models are largely black boxes that train on past patterns and spit out predictions of future patterns. But WeatherNext 3 is adding a tiny bit of physical information to calculate surface temperature and dew point at any specific location you want to pull up. It checks whether that point is land or ocean and uses its surface elevation. By training on past weather station data tagged with that information, the team says they get better forecast predictions.
与利用特定地点的物理属性来模拟物理过程的传统模型不同,机器学习模型在很大程度上是“黑箱”,它们通过训练历史模式来输出对未来模式的预测。但 WeatherNext 3 加入了少量的物理信息,用于计算用户查询的任何特定地点的地表温度和露点。它会检查该点是陆地还是海洋,并利用其海拔高度。团队表示,通过使用带有这些标签的历史气象站数据进行训练,他们获得了更好的预报结果。
Some oddities
一些异常现象
The white paper shows some results to document forecast performance improvements over WeatherNext 2, as well as the European Centre for Medium-Range Weather Forecasts (ECMWF) AI model. They note a roughly 5 percent improvement in upper atmosphere condition accuracy over their previous model, for example, which they say equates to about six more hours of accurate forecast lead time. And their change to calculating surface temperature for a specific location improved accuracy by up to 30 percent. They’re generally beating the ECWMF model on these metrics as well.
白皮书展示了一些结果,记录了其预报性能相较于 WeatherNext 2 以及欧洲中期天气预报中心(ECMWF)AI 模型的提升。例如,他们指出高层大气状况的准确性比之前的模型提高了约 5%,他们称这相当于增加了约 6 小时的准确预报提前期。此外,他们针对特定地点计算地表温度的改进,使准确度提高了多达 30%。在这些指标上,他们通常也优于 ECMWF 模型。
There is one curious exception that the paper doesn’t even guess at the cause of. For a number of variables, their comparison to the initial six-hours-ahead forecast from the other models shows WeatherNext 3 doing worse before pulling ahead for the rest of a 15-day forecast. Larger-scale patterns are also not without some weirdness. You can see the shape of the model’s grid in some predictions, like the map of precipitation showing some distinctly hexagonal blobs. Their method of generating multiple surface temperature forecasts to represent the range of possible outcomes also has a habit of producing snapshots where the global average temperature is higher or lower. Normally, you would want to see the average be consistent, with local-scale variability that averages out across the globe.
有一个奇怪的例外,白皮书甚至没有推测其原因。对于多个变量,他们与来自其他模型的初始 6 小时预报进行对比时发现,WeatherNext 3 在初期表现较差,但在 15 天预报的后续阶段则处于领先地位。更大尺度的模式也并非没有怪异之处。你可以在一些预测中看到模型网格的形状,例如降水图显示出一些明显的六边形斑块。他们生成多个地表温度预报以代表可能结果范围的方法,也容易产生全球平均温度偏高或偏低的快照。通常情况下,人们希望平均值保持一致,而局部尺度的变异性在全球范围内相互抵消。
Overall, the team says their new model “represents a major step forward for AI-based weather predictions by going beyond relying purely on analysis and utilizing information-dense, low-latency observation data.” WeatherNext 3 is now the source of forecast information across Google services, including Search, Gemini, and Maps.
总的来说,团队表示他们的新模型“通过超越单纯的分析依赖,利用信息密集、低延迟的观测数据,代表了基于 AI 的天气预报向前迈出的重要一步。” WeatherNext 3 目前已成为谷歌各项服务(包括搜索、Gemini 和地图)的天气预报信息来源。