Storm Stream.

data

What Radar Archives Hold, and What They Miss

A dataset by dataset look at what NOAA's radar archive actually stores, what its filters drop, and where MESH and ground reports diverge.

The radar record is the only part of the severe weather archive that covers ground where nobody was standing. Ground reports come from people, and people are not evenly distributed, so the report record thins out exactly where the land is empty. Radar does not care who lives below it. That makes it uniquely useful, and it is also why people most often read it as more than it is. What follows is what is actually in it, dataset by dataset, and what each piece can be asked to support.

What is actually archived

NCEI's Severe Weather Data Inventory holds a specific set of algorithm outputs, not raw radar volumes. The datasets it lists are these:

Access is through an interactive map, a bulk download of the whole database as CSV over HTTP, and a REST service, in Shapefile, KMZ, CSV and XML. Every row in those files is a detection: something an algorithm flagged from a radar scan at a time and a place. A detection is not a measurement of a hailstone, and the page is direct about that.

What a filter like "probability 100 percent" costs you

The two variants of each dataset are the most consequential detail on the page, and the consequence is not spelled out for you. The filtered hail file keeps only signatures whose maximum size is greater than zero and whose probability is 100 percent. Everything the algorithm flagged with less than full confidence, and everything it flagged with a maximum size of zero, is in the all-signature file and absent from the filtered one.

So the filtered file is not the archive. It is a subset selected for confidence. If you query it over an address and get nothing back, what you have learned is that no signature there cleared that bar, which is a narrower statement than the one people tend to write down. The same logic applies to storm structure, where the filtered cell file imposes a 45 dBZ floor and the all-cell version does not.

The mirror mistake is just as easy. Pull the all-signature file and treat every row as an event, and you will count low-probability flags as hail. Neither variant is wrong. What is wrong is not saying which one you used. Two people can quote the same archive, pick different variants, and disagree completely without either of them making an error.

MESH is a separate product, and there is more than one of it

Hail size estimates from radar usually come from MESH, the maximum estimated size of hail, which is not in the list above. MESH estimates hail size from the radar reflectivity properties of a storm above the environmental 0 degree C level. It is produced by the Multi-Radar Multi-Sensor system and is commonly used across the NWS, including at SPC, to diagnose expected hail size.

There is not one MESH. The original algorithm was published by Witt and colleagues in 1998. Later work produced two further power-law relationships: MESH75, fit to the 75th percentile of the sampled hail size distribution, and MESH95, fit to the 95th percentile. That is the whole reason the same storm yields different hail sizes from different products. A fit to the 75th percentile and a fit to the 95th percentile of the same distribution are different answers by construction.

Which one is better depends on the size you care about. Near the 1 inch threshold, which is 25.4 mm and also the NWS criterion for severe hail, MESH75 and MESH95 predict hail size more reliably than the Witt formulation, but they are more likely to overestimate between 1 inch and 2 inches, 2 inches being the significant severe threshold. At the exact 25.4 mm and 50.8 mm thresholds, the Witt version has a higher critical success index than both. Murillo and Homeyer found in 2019 that MESH75 performs best at 40 mm for severe hail and 47 mm for significant, and MESH95 at 64 mm and 83 mm. In other words the threshold is part of the product. A MESH number without the formulation and the threshold attached is not a usable number.

The quality control, stated plainly

The 2021 Wendt and Jirak study that compared MESH against ground reports across the contiguous United States for 2012 to 2019 applied two screens worth knowing. A MESH pixel counted only if it sat within 40 km of a detected cloud-to-ground lightning flash in the same hour, which removes radar signatures with no nearby evidence of convection. And MESH values above 127 mm, five inches, were removed as likely spurious.

That second screen is the honest part. A product whose top end has to be discarded as probably false is telling you where it stops being trustworthy. The study also used a severe-MESH threshold of 29 mm, which is 1.14 inches, for the Witt formulation, not the 1 inch that defines a severe hail report. Thresholds in radar products and thresholds in the report record are not the same thresholds.

How far radar and reports disagree

The gap between the two records is large and it has been measured. Even in plains areas where population density is relatively high, severe hail hours estimated from MESH can be 2 to 4 times greater than those estimated from Storm Data, and where population density is low the differences can be larger. Significant-hail hours show the same 2 to 4 times pattern. MESH estimates of severe hail hours per year are higher than report-based estimates at every location in the contiguous United States, and across much of the High Plains MESH puts the figure on the order of 20 hail hours per year above what Storm Data records.

The paper does not resolve that gap in radar's favor, and neither should anyone quoting it. It says severe hail is plausibly going underreported in the High Plains because of low population density, while allowing that MESH likely overestimates severe hail to some degree. Both of those are true at once. The report side of the comparison is built from damage surveys, emergency managers, and SKYWARN spotters, which is a good process and still a human one. Neither number is ground truth. The difference between them is the size of the uncertainty, not the size of the error in one of them.

Where radar simply cannot see

Radar coverage is not uniform, and the same study shows where it falls off. The difference between MESH and reports reaches relative minima across much of the Intermountain West, the West Coast, the Appalachians and the Northeast, and a lack of radar coverage, from beam blockage or widely spaced radars, accounts for part of that.

Reliability also degrades at the top of the size range. As hail approaches 4 inches, all of the MESH formulations become less reliable, which is awkward because the largest stones are the ones that end up in the most expensive disputes.

The archive's own statements are the shortest version of this article. SWDI adds no quality control beyond archival processing, missing data does not mean no severe weather occurred, and much of the automatically derived data is radar based and represents probable rather than confirmed conditions.

What to ask a radar archive for

Ask it whether a storm with the right structure was over a place at a time. It answers that well, and it answers it for empty country where the report record has nothing to say.

Do not ask it to rule anything out. Missing data does not mean no severe weather occurred, and in the terrain where coverage is weakest the absence means least. Do not ask it for a hail size you intend to present as fact, unless you also name the formulation, the threshold and the variant you queried. And do not treat it as a replacement for the settled report record, which runs 90 to 120 days behind the event and is the thing a certified copy is drawn from. The radar archive and the report archive answer different questions. Used together, with each one's limits stated, they are a reasonable account of what happened. Used as substitutes for each other, they are a way to be confidently wrong in two directions.

Sources

More on this