Storm Stream.

science

MESH Versus Ground Truth on Hail Size

What MESH measures, why radar hail hours run 2 to 4 times above Storm Data, and why two vendors quote different sizes for one storm.

Somebody tells you the radar showed two inch hail over a subdivision. Somebody else tells you the official record has no hail report within miles of it. Both statements can be true at once, and the gap between them is large, measured, and published. The number that closes the gap is not a secret. It is in a peer reviewed paper that SPC hosts itself.

What MESH is looking at

MESH stands for maximum estimated size of hail. It estimates hail size from radar reflectivity properties of a storm above the environmental 0 degree C level, and it is produced by the Multi-Radar Multi-Sensor system. It is used widely across the NWS, including at SPC, to diagnose expected hail size.

Read the definition again and notice what is absent. Nothing in that chain touches a hailstone. The radar measures returned energy from a volume of air above the freezing level, and an equation converts that into a diameter. The output is in inches, which makes it look like a measurement. It is a conversion, and the quality of the conversion is the whole subject.

The gap is not a rounding error

The clearest comparison of radar estimates against the official record is Wendt and Jirak, published in Weather and Forecasting in April 2021, covering 2012 to 2019 over the contiguous United States. Their headline comparison is blunt. Even in plains areas where population density is relatively high, severe hail hours estimated from MESH can be 2 to 4 times greater than those estimated from Storm Data. Where population density is low, the differences can be higher.

The same multiple shows up at the larger size class: MESH significant-hail hours are similarly 2 to 4 times larger than estimates built from significant hail reports. For context, severe hail is 1 inch or larger and significant severe hail is 2 inches or larger.

Two more findings from the paper are worth memorizing, because they come up in arguments. MESH estimates of severe hail hours per year are higher at every location in the contiguous United States than estimates from reports. And across much of the High Plains, MESH puts the figure on the order of 20 more hail hours per year than Storm Data records.

Both numbers are wrong, in opposite directions

This is the part that usually gets dropped when somebody quotes the 2 to 4 figure at you. The paper says severe hail is plausibly going underreported in the High Plains because of low population density, while allowing in the same breath that MESH likely overestimates severe hail to some degree.

That reading lines up with what SPC publishes about its own database: reports cluster around population centers, and the data are used for warning verification and may not capture all storm events. So you have a report record that is thin where nobody lives and a radar record that leans high. Neither one is the truth. The difference between them is not a correction factor you can apply.

There is also a definitional seam. The study used a severe MESH threshold of 29 mm, which is 1.14 inches, for the original algorithm. The report threshold is 1 inch. Those are not the same line, and a comparison of counts across them carries that difference inside it.

MESHWitt, MESH75, MESH95: three fits, three answers

The original algorithm came from Witt and colleagues in 1998, and the paper calls it MESHWitt. Later work produced two new power-law relationships: MESH75, fit to the 75th percentile of the sampled hail size distribution, and MESH95, fit to the 95th percentile.

That is the mechanical reason two vendors can quote different hail sizes for the same storm over the same address. They are not disagreeing about the radar. They are reporting different percentiles of a distribution. A 95th percentile fit is answering a question about the largest stones that plausibly fell. A 75th percentile fit is answering a question closer to the typical large stone. Both are defensible. Neither is "the hail size."

The performance details complicate it further. Near the 1 inch threshold, 25.4 mm, MESH75 and MESH95 predict hail size more reliably than MESHWitt, but they are more likely to overestimate hail size between 1 inch, 25.4 mm, and 2 inches, 50.8 mm, than MESHWitt is. MESHWitt has a higher critical success index than both of the newer fits at the exact 25.4 mm and 50.8 mm thresholds. And as hail approaches 4 inches, 101.6 mm, all of the formulations become less reliable.

Thresholds move too. The paper cites Murillo and Homeyer (2019), who found MESH75 performs best at 40 mm for severe hail and 47 mm for significant hail, and MESH95 at 64 mm and 83 mm. Different studies, different operating points, same word on the label.

The practical consequence is that "radar-estimated hail size" is not a specification. The specification is the formulation plus the threshold.

The quality control shows you where the noise lives

The study's filtering is informative in its own right. MESH pixels had to fall within 40 km of a detected cloud-to-ground lightning flash in the same hour to be counted. Values above 127 mm, 5 inches, were removed as likely spurious.

Both screens are admissions. The lightning proximity filter exists because the algorithm will return hail-like values in places where there is no deep convection to make hail. The 5 inch cutoff exists because the top of the output range produces values that researchers do not believe. If you are handed a raw radar hail product and it contains values above 5 inches, you are looking at something the literature throws away.

The archival side carries the same warnings. NCEI's Severe Weather Data Inventory hosts NEXRAD Level-III hail signatures in a filtered version, keeping maximum size greater than zero and probability 100%, and an unfiltered all-signature version. The product page states that SWDI adds no quality control beyond archival processing, that missing data does not mean no severe weather occurred, and that much of the automatically derived data is radar-based and represents probable rather than confirmed conditions.

Where the radar itself cannot see

Radar coverage is not uniform, and the paper is explicit about it. Relative minima in the MESH-minus-reports difference show up across much of the Intermountain West, the West Coast, the Appalachians and the Northeast, and lack of radar coverage, from beam blockage or widely spaced radars, accounts for part of it.

In mountain country, a low radar hail estimate can mean the beam was blocked or the nearest radar was far enough away that it was sampling well above the storm. That is a geometry problem, not a weather observation. The same quiet radar output that looks like good news for a denial is, in those regions, partly an artifact of where the radars sit.

Knowing which number you are holding

Neither number is ground truth. The report record tells you a person was there. The radar estimate tells you an algorithm converted reflectivity aloft into a diameter. Useful work starts with knowing which one is in your hand and what its known failure mode is.

When someone cites a radar hail size, three questions settle most of it. Which formulation produced it, MESHWitt or a percentile fit, and at what threshold. Was it screened for lightning proximity and for implausibly large values. And is the location one where coverage is known to be weak.

When someone cites the absence of a report, the question is simpler: who would have been there to make it. And when someone cites a report's size, remember that hail is often estimated by the public by comparison to a reference object of known size, against pairings like a quarter at 1 inch and a golf ball at 1 3/4 inch. That is the precision available, and it is worth more than a false decimal.

Sources

More on this