data
Geocoding Addresses Against Storm Polygons
The engineering of point-in-polygon address matching, the real limits of the free geocoders, and the failures that return a wrong answer silently.
The question looks simple. Was this address inside that polygon? In production it is two questions bolted together, and both of them have failure modes that return a confident answer instead of an error.
The first question is where the address is. That is a geocode: a string in, a coordinate pair out. The second is whether that coordinate pair falls inside a given shape. That is a point-in-polygon test, and it is pure geometry. Each step is cheap and well supported. The trouble is that the geometry step has no way to know how good the coordinate was, and the geocode step has no way to know how close to a boundary it landed.
The Census Geocoder
The US Census Geocoder offers interactive and REST access, and splits into two groups. Find Locations returns coordinates. Find Geographies returns the census geography a location falls in as well. Within those groups the options include a one-line address, a stateside parsed address, a Puerto Rico parsed address, batch addresses, geographic coordinates and batch coordinates.
For working a list of properties, the batch paths matter most. Batch address processing accepts a batch of addresses up to 10,000, and batch geographic coordinates accepts up to 10,000 coordinate pairs. Accepted file formats are CSV, XLS, XLSX, Text and Data files. The same guide describes the two settings people skip: you select a benchmark and a vintage, and the vintage options depend on the benchmark you chose. Those two settings determine which snapshot of the reference data your coordinates came out of, which is worth writing down next to any result you may have to defend.
One hard boundary is worth knowing before you build on it. The service only geocodes addresses that are within the United States, Puerto Rico, and the U.S. Island Areas. For US storm work that is usually fine. It is not a general-purpose geocoder.
Nominatim, and the policy that comes with it
Nominatim is the other common free option, and its usage policy reads as an engineering specification rather than legal boilerplate.
The OpenStreetMap Foundation's policy sets an absolute maximum of 1 request per second. The unit is the important part. The limit applies per website or application, so the traffic of all your users combined must stay under it. One request per second is not one request per second per customer.
Scripts that run longer than a day, or on a schedule, are capped at 4 requests per minute. A valid HTTP Referer or User-Agent identifying the application is required, and the policy is blunt about what does not count: stock user agents as set by HTTP libraries will not do.
Then come the parts that decide whether Nominatim belongs in your pipeline at all. The policy says bulk geocoding of larger amounts of data is not encouraged. Small one-time jobs may be allowed, using a single thread on one machine, with no distributed scripts. Results must be cached on your side, and clients that repeatedly send the same query may be treated as faulty and blocked. Heavy use may lead to access being withdrawn, and the policy may change without notice.
Read as requirements, those lines are clear. Caching is mandatory, not an optimization. Your request rate is a global property of the whole application, not of any one user. And a service whose policy may change without notice is not a dependency to put underneath a billing process.
The geometry that decides the answer
Once you have a coordinate, the rest is geometry, and GeoJSON specifies it tightly.
A position is an array whose first two elements are longitude then latitude, with an optional third element for altitude in meters. The default coordinate reference system is the WGS 84 datum in decimal degrees, and alternative coordinate reference systems were removed in this version of the specification, so there is no projection to negotiate.
A linear ring is a closed LineString with four or more positions, and its first and last positions must have identical values. The first ring of a polygon is the exterior ring, and additional rings are interior rings, meaning holes. Exterior rings run counterclockwise and holes clockwise, by the right-hand rule, though parsers should not reject polygons that ignore it for backward compatibility, which means winding order cannot be used to identify which ring is which.
And there are seven geometry types: Point, MultiPoint, LineString, MultiLineString, Polygon, MultiPolygon and GeometryCollection. Feature and FeatureCollection are GeoJSON types but not geometry types.
Where it quietly goes wrong
This is the part worth keeping next to the code.
Swapped coordinate order. Longitude comes first in GeoJSON, and almost everything a person writes by hand puts latitude first. Swap them and no parser complains. The point moves to a different part of the world, every test returns false, and the run reports that no addresses were affected. A universal negative is the signature failure of this pipeline, and it is indistinguishable from good news.
A ring that is not closed. Four or more positions, first and last identical. Build a ring from a vertex list without re-appending the first vertex and you do not have a polygon. Round coordinates on serialization and a ring that was closed can stop matching exactly. Either way the shape is invalid, and what happens next depends on how forgiving your library is.
A geocode that is not the parcel. A geocoder may resolve an address to a point interpolated along a street, or to the centroid of a larger area such as a ZIP code, rather than to the building. Inside a wide polygon that difference rarely changes the answer. Near a boundary it decides the answer. The test will still return a crisp true or false, about the centroid.
Points on or near the edge. A point exactly on a boundary is ambiguous in principle and library-dependent in practice. More common, and more dangerous, is a point just inside or just outside, where the geocode error and the polygon error are both larger than the distance to the edge. The output of the test carries no uncertainty. The inputs carry plenty.
A polygon that was never a measurement. If the shape came from a radar-derived product, the shape itself is an estimate. NCEI's Severe Weather Data Inventory states that it adds no quality control beyond archival processing, that missing data does not mean no severe weather occurred, and that much of the automatically derived data is radar-based and represents probable rather than confirmed conditions. Radar hail estimates work that way by construction: MESH, the maximum estimated size of hail, estimates hail size from the reflectivity properties of a storm above the environmental 0 degree Celsius level. Ground reports have their own softness, and the same paper notes that hail is often estimated by the public by comparison to a reference object of known size.
A true that came from an interpolated coordinate tested against a probable-conditions polygon is an exact statement about two estimates. The exactness belongs to the arithmetic, not to the roof.
Making the result defensible
None of this argues against running the test. It argues for recording enough that the answer can be examined.
Cache the geocode and store the coordinate rather than re-deriving it on each run. Nominatim requires caching outright, and the reason holds for every geocoder: a result you cannot reproduce is not a result. Store the coordinate next to the exact address string you sent.
Record the match quality. Geocoders tell you something about how they resolved an address, whether that was a direct match, an interpolation along a street, or a fallback to a larger area. A row that says the point came from an area centroid is a row you can triage. A bare latitude and longitude is not. If you used the Census Geocoder, record the benchmark and the vintage with it.
Record which polygon and which pull. Store the geometry you tested against and the time you retrieved it, not just the name of the source. Live feeds change and archives get reprocessed, and the question somebody asks later is what you tested, not what the source says today.
Treat edge results differently from interior results. A point well inside a large polygon tolerates a sloppy geocode. A point near the boundary does not, and it belongs in a separate bucket that routes to a better geocode or to a person, not into the same column as the confident cases. You need to stop presenting two different kinds of answer as one kind.
That is the whole discipline: a reproducible coordinate, a recorded match quality, a stored shape with a timestamp on it, and a different path for the cases that sit on a line. Storm Stream is an API over these same public feeds: its point-in-swath query accepts a list of the caller's own points and returns which of them each swath contains.
Sources
- https://geocoding.geo.census.gov/geocoder/
- https://www2.census.gov/geo/pdfs/maps-data/data/Census_Geocoder_User_Guide.pdf
- https://operations.osmfoundation.org/policies/nominatim/
- https://www.rfc-editor.org/rfc/rfc7946.html
- https://www.ncei.noaa.gov/products/severe-weather-data-inventory
- https://www.spc.noaa.gov/publications/wendt/meshwaf.pdf