Extraction accuracy methodology
How accurate is ZoneSheet, exactly?
Two rounds of testing, on documents we can name, with the field-by-field results, including the ones ZoneSheet got wrong and how those misses actually surface to a user. Last run and verified: 2026-06-16.
What was measured
Field-level accuracy: for a given zoning PDF, did the extracted value for each field (zoning district, front/side/street-side/rear setback, max building height, floor area ratio, lot coverage, allowed uses, parking minimums, key fees) match the source document, or come back null when it genuinely wasn’t there, with zero fabricated values. Two rounds, on two different kinds of documents.
Round 1: two clean sample PDFs
Tested 2026-06-16 with claude-sonnet-4-5 (the OpenAI gpt-5.5 path rejected temperature=0 and json_object at the time). These two documents were built for testing, not pulled from a live government server, which is exactly why Round 2 below exists: clean, well-formatted samples answer "does the extraction logic work in principle," not "what happens with the PDF a real user actually uploads."
| Document | Fields tested | Correct | Accuracy |
|---|---|---|---|
| Austin, TX — LDC dimensional standards (SF-2) | 13 | 12 correct, 1 partial | 92.3% |
| Denver, CO — Zoning Code (S-SU-A) | 12 | 12 | 100% |
| Combined | 25 | 24 correct, 1 partial | 96% |
The one partial: Austin has two side setbacks (5 ft interior, 10 ft street-side) and the schema at the time had a single setback_side_ft field. The model picked the interior value and flagged the street-side value in its notes, which is correct behavior for a schema gap, not an extraction error. The schema now has a dedicated setback_street_side_ft field (used correctly in Round 2, Test 2 below).
Round 2: real government PDFs
Four documents pulled from live municipal servers, tested 2026-06-16. Honest note on sourcing: most major city PDF servers tested (Nashville, Portland, Charlotte, Raleigh, Phoenix, Fort Collins) returned HTML login pages or 403 errors to a direct fetch. Albuquerque, NM was the reliable source found, so this round is one city’s ordinance tested at two granularities, plus two edge-case robustness tests on non-zoning documents. That is a real limitation of this test corpus, stated plainly rather than dressed up as broader coverage.
| Test | Document | Before the fix | After the fix |
|---|---|---|---|
| 1 | Full 672-page Albuquerque IDO 2023 | 54.5%: all 5 dimensional fields null | Front 10, side 5, street-side 10, rear 10, height 26 ft — all high confidence, verbatim-cited |
| 2 | 22-page zone-districts subset of the same ordinance | 90.9% (10 of 11 fields, 1 partial) | Unchanged: R-1B exact match — front 15, side 5, street-side 10, rear 15, height 26 ft |
| 3 | Scanned PDF, no text layer (unrelated 76-page housing study, not a zoning code) | JSON of nulls, no explanation | Structured error scanned_pdf_no_text_layer, no wasted LLM call |
| 4 | Wrong document: a construction-progress report a URL implied was zoning | 11 of 11 nulls, 0 hallucinations | Unchanged: 11 of 11 nulls, 0 hallucinations |
The Albuquerque walkthrough
The centerpiece case, because it is the one that actually broke something and shows how the fix worked.
A user uploading the full City of Albuquerque, NM IDO 2023 zoning ordinance is a completely realistic case: a typical municipal ordinance runs 200 to 800 pages. Before the fix, ZoneSheet truncated every document to the first 30,000 characters before sending it to the model, roughly the first four to five pages: a table of contents, an amendment log, an adoption page. The Albuquerque IDO’s per-district dimensional-standards table starts around character 180,000. Every dimensional field came back null. Zero hallucinations, but zero usable data.
The fix: replace first-30k truncation with section targeting. The document is split into line-aligned windows and scored by zoning-field diversity, so a window holding front, side, rear, height, and coverage together (the signature of a per-district summary table) outranks verbose prose that just repeats the word "setback." The highest-scoring windows fill a 120,000-character budget, reassembled in document order. On the full Albuquerque IDO, this pulls the real R-1 dimensional table (window 22 of 280) into what the model actually sees. Re-run after the fix: front 10 ft, side 5 ft, street-side 10 ft, rear 10 ft, height 26 ft, every value with a verbatim citation, zero hallucinations.
"the actual numerical values for setbacks, height limits, FAR, lot coverage, specific allowed uses, parking requirements, and fees are not included in this excerpt — they would be found in the body sections referenced." — the model’s own notes field, Test 1, before the fix. This is the behavior we want when data genuinely isn’t visible: an honest explanation, not a fabricated number.
What failed, and how failures surface
A failure a user can see and understand is worth more than a slightly higher accuracy percentage. Here is every failure mode found, and what ZoneSheet actually does when it happens.
| Failure mode | What the user sees | Status |
|---|---|---|
| First-30k truncation missed mid-document tables | All dimensional fields null on a full-length ordinance, with a low-confidence flag and an honest note | Fixed 2026-06-16 (section targeting) |
| No OCR fallback for scanned PDFs | Previously a silent JSON of nulls; now a structured scanned_pdf_no_text_layer error before any LLM call is spent | Fixed 2026-06-16 (detection, not OCR) |
| OpenAI gpt-5.5 rejected temperature=0 and json_object, forcing every call onto the Anthropic fallback | No visible difference, but the intended primary model was never actually used | Fixed: pinned primary path to gpt-4o |
| Allowed-uses sometimes cited from a district’s purpose paragraph rather than the actual use table | A directionally correct value with a paraphrased, not verbatim, citation | Open: needs table-vs-paragraph citation discipline |
| Schema returns one district row even when a jurisdiction has sub-zones with different setbacks | Two correct-but-different extractions of the same ordinance, as in the Albuquerque walkthrough above | Open: needs multi-district output mode |
Confidence scores are calibrated, not decorative
A null value on the full untargeted Albuquerque document came back low confidence (data might exist, the model couldn’t see it). A null value on the wrong-document test came back high confidence (this field genuinely doesn’t apply to a construction report). That distinction is the point: "I don’t know" and "this doesn’t exist" are different findings and ZoneSheet is built to tell them apart rather than collapsing both into the same blank field.
Frequently asked
Where does the 96% number come from?
From a controlled test on two agent-constructed sample PDFs (an Austin, TX dimensional-standards table and a Denver, CO zoning code excerpt), tested 2026-06-16: 24 of 25 fields correct, zero hallucinations. It measures whether the extraction logic works when the document is clean and well-formatted. It is not a claim about every real-world upload, which is why this page also shows the real-municipal-PDF numbers below.
What happens when ZoneSheet can’t find a value?
It returns null with a confidence score rather than guessing. Across every test run to date, zero values have been hallucinated: when the model could not see a field, it said so, whether that meant "this genuinely isn’t in the document" (high confidence) or "this might exist somewhere I couldn’t see" (low confidence).
What happens if I upload a scanned PDF with no text layer?
The pipeline detects it before calling the LLM (by checking characters extracted per page) and returns a structured error, scanned_pdf_no_text_layer, instead of a silent JSON of nulls. There is no OCR step yet: a scanned document needs to be run through OCR first before ZoneSheet can extract from it.
Should I rely on ZoneSheet’s numbers without checking the source?
No. ZoneSheet is a research aid, not legal or land-use advice. Every extracted value carries a citation to the exact line of your source document specifically so it can be checked in two clicks, and it should be, before it goes into a permit submission, a bid, or a design decision.
What this page is, and isn’t
This is our own internal testing, run and written up by us, not a certified third-party audit. The real-world corpus is currently one city’s ordinance (Albuquerque, NM) tested at two granularities, plus two edge-case documents chosen to test failure handling rather than typical accuracy. That is a real limitation: it shows the pipeline works and fails the way we say it does on the documents we tested, not a statistically representative sample of all 30,000-plus US jurisdictions. We are expanding the real-document test set over time and this page will be updated when we do. In the meantime: verify every extracted value against your source ordinance before relying on it.
Run it on your own PDF
Paste zoning text into the free demo and check the citations yourself.
- Internal researchZoneSheet Extraction Accuracy Report: two agent-constructed sample zoning PDFs (Austin, TX and Denver, CO), tested 2026-06-16. 24 of 25 fields correct (96%), zero hallucinations.
- Internal researchZoneSheet Real-World Extraction Accuracy Report: four real government PDFs, including the 672-page City of Albuquerque, NM IDO 2023, tested 2026-06-16; fixes re-verified same day. Re-verification scripts: claude_scripts/2026-06-16-zonesheet-verify-extraction.py, -verify-regressions.py, -verify-logic.py.