ProjectSidewalk / ProjectSidewalk/RampNet

Assess location precision for the unassessed candidate cities — the gate every route to 500k depends on

Open
#96 11 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

documentation
Dominant language
Python
Stars
7
Forks
1
Avg merge
4d 11h
Merged PRs (30d)
7

Description

Why this is the critical path

docs/curb_ramp_data_sourcing.md §6 prices the routes to a 500,000-ramp corpus and lands on
one blocking result:

Pool Ramps Cumulative
Good (NYC + Portland + Bend) — already fully used 276,615 276,615
+ every OK city (LA + Austin + DC + Nashville) 193,898 470,513
+ unassessed cities ~236,000 ~706,500
+ state DOTs ~198,800 ~905,300

500,000 is not reachable on assessed data alone. Good plus every OK city reaches 470,513
— ~30k short — and that is already after accepting a quality tier the paper deliberately
rejected. So every route to 500k depends on cities whose location precision nobody has
checked
, which makes this assessment the gate, not discovery. Discovery is done (§3).

The work is cheap and unblocked: a few days of visual review, no GPU, no pipeline run, no
sourcing commitment
. It determines whether 500k is a real target or an arithmetic one. If
roughly 60% of the ~236k unassessed pool passes at Good, 500k clears on city-inventory data
alone with no state-DOT tail.

What "assessment" means here

Paper §3.1 / Table 1 assessed eight cities by manually overlaying curb-ramp coordinates on
aerial imagery
and judging whether they land on the physical ramp (Fig. 2), bucketing Good /
OK / Poor. No thresholds were published. The selection rule was "Good only" — the corpus is
three cities because only three passed, not because only three were available.

docs/curb_ramp_data_sourcing.md §5 proposes quantifying it while extending it, so the tiers
stop being a judgment call:

Check Why it matters Test
Positional offset The Stage 1 mechanism — a misplaced coordinate produces a misplaced label. Fig. 2b is the failure Sample ~50 points/city, measure metres from the true ramp on aerial imagery; report a distribution, not a bucket
Per-ramp vs per-corner One point per corner collapses paired ramps to a single label — the supervision gap behind Paterson's failure Records per intersection; inspect a paired corner. NYC's ~1.8/intersection implies per-ramp
Staleness Ramps built since the capture are missing (recall loss); removed ramps are phantom labels (precision loss) Compare capture date to GSV capture date
Completeness Sets the label ceiling. Stage 1 agreement is P .9403 / R .9245 Spot-check N intersections in GSV for ramps absent from the inventory
CRS / datum A wrong projection silently shifts a whole city Round-trip known points
Active vs retired Seattle publishes 46,386 total but 38,468 active Prefer the publisher's active filter

A quantitative offset distribution would also let the OK tier be re-examined: "OK" may mean
2 m or 8 m, and that difference plausibly decides whether LA's 91,759 records are usable.

What is already done — and what it does not cover

The second, independent gate (temporal distance: did the ramp and the pixels exist at the
same time?) is implemented as scripts/analysis/temporal_gap.py and has been run on seven
cities (§5c, 2026-07-30). That gate is not this one, and a city can pass one and fail the
other — Denver is the clearest case, where "delineated from 2022 aerial imagery" is a
positional concern but a temporal strength.

Temporal verdicts so far: Denver ✅, Sioux Falls ✅, Minneapolis ✅, Austin ⚠️ (84.6% undated),
Nashville ⚠️, Boston ❌ (~12-yr gap), DC ❌ (~6-yr gap).

No city in the list below has had any positional assessment.

Checklist

Unassessed candidate cities, ordered by size (counts are a 2026-07-30 snapshot; endpoints in
docs/curb_ramp_data_sourcing.md §3). Each needs the six checks above.

  • Denver, CO — 72,770. GOOD (median 0.29 m, 92.3% within 1 m, phantom 5.5%; §5f).
    Per-ramp confirmed on imagery — its low 1.21 rec/corner is design vocabulary, not missing
    records. ⚠️ temporal ✅ downgraded to unconfirmed: 74.4% carry a 2015 CREATEDATE and
    Modify is used zero times, so the "2022 aerial imagery" claim does not survive the data
  • San Francisco, CA50,096 DISQUALIFIED. Only 7,553 distinct coordinates, 1:1
    with cnn (intersection node id) ⇒ intersection centroids, ~4.7 identical labels per pixel.
    No imagery needed. 14,414 rows are crexist=0 — wrong polarity for Stage 1, right for #86
  • Charlotte, NC — 40,601. From Charlotte's ADA self-evaluation. Sunbelt/Southeast
    vocabulary, which is where Paterson and Gainesville fail
  • Boston, MA — 24,022. Temporal ❌ (~12-yr gap) — assess only if the temporal verdict is
    contested
  • Sioux Falls, SD — 19,977. Best install-date coverage found (37.6% undated)
  • Minneapolis, MN — 18,447. Richest attributes found, relevant to #86
  • Arlington, VA — 10,342

State DOTs (~199k), lower priority — partial per city and skewed toward wide arterials:

  • VDOT (83,000) — ⚠️ statewide Virginia includes Richmond, a benchmark split and the
    Mapillary OOD flagship. Must be clipped to exclude the deployment footprint before ingest;
    the failure mode is silent
  • NYSDOT (42,297) — needs the same treatment for NYC overlap (already in training)
  • CDOT (24,549) · WisDOT (~49,000, 2014/15 desktop inventory — assess the method
    before the coordinates)

Table 1 re-checks worth folding in:

  • Resolve the LA discrepancy. Table 1 rates LA "OK" at 91,759, but LA's published Access
    Ramps layer is documented as "the geographic center of corner polygon features" derived
    from 2014 aerial imagery — i.e. corner-centroid, not ramp-located. Either a different layer
    was assessed or "OK" tolerates corner-level placement. Settle before counting the 91,759
  • Re-assess Seattle? In progress — partially assessed, not resolved. Rated Poor,
    which disqualifies it, and that is the one axis where its otherwise-ideal profile fails
    (weekly refresh, rich attributes, already contamination-burned so free to train on). Its
    pairing density is the surprise: 1.34 records/corner, above Bend, which is in training.
    Attempted 2026-07-31 and the anchor did not come out. 19 of 60 chips attempted, 11
    measurable, median 2.33 m — but 36.8% of attempted chips were unjudgeable under leaf-on
    King County canopy, with a selection effect toward un-treed corners, on a basemap §5h
    declares adequate only to size a large error. The apparent 87% "systematic shift" was
    investigated and is NOT a registration error
    (§5i): coordinates read 0.00 m against
    Seattle's own street network over 31,430 samples, and the imagery clears at ≤0.32 m at all
    eleven reviewed chips. So there is no constant to subtract, and no rehabilitation by that
    route. Still open, and any argument for Seattle has to start by contesting Table 1:
    a better basemap (Seattle's own EPSG:2926 caches need the sheet's tile math generalised),
    more chips, and the 62.9%-installed-after-2019 confound handled — against 2019 imagery
    a record for a ramp built later has no correct answer

Selection rule (§8) — applies to whatever passes

Train on cities you would never want as a benchmark split. Every city added to training is
permanently disqualified as clean evaluation ground.

  • Already contamination-burned, so free — Seattle, Columbus, Chicago, Pittsburgh, St. Louis,
    Knoxville. Painfully, Seattle is rated Poor and the others publish no inventory, so this
    category currently yields nothing.
  • Registry-clean, so costly — everything in the checklist above. Acceptable, since none is
    among the current nine splits, but it should be a deliberate, recorded decision.
  • Never add Paterson or Gainesville — two of only three GSV benchmark splits, i.e. nearly
    all in-domain-imagery evaluation.

Precision gates diversity. A Good-rated bland city beats a Poor-rated diverse one, because
the Poor city's labels are wrong wherever they land.

Also do at ingest (§9)

Commit a dated snapshot of every inventory assessed, with the fetch URL and query, the way
benchmark/*/records.jsonl already pins benchmark inputs. These are point files — a few MB
gzipped. Bend has drifted +8.7% since the paper, so anyone re-running Stage 1 from the
README links today builds a measurably different dataset and has no way to detect the
difference. Every count in §3 is a snapshot of a moving target.

Out of scope

This issue does not decide whether more data helps — that is #59's E1–E3 (E1 is answered:
the harm hypothesis is not supported). It establishes only which cities are usable if the
answer is yes.

Refs: #59 (does scaling help — parent), #86 (attribute/condition supervision; several of these
inventories carry it), #46 (miss taxonomy — sizes the near-field prize the sourcing programme
targets), #11 (the temporal ordering filter, which is not either gate here).

🤖 Generated with Claude Code (claude-opus-5[1m])

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with docs/curb_ramp_data_sourcing.md §§3 and 5, then review scripts/analysis/temporal_gap.py only to distinguish the existing temporal gate from this positional assessment. Review the listed city inventories against aerial imagery, perform the six checks, and record dated snapshots, offset distributions, per-city verdicts, and resolved LA or Seattle questions in the sourcing documentation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
computer-vision, data-engineering, documentation
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.