· architecture, vision

Signals on a grain

Scripts and address lookup could inspect a parcel. They could not make one list of parcels with a court filing, an assessor column, and an absentee-owner flag. That mixed list needs a grain.

Problem

The system I built captures public real-estate records and turns them into lists. A buyer wants one file: properties that are absentee, and have a court filing, and sit in a value band. Those facts already exist. They show up as portals, PDFs, and county files. They do not show up as a list.

For context, I started with scripts. One address, one screen per source. That can inspect a parcel. It cannot answer the mixed question. The product is the combined list, not a stack of document screens. A filing I cannot match to a parcel stays on disk. I do not sell it as a property.

I locked this in when the lookup died on the first mixed list, and a frozen “one schema on day one” died on the first ugly PDF.

Grain: one list row is one property, so signals can be layered. Match before the list: the PDF has no shared key, so I resolve identity. Unmatched stays out. Collectors: land the file, do not decide inventory. A bad parse does not stop ingest. A new source does not copy list rules. Wide property row (OBT): the buyer filters one row instead of joining a maze of ugly tables. Preview, export, and masking all read that same row.

I build property lists today. Shared landing: the file is not owned by this list. A later product would be another grain and its own spine. Not a second scrape. Not this post.

Options considered

  1. One product per source — a courts screen, a roll screen, an export per file type. Fast while the script count is small. It cannot answer absentee and filing and value on one list.
  2. One schema on day one — force every feed into the same tables before I have seen the PDFs. Looks tidy. Breaks on the first ugly docket.
  3. Signal layering on a grain — land raw records, extract signals, match them to the property, query one serving row. Unmatched stays out.

Decision

I went with option 3.

Option 1 was the lookup. It inspects a parcel. It does not make a mixed list. Option 2 froze the schema too early. The first ugly PDF had no place to land.

Collectors and the wide row are not a second product. They are how option 3 survives more sources. Unmatched stays out: no parcel match, no list filter.

How it works

Land, then extract, then match to a parcel, then one serving row, then the list.

A collector lands the PDF. I extract fields, then match. A hit writes flags onto the property row. Preview, export, and masking all read that same row. No match: I keep the file. I do not put it on the list. Buyers get properties, not documents.

flowchart LR
  pdf[Court PDF] --> land[Land the file]
  land --> extract[Extract signal]
  extract --> ok[Match to parcel]
  extract --> no[Unmatched]
  ok --> row[One property row]
  row --> list[Layered property lists]

Hard parts

Disagree on the grain and I invent properties that do not exist. Identifiers fight (address, APN, legal description, parties): a loose match duplicates; a strict match drops real filings. Two sources disagree on the same parcel. The list needs a rule, not last-write-wins. The pressure to “just show the filing” is the lookup again. Documents, not properties. A new extractor with no match and no unmatched state fills the lists with junk. Slowly.

What we’d change

I would write the grain contract sooner: which signals are list filters, and what a matched event means. Extractors can stay messy. If flags on the wide row are not enough later, I would add satellites, still projected onto the grain for list paths. A later domain gets its own matching rules on the same landed files. Not a second philosophy.

References

← All notes