· architecture, operations

Scale sources, not requests

When sources multiply, I schedule one polite job per source. I do not speed crawls, and I do not put raw HTML on a queue.

Problem

A source is a place I collect from. A government app, a vendor listing screen, a file drop. It is not a URL, and it is not the raw HTML. I need court PDFs and tax files from a lot of those places without rewriting extract, match, or the lists. A filing still joins a list of properties only after I match it to a parcel. That stays. It is not the fork.

Collection already waits between calls so I do not get blocked. It must not run as fast as it can. Scale is more sources, not more requests per source.

Cron around a handful of scripts is honest while the job count is tiny. At hundreds of sources, one hung job stalls the night, and the crontab file becomes the catalog. The fork is how I schedule that work.

I build property lists today. Shared landing: the same file could match another grain later. Not this post.

Options considered

  1. One crontab line per source — cheapest while the set is small. Hundreds of lines are not something I can read. One hung source stalls the whole night.
  2. A queue of raw HTML — every page body or PDF as a message, so workers crawl faster. That scales requests. I cap requests on purpose. Unmatched filings would ride the same bus as inventory.
  3. One job per source — something starts gather plus extract for that source, in process. The waits stay. A failure stays in that job. The tables the app already reads still get rows.

Decision

I went with option 3.

Option 1 worked until the job list became a second codebase. Option 2 looked like a pipeline and optimized the wrong thing: requests per source, not sources. Loaders such as Airbyte, dlt, or Singer taps help for documented APIs and file drops. They do not replace portal crawlers, so they are not the move. A full hosted integrator fails the same way on clerk sites that keep state.

The work unit is a source (or a source and a day), not a URL. I can run a lot of polite jobs at once. I cannot run a lot of faster request loops. The site will block me. Match stays in the same job. Unmatched stays out of lists and off any bus. The qualification gate does not move. The app stays as it is.

I would do this in steps. I keep the collectors. That is the advantage. A thin runner, with retries and a note when a job fails, can make collect then extract then transform one graph. Cron still works while the job count is small. I do not need a huge scheduler deploy, a crawler fleet, or a message bus to get there.

How it works

A new court portal is one more source job, not a faster crawl of an old one. The runner starts that job. Collection waits between requests, stores the PDF, and does not decide the parcel. Extract and match run in the same job. If the identifiers match a parcel, the filing can join a property list. If they do not, I keep the file. I do not enqueue the unmatched filing. A hung neighbor retries on its own. It does not stall this job.

flowchart LR
  src[Public sources] --> orch[Runner]
  orch --> job[One source job]
  job --> gather[Collect then extract]
  gather --> db[Same tables]
  db --> compose[Transform]

Hard parts

  • It is easy to treat “scale” as faster crawls because the machine is idle. That burns the source.
  • A source job that also writes “ready for lists” rows is a one-stage scraper again. Landing still does not mark inventory.
  • Wrapping the scripts I have is a smaller step than renaming everything. Either way, the noun stays the source, not the URL.
  • File-shaped sources already import into tables. They can keep doing that. Where the raw layer lives is a later fork.
  • A large scheduler deploy and a crawler fleet cost more ops than a product that is fine with week-old data.

What we’d change

I would name the source as the thing I retry, before the crontab file became the catalog. I would add retries, run ids, and a count of failed jobs before I invent a bus. I would not put raw HTML on a queue to get there. When files want object storage and match wants its own step, that is a different post: queue grain, not raw HTML.

References

← All notes