The shell scripts are already the pipeline
The data path is already a medallion pipeline. An orchestrator would rewrite the scheduler, not the system.
The system I developed for (proptech data hub)[https://proptechdatahub.com/] captures all kinds of realestate datasets from public sources. It evolved from a bunch of highly specific python scripts, and as a results still has lots of manually written code.
For context I started producing those scripts when AI agents for coding were not a thing. On top of that I wanted to dump as much as possible of my professional experiences in it when I decided it was time for a proper application setup, thus many manually written code ;).
While it is currently in a good stable state in production: nice automated data collection, expandable and flexible enrichment module and finally a nice UI that allows to browse and create poperty lists selecting features, signals, etc and then dump it in a file of your choosing format; I still got that scratch I need to itch, and formally, there is always the need in the lifecycle of any app to make sure there are no missed opportunities to improve it, So I prompted the ai (Grok 4.7) with the following prompt:
> Can the whole system be upgraded to follow a modern data engineering application pattern such as data pipelines?
Context: At this point the agent already have a good handle of my aims when I interact with it.
And this is the response I got:
The data path already is a pipeline. A full move onto an orchestrator would rewrite the scheduler, not the system, and that rewrite is not justified by how the jobs run today.
What you have is the usual medallion layout:
- **Ingest.** Collectors pull public-record sources into raw warehouse schemas.
- **Enrich.** Enrichment jobs extract signals and write the property-search index.
- **Transform.** The transform project builds bronze reference tables, silver entities, then gold tables the API serves, plus a feature layer. Incremental models advance on a load cursor.
- **Serve.** The API and the marketplace read gold. They are not pipeline steps.
The pipeline shell scripts are the DAG. Court work is gather, then process, then enrich. Assessor loads on the batch machine are supposed to finish before the transform run on the app machine. The two crontab files are the orchestrator. An earlier in-app scheduler was removed on purpose. Crontab is the only supported runner.
A “modern data engineering” upgrade usually means Airflow, Dagster, or Prefect: one graph, a run starts only after the upstream task succeeds, retries, a run history, and a UI. The tasks would stay the same Python modules and the same transform project. The shell scripts and the two crontabs would become that graph.
That graph is small and awkward for those tools:
- About a dozen jobs, on two machines, with the batch machine asleep outside the batch window.
- The cross-node handoff is a clock, not a dependency. Weekly assessor loads are expected to finish shortly before the transform starts on the app machine. An orchestrator would replace that guess with “the transform starts after every assessor load exits 0.” That is the one real gain.
- Collectors need a checkout, a virtualenv, local caches, and a local model server on the batch machine. They are not small container tasks you can ship to a worker pool.
- The app machine has little spare memory. A scheduler process there sits next to the API. Putting it on the batch machine means the scheduler is suspended every afternoon, so nothing can notice a missed run until the box wakes.
The marketplace, login, cache, and public edge should stay an application. Folding them into a pipeline framework does not make the product more modern.
Leave the collectors, enrichment jobs, and transform project where they are. If the transform ever starts before the assessor loads have landed, or one court job fails silently inside the shared log, the fix is a completion signal between those existing scripts. A new orchestrator is the step after that signal is not enough.
I mean, it could’ve sold me on the newest, the latest, and the bestest; instead I got a very conservative response, with which I aggree with wholely.
I personlally try to avoid touching working systems, but I am open to the idea of upgrading the system to a more modern data engineering pattern. One of the reasons I find modern AI tools for development so powerful.