Meridian

Architecture · Runbook

From a new data drop to a deployed site

One command runs the whole thing in the right order:

python3 run_pipeline.py

--dry-run prints the plan without running it. --from <step> resumes after a failure — the script tells you the exact command when a step fails. --only payloads rebuilds just the JSON the app reads, which is the common case when a builder changed but the source data did not. --with-lanxess adds the catalog steps, skipped by default because that spreadsheet changes on a different cadence.

Order is not a style preference, which is why it is enforced in code

scripts.prospect.index reads public/dashboard-data.json to decide which companies count as covered accounts. Run it before scripts.opportunity.build and it silently indexes against the previous run's account list — no error, just a wrong account flag on every company, which surfaces much later as a supplier page claiming it reaches no covered accounts. The dependency is declared in run_pipeline.py, and selecting a subset that leaves a dependency out prints a warning naming it rather than failing quietly.

The run checks itself, but read the check correctly

Every step is idempotent, so after a logic-only change the script should finish by reporting public/ is unchanged — the proof the change was a no-op on the full corpus rather than only on the sample you tested. A diff does not mean your logic changed something. It means the payloads differ from the last commit, which is equally consistent with the inputs having changed underneath you. Check input/ before blaming the code.

Two exports of the same data double-count everything

The most expensive mistake this pipeline permits, and it has already happened: two files landed in input/ holding the same 24,059 Owens Corning declarations — one 20-column with ISO dates and plain numbers, one 257-column with day-first dates and comma-formatted values. Nothing matched textually, so no duplicate flag fired. Both Owens accounts came out at exactly 2.00× — 4,601 rows became 9,199, $150.7M became $301.1M — and the inflated pair displaced five genuine accounts out of the covered set.

exact_duplicate could never catch it: it is computed per file. The consolidation step now compares declarations after normalization — parsed dates, canonical values, resolved company names, the only level at which two differently-formatted exports of one declaration look alike — and writes output/cross_file_duplicates.json naming the overlapping files. It reports and never drops: which export is authoritative is a question about provenance the pipeline has no basis to answer.

The rest of this page is what the script runs, in order, and what each step can get wrong. Two of these steps did not exist as code until recently — the Excel conversion and the Lanxess catalog were run by hand, which meant lanxess-catalog.json was committed with no way to rebuild it. Both are now in the repository and both reproduce the committed payloads byte-for-byte.

1 · Convert the drop to CSV

Drop .xlsx files into input/excelfiles/. Existing CSVs are skipped unless --force is passed.

python3 -m pipeline.excel_to_csv

~30s for 26 files. pandas writes dates as ISO YYYY-MM-DD while older hand-made CSVs are day-first; both parse only because "%Y-%m-%d" is in date_formats. A source with a third date format will parse to NaT silently — the review queue in step 3 is where you would notice.

2 · Rebuild the knowledge base

python3 manage_knowledge_base.py build --input-dir input --knowledge-base knowledge_base

Generates observed country aliases, entity candidates, and the observed HS hierarchy. Candidates are not used by processing. If new companies arrived in this drop, review knowledge_base/entity_alias_candidates.csv and promote what is correct into entity_aliases.csv before step 3 — afterwards is too late for this run.

3 · Normalize

python3 process_data.py --all

The long step — it writes output/<file>/ per source plus the concatenated output/trade_clean_all.csv, currently ~540 MB. Then read the review queues before trusting anything downstream: a queue that suddenly jumps from 2% to 100% of rows means a format changed, not that the data got worse.

4 · Build the payloads

python3 -m scripts.opportunity.build     # must be first
python3 -m scripts.prospect.index        # reads dashboard-data.json
python3 -m scripts.analytics.cube
python3 -m scripts.analytics.entities

The last three are independent of each other; only the first two are ordered.

5 · Rebuild the Lanxess catalog

Only needed when input/Lanxess Products.xlsx changes. It does not read the customs pipeline at all — it is the one payload built from a customer spreadsheet.

python3 -m scripts.lanxess.catalog        # ~0.5s — writes output/ and public/
python3 -m scripts.lanxess.analysis      # ~1s — every Lanxess figure
python3 -m scripts.check_payloads        # ~0.4s — fails the run on a wrong number
python3 -m scripts.check_selftest        # ~5s — proves those checks still fire

The second is where the Lanxess numbers are computed — not a check on them. It used to be an oracle that recomputed the figures in Python while the browser computed its own in TypeScript, with agreement between the two maintained by hand; it now writes lanxess-analysis.json and lib/lanxess.ts derives nothing. It currently reports 709 clean opportunities worth $911,262,098 across 24 accounts — which is what /lanxess displays. The count fell from 803 when MIN_OPPORTUNITY_VALUE put a $1,000 floor under a demand row, taking $17.6K of one-declaration noise with it.

6 · Check, then commit the payloads

npx tsc --noEmit
npm run build
git add public/ && git commit

npm test is an alias for npm run build — it proves the code compiles, not that a number is right. What checks the numbers is scripts/check_payloads.py, which runs as the last pipeline step and exits non-zero on a wrong one: the opportunity floor, surviving cross-file duplicates, a company buying from itself, the hierarchy trees agreeing with the heat matrix, and a ±25% collapse guard on every headline count. Payloads only reach production when someone commits public/, so a failed run is a gate. Three defects that shipped for months — $17.6K of sub-dollar rows, $35.9M of duplication, $59.6M of intra-group trade — are each covered by an assert that has been seen to fail. That last part is enforced: python3 -m scripts.check_selftest reintroduces every defect against a temporary copy of the payloads and requires the checker to catch it, because a check nobody has seen fail is decoration.

The payloads are committed on purpose

Vercel builds from git and cannot regenerate them: that needs Python, pandas and the 540 MB intermediate CSV, none of which exist in the build image. Ignoring them would deploy a site where every page hangs on “Loading…”. The contracts page tracks their total against the 50 MB ceiling →

7 · Deploy

Push. Vercel builds from the repository root — framework preset Next.js, Node 22.x, no environment variables, no database. Root Directory must stay at the repository root; the Python package is named pipeline/ rather than src/ precisely because Next.js would treat src/app as a source root and take over the build from app/.

Measured timings

StepWall clock
excelnot in the last run
knowledge-basenot in the last run
normalizenot in the last run
opportunitynot in the last run
prospectnot in the last run
cubenot in the last run
entitiesnot in the last run
lanxess-catalognot in the last run
lanxess-analysisnot in the last run

No timings yet. run_pipeline.py records each step's wall clock to output/pipeline_run.json; this table reads it when the file exists. It never will on Vercel — output/ is gitignored — so a deployed copy of this page shows blanks rather than figures measured on somebody else's laptop.