From a new data drop to a deployed site
One command runs the whole thing in the right order:
python3 run_pipeline.py--dry-run prints the plan without running it. --from <step> resumes after a failure — the script tells you the exact command when a step fails. --only payloads rebuilds just the JSON the app reads, which is the common case when a builder changed but the source data did not. --with-lanxess adds the catalog steps, skipped by default because that spreadsheet changes on a different cadence.
scripts.prospect.index reads public/dashboard-data.json to decide which companies count as covered accounts. Run it before scripts.opportunity.build and it silently indexes against the previous run's account list — no error, just a wrong account flag on every company, which surfaces much later as a supplier page claiming it reaches no covered accounts. The dependency is declared in run_pipeline.py, and selecting a subset that leaves a dependency out prints a warning naming it rather than failing quietly.
Every step is idempotent, so after a logic-only change the script should finish by reporting public/ is unchanged — the proof the change was a no-op on the full corpus rather than only on the sample you tested. A diff does not mean your logic changed something. It means the payloads differ from the last commit, which is equally consistent with the inputs having changed underneath you. Check input/ before blaming the code.
The most expensive mistake this pipeline permits, and it has already happened: two files landed in input/ holding the same 24,059 Owens Corning declarations — one 20-column with ISO dates and plain numbers, one 257-column with day-first dates and comma-formatted values. Nothing matched textually, so no duplicate flag fired. Both Owens accounts came out at exactly 2.00× — 4,601 rows became 9,199, $150.7M became $301.1M — and the inflated pair displaced five genuine accounts out of the covered set.
exact_duplicate could never catch it: it is computed per file. The consolidation step now compares declarations after normalization — parsed dates, canonical values, resolved company names, the only level at which two differently-formatted exports of one declaration look alike — and writes output/cross_file_duplicates.json naming the overlapping files. It reports and never drops: which export is authoritative is a question about provenance the pipeline has no basis to answer.
The rest of this page is what the script runs, in order, and what each step can get wrong. Two of these steps did not exist as code until recently — the Excel conversion and the Lanxess catalog were run by hand, which meant lanxess-catalog.json was committed with no way to rebuild it. Both are now in the repository and both reproduce the committed payloads byte-for-byte.
1 · Convert the drop to CSV
Drop .xlsx files into input/excelfiles/. Existing CSVs are skipped unless --force is passed.
python3 -m pipeline.excel_to_csv~30s for 26 files. pandas writes dates as ISO YYYY-MM-DD while older hand-made CSVs are day-first; both parse only because "%Y-%m-%d" is in date_formats. A source with a third date format will parse to NaT silently — the review queue in step 3 is where you would notice.
2 · Rebuild the knowledge base
python3 manage_knowledge_base.py build --input-dir input --knowledge-base knowledge_baseGenerates observed country aliases, entity candidates, and the observed HS hierarchy. Candidates are not used by processing. If new companies arrived in this drop, review knowledge_base/entity_alias_candidates.csv and promote what is correct into entity_aliases.csv before step 3 — afterwards is too late for this run.
3 · Normalize
python3 process_data.py --allThe long step — it writes output/<file>/ per source plus the concatenated output/trade_clean_all.csv, currently ~540 MB. Then read the review queues before trusting anything downstream: a queue that suddenly jumps from 2% to 100% of rows means a format changed, not that the data got worse.
4 · Build the payloads
python3 -m scripts.opportunity.build # must be first
python3 -m scripts.prospect.index # reads dashboard-data.json
python3 -m scripts.analytics.cube
python3 -m scripts.analytics.entitiesThe last three are independent of each other; only the first two are ordered.
5 · Rebuild the Lanxess catalog
Only needed when input/Lanxess Products.xlsx changes. It does not read the customs pipeline at all — it is the one payload built from a customer spreadsheet.
python3 -m scripts.lanxess.catalog # ~0.5s — writes output/ and public/
python3 -m scripts.lanxess.analysis # ~1s — every Lanxess figure
python3 -m scripts.check_payloads # ~0.4s — fails the run on a wrong number
python3 -m scripts.check_selftest # ~5s — proves those checks still fireThe second is where the Lanxess numbers are computed — not a check on them. It used to be an oracle that recomputed the figures in Python while the browser computed its own in TypeScript, with agreement between the two maintained by hand; it now writes lanxess-analysis.json and lib/lanxess.ts derives nothing. It currently reports 709 clean opportunities worth $911,262,098 across 24 accounts — which is what /lanxess displays. The count fell from 803 when MIN_OPPORTUNITY_VALUE put a $1,000 floor under a demand row, taking $17.6K of one-declaration noise with it.
6 · Check, then commit the payloads
npx tsc --noEmit
npm run build
git add public/ && git commitnpm test is an alias for npm run build — it proves the code compiles, not that a number is right. What checks the numbers is scripts/check_payloads.py, which runs as the last pipeline step and exits non-zero on a wrong one: the opportunity floor, surviving cross-file duplicates, a company buying from itself, the hierarchy trees agreeing with the heat matrix, and a ±25% collapse guard on every headline count. Payloads only reach production when someone commits public/, so a failed run is a gate. Three defects that shipped for months — $17.6K of sub-dollar rows, $35.9M of duplication, $59.6M of intra-group trade — are each covered by an assert that has been seen to fail. That last part is enforced: python3 -m scripts.check_selftest reintroduces every defect against a temporary copy of the payloads and requires the checker to catch it, because a check nobody has seen fail is decoration.
Vercel builds from git and cannot regenerate them: that needs Python, pandas and the 540 MB intermediate CSV, none of which exist in the build image. Ignoring them would deploy a site where every page hangs on “Loading…”. The contracts page tracks their total against the 50 MB ceiling →
7 · Deploy
Push. Vercel builds from the repository root — framework preset Next.js, Node 22.x, no environment variables, no database. Root Directory must stay at the repository root; the Python package is named pipeline/ rather than src/ precisely because Next.js would treat src/app as a source root and take over the build from app/.
Measured timings
| Step | Wall clock |
|---|---|
excel | not in the last run |
knowledge-base | not in the last run |
normalize | not in the last run |
opportunity | not in the last run |
prospect | not in the last run |
cube | not in the last run |
entities | not in the last run |
lanxess-catalog | not in the last run |
lanxess-analysis | not in the last run |
No timings yet. run_pipeline.py records each step's wall clock to output/pipeline_run.json; this table reads it when the file exists. It never will on Vercel — output/ is gitignored — so a deployed copy of this page shows blanks rather than figures measured on somebody else's laptop.