Skip to main content

Deployment & Build

Docker image, Vite/Rollup externalization gotchas, staging deploy workflow.

Deep-dive doc split out of .claude/rules/architecture.md (which is now an index). Append new findings about this area here, not to the index.

Docker image​

Dockerfile is minimal: FROM node:24-alpine, COPY . ., RUN npm ci, RUN npm run build, CMD npm run start. Two consequences:

  • npm ci does NOT use --omit=dev — devDependencies (tsx, prisma, vitest, etc.) ship in the production image. This is what makes tsx scripts/...ts work from the DO console without further changes. If the image is ever slimmed via --omit=dev, all tsx-based scripts will break unless tsx is moved to dependencies.
  • COPY . . copies the entire working tree, including scripts/ and src/. So scripts that import from ../src/... continue to resolve at runtime; nothing is pruned.

Vite/Rollup Build — dependencies Are Externalized by Bare Specifier Only (INFRA-611)​

vite.config.ts's rollupOptions.external array externalizes third-party packages so npm run build (vite build, emitting dist/index.js as pure ESM per build.lib.formats: ['es']) doesn't bundle them — mostly via ...Object.keys(dependencies) (every package.json dependencies entry, by exact bare specifier), plus a few explicit extra strings/regexes ('express', /^node:/, etc.). A regex entry is required, not the bare-string form, for any package whose code is imported via a deep subpath (e.g. 'nodemailer/lib/mail-composer', not 'nodemailer') — Object.keys(dependencies) only ever contributes exact bare package-name strings, and Rollup's external matching is exact-string (or regex/function) against the full imported specifier, so a subpath import silently falls through and gets bundled even though the bare package name is nominally "external". This was already hit once before and fixed for @shopify/shopify-api/adapters/node-style subpaths via a dedicated /^@shopify\/shopify-api.*$/ regex entry (predates INFRA-611, comment explains the original tree-shaking/dynamic-import symptom).

INFRA-611 gotcha, found live in Staging after INFRA-578 merged: src/clients/gmailClient.ts added import MailComposer from 'nodemailer/lib/mail-composer'; (deep subpath) for the new Gmail-notification client. Since only the bare 'nodemailer' string was externalized (via the dependencies spread), Rollup bundled nodemailer's internal lib/base64 submodule — which does class Base64Encoder extends Transform against Node's stream built-in — and something in Rollup's bundling of that nested CJS chain broke the Transform reference, producing TypeError: Class extends value undefined is not a constructor or null at module-load time, i.e. before Express ever calls app.listen(). The container therefore never binds to its port, which is why DigitalOcean's automated deploy-failure diagnosis reported a "Port binding issue" — a real symptom, but a red herring for the actual root cause. The GMAIL_* env vars being present/absent in DO was a plausible-looking but incorrect first hypothesis; the crash reproduces identically even with all required env vars correctly set, since it happens at import-time, before any env-dependent code runs.

Fix requires two changes together (verified by checking out the merged commit in an isolated git worktree, running the real npm run build + node dist/index.js, and reproducing then resolving the exact stack trace):

  1. Add a regex externalizing the whole package's subpaths — /^nodemailer/ — alongside the existing /^@shopify\/shopify-api.*$/ entry, so Rollup stops bundling any of nodemailer's internals.
  2. Change the import to the fully-specified file path: 'nodemailer/lib/mail-composer/index.js'. This is necessary in addition to (1) — once truly externalized, the specifier is left as-is in the ESM output and Node's native ESM resolver (unlike CJS require) does not auto-resolve directory imports; without the explicit /index.js, boot instead fails with ERR_UNSUPPORTED_DIR_IMPORT.

Rule of thumb: any time a new dependency is consumed via a deep subpath import (rather than the package's main entry), check whether it needs its own regex external entry — the blanket Object.keys(dependencies) spread does not cover this case, and the failure mode (crash at container boot, before the port binds) is easy to misattribute to something else since it only manifests in the real vite build output, not in tsx/dev-mode runs (which don't bundle node_modules at all, so the bug is invisible locally under npm run dev).

Staging Deploy Workflow (.github/workflows/deploy-staging.yml)​

Push-to-main: biome→tsc→test→build (migrate DB, build/push image, rotate tags candidate→staging, prior :staging digest→:previous for rollback; registry ops via DO REST API not doctl, which mishandles multi-registry)→deploy. GC gotcha (INFRA-503): manual doctl registry garbage-collection start step is commented out — DO automated GC now enabled on manufacturing-admin-staging, manual call 412s (manual GC not available while automated GC enabled) and failed build (last step). Left commented (not deleted) for possible revert; don't re-enable while automated GC on. Prod workflow has no GC step.

AdminJS Components Bundle — Build Pipeline & the "Bad Bundle" Failure Mode​

Pipeline. npm run build (vite build) runs the adminJSBundler() plugin (scripts/build/vite-plugin-adminjs.ts) in closeBundle, which spawns npx tsx scripts/build/bundle-components.ts. That script calls bundle({ componentLoader, destinationDir: 'dist/adminjs-bundles' }) from @adminjs/bundler, emitting four files: global.bundle.js, design-system.bundle.js, app.bundle.js (copied from node_modules) and components.bundle.js (our custom components). Because this runs inside RUN npm run build in the Dockerfile, the bundle is baked into the image — a DO restart or DO "Force rebuild and deploy" re-serves the same broken file; only re-running the GitHub build job regenerates it.

How the bundle registers components. components.bundle.js is one IIFE taking globals (React, AdminJS, PropTypes, ReactDOM, ReactRouterDOM, AdminJSDesignSystem, ReactRouter, ReactRedux). Vendored third-party code is inlined at the top; the AdminJS.UserComponents.X = ... assignments all happen in a single statement at the very end. So a throw anywhere during vendor-module initialization registers zero components, and every custom property/action renders "Component X has not been bundled". Object.keys(componentLoader.getComponents()) (both .add and .override entries — 48 + 15 = 63 as of INFRA-710) is exactly the set that should end up on window.AdminJS.UserComponents.

Observed failure (2026-07-29, 2026-09-28). CI builds that passed produced a components.bundle.js that throws TypeError: Object.defineProperty called on non-object from a lazily-initialized CJS-interop module (d3-interpolate-style var r0={}; function tk(){ ... Object.defineProperty(r0, ...) }, called before its var initializer ran). The file was not truncated — it contained all 63 registrations and was within ~30 bytes of a good build — and a clean rebuild of the identical commit was fine. Treat it as non-deterministic module ordering in that build, not a code regression in whatever PR just shipped.

Reproducing / verifying without a browser. jsdom (already a dependency) reproduces the failure faithfully: evaluate the four bundles in order in a new JSDOM(..., { runScripts: 'outside-only' }) window via window.eval(...), then read window.AdminJS.UserComponents. Against the saved bad staging bundle this gives 0 components + the same TypeError in ~1s; against a good build, 63 + no errors. Chromium is not in the image (it would add ~821 MB — see work-order-pdf-export.md), so jsdom is the practical in-build check. To test a candidate bundle against real staging instead, Playwright page.route('**/admin/frontend/assets/components.bundle.js', ...) can swap in a locally served file on the unauthenticated /admin/login page (it loads all four bundles).

Post-build check (INFRA-710). scripts/build/bundle-components.ts now runs verifyComponentsBundle (scripts/build/verify-components-bundle.ts) right after bundle(): it loads the four bundles into jsdom and requires every name in componentLoader.getComponents() to be on window.AdminJS.UserComponents with no load errors, otherwise it process.exit(1)s — failing npm run build, and therefore docker build for DO staging, DO production and Dockerfile.gcpJobs, before any image is pushed. Adds ~0.5s. Deliberately no automatic retry: the root cause of the non-deterministic bad bundle is unknown, and a retry would hide how often it happens. If it fires on a commit that builds cleanly elsewhere, re-run the workflow; if it fires on every build, a component (or a new dependency it pulls in) throws at module-load time — the log names the error.