Skip to content
Back to all work
Full-Stack / Distributed Systems2026Client work

Dispatch

A self-hosted bulk email platform sending a research desk's daily market reports to ~12,000 subscribers — tenant isolation enforced by the database, a queue that cannot outrun the provider's quota, and bounce handling that fails safe.

Role
Architect & sole engineer
Client
a Lagos stockbroking firm
Scale
191 commits across 14 active days · ~35,500 lines · 62 API routes · 20 migrations
Next.js 15React 18TypeScriptPostgreSQL 16DrizzleBullMQRedisReact FlowAmazon SESDockerCaddy

The application

Screens, captured from the running app

Booted against its seeded database and walked route by route — these are the real pages, not mockups.

Dispatch overview showing sending health, recent campaigns and audience growth.
Sending health first. For a desk that mails three times a day, the question is never “how many campaigns” — it is whether the last one landed.
Dispatch campaign monitor with delivery rate, an opens-per-hour curve and a click breakdown by link.
The monitor reads denormalised counters and per-second throughput samples, so watching a send in progress never scans a table of millions of rows.
Dispatch block-based email builder with formatting bar, block library and styling panel.
The builder constrains rather than liberates: a 600px content width and a fixed block set, because the failure mode of an email editor is a layout that only works in one client.
Dispatch journey builder: a React Flow canvas branching on whether the recipient opened.
Journeys advance on a timer rather than through a queue — a contact can sit on a wait node for weeks, which a job queue models badly.
Dispatch deliverability page showing DKIM, SPF and DMARC status alongside reputation metrics.
Authentication state is a page in the product, not a runbook step — the firm's compliance officer can read it without asking an engineer.
Dispatch cross-campaign reporting with trends over time.
Cross-campaign trends are where list decay shows up — a bounce rate creeping up over nine sends is the signal the hygiene rule was built for.

The idea

A stockbroking research desk was paying Mailchimp per subscriber to send a market report three times a day. The bill scaled with the list; the feature set they actually used did not. They needed perhaps a tenth of what they were renting.

Dispatch is that tenth, self-hosted. It is not a Mailchimp clone — it is the subset a financial research desk genuinely uses, built so the running cost is a server and metered delivery rather than a per-subscriber licence. It went live on 1 September 2026 and has sent on every trading day since.

Every screen below comes from a local instance seeded with invented data — the account is "Acme Retail" and all 500,000 addresses are on the reserved example-mail.test domain. The client's own instance holds real subscribers and is never screenshotted.

Isolation is the database's job, not the application's

Every tenant table has row-level security keyed on current_account_id(). Application queries run as a role that cannot see past the policies, and the tables are marked FORCE ROW LEVEL SECURITY so ownership does not quietly exempt anyone.

The point is what a mistake costs. A query that forgets WHERE account_id = … returns no rows rather than somebody else's. The account context is carried in AsyncLocalStorage rather than threaded through forty call sites as an argument — not to save typing, but because an argument someone forgets to pass produces a query that silently returns another tenant's data, whereas a missing context produces nothing at all.

The send worker and the sign-in path legitimately span accounts, so they run as a separate BYPASSRLS role and carry the responsibility of setting the context as soon as they know it. Two roles, one of which is deliberately weak.

A send pipeline that cannot outrun the provider

Fanout resolves the audience once and streams it into the queue using keyset pagination, so a 500,000-recipient campaign never runs an OFFSET query. Delivery is a second stage.

Rate limiting is a Redis token bucket shared by every worker process rather than a per-process limit. That distinction is the whole design: with per-process limits, adding a worker to go faster is also how you breach the provider's quota and get your account reviewed. With a shared bucket, scaling horizontally raises throughput and cannot exceed the ceiling.

The live monitor reads per-campaign counters and per-second throughput samples. Watching a blast in progress never makes the dashboard scan a table of millions of send rows.

Compliance in the render path

Every message gets RFC 8058 one-click unsubscribe, a List-Unsubscribe header, a visible footer link, a preference page, the sender's postal address, and a line explaining why the recipient is receiving it. Unsubscribes are honoured platform-wide and permanently, and there is deliberately no mechanism in the product to put someone back.

This is in the render path rather than a checklist because a checklist is a thing you can skip on the send that matters.

What happened when it met reality

The parts worth writing about are the ones that only showed up under load and under scrutiny.

Amazon paused the account mid-flight

Six full-audience campaigns went out in 48 hours from an SES account with no sending history. Amazon's trust and safety team paused sending with no warning, in the middle of the client's daily schedule.

Recovery meant identifying the real causes honestly — legacy links inherited from the previous provider, and a sending pattern that looked like a burst — correcting them, documenting the changes with dates, and answering Amazon's four questions with evidence rather than assurances. Reinstated the same day.

The bug the pause exposed

While sending was paused, SES rejected every message with 554 Message rejected: Sending paused for this account. The bounce classifier saw a 554, called it permanent, and suppressed 3,525 real customers as hard bounces — removed from every future campaign, with no error raised anywhere. The next send would have silently skipped a third of the list.

There was already a guard for exactly this case. Its phrase list contained "account is paused". Amazon says "Sending paused for this account". No substring match, no warning, no failure — just a quietly shorter list.

Two fixes, and the second is the one that matters. The phrase list was corrected, which handles the case that actually happened. Then a guard was added that does not depend on guessing anyone's wording: a batch where every single recipient fails with one identical message and none succeed is a fault on our side, not a coincidence of dead mailboxes. All 3,525 were restored from the send records.

The lesson is not "add the missing string". It is that a rule keyed on a provider's exact phrasing will fail the moment the provider rewrites it, so the real guard has to be shaped like the anomaly rather than like the message.

Hygiene that does not remove readers

Analysis showed 17 addresses accounted for every bounce across nine campaigns — the same 17 every time, none ever delivered to. Nothing in the product would have removed them, because they were soft bounces, and soft bounces are deliberately never suppressed.

The obvious rule — suppress after N consecutive soft bounces — would have removed the most engaged reader on the list: a corporate mail server with a forwarding loop that bounces one copy while delivering another, whose owner had clicked sixteen times.

So the rule is a bounce streak and no sign of a human. Clicks always count. Opens count only when they were not flagged as machine prefetch — which mattered more than expected: of five addresses that looked engaged, three had nothing but Apple Mail and Gmail fetching the tracking pixel. Fifteen addresses retired, two readers kept, bounce rate down from 0.12% across the preceding thirty days to 0.02% on the send that followed.

Where it stands

In production since 1 September 2026. That is six days at the time of writing, so what follows is an early reading rather than a track record.

On the most recent send, on 4 September, 12,040 of 12,042 messages were accepted and not bounced — 99.98%, against two bounces, one unsubscribe and no spam complaints. Across the thirty days to that date, 156,693 messages were delivered at a 0.000% complaint rate and a 0.12% bounce rate; the gap between that and the 0.02% above is the fifteen addresses retired in between.

Delivered here means the receiving server accepted the message. Whether it landed in an inbox or a spam folder is not measurable from the sending side, and nothing above should be read as claiming it. A full 12,000-recipient send clears in about eight minutes at 25 messages a second, measured rather than projected.

The engineering that earns its place here is not the CRUD. It is the four things that only matter when something goes wrong: isolation the application cannot bypass, a queue that cannot exceed its quota, a classifier that fails safe, and compliance that cannot be skipped.

You have a process thatshould be a system.

Tell me what arrives, who has to act on it, and where it currently falls over. That conversation is usually enough to scope the build.