FuturumBecome a Client
Insight

Inside the new publishing pipeline: from submit to live in seconds

Futurum Web Team
The short answer

A one-submission core moves an insight from a validated request to a cache-purged public page in single-digit seconds, with sanitize, scan and idempotency doing the parts that don't get to fail quietly.

Futurum's Futurum Web Team,

Cite this

Futurum Web Team, The Futurum Group, "Inside the new publishing pipeline: from submit to live in seconds," September 24, 2026. https://preview.erikbethke.com/insights/sample-insight-publishing-pipeline/

A publishing pipeline usually earns attention only when it breaks. This one is worth a look while it is working: on 2026-09-24 this site's own publishing core went from a validated submit() call to a purged, indexable public page in under three seconds, measured through the live URL with no cache-busting. That is not a benchmark run in a lab. It is the ordinary path an insight now takes to get in front of a reader.

One core, several doors

Before this phase, the API lane's biggest liability was not code quality; it was surface area. A raw curl could reach the database, a retried request could duplicate a post, and the credential scan that was supposed to catch a stray API key in a pasted paragraph was a convention rather than a gate. The fix was to collapse every entry point — a CLI, a WordPress-shaped adapter for the internal content platform, and the plain HTTP API — onto one function that does the same ten things in the same order, every time, regardless of who is calling.

  1. Authorize the caller against its capabilities.
  2. Resolve idempotency from the request key or an external id, and stop here on a replay.
  3. Normalize the body: Markdown to HTML, or the WordPress import cleaner for raw HTML.
  4. Sanitize the HTML against a fixed element and attribute allowlist.
  5. Scan for credentials, internal hostnames and account-shaped identifiers, and block on a hit.
  6. Resolve taxonomy: companies must already exist, tags are created on first use.
  7. Persist a revision — metadata in DynamoDB, the body in S3, keyed by content hash.
  8. Transition the post's state machine according to what the caller's capabilities allow.
  9. Queue effects in the same transaction as the write, so a crash cannot lose one.
  10. Append an audit record: principal, action, revision hash, scan result.

The order matters more than any single step. Sanitize runs before scan, not after, because a scan that reads unsanitized HTML would be scoring markup that never ships. Idempotency runs before normalize, because normalizing twice on a retried request would hash two different byte strings for what is supposed to be one write. The sanitizer itself follows the allowlist approach industry guidance recommends over trying to blocklist every dangerous tag [5]: a fixed set of elements and attributes survive, everything else is dropped, and no class or style reaches the page.

Why a revision, not an overwrite

Every accepted submission becomes a new revision rather than a mutation of the last one. The post's pointer moves; the old body stays in the content-addressed store at its own hash. That makes a bad edit a one-line revert instead of a support ticket, and it means the credential scan's verdict on revision N is never invalidated by what happens in revision N+1.

The Futurum wordmark in white on a near-black background, with a blue rule beneath it
The site's own fallback share-card art — the image every open-graph preview uses when a post has no photograph of its own, including, fittingly, a post about the pipeline that serves it.

What the numbers say

Phase 1's acceptance test is blunt: submit a post, then poll the public URLs exactly as a browser would, with no cache-busting query string. The measurements below are what that run recorded against the preview stage.

Phase 1 timings on the preview stage, from a plain public request to a changed page:

Step Elapsed
New article live at its own URL 2.9 s
First listing page shows it 6.1 s
Second listing page shows it 11.4 s
feed.xml includes it 4.3 s
An edited title live on the article 2.4 s
A soft-deleted post gone from its URL 2.1 s

Nothing here required a rebuild. The site is served from a CDN in front of a server-rendered app; a publish writes to a single DynamoDB transaction [2], and an outbox row from that same transaction drives a Lambda that calls the framework's own tag-based revalidation [1] for the post, its listings and its feed, then issues one CDN invalidation [3] for the whole batch rather than one per path.

There is one publishing core and several doors into it. Every door translates its own dialect into one validated Submission, and the core does everything else: validate, sanitize, scan, resolve taxonomy, store a revision, move the state machine, and fan out the side effects.

— from this site's publishing architecture plan, “The idea in one paragraph”

Replays are free

Every write carries an idempotency key, checked with a conditional write against its own partition before anything else runs. A client that times out and retries the exact same request gets the exact original response back, byte for byte, with a header saying so, and no second revision is written. The IETF has a draft standardizing this exact header [4], and the behavior here matches its intent: a retry is a read of the past, not a new write.

What's still open

Phase 1 is the core and the API door; it is deliberately not the whole plan.

  • Scheduling a future publish time is refused outright rather than half-supported, until a real scheduler exists to fire the revalidation at the right moment.
  • The signed webhook back to the internal content platform, and the ID registry that lets it keep using integer post IDs it already has stored, are not built yet.
  • There is no rate limit in front of the write API beyond a per-principal daily quota.
  • Listings still read full post bodies rather than a lighter summary projection, and stop at a fixed page size.

Key takeaways

  • One validated shape, one code path: every door into the site's content ends up calling the same function, so a security fix or a new taxonomy rule applies everywhere at once.
  • Sanitize and scan are gates, not conventions: a submission that fails either is rejected, not flagged for later cleanup.
  • Idempotency is structural, not best-effort: a conditional write closes the retry race that a check-then-write cannot.
  • Live in seconds, not a deploy: revalidation and a targeted CDN purge replace a rebuild for every publish, edit and delete.

Sources

  1. Next.js docs: revalidateTag
  2. AWS docs: DynamoDB TransactWriteItems
  3. AWS docs: CloudFront invalidation
  4. IETF draft: The Idempotency-Key HTTP Header Field
  5. OWASP Cross Site Scripting Prevention Cheat Sheet

Published by Futurum.