← writing

Bytes through the API, not around it: designing a document analytics service

August 24, 2026

100 KB files. 100,000 uploads a day. Bursts of 10,000 per minute. Extract a page count, a word count, a character count from each file, and show them in a dashboard. Sounds simple. The design space is wider than it looks.

That combination — small files, high burst, mixed formats, a couple of variable-cost failure modes — is where a lot of ingest pipelines live. Five decisions shape this design, in the order they matter.

What most engineers would build

The default architecture is well-worn: presigned URLs for uploads, a Lambda for parsing, an LLM to extract counts because “AI is what we do now,” and a synchronous API because the file is small — how long could it take?

Each of those defaults is defensible in the general case and wrong for this specific one. Presigned URLs win on throughput above ~1 MB; at 100 KB a SAS handshake is a wasted round-trip that also moves validation off the write path. An LLM for counting pages and words costs two to three orders of magnitude more than pypdf and returns the same integer. Sync API works at low burst; at 10K/minute it turns the tail latency of OCR into everyone’s problem.

The five decisions below are what happens when you take the standard answer, ask what breaks at this scale, and change the ones that don’t survive contact.

Async with push, not sync

Sync feels simpler. It isn’t.

At 10K uploads/minute burst, a synchronous API has to keep parser capacity pre-warmed for the peak or time out under load. Scanned PDFs behind OCR can take 3–5 seconds; a synchronous upload holds that connection open, occupies an API pod for the duration, and blows the p95 budget on every OCR case. The tail latency of one job type dictates the entire service.

Async breaks the coupling. Upload returns 202 {status: QUEUED} in under 200 ms. A queue absorbs the burst; workers scale independently on queue depth. The user experience stays synchronous through a Web PubSub push — the browser joins group user:{id} with a scoped token from the 202, and the dashboard row updates live when the parser completes.

The honest tradeoff: async means a state machine. QUEUED → PROCESSING → COMPLETED | FAILED | RETRYING, retry with exponential backoff, DLQ after N attempts, user-initiated retry with a new attempt_id. Sync would have skipped all of that. But sync at this burst shape is a load-test failure waiting to happen.

When sync is fine: every document parses in under 500 ms and burst is well under 100/min. The design collapses to a single API tier and the state machine becomes overhead.

Bytes through the API, not around it

The default answer is presigned upload — a SAS URL, bytes go client-to-storage direct, API never touches the payload. For 100 KB files, that’s exactly the wrong call.

The argument for presigned URLs is throughput: bytes bypass the API tier. That win is real when files are big. At 100 KB, one HTTP round-trip carries the entire file — a SAS handshake adds a request without saving bandwidth that matters.

More importantly, it moves three checks off the golden path: MIME sniffing (magic bytes must match the declared type), dedupe against sha256, and — as soon as it’s added — antivirus scanning. All three want to be done where the bytes are. Doing them in a webhook after presigned upload creates a race window where the row exists but hasn’t been validated.

The honest tradeoff: the API tier becomes a proxy. At 100 KB × 167 rps peak, that’s ~16 MB/s inbound. Serve it with async I/O so bytes stream rather than buffer; the pod count scales the same as any request-serving tier. It’s still cheaper than the operational surface of a two-phase upload protocol.

The rule I’d apply generally: presigned upload above ~1 MB, bytes through below ~200 KB. Between those, use whichever keeps validation on the write path.

When you’d use presigned: files above ~1 MB, or when the client sits on a fat pipe you don’t want to hairpin through your API. The rule is the file size, not the pattern.

Service Bus, not Event Hubs / RabbitMQ / Celery

Choosing the wrong queue shape is the most common mid-level mistake.

Event Hubs is a partitioned append log. It’s built for telemetry streaming — high-throughput ingest with consumer group offsets. It isn’t built for per-message ack, scheduled retry with exponential backoff, or a native dead-letter queue. You can force those on top, but you’re building on the wrong shape.

RabbitMQ is the right shape and the wrong operational cost. Managed RabbitMQ still needs someone patching, tuning the cluster, and rebuilding a node when it dies. If a managed cloud-native option is shape-correct, prefer it.

Celery collapses queue and worker framework into one bundle. That’s language lock (Python-only) and a set of silent failure modes — tasks that die mid-execution, redelivered without idempotency guarantees, and worker pools that quietly drift out of sync with queue backpressure.

Azure Service Bus (or its equivalent on any cloud) gives you what a job queue actually needs: message dedup at the transport layer, scheduled delivery for backoff without an in-worker sleep, and a per-queue DLQ. Autoscale (KEDA on Kubernetes, or equivalent) scales workers on queue depth. Nothing exotic; everything shape-correct.

When Event Hubs is the right call instead: the pipeline is telemetry-shaped — high-throughput, ordered per partition, no per-message retry semantics. Different problem, right tool for it.

Deterministic parsers primary; AI only for OCR

At an AI-first company, the temptation is to reach for AI first. On this problem, that’s the wrong instinct — and it’s worth being explicit about why, because gratuitous LLM use is an anti-pattern that hides behind AI-first branding.

Counting pages, words, and characters in text-native documents is a solved library problem. pypdf, python-docx, python-pptx — one library per format, deterministic, fast, free, boring. Sending a text PDF through OCR (Azure AI Document Intelligence Read, or equivalent) would cost ~$1.50 per 1,000 pages and add 2–4 seconds of tail latency for exactly the same answer. An LLM prompted to “count pages” would be worse still — non-deterministic, unbounded cost, and wrong just often enough to be unusable.

The place OCR earns its slot is scanned PDFs — pixels, not text. The heuristic to detect them is embarrassingly simple: extract text with the library parser, count characters per page, branch to OCR if it’s under 50 chars/page. Above the threshold, deterministic wins.

The framing I’d use: AI where the input is a modality libraries can’t handle. Libraries where they can.

When AI earns a seat at the table: stage-2 enrichment — summary generation, topic extraction, entity linking, semantic classification, retrieval over the corpus. These are tasks where the answer isn’t a well-defined function of the input, so a library can’t produce it. That’s the moment to add an LLM worker. The architecture already accommodates it as a second worker deployment fanning off the same Service Bus — no re-plumbing, and the deterministic path stays cheap and predictable underneath.

The honest tradeoff on the OCR threshold: it isn’t universal. Partial-OCR PDFs — where some pages are text and some are scanned — misfire. A shadow-parse metric in production (run both paths on a sample, compare outputs) is the right way to tune it.

Postgres is the dedupe truth. Redis is the cache.

The design has two dedupe layers. Only one is authoritative.

Redis holds a sha256 → file_id lookup. It’s fast, it handles the 99% hot path, and it’s not the source of truth. A cache is a probabilistic optimizer, not an invariant enforcer.

The invariant lives in Postgres: UNIQUE (user_id, sha256). If two uploads race past Redis with the same hash, the second insert fails at the database constraint, and the API returns the original file_id. The race becomes impossible, not just unlikely.

The rule generally: never use a cache as a correctness boundary. If you need “exactly one,” it lives in a store with real constraints, and the cache accelerates the common path in front of it.

When Redis-only is defensible: if duplicates are cheap (idempotent side effects) and the race window is short enough not to matter — analytics ingestion, some cache-warming workloads. Never for anything where a user gets billed or a document surfaces twice.

What I’d change on v2

No design ships done. Five things I’d change with real traffic under it:

Add attempt_id and a retry-history table. The current schema overwrites failure_code on each retry. That’s fine for the happy path and terrible for a poison message that’s been retried five times with three different failure modes. An attempts table with one row per attempt is the fix.

Content-addressed global dedup, opt-in per tenant. Current dedup is per-user. If the same document is uploaded by 5,000 users of the same enterprise tenant, it’s parsed 5,000 times. A global sha256 index — with tenant-scoped ACLs on the result, not the file — collapses that to one parse.

Antivirus on ingest, not post-hoc. The failure taxonomy already reserves a FILE_MALICIOUS code. The scan needs to run before the row transitions to PROCESSING, not as a background sweep. Azure Defender for Storage (or equivalent) is the cheap first pass.

Per-tenant queue depth cap. The API rate-limits uploads. The queue doesn’t. One tenant can push 10K messages and starve everyone else. A per-tenant partition — or at minimum a soft cap on inflight jobs — is the fix.

Real DOCX pagination. Currently words / 275 as an estimate, because true pagination needs a rendering engine. LibreOffice-headless is the roadmap fix; it changes worker resource requirements meaningfully (memory-heavy, needs an isolated node pool like OCR), so it’s deferred rather than bolted on.

How I used Claude to write this

Full transparency: I used Claude to draft the HLD and iterate on the diagram. What matters is the shape of the collaboration, not the fact of it.

Prompt iteration, not one-shot generation. The first pass produced a competent generic doc — presigned URLs, a standard queue, sync API. That’s the median engineer’s answer, because it’s the mean of the training data. I rewrote the brief to lock the five contrarian decisions before Claude drafted, forcing the design into a specific shape. The output only becomes yours when you’re the one deciding what it should say.

Direct the tool when it’s wrong. The architecture diagram went through three renderers. Mermaid’s default (dagre) with LR gave a good aspect ratio; switching to ELK for cleaner subgraph grouping reversed the flow — User ended up on the right, Front Door on the far right, reading right-to-left. Switching back to dagre while keeping the ELK-style subgraph grouping was the fix. Separately, the Observability node initially rendered as a small floating box; I forced it out into an HTML bar strip beneath the SVG, because a wide bottom bar is what an observability plane should look like. The tool’s default was wrong for the reader.

Adversarial review before shipping. After the first defensible draft, I ran a critical review with a “poke every hole” instruction. It surfaced real gaps: no attempt_id for retry history, no idempotency key on the write endpoint, no malware-scan plan, an arbitrary DLQ alert threshold, no OCR cost math. I patched the doc against each attack before considering it done. Better to find your own holes than to be found by them.

Where the tool got it right. Compression — cutting a four-page draft to two, and forcing every sentence to earn its place. Table shapes — the failure-code table and format-handling table landed cleanly on the first pass. The “honest tradeoff:” pattern per decision came from the model; I kept it because it’s the right mature-engineer shape.

The AI-fluency signal isn’t whether you used the tool. It’s whether the output reads as yours and whether you’d defend every line.

The artifact

The full HLD — architecture diagram, state machine, failure code table, data model, observability plan — is written up as a two-page reference doc: document-analytics-hld.pdf.

The two-page constraint is worth calling out. It forces every sentence to earn its place, and it forces trade-offs to be named rather than hidden in prose. A useful exercise for any engineer working on ingest pipelines: take a system you already own, write it down in two pages, and see which decisions survive the compression.