← Articles

One rule, three sites

I recently built a small internal utility on Cloudflare Workers. Drop a file in, get back a link that renders it. Early on it had an option called “never expire.” I took it out.

That’s a smaller decision than it sounds, and also a bigger one. Removing an option is rarer than adding one. Someone has to decide the convenience isn’t worth what it costs, and then live with the complaints. In this case the replacement was simple: everything expires, one year at most. No exceptions, no admin override, no “just this once.”

The rule reads clean. Getting it to actually be true was the interesting part.

Why “everything expires” wasn’t automatically true

A rule like this sounds like a UI setting: remove the dropdown option, ship it, done. But “everything expires” isn’t something you assert once. It’s something every part of the system has to agree on, every time it makes a decision that touches an expiry. And there were three separate places in this codebase that each got to decide what an expiry is:

  1. The function that checks, at read time, whether a given share has expired.
  2. The scheduled job that walks the database and deletes anything expired.
  3. The endpoint that lets someone renew a share and push its expiry out.

Each of these had its own idea of “no expiry set.” One of them treated a missing value as immortal. Another only looked at rows where an expiry was explicitly recorded, so a null-expiry row was invisible to it, permanently. The third just clamped forward from “now,” which sounds harmless until you notice what that does over many renewals.

Fix one of these in isolation and the system looks correct. Fix two of three and you get something worse than doing nothing: a silent, opposite-direction failure. A share can return a “this link has expired” page while its files sit in storage forever, because the delete job never learned to look at rows without an explicit expiry. Or a share can vanish from storage while the page that’s supposed to show it as gone is still serving it fine, because the two functions disagree about what “expired” means for the same row. Either way, nothing in the logs looks wrong. The system just quietly does the opposite of what the product promised.

The renewal clamp, which is subtler than it looks

The renewal endpoint is the one that’s easy to get wrong without noticing, because a naive fix reads correct. The obvious version clamps every renewal to “now plus one year.” That preserves one property (no renewal exceeds the ceiling) and destroys another: repeated renewals can push a share’s effective lifetime out indefinitely, one small nudge at a time, which is exactly the “never” behaviour the rule was supposed to remove.

The fix clamps to creation time plus the ceiling, not to renewal time plus the ceiling. That single change has to hold two properties at once: a renewal should never expire a currently-live share as a side effect of running the clamp math, and a genuine shortening of the expiry should still land correctly even when the share is already older than the ceiling. Neither property is hard alone. Getting both at once out of one clamp is where it’s easy to quietly break one while fixing the other.

The migration that nearly deleted live data, twice

Backfilling old rows to the new rule meant a migration: for every existing share with no expiry recorded, write one in. Two SQL statements, straightforward on paper. Almost destructive in practice, and in two different ways at once.

The purge query, the one that decides what to delete, had its own protection against a row with no recorded expiry: fall back to creation time plus the ceiling, rather than treat a missing value as unreachable. That protection only fires when the expiry is missing. The migration’s first statement wrote a real, non-null expiry into every row it touched, which means the fallback never triggers again for that row. Nothing about the write statement was wrong on its own terms. It just routed around a protection built for a different case, by construction.

The second hazard was sharper. If created_at had ever landed as text instead of a clean integer, the backfill’s arithmetic would still run, because SQL happily does math on strings that look numeric. It doesn’t fail. It produces a real-looking timestamp.

In this case, one from 1971.

Hand that number to the next scheduled purge and it reads as a share that expired decades ago, live future expiries and all, marked for deletion on the next run.

Neither of those was visible reading either statement on its own. Both statements looked fine. The problem only showed up when the read guard and the write statement were checked against each other, which is a different exercise than reviewing either one for correctness in isolation. The fix was to put the same type guard on both statements, so the write can’t produce anything the read wouldn’t also trust. I reproduced the failure against a real SQLite engine first, to confirm it was real and not theoretical, then reproduced it clean after the fix.

The generalisable version of this, if there is one: I guarded the query that reads and left the statement that writes unguarded, and each half looked correct alone. That’s the pattern worth carrying into other codebases, more than the specific SQL is.

The residual, said out loud

The guard added to the migration checks that a column’s stored type is integer. The read-time predicate elsewhere in the app checks, in JavaScript, whether a value is a number. Those aren’t quite the same test, and it’s worth saying plainly rather than claiming the invariant is airtight: a database column that ended up storing something like a floating-point number instead of a clean integer would be skipped by the SQL guard, but accepted as valid by the JavaScript check. Two rules, one column, still not fully aligned.

The failure direction, at least, is the safe one: a mismatched row gets skipped, not deleted. Nothing in that gap can destroy data on its own; it just means a row might sit un-purged longer than it should. So “one rule, three sites” is really “one rule, three sites, assuming the underlying column always holds the type it’s declared to hold.” That’s a real assumption, not a closed case, and it’s the kind of thing worth writing down rather than letting a green test suite imply it’s handled.

Deploy order as a correctness property

There’s one more wrinkle worth keeping, because it’s less about expiry logic and more about how systems ship. The old “never expire” value still exists in the code, as a legacy alias that resolves to the one-year ceiling instead of being rejected. It exists because the client interface and the server ship from two different repositories, on two different schedules. An admin interface on an old build still offers “never,” and a user who picks it needs the server to translate the value rather than reject the request. Reject it, and a legitimate save fails, stranding files the user already uploaded to storage with nothing to point at them. Silently accept it as “no expiry,” and you’ve reopened the exact hole the rule was built to close.

So the deploy order becomes load-bearing, in one direction. The server has to ship first, understanding both the new value and the legacy one. The client can lag behind safely, because the alias covers the gap. Reverse that order, shipping a client that offers a new option the server doesn’t understand yet, and the failure is different and worse: a request the server can’t interpret falls back to some default nobody chose on purpose. Neither system is “wrong” in isolation. The correctness lives in the sequence they arrive in, not in either one alone.

The shape of it

None of this is really about expiry dates, or SQLite type affinity, or Cloudflare Workers specifically. It’s about what “invariant” actually means once you look for it. An invariant isn’t a rule you write in one place and enforce with a validator at the door. It’s a claim about every path that writes, and every path that reads, agreeing with each other — not just agreeing with the rule as it’s documented. A validator can be flawless and the invariant can still be false, because the validator was never the only door.

Worth checking, the next time a rule in your own system sounds settled: how many places get to decide what that rule means, and do they still agree with each other today?