In my last post, I introduced Durable Execution (or “Code-First Workflows”) as a fast-emerging architectural primitive. A paradigm only really clicks when you watch it solve a messy, real problem, so let’s look at one.
Back in 2018-19, a friend at Uber walked me through how they handled ride-end payment notifications in India. It ran on Cadence, the engine that Temporal later grew out of, built by the same people. Below is the problem, and how I’d model it with Temporal today.
One note: this is my understanding from that conversation, and the code is a simplified sketch, not production code from Uber.
The Product Challenge: Cash vs. Digital
Cash was still heavily used for ridesharing in India back then, and it created real friction at the end of a trip:
- The double-pay confusion. A rider picks Digital (UPI, credit card, Paytm), but the driver asks for cash anyway, out of habit or confusion. The app needed to say it plainly: “Digital payment selected. Do NOT pay cash to the driver.”
- The default payment trap. Riders carried over a last-used method, often Cash, and only realized near the end of the trip that they had no cash on them.
- No post-ride switching. Changing the payment method after the ride ended hadn’t shipped yet, so a trip that ended on Cash left rider and driver in an awkward standoff. ## The Requirement
To give riders one last window to switch payment methods before the ride locked, product wanted a notification with three rules:
- Timing: send it 5 minutes before the ride ends.
- Cap: at most 3 notifications per ride.
- Moving target: the ETA keeps changing with traffic, so “5 minutes before the end” is never a fixed timestamp. Sending a notification is trivial. Sending it at the right moment, the right number of times, on a trip whose end time keeps shifting, is the actual problem.
Option A: Build It Yourself
With traditional building blocks (DB status flags, SQS delay queues, Redis TTL keys, cron scripts), this is possible, but painful for what it delivers. You end up writing a bespoke distributed state machine just to manage time-based logic. The glue code you own looks like this:
- Timer engine. Thousands of ephemeral delayed jobs in SQS or RabbitMQ, or a polling mechanism, just to track “5 minutes before the end” while ETAs keep shifting with traffic.
- Locking and state sync. Guarding against double notifications when an ETA update and a timer expiry hit two different workers at once, and keeping a Redis counter for the 3-notification cap in sync with the ride lifecycle across services.
- Cleanup. Purging or skipping scheduled messages when a ride cancels, ends abruptly, or completes early. Most of your code ends up being plumbing (retries, locks, timer offsets, state cleanup) rather than the product logic of sending one notification. You’ve built a fragile, miniature workflow engine inside your own application.
Option B: Durable Execution
With a code-first workflow engine like Temporal, you don’t build a state machine out of tables and timers. Your code IS the state machine.
Temporal gives you durable timers and signals. You write a plain loop with a sleep in it, and signals feed in ETA updates and ride-end events. If the worker running it dies, the cluster restores the exact loop state, timer, and variables and carries on as if nothing happened.
Here’s a simplified sketch of how this could look in Temporal (pseudocode, not production code):
async function rideEndNotificationWorkflow(rideId: string) {
let eta = await getInitialEta(rideId);
let rideActive = true;
let sent = 0;
onSignal('etaUpdated', (newEta) => { eta = newEta; });
onSignal('rideEnded', () => { rideActive = false; }); // ended or cancelled
while (rideActive && sent < 3) {
const waitTime = eta - now() - minutes(5);
if (waitTime > 0) {
// durable timer: wakes early if a signal arrives
await sleepOrSignal(waitTime);
continue; // re-check: ETA may have changed, or the ride may be over
}
await sendPaymentReminder(rideId);
sent++;
await sleepOrSignal(minutes(1)); // brief pause before any re-send
}
}
Enter fullscreen mode Exit fullscreen mode
Read it top to bottom and it’s just the business rule: wait until 5 minutes before the end, notify, stop after 3 or when the ride is over.
Why This Wins
- No external state. No Redis counter, no delay queues, no status flags. The counter and the timer live in the workflow itself, so there’s nothing to keep in sync.
- Cancellation and cleanup come almost free. When the ride ends or cancels, the signal flips a flag and the loop exits. There are no orphaned delayed messages to hunt down.
- The code reads like the business rule. Product can look at that loop and recognize their own spec. That’s a big deal for maintenance. None of this is free, and I’ll cover where it bites in the next post.
In the next one, I’ll get into why some engineering organizations evaluated Temporal and walked away, and how those evaluations played out.
Cheers!!