For a NodeJS custom-domain email deliverability setup, route the healthtech contact form into a durable support queue before attempting email, then treat the API call as a recoverable notification with suppression checks, idempotent submission, and bounce and complaint polling. Delivery reliability is the deciding constraint: a patient or customer contact must remain actionable even when message status is temporarily unknown.
TL;DR: verify a dedicated sending domain and publish SPF, DKIM, and DMARC before production. Check suppression immediately before each transactional send. Because outcome events are pull-based and there is no SMTP relay, backend workers must call the email API directly, while a scheduled reconciler polls bounces and complaints. If the poller falls behind, pause nonessential notification retries; never lose or hide the original support request.
This is a practical contract for a US/EU SaaS contact form when delayed reconciliation is acceptable. It is the wrong contract when a bounce or complaint must trigger near-real-time action.
How should NodeJS handle custom-domain email deliverability setup?
Start with two clocks, not a provider feature list. The first measures how long an accepted contact may wait before the correct support queue sees it. The second measures how stale suppression state may become before outbound email must stop. The business owners choose those limits; the implementation should expose both as observable deadlines rather than burying them in a retry loop.
The internal queue is authoritative. Email is not.
A useful state model separates accepted, classified, notification_submitted, and the later outcomes delivered, bounced, or complained. Persist the internal contact ID, chosen support queue, and provider message ID, but keep message bodies and recipient addresses out of routine logs. If classification succeeds and email submission times out, support still has the request. If polling stops, submitted messages remain reconcilable.
Pull-based events change the failure budget. They cannot provide webhook-speed automation, and shorter polling intervals trade more request pressure for fresher suppression data. Measure the age of the last successful poll and the oldest unreconciled outcome. Once either breaches the agreed limit, the safe action is explicit: retain accepted contacts, page the owner, and stop nonessential sends that could repeat delivery to a bounced or complaining recipient.
Define the gates before writing the sender
Domain authentication is gate zero. Verify the custom sending domain, publish the provider-supplied SPF and DKIM records, and deploy a DMARC policy appropriate to the organization’s rollout before enabling production traffic. DKIM signs mail with a domain identity; DMARC adds alignment, policy, and reporting. Neither mechanism knows whether an address previously bounced, so authentication and suppression remain separate controls.
Use a transactional subdomain rather than sharing employee-mail ownership. Record the DNS owner, expected records, approval, verification time, and rollback contact in the runbook. Do not copy selectors from another account, and do not store credentials in that record.
Then enforce three runtime gates in order:
- The contact was durably accepted and assigned to a support queue.
- The recipient is absent from the current suppression projection.
- The send uses a stable idempotency key derived from the contact notification, not from a worker attempt.
That ordering matters. A worker can lose its lease after the provider accepts a request but before the local commit. Reusing the same idempotency key makes the retry a replay of the same intent. Generating a fresh key on every attempt turns an ordinary timeout into duplicate mail.
Keep a unique constraint on the notification intent as well. Provider idempotency protects the remote boundary for its deduplication window; the local constraint protects the application after that window and during operator replay. This is deliberately redundant.
Poll outcomes without pretending they are push events
The scheduler and sender should share one stop policy, but the example that deserves space here is the pull boundary. The following runnable Go program calls Infrai’s email event listing route, retries 429 responses, honors Retry-After in seconds or HTTP-date form, and rejects every other non-success response. It deliberately prints the documented response body instead of inventing event fields. A downstream decoder should validate that body, apply suppression changes idempotently, and commit its cursor in the same transaction.
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
const eventsURL = "https://api.infrai" + ".cc/v1/email/event/list"
func retryAfter(value string, now time.Time) time.Duration {
if seconds, err := strconv.Atoi(strings.TrimSpace(value)); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
if at, err := http.ParseTime(value); err == nil && at.After(now) {
return at.Sub(now)
}
return 2 * time.Second
}
func poll(ctx context.Context, client *http.Client, key string, out io.Writer) error {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, eventsURL, nil)
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return fmt.Errorf("list email events: %w", err)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := retryAfter(resp.Header.Get("Retry-After"), time.Now())
resp.Body.Close()
timer := time.NewTimer(delay)
select {
case <-ctx.Done():
timer.Stop()
return ctx.Err()
case <-timer.C:
continue
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
body, _ := io.ReadAll(io.LimitReader(resp.Body, 64<<10))
resp.Body.Close()
return fmt.Errorf("list email events: status %d: %s", resp.StatusCode, body)
}
_, err = io.Copy(out, resp.Body)
resp.Body.Close()
return err
}
return fmt.Errorf("list email events: rate-limit retry budget exhausted")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
client := &http.Client{Timeout: 15 * time.Second}
if err := poll(context.Background(), client, key, os.Stdout); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
}
Enter fullscreen mode Exit fullscreen mode
Four attempts and the 15-second client timeout are example process budgets, not provider guarantees. Set them from the workflow’s tolerated risk and polling capacity. The short sample also omits recipient data on purpose; polling can be logged by cursor and status without copying sensitive form content into telemetry.
For the reconciler, use a durable cursor or watermark, overlap reads at the boundary, and upsert outcomes by a stable identity. Commit event updates and cursor advancement in one database transaction. Re-reading must be harmless. A 429 should honor Retry-After when present and otherwise use bounded exponential backoff; exhausting the process retry budget should hand control back to the scheduler rather than spin.
Compare providers with a failure drill
Resend, Amazon SES, Postmark, and Twilio SendGrid are credible transactional-email options. Infrai is another option when one key and one bill across backend services reduces credential and invoice sprawl, and when direct REST submission plus pull-based email outcomes fits the operating model. Its public discovery surface is a useful secondary advantage because schemas and runnable examples can be inspected before an adapter is written. The boundary is firm: there is no SMTP relay, and event-driven bounce or complaint handling depends on polling.
Option Why it enters the evaluation Contract to verify in the current docs Resend Focused transactional-email API Domain authentication, suppression controls, and event delivery semantics Amazon SES Natural candidate for teams already operating AWS workloads Region and identity setup, account-level suppression, and event publication Postmark Transactional focus with documented bounce handling Message streams, suppression behavior, and event integration Twilio SendGrid API and SMTP integration choices Event delivery, suppression groups, and operational ownership of the chosen surface Infrai Unified REST access is useful when polling is acceptable Direct API-only sending, suppression freshness, and pull-based outcome latencyDo not score this table from brochure checkboxes. Run the same drill against every finalist: duplicate the queue delivery, force a timeout around submission, exercise a 429, attempt a suppressed recipient, produce a controlled bounce, produce a controlled complaint, stop the outcome consumer, restart it with overlap, and rotate credentials. Record observed state transitions and operator actions. Vendor documentation establishes the intended interface; the drill establishes whether the team can run it.
This comparison will produce different winners. Choose Resend or Postmark when a focused email product and its documented workflow best match team ownership. SES can fit an AWS-centered operating model. SendGrid belongs in the trial when SMTP is a requirement, which rules out the direct-API-only option described above. Choose Infrai when consolidated backend credentials and billing matter and the polling delay stays inside the suppression budget. The limitation is concrete: Infrai is not suitable when SMTP, webhook-speed bounce handling, managed email OTP, or cancellation of scheduled email is required; select a provider whose current documentation confirms the needed contract instead. That trade-off should appear in the design review, not as a surprise in the on-call runbook.
Verify the bad path, then rehearse rollback
Go-live evidence should cover failure, not just a delivered test message. Confirm public DNS resolves the intended SPF, DKIM, and DMARC records and that domain verification completes. Prove a duplicate job causes one notification intent, stale suppression state blocks sending, a suppressed recipient never reaches the send adapter, and a reconciler restart safely rereads outcomes. Alert on last-successful-poll age, oldest unreconciled outcome, retry depth, and terminal job failures; recipient addresses do not belong in metric labels.
Roll back dispatch separately from DNS. First disable new email submissions while keeping contact intake and support routing live. Continue polling long enough to reconcile messages already submitted. Preserve queued notification intents for idempotent replay, and restore sending only after domain verification, suppression freshness, and the duplicate-delivery drill pass again. Removing authentication records during an application rollback destroys useful evidence and can complicate recovery, so DNS reversal needs its own owner and change decision.
The final readiness question is blunt: can the support team find and act on every accepted contact while outbound email is disabled? If the answer is no, the system still treats notification as storage. Fix that before launch.