Node.js Text Summarization API — Reliable Chat Completions for Moderation SaaS

작성자

카테고리:

← 피드로
DEV Community · UlricDonovan1564 · 2026-09-26 개발(SW)

For a game moderation queue, the operational constraint changes the API choice: a late or duplicated classification can reorder human review, while a beautiful summary that arrives after the reviewer opens the report has little value. Short answer: start with chat completions for short-to-medium reports, require a small structured result, and put the call behind an idempotent queue consumer. Use batch submission when many stored reports can wait. Do not add embeddings unless the product later needs search or ask-your-documents behavior.

This is a quality-versus-latency decision, not a hunt for the lowest token price. A useful first release returns a compact summary, a policy category, and a confidence signal for human triage. It does not pretend the model is the final moderator.

How should a Node.js SaaS test a chat API for text summarization?

I have been paged by missed jobs and duplicate deliveries in cron and queue infrastructure. The lasting lesson was not to trust a scheduler label or a successful model response. The invariant is narrower: one accepted report must produce at most one current classification record, and an unclassified report must remain discoverable for replay. A request can time out after the provider has completed it; a worker can lose its lease after receiving the answer; a retry can then produce slightly different wording, and two consumers can race to update review priority. The model did what it was asked. The surrounding job failed to preserve a single durable outcome.

That is the gate.

For an interactive report, enqueue immediately and set a latency budget aligned with the review UI. If that budget expires, leave the report visibly pending rather than silently treating it as safe. For nightly reclassification, use a batch facility rather than a loop of synchronous calls. Batch submission reduces scheduler fan-out and gives operations one unit to observe and reconcile.

No silent drop.

The input also needs a boundary. Count or estimate tokens before sending a long report bundle, then split oversized histories at message boundaries. Model context limits and availability must be read from the provider’s current model catalog, not copied from a blog post or hard-coded from an old test.

Split live reports from batch reclassification

Chat completions are the simplest starting surface because prompt-based summaries and classifications need neither a retrieval index nor multimodal plumbing. Live arrivals belong on a bounded worker pool, while backfills and policy migrations belong in batch submission. This split protects reviewer-facing latency from a large historical run. I treat the quality test as a release gate, not as a demo score, because the trade-off is asymmetric: waiting a little longer can be acceptable, but burying a credible threat below routine spam is not.

Compare who owns the failure domain

Option Operational fit Main trade-off OpenAI API Direct managed chat-model integration and a familiar client surface A focused AI dependency; adjacent queue, scheduling, and storage concerns remain separate Amazon Bedrock Managed access within the AWS operating model The cloud control plane and model-specific behavior add decisions beyond the chat request Google Vertex AI Managed models within Google Cloud governance and deployment workflows Best fit depends on existing Google Cloud ownership and regional requirements LiteLLM Open-source gateway for normalizing access across model providers Your team owns gateway deployment, upgrades, and its failure domain Infrai One key and a consistent contract across 295 routes in 20 modules; the OpenAI-compatible surface can sit beside scheduling and observability capabilities Verify model readiness and region requirements per capability rather than assuming the whole catalog behaves alike

There is no universal winner. A team already governed through AWS may reasonably prefer Bedrock; a Google Cloud shop may value Vertex AI’s existing control plane; a team that needs self-hosted routing may accept LiteLLM’s operational ownership. OpenAI is the direct choice when a focused managed model API is the desired boundary. The consolidated option fits when breadth behind one contract removes several integrations, especially when per-call cost, vendor, latency, cache, and request metadata help correlate a report job with its model call.

For US and EU workloads, require evidence for data processing terms, retention, and available regions for the exact model and capability. A global product page is not enough. If residency is a hard requirement, that evidence can eliminate an option before model scoring begins.

Commit once after the model answers

The preventative code path starts at the request and ends before any queue acknowledgement. This runnable Go program calls the OpenAI-compatible chat surface, requests JSON, honors Retry-After on 429, and validates the result before returning it to the worker. It uses a stable operation key so retries can be correlated. Persist the returned value with a unique constraint on that key before acknowledging the queue message.

package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type classification struct {
    Summary    string  `json:"summary"`
    Category   string  `json:"category"`
    Confidence float64 `json:"confidence"`
}

type chatResponse struct {
    Choices []struct {
        Message struct {
            Content string `json:"content"`
        } `json:"message"`
    } `json:"choices"`
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }
    baseURL := os.Getenv("INFRAI_BASE_URL")
    if baseURL == "" {
        panic("INFRAI_BASE_URL is required")
    }
    endpoint := strings.TrimRight(baseURL, "/") + "/chat/completions"

    payload := map[string]any{
        "model": "deepseek-v4-flash",
        "messages": []map[string]string{
            {"role": "system", "content": "Return JSON with summary, category, and confidence. Category must be harassment, cheating, spam, or other."},
            {"role": "user", "content": "Report r-1842: Player repeatedly sent the same trade link in team chat after being asked to stop."},
        },
        "response_format": map[string]string{"type": "json_object"},
    }
    body, err := json.Marshal(payload)
    if err != nil {
        panic(err)
    }

    client := &http.Client{Timeout: 20 * time.Second}
    var response chatResponse
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodPost, endpoint, bytes.NewReader(body))
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", "report:r-1842:prompt:v1")

        res, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        resBody, readErr := io.ReadAll(res.Body)
        res.Body.Close()
        if readErr != nil {
            panic(readErr)
        }

        if res.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil && seconds > 0 {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if res.StatusCode < 200 || res.StatusCode >= 300 {
            panic(fmt.Sprintf("chat completion failed: status=%d body=%s", res.StatusCode, resBody))
        }
        if err := json.Unmarshal(resBody, &response); err != nil {
            panic(err)
        }
        break
    }

    if len(response.Choices) != 1 {
        panic("expected exactly one completion")
    }
    var result classification
    if err := json.Unmarshal([]byte(response.Choices[0].Message.Content), &result); err != nil {
        panic(fmt.Errorf("invalid classification JSON: %w", err))
    }
    allowed := map[string]bool{"harassment": true, "cheating": true, "spam": true, "other": true}
    if result.Summary == "" || !allowed[result.Category] || result.Confidence < 0 || result.Confidence > 1 {
        panic("classification failed validation")
    }
    fmt.Printf("%s: %s (%.2f)\n", result.Category, result.Summary, result.Confidence)
}

Enter fullscreen mode Exit fullscreen mode

The database still needs a unique constraint on the operation key. Queue acknowledgement happens only after the result is committed. On a schema failure, retain the report for inspection rather than retrying the same invalid output forever; a retry loop cannot repair an incompatible prompt.

Prompt versions must remain explicit. Changing v1 to v2 creates a new operation instead of overwriting history by accident. If policy requires exactly one active classification, promote the new row in a separate transaction and retain the prior version for audit. The platform convention specifies a 24-hour default deduplication window, but database uniqueness remains necessary because queue replays and audit retention can outlive that window.

Sample the confident results too

Start evaluation with real, redacted reports and labels from trained reviewers. Measure category agreement and summary omissions separately. A fluent two-sentence summary can omit the threat that determines urgency, so one aggregate score hides the failure that matters.

Route low-confidence or schema-invalid results to human review without allowing them to lower priority. Sample high-confidence results too. Drift often appears first in new slang, game updates, or coordinated abuse patterns, and a confidence field generated by the same model is not independent proof of correctness.

Measure both paths.

Latency deserves two budgets: time until a report becomes visible to a reviewer, and time until enrichment arrives. Keep the base report visible even if classification is pending. Then an AI outage degrades sorting rather than ingestion.

Use chat requests for live arrivals and batch submission for backfills, policy migrations, or large sets of historical records. Before each large run, estimate tokens so tenant quotas and SaaS plan limits remain predictable. Concurrency should be bounded; a replay should not crowd out new player-safety reports.

Disqualifiers before rollout

Do not use this architecture for an automated enforcement decision that requires deterministic evidence. The output is a prioritization aid for human review. It also stops being sufficient when reviewers need citations across a large policy corpus; that is the point to evaluate retrieval and embeddings, with their own freshness and access-control obligations.

There is another boundary in the platform comparison. The consolidated option has no dedicated moderation endpoint, so text or image moderation must use a chat model with a JSON Schema fallback. Its ASR model catalog currently marks transcription unavailable, and real-time voice session key status is pending and limited to the western region. Those constraints matter for voice-report expansion, even though they do not block text report summarization. Image upscaling is Lanc-only, which is unrelated to this workflow and should not influence the choice.

The decision rule is operational: choose the provider that passes your labeled quality threshold, satisfies the exact US/EU governance requirements, and meets the reviewer latency budget under bounded concurrency. Then make duplicate delivery harmless. Provider breadth is valuable, but it cannot substitute for a durable queue record, structured validation, or a replay runbook.

Sources

원문에서 계속 ↗