Contributing a GCP Cloud Spanner Scaler to KEDA

작성자

카테고리:

← 피드로
DEV Community · Petr Petrenko · 2026-08-05 개발(SW)
Cover image for Contributing a GCP Cloud Spanner Scaler to KEDA

Petr Petrenko

How we added native Spanner support to Kubernetes autoscaling — and what we learned along the way

The problem

We run several workloads on Kubernetes that process jobs stored in Cloud Spanner tables. The pattern is simple: a producer writes rows with status = 'pending', workers pick them up and mark them done. The question is — how many workers do you run?

Fixed replica counts mean either wasted money during quiet periods or dropped throughput during spikes. We needed autoscaling based on actual queue depth, not CPU or memory.

KEDA (Kubernetes Event-Driven Autoscaling) is the standard answer for this — it scales workloads based on external metrics like queue lengths, database counts, and custom queries. It already had scalers for GCP Pub/Sub, Cloud Tasks, and Cloud Storage. But not Spanner.

So we built one.

How KEDA works

KEDA polls Spanner on every interval, translates the query result into a replica count, and tells HPA to scale the worker deployment accordingly.

KEDA sits between your workload and the external system. On every polling interval it runs your query, gets a number back, and tells Kubernetes HPA how many replicas to run based on ceil(currentValue / targetValue).

The scaler

The gcp-spanner trigger takes any SQL query that returns a single INT64 value:

triggers:
  - type: gcp-spanner
    metadata:
      projectId: my-project
      instanceId: my-instance
      databaseId: my-database
      query: "SELECT COUNT(*) FROM jobs WHERE status = 'pending'"
      targetValue: "5"       # one replica handles 5 pending jobs
      activationValue: "2"   # stay at 0 replicas below this threshold
      credentialsFromEnv: GOOGLE_APPLICATION_CREDENTIALS_JSON

Enter fullscreen mode Exit fullscreen mode

With targetValue: 5 and 20 pending jobs, KEDA will maintain 4 worker replicas. When the queue drains to 0, it scales back to zero.

Scaling behaviour

KEDA keeps workers at zero until the queue exceeds the activation threshold, scales out proportionally to queue depth, and scales back to zero when the queue drains.

What we learned contributing to KEDA

1. Schema generation matters

KEDA auto-generates its scaler schema from Go struct tags. We needed to add Credentials and CredentialsFromEnvFile fields to the metadata struct purely so they appear in the schema — even though GetGCPAuthorization reads them directly from the config. This is the same pattern used by other GCP scalers.

2. ParseCommand has a quirk

KEDA’s e2e test helper splits commands on spaces, honouring single-quotes only. So --ddl="value with spaces" breaks — "value" doesn’t get treated as a quoted string. The fix is --ddl 'value with spaces' — space-separated flag and value, with single quotes.

3. Cleanup must be independent of Kubernetes

Our first attempt deleted the Spanner test instance via gcloud running in a pod. But DeleteKubernetesResources removes the pod before t.Cleanup fires — leaving orphaned (and billing) Spanner instances. The fix: delete via the Go Spanner Admin API directly from the test binary, completely independent of any pod.

Registering t.Cleanup immediately after instance creation ensures the Spanner instance is always deleted — even if later setup steps fail or Kubernetes resources are removed first.

4. targetValue: 0 causes HPA divide-by-zero

We added a Validate() method to reject targetValue <= 0 — caught by a Copilot review comment. KEDA’s TypedConfig calls Validate() automatically via the CustomValidator interface.

The PR

The contribution includes:

  • Scaler implementation with all three GCP auth methods (inline JSON, env file, Workload Identity)
  • 18 unit tests
  • 6 integration tests against the Cloud Spanner emulator
  • e2e test that provisions a real Spanner instance, runs scaling scenarios, and cleans up

Using it

kubectl apply -f - <<EOF
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: spanner-worker-scaler
spec:
  scaleTargetRef:
    name: job-processor
  minReplicaCount: 0
  maxReplicaCount: 20
  triggers:
    - type: gcp-spanner
      metadata:
        projectId: my-project
        instanceId: my-instance
        databaseId: my-database
        query: "SELECT COUNT(*) FROM jobs WHERE status = 'pending'"
        targetValue: "5"
        activationValue: "2"
        credentialsFromEnv: GOOGLE_APPLICATION_CREDENTIALS_JSON
EOF

Enter fullscreen mode Exit fullscreen mode

That’s it. KEDA will handle the rest.

원문에서 계속 ↗

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다