GCP

Scale GKE workloads using PromQL metrics in HorizontalPodAutoscaler

GKE now lets you use PromQL queries directly in AutoscalingMetric resources to drive HPA decisions without external adapters.

E

Everything Cloud

Everything Cloud

Scale GKE workloads using PromQL metrics in HorizontalPodAutoscaler

GKE’s built-in support for Prometheus metrics enables HorizontalPodAutoscaler to consume PromQL queries from Cloud Monitoring or Google Managed Prometheus, eliminating the need for third-party adapters. You define metrics via AutoscalingMetric custom resources and reference them in HPA using the autoscaling.gke.io|<resource>|<metric> format, achieving low-latency scaling with minimal operational overhead.

https://storage.googleapis.com/gweb-cloudblog-publish/images/1_zycrwiE.max-900x900.jpg

https://storage.googleapis.com/gweb-cloudblog-publish/images/2_5IffDdj.max-800x800.jpg

Define a Prometheus metric for autoscaling

Create an AutoscalingMetric resource to expose a PromQL query as a custom metric. Specify the query under the promql block, including any required labels or aggregations. For example, to scale based on Pub/Sub undelivered messages, use a query that filters subscription_id and returns the num_undelivered_messages metric. The controller runs on the GKE control plane and only activates when a metric is requested.

The AutoscalingMetric resource uses apiVersion: autoscaling.gke.io/v1beta1 and kind: AutoscalingMetric. In the spec.metrics section, define each promql metric with a name and query string. The query can include PromQL functions like sum, avg_over_time, or rate. If the metric returns per-pod values, set type: Pods; for global values, omit the type or use External in HPA.

This approach removes the need for adapters like the Prometheus Adapter or Stackdriver Custom Metrics Adapter. Since the controller is managed by GKE, there are no pods to install, upgrade, or monitor. Metrics are read directly from Cloud Monitoring or Google Managed Prometheus with low latency, reducing complexity in the autoscaling pipeline.

Link the metric to HorizontalPodAutoscaler

Reference the AutoscalingMetric in your HPA using the external metric type and the format autoscaling.gke.io|<AutoscalingMetric-name>|<metric-name>. For example, if your AutoscalingMetric is named gmp-metric and exposes pubsub-queue-depth, the metric name in HPA becomes autoscaling.gke.io|gmp-metric|pubsub-queue-depth. This format mirrors how custom metrics are referenced in standard HPA configurations.

In the HPA spec, set maxReplicas to define the upper bound and configure the target value under averageValue for External metrics. For instance, to maintain ~100 messages per pod, set averageValue: 100. The HPA controller fetches the metric value every 15 seconds via the AutoscalingMetric system, ensuring timely scaling decisions without adapter-induced delays.

This method works for both global and per-pod metrics. For per-pod scaling (e.g., average memory usage), ensure the PromQL query includes a pod label and set type: Pods in AutoscalingMetric. The HPA then uses the pod metric source type to scale based on individual pod values, enabling fine-grained autoscaling based on workload-specific signals.

Benefits and operational considerations

The built-in integration eliminates adapter maintenance, reduces IAM complexity, and improves reliability by removing failure points in the autoscaling loop. Security is streamlined because the GKE Default Node Service Agent already has read access to Cloud Monitoring and Google Managed Prometheus within the same project, requiring no additional service accounts or keys.

Latency is low: the AutoscalingMetric controller polls the backend every 15 seconds, enabling fast reactions to changing conditions. Resource usage is minimal—the controller scales to zero when no PromQL metrics are active, avoiding overhead. This supports cost efficiency, especially in development or intermittent workloads.

You can combine this feature with HPA scale-to-zero and CapacityBuffers to scale workloads to zero replicas during idle periods and restore them quickly when demand returns. This is effective for event-driven systems like Pub/Sub consumers or batch jobs, reducing spend while maintaining responsiveness.

What to do next

Start by defining an AutoscalingMetric resource with your PromQL query, then reference it in your HPA using the autoscaling.gke.io|<resource>|<metric> format. Test with a simple metric like queue depth or request rate, monitor scaling behavior, and expand to per-pod or advanced PromQL expressions as needed. Refer to the GKE documentation for full examples and limitations.

FAQ

Do I need to enable any APIs or install additional components to use Prometheus metrics in GKE HPA?

No. The feature uses the managed AutoscalingMetric controller in the GKE control plane and relies on existing Cloud Monitoring or Google Managed Prometheus access. No extra APIs, adapters, or installations are required.

Can I use this with self-hosted Prometheus servers?

Not yet. Currently, only Cloud Monitoring and Google Managed Prometheus are supported. Support for self-hosted Prometheus is planned for general availability after the preview phase.

Source: Scale your own way, using HPA with built-in support for PromQL metrics queries in GKE (GCP).

Share:TwitterLinkedIn

Related Articles