Skip to main content
This guide describes how to scale the number of Limina Pods automatically with the Kubernetes Horizontal Pod Autoscaler (HPA), using the Prometheus metrics that the container exposes on its /metrics route. The example below scales on text_processing_queue_depth, the per-pod average backlog of in-flight text processing calls.

How autoscaling works

The HPA is a built-in Kubernetes controller that adjusts a Deployment’s replica count to keep an observed metric near a target value. On its own it can only read CPU and memory. To scale on an application metric such as text_processing_queue_depth, the Prometheus Adapter is needed to publish that metric through the Kubernetes custom metrics API. The steps in this guide build that chain end to end:
Metrics Pipeline

Available metrics

Limina exposes three backlog depth gauges, one per pipeline layer. Any of them can be used as the scaling signal. The queue_depth meter covers every workload type, while text_processing_queue_depth counts only the text processing workload which may include the NER, the coreference resolution, the synthetic entity generation and the relation extraction.
This guide uses text_processing_queue_depth, which best tracks the bottleneck for text processing workloads. For image or audio files, the inference work (OCR, object detection, audio transcription) is not counted by this metric — use queue_depth as the scaling signal instead. For more information on the /metrics route, see Deploying into Production.

1. Prerequisites

Before you begin, make sure that your cluster and nodegroup have been created and that Limina Pods are running. See the Kubernetes Setup Guide for the Deployment and Service manifests, and the AWS EKS Setup Guide if you are running on EKS. The examples in this guide assume a nodegroup defined as follows.
eksctl Command
Set --nodes-max to at least the HPA’s maxReplicas. The Limina GPU container occupies one GPU per Pod, so the number of available GPUs caps the Pod count.

2. Install Prometheus

Helm Command
expected output
Output

3. Configure the Prometheus Adapter

The adapter is what makes Prometheus metrics visible to the HPA. Create a values file as follows.
Prometheus Adapter Values
Always set prometheus.url explicitly as shown. The chart’s default value does not resolve.
Now install the adapter with this values file.
Helm Command
expected output
Output

4. Create an internal ClusterIP Service

This step is required when Limina is exposed through an ELB. It provides a monitoring-only Service for the ServiceMonitor to select. If Limina is only exposed internally, skip this step and see Without an ELB below.
Monitoring Service Manifest
Apply it with this kubectl command.
Kubernetes Command

5. Create the ServiceMonitor

The ServiceMonitor tells Prometheus which Pods to scrape.
ServiceMonitor Manifest
Apply it with this kubectl command.
Kubernetes Command
expected output
Output

Without an ELB

If the app already has a ClusterIP Service, step 4 is unnecessary — point the ServiceMonitor at that Service instead. Match selector to its labels and port to its port name.
ServiceMonitor Manifest

6. Create the HPA

HPA Manifest
Apply it with this kubectl command.
Kubernetes Command
expected output
Output
averageValue is the per-pod backlog that the HPA holds the Deployment at, so a lower value scales out earlier. Pick it relative to the number of simultaneous requests each container is expected to handle — see Concurrency for the recommended levels — then tune it against your own traffic patterns.

7. Verify the setup

First, check that all Pods have started successfully.
Kubernetes Command
expected output
Output
Then check that Kubernetes recognizes the custom metric.
Kubernetes Command
expected output
Output
Once pods/text_processing_queue_depth_avg2m is returned, the metric is available to the HPA.

8. Test autoscaling

Replace <your-load-balancer> in the commands below with your own ELB hostname. Before submitting any requests, the metric reads zero.
Command
expected output
Output
After submitting 5 concurrent redaction requests, the backlog is reflected in the gauge.
Command
expected output
Output
The averageValue: "4" threshold has been exceeded, so the HPA starts a second Pod.
Kubernetes Command
expected output
Output
Once the backlog clears, the replica count returns to minReplicas. This takes at least 5 minutes, because scaleDown.stabilizationWindowSeconds: 300 holds the count steady until the metric has stayed low for that long.
Kubernetes Command
expected output
Output

Additional Resources

  • Kubernetes Setup Guide — the Deployment and Service manifests that this guide scales.
  • Concurrency — recommended simultaneous requests per container, for choosing your HPA target.