A Fully Working Monitoring Cluster for All Your Kubernetes Clusters
Date Published

I wanted one place to look when something is wrong in any of our Kubernetes clusters. Not a Prometheus per cluster, not a Grafana per cluster, but a single monitoring cluster that every other cluster reports to. This post walks through how I built it, from the collectors on each node to the GitOps setup that deploys all of it.
The goal
Before choosing any tools, I wrote down what the setup had to do:
- Collect metrics and logs from every cluster and keep them in one place.
- Use a single tenant,
central-org, so one Grafana can query everything. - Keep recent data on fast local disks and long-term data in cheap object storage.
- Add a new cluster without editing anything by hand on that cluster.
- Keep every piece of configuration in Git.
The architecture at a glance
There are two kinds of clusters. Each source cluster runs a small collector layer. The monitoring cluster receives everything and stores it. A third cluster, the devops cluster, runs ArgoCD and the GitOps repository, and deploys to both.
- Source clusters: an Alloy DaemonSet, kube-state-metrics, and node-exporter.
- Monitoring cluster: Gateway API, an Alloy gateway, distributed Mimir and Loki, Kafka, Rook-Ceph, and Grafana.
- Devops cluster: ArgoCD and the GitOps repo.
The rest of the post follows the path of the data, then covers how Git and ArgoCD put it all in place.
Collecting at the source
Every node in every source cluster runs one Alloy collector, deployed as a DaemonSet. It gathers four kinds of data:
- Kubelet and cAdvisor metrics, scraped directly from the host. The kubelet exposes node and pod metrics, and cAdvisor, which is built into the kubelet, exposes container CPU, memory, network, and disk usage.
- Node metrics from node-exporter, which also runs as a DaemonSet, so there is one per node.
- Cluster object metrics from kube-state-metrics (KSM). This one is a Deployment, because it reads the state of Kubernetes objects from the API and a single copy per cluster is enough.
- Logs from the containerd log files under
/var/log. Alloy tails them from the node, so applications do not need to know anything about logging.
Because Alloy runs on every node, each collector only needs to look at its own node. That keeps the scrape load spread out and avoids one big collector that has to reach everything.
Getting data in: Gateway API and the Alloy gateway
Everything the collectors gather is sent to the monitoring cluster, which is where the Gateway API comes in. I use two kinds of routes:
- An
HTTPRoutecarries metrics and logs today. - A
GRPCRouteis reserved for Tempo and Pyroscope, which I plan to add later and which speak gRPC. Having the route type in mind from the start means I do not need to redesign the entry point when traces and profiles arrive.
Behind the Gateway sits an Alloy gateway. It is the single receiver for the whole fleet. It takes in metrics and logs and forwards metrics to Mimir and logs to Loki. The tenant is central-org all the way through.
One consequence of a single tenant is that clusters are separated by labels, not by tenant. Every collector adds a cluster label to everything it sends, taken from the cluster name in ArgoCD. If you copy this design, decide on that label on day one, because adding it later means gaps in your history.
Mimir and Loki, and why Kafka shows up
For storage I use the distributed Helm charts for Mimir and Loki. In distributed mode, each part of the system (distributors, ingesters, queriers, compactors and so on) runs as its own workload and scales on its own.
The charts I deploy need an external Kafka. I did not want to run Kafka by hand, so I installed the Strimzi operator and describe the Kafka cluster as Kubernetes resources. That means Kafka is deployed and kept in sync the same way as everything else, through ArgoCD. The details are in the GitOps section below.
Local PV first, then S3
This is the part of the design I like most. Each ingester keeps its recent data on a 10Gi local PV. For Mimir that is the TSDB data, and Loki ingesters hold their recent data locally in the same way. Recent queries are answered from there, quickly, without touching object storage.
Then the data moves on:
- Data is compacted. Small blocks are merged into larger ones, which makes long-term storage and queries much more efficient.
- It is shipped to S3 buckets served by the Rook-Ceph RGW (the RADOS Gateway, which speaks the S3 API).
Each bucket is requested through an ObjectBucketClaim, so Rook creates the bucket and the credentials, and I point the Mimir and Loki configuration at them. The result is that the local disks can stay small, and long-term data lives in object storage that I can grow independently.
Reading it back with Grafana
Grafana runs as a Deployment in the monitoring cluster and has two data sources: Mimir for metrics and Loki for logs. When you run a query, each query path combines two sources of data: recent data from the ingesters, and older data read back from S3. For Mimir the store-gateway serves the older blocks, and for Loki the index gateway serves the index. From the dashboard's point of view this is invisible, which is exactly the point.
Everything lives in Git
All of the above is deployed with ArgoCD. ArgoCD and the GitOps repository both live in the devops cluster. The monitoring cluster and every source cluster are registered in ArgoCD as deployment targets.
The manifests in this section are simplified versions to show the structure. Names, URLs, and versions are placeholders, so check each chart's documentation for the current values.
The shape of the repo
The repository holds only values and small manifests. The charts themselves come from their upstream Helm repositories.
gitops/
├── monitoring/
│ ├── mimir/values.yaml
│ ├── loki/values.yaml
│ ├── grafana/values.yaml
│ ├── alloy-gateway/values.yaml
│ └── kafka/ # Strimzi resources
├── collectors/
│ ├── alloy/values.yaml # includes the Alloy config
│ ├── kube-state-metrics/values.yaml
│ └── node-exporter/values.yaml
└── argocd/
├── applications/ # Mimir, Loki, Grafana, Kafka, gateway
└── applicationsets/ # collectors for every cluster
Charts from upstream, values from Git
Each part of the monitoring cluster is an ArgoCD Application with two sources. The first is the upstream chart. The second is my GitOps repository, which is referenced by name so that its values files can be used by the first source.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: mimir
namespace: argocd
spec:
project: monitoring
sources:
- repoURL: https://grafana.github.io/helm-charts
chart: mimir-distributed
targetRevision: <CHART_VERSION>
helm:
valueFiles:
- $values/monitoring/mimir/values.yaml
- repoURL: https://git.example.com/platform/gitops.git
targetRevision: main
ref: values
destination:
name: monitoring
namespace: mimir
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
- ServerSideApply=true
Loki, Grafana and the Alloy gateway follow the same pattern with their own charts and values files. The upside is that upgrading a chart is a one-line change to targetRevision, and changing configuration is a normal pull request on a values file.
The local PV size and the bucket live in the values file. A simplified Mimir example looks like this:
ingester:
persistentVolume:
enabled: true
size: 10Gi
mimir:
structuredConfig:
blocks_storage:
backend: s3
s3:
endpoint: <RGW_SERVICE>:80
bucket_name: mimir-blocks
insecure: true
With an ObjectBucketClaim, the S3 credentials land in a Secret next to the bucket, so no keys need to be stored in Git.
Kafka with Strimzi
Because Mimir and Loki need Kafka, the Strimzi operator is its own Application, and the Kafka cluster is described with the operator's custom resources in the monitoring/kafka folder of the repo. A simplified version of the Kafka resource looks like this:
apiVersion: kafka.strimzi.io/v1beta2
kind: Kafka
metadata:
name: monitoring-kafka
namespace: kafka
spec:
kafka:
version: <KAFKA_VERSION>
listeners:
- name: plain
port: 9092
type: internal
tls: false
entityOperator:
topicOperator: {}
The exact fields, including how broker nodes are defined with node pools, depend on your Strimzi version, so follow its documentation for your release.
Two ArgoCD details matter here:
- Order. The Kafka resource cannot be applied before the operator has installed its CRDs, and Mimir and Loki should not start before Kafka exists. A clean way to handle it is to deploy the operator, the Kafka cluster, and then the charts as separate Applications, and to set the
SkipDryRunOnMissingResource=truesync option on the Application that holds the custom resources, so ArgoCD does not fail while the CRDs are not there yet. - Large CRDs. Some CRDs are too big for client-side apply, which stores the whole object in an annotation.
ServerSideApply=trueavoids that.
One config for every collector
The Alloy collector is configured with River, the syntax Grafana now calls the Alloy configuration syntax. The config lives in collectors/alloy/values.yaml in Git, and the same file is used for every cluster. A simplified version looks like this:
discovery.kubernetes "nodes" {
role = "node"
}
// Only look at the node this collector runs on.
discovery.relabel "this_node" {
targets = discovery.kubernetes.nodes.targets
rule {
source_labels = ["__meta_kubernetes_node_name"]
regex = sys.env("NODE_NAME")
action = "keep"
}
}
prometheus.scrape "kubelet" {
targets = discovery.relabel.this_node.output
scheme = "https"
bearer_token_file = "/var/run/secrets/kubernetes.io/serviceaccount/token"
forward_to = [prometheus.remote_write.central.receiver]
tls_config {
insecure_skip_verify = true
}
}
prometheus.scrape "cadvisor" {
targets = discovery.relabel.this_node.output
scheme = "https"
metrics_path = "/metrics/cadvisor"
bearer_token_file = "/var/run/secrets/kubernetes.io/serviceaccount/token"
forward_to = [prometheus.remote_write.central.receiver]
tls_config {
insecure_skip_verify = true
}
}
local.file_match "pod_logs" {
path_targets = [{ "__path__" = "/var/log/pods/*/*/*.log" }]
}
loki.source.file "pod_logs" {
targets = local.file_match.pod_logs.targets
forward_to = [loki.process.cri.receiver]
}
loki.process "cri" {
stage.cri {}
forward_to = [loki.write.central.receiver]
}
prometheus.remote_write "central" {
external_labels = { cluster = sys.env("CLUSTER_NAME") }
endpoint {
url = "https://<GATEWAY_HOST>/api/v1/metrics/write"
headers = { "X-Scope-OrgID" = "central-org" }
}
}
loki.write "central" {
external_labels = { cluster = sys.env("CLUSTER_NAME") }
endpoint {
url = "https://<GATEWAY_HOST>/loki/api/v1/push"
tenant_id = "central-org"
}
}
A real config has a few more pieces. It scrapes the node-exporter and kube-state-metrics pods in the same way, discovering them by label, and it extracts the namespace and pod name from the log file path. The shape stays the same.
Notice that the config contains no cluster name. It reads CLUSTER_NAME and NODE_NAME from environment variables. That keeps the file identical everywhere, and it means the file in Git never has to be edited when a cluster is added. It also keeps the config out of the ApplicationSet itself, where its curly braces could clash with ApplicationSet templating.
The values file around it is short: run the chart as a DaemonSet, mount the host's /var/log, and put the config above in the chart's config map.
The ApplicationSet that deploys it uses the clusters generator, which produces one entry for every cluster registered in ArgoCD. It sets the two environment variables, and the cluster name comes straight from ArgoCD:
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: collector-alloy
namespace: argocd
spec:
goTemplate: true
goTemplateOptions: ["missingkey=error"]
generators:
- clusters: {}
template:
metadata:
name: '{{.name}}-alloy'
spec:
project: collectors
sources:
- repoURL: https://grafana.github.io/helm-charts
chart: alloy
targetRevision: <CHART_VERSION>
helm:
valueFiles:
- $values/collectors/alloy/values.yaml
valuesObject:
alloy:
extraEnv:
- name: NODE_NAME
valueFrom:
fieldRef:
fieldPath: spec.nodeName
- name: CLUSTER_NAME
value: '{{.name}}'
- repoURL: https://git.example.com/platform/gitops.git
targetRevision: main
ref: values
destination:
server: '{{.server}}'
namespace: monitoring-agents
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
Exporters everywhere with the matrix generator
kube-state-metrics and node-exporter need to be in every cluster as well. For these I use an ApplicationSet with a matrix generator. It combines two generators and produces every pairing of them: the first generator lists the registered clusters, and the second is a list of the apps to deploy.
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: collector-exporters
namespace: argocd
spec:
goTemplate: true
goTemplateOptions: ["missingkey=error"]
generators:
- matrix:
generators:
- clusters: {}
- list:
elements:
- app: kube-state-metrics
chart: kube-state-metrics
repoURL: https://prometheus-community.github.io/helm-charts
version: "<CHART_VERSION>"
- app: node-exporter
chart: prometheus-node-exporter
repoURL: https://prometheus-community.github.io/helm-charts
version: "<CHART_VERSION>"
template:
metadata:
name: '{{.name}}-{{.app}}'
spec:
project: collectors
sources:
- repoURL: '{{.repoURL}}'
chart: '{{.chart}}'
targetRevision: '{{.version}}'
helm:
valueFiles:
- '$values/collectors/{{.app}}/values.yaml'
- repoURL: https://git.example.com/platform/gitops.git
targetRevision: main
ref: values
destination:
server: '{{.server}}'
namespace: monitoring-agents
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true
With three registered clusters and two apps, ArgoCD generates six Applications. Adding a third app to the list adds one more to every cluster. If you ever need to leave a cluster out, a label selector on the clusters generator does it.
Adding a new cluster
This is the payoff of the whole setup. To bring a new cluster into monitoring, I register it in ArgoCD, either with the argocd cluster add command or with a declarative cluster Secret. The clusters generator picks it up, and the ApplicationSets create its Alloy, kube-state-metrics, and node-exporter Applications. The collectors start, and its metrics and logs show up in Grafana with their own cluster label.
There is nothing to configure by hand on the new cluster, and nothing to edit in Git.
What to watch out for
A few things I would plan for if you build something similar:
- Kafka is a real dependency. If Mimir and Loki need it, then it needs to be as reliable as they are. Treat it as part of the monitoring stack, not as an add-on.
- Single tenant means shared limits. Ingestion limits apply to the tenant as a whole, so one noisy cluster can use up the budget of all the others. Watch the per-cluster volumes through the
clusterlabel. - Size the local PV deliberately. It holds the recent window before data is shipped to S3, so watch its usage and set the size with that in mind.
- Keep the ordering explicit. The operator, then Kafka, then the charts. Letting ArgoCD figure that out on its own leads to confusing failures on the first sync.
What's next
The route for gRPC is already reserved, so the next step is adding Tempo for traces and Pyroscope for profiles, using the same pattern: charts from upstream, values in Git, and a collector change in the same Alloy config. If you have built something similar, I would like to hear what you would do differently.

How pod QoS classes, the scheduler's blind spot, and vertical autoscaling work together illustrated with real overcommitment data from a prod cluster.