A
Ahmed Tariq
Guest
Two AKS clusters used the equivalent Fleet processor configuration and produced different cluster metadata. One cluster produced an empty
The difference came from Azure identity. The cluster that looked correct had two user-assigned identities on its node pool. Elastic Agent could not choose one, so its Azure token request failed, and the cloud metadata processor skipped the update. The cluster with one identity obtained a token, completed an Azure Resource Manager request, found no matching cluster, and wrote an empty string. The failed credential path hid the metadata defect.
I began with the rendered agent configuration. Running
Both policies resolved to the same relevant structure. Each had 24 Kubernetes streams with a
Elastic Agent's Azure cloud metadata provider reads instance metadata and then tries to add AKS cluster metadata. The provider creates an Azure Resource Manager client, lists managed clusters in the subscription, and compares each cluster's node resource group with the resource group that the Azure Instance Metadata Service reports.
The Elastic Beats v9.5.3 source shows two outcomes that look similar to an operator but produce different documents.
That small distinction explained the opposite results.
Microsoft's managed identity guidance requires a client ID, object ID, or resource ID when a VM has several user-assigned identities. The agent's default credential path did not select one. Its token request therefore failed on the production node pool before the managed-cluster query could run.
On the non-production node pool, the token request succeeded. The managed-cluster list returned an empty
I also checked the role assignments on both node identities. Neither identity could read managed cluster resources. The lookup could not resolve a cluster name in either environment. One path failed before the lookup, while the other path reached the lookup and returned no match.
Elastic's AKS guidance says that Elastic Agent does not populate
The configuration followed that guidance. In this deployment, the field sometimes existed with an empty value. That difference changed the behaviour of the workaround.
A stored document contained both the configured cluster name and the empty value:
Elasticsearch supports arrays without a dedicated array type. A terms aggregation therefore saw two values in the same field and placed one stored document in two buckets. The ingestion path had created a multi-valued field. The dashboard displayed the consequence.
Kibana's blank label hid another distinction. Some agent documents contained an empty string. Other documents had no
I separated the cases with field-existence and exact-value conditions before changing the pipeline. A reduced query for the explicit empty-value case is:
A second query was used
The fix needed two changes. The ingest path first had to remove empty cluster metadata without damaging valid values. Fleet then had to add the configured cluster name to the System streams as well as the Kubernetes streams.
Elastic provides a
The pipeline normalized
The script below is a generalized version of that pattern for Kibana Dev Tools:
The simulate API can test the behaviour before the pipeline processes live data:
The first simulated document kept
I activated the cleanup pipeline, then added the existing cluster-name processor to the System streams. Applying the cleanup first matters because extending
I evaluated the following changes against the stored documents, the rendered policy, or the identity model.
The node identity did not need broader access for this repair. The configured
Agent startup logs can roll away before anyone investigates. The on-disk agent logs preserved the cloud metadata error after the current console output no longer showed it. The surprising part of this incident was simple once the two paths were visible. A successful Azure lookup with no matching cluster made Elastic Agent write an empty name. The identity failure stopped the lookup earlier, so the configured name remained untouched.
orchestrator.cluster.name. The other retained the configured name and looked correct in Kibana. I observed this behaviour with Elastic Agent 9.5.3.The difference came from Azure identity. The cluster that looked correct had two user-assigned identities on its node pool. Elastic Agent could not choose one, so its Azure token request failed, and the cloud metadata processor skipped the update. The cluster with one identity obtained a token, completed an Azure Resource Manager request, found no matching cluster, and wrote an empty string. The failed credential path hid the metadata defect.
The Fleet processor configuration did not explain the difference
I began with the rendered agent configuration. Running
elastic-agent inspect inside an agent pod shows the policy that the agent actually received, including the processors attached to each stream. This avoids concluding what the Fleet interface claims it sent.Both policies resolved to the same relevant structure. Each had 24 Kubernetes streams with a
add_fields processor that supplied the cluster name, plus 18 System streams without that processor. The processor structure matched. That comparison ruled out the most obvious explanation. The remaining difference was the identity configuration of the AKS node pools.Elastic Agent treats no match and lookup failure differently
Elastic Agent's Azure cloud metadata provider reads instance metadata and then tries to add AKS cluster metadata. The provider creates an Azure Resource Manager client, lists managed clusters in the subscription, and compares each cluster's node resource group with the resource group that the Azure Instance Metadata Service reports.
The Elastic Beats v9.5.3 source shows two outcomes that look similar to an operator but produce different documents.
| Outcome | Function result | Document result |
|---|---|---|
| The list succeeds but finds no match | Empty name, empty ID, and no error | The caller writes empty cluster fields |
| Token acquisition, client creation, or the list fails | An error | The caller returns without adding the cluster fields |
getAKSClusterNameID reaches the end of the list and returns two empty strings with a nil error when it finds no match. The caller checks only the error. A nil error causes it to write both values:
Code:
clusterName, clusterID, err := getAKSClusterNameID(...)
if err == nil {
result.metadata.Put("orchestrator.cluster.id", clusterID)
result.metadata.Put("orchestrator.cluster.name", clusterName)
}
That small distinction explained the opposite results.
| Environment | Node-pool identities | Token request | Managed-cluster list | Cloud metadata contribution |
|---|---|---|---|---|
| Production | Two user-assigned identities | Fails because the request does not select an identity | Never runs | Cloud metadata processor adds no cluster field |
| Non-production | One user-assigned identity | Succeeds | Returns an empty list | Cloud metadata processor writes an empty string |
Microsoft's managed identity guidance requires a client ID, object ID, or resource ID when a VM has several user-assigned identities. The agent's default credential path did not select one. Its token request therefore failed on the production node pool before the managed-cluster query could run.
On the non-production node pool, the token request succeeded. The managed-cluster list returned an empty
value array. That response proves that the call found no candidate. It does not, by itself, prove why Azure returned an empty list.
I also checked the role assignments on both node identities. Neither identity could read managed cluster resources. The lookup could not resolve a cluster name in either environment. One path failed before the lookup, while the other path reached the lookup and returned no match.
The documented add_fields workaround exposed the empty value
Elastic's AKS guidance says that Elastic Agent does not populate
orchestrator.cluster.name on AKS. It recommends a add_fields processor that supplies the cluster name for each Kubernetes integration component.The configuration followed that guidance. In this deployment, the field sometimes existed with an empty value. That difference changed the behaviour of the workaround.
A stored document contained both the configured cluster name and the empty value:
Code:
{
"orchestrator": {
"cluster": {
"id": "",
"name": ["", "<cluster-name>"]
}
}
}
Elasticsearch supports arrays without a dedicated array type. A terms aggregation therefore saw two values in the same field and placed one stored document in two buckets. The ingestion path had created a multi-valued field. The dashboard displayed the consequence.
A blank bucket combined empty and missing fields
Kibana's blank label hid another distinction. Some agent documents contained an empty string. Other documents had no
orchestrator.cluster.name field at all. The missing-field group included application traces that never passed through Elastic Agent and System streams that had no cluster-name processor. Those documents could not support the same conclusion as an agent document with an explicit empty value.I separated the cases with field-existence and exact-value conditions before changing the pipeline. A reduced query for the explicit empty-value case is:
Code:
GET logs-*,metrics-*,traces-*/_search
{
"size": 0,
"query": {
"bool": {
"filter": [
{ "exists": { "field": "orchestrator.cluster.name" } },
{ "term": { "orchestrator.cluster.name": "" } }
]
}
},
"track_total_hits": true
}
A second query was used
must_not with the exists clause to count documents where the field did not exist. Splitting the two populations by data_stream.dataset showed which integrations created each one. This also explained why dashboard bucket totals could exceed the number of stored documents. A terms aggregation counts each value in a multi-valued field. One document with an empty string and a cluster name contributes to both terms.I repaired the field in a custom ingest pipeline
The fix needed two changes. The ingest path first had to remove empty cluster metadata without damaging valid values. Fleet then had to add the configured cluster name to the System streams as well as the Kubernetes streams.
Elastic provides a
global@custom hook for custom data processing. Its data-stream documentation says that default integration pipelines call this hook and that custom pipelines persist across version upgrades. Elastic also warns users to test custom processing carefully.The pipeline normalized
orchestrator.cluster.name and orchestrator.cluster.id in four cases:- It left a valid scalar unchanged.
- It removed empty members from a list.
- It collapsed a one-item list to a scalar.
- It removed the key when no valid value remained.
The script below is a generalized version of that pattern for Kibana Dev Tools:
Code:
PUT _ingest/pipeline/global@custom
{
"processors": [
{
"script": {
"lang": "painless",
"source": """
if (ctx.orchestrator == null || ctx.orchestrator.cluster == null) {
return;
}
def cluster = ctx.orchestrator.cluster;
for (def key : ['name', 'id']) {
if (!cluster.containsKey(key)) {
continue;
}
def original = cluster[key];
def cleaned = new ArrayList();
if (original instanceof List) {
for (def item : original) {
if (item != null && item.toString().length() > 0) {
cleaned.add(item);
}
}
} else if (original != null && original.toString().length() > 0) {
cleaned.add(original);
}
if (cleaned.isEmpty()) {
cluster.remove(key);
} else if (cleaned.size() == 1) {
cluster[key] = cleaned.get(0);
} else {
cluster[key] = cleaned;
}
}
"""
}
}
]
}
The simulate API can test the behaviour before the pipeline processes live data:
Code:
POST _ingest/pipeline/global@custom/_simulate
{
"docs": [
{
"_source": {
"orchestrator": {
"cluster": {
"name": ["", "example-cluster"],
"id": ""
}
}
}
},
{
"_source": {
"message": "document without cluster metadata"
}
}
]
}
The first simulated document kept
example-cluster as a scalar and lost the empty ID. The second document passed through unchanged. The ingest pipeline corrected new events. Existing indices kept their previous field shape unless they were reprocessed through a pipeline.I activated the cleanup pipeline, then added the existing cluster-name processor to the System streams. Applying the cleanup first matters because extending
add_fields alone can turn a bare empty string into the same two-value array seen on the Kubernetes streams.Repairs I rejected
I evaluated the following changes against the stored documents, the rendered policy, or the identity model.
| Candidate | What the check showed |
|---|---|
| Drop the field in each Fleet stream | The resolved agent policy contained the drop processor, but indexed documents still gained the empty value later in the agent path |
| Add global data tags in Fleet | Another field-writing processor would leave both the configured value and the empty value in the document |
| Disable Azure in BEATS_ADD_CLOUD_METADATA_PROVIDERS | This removes useful cloud.* enrichment along with the defective cluster fields |
| Grant managed-cluster read to the node identity | This broadens node-identity permissions and does not resolve the ambiguous-identity failure on a node pool with several identities |
| Add another identity to the working token path | This copies the authentication failure, so the processor aborts and the field disappears |
The node identity did not need broader access for this repair. The configured
add_fields processor already knew the cluster name. The ingest pipeline only removed the invalid value that the cloud metadata path added.The checks that now catch this failure
elastic-agent inspect provide the resolved Fleet policy. I use it to compare the processors for each stream. An empty string and an absent field need separate queries. A dashboard can render both as blank, even though they describe different execution paths. The node pool's managed identities belong in an observability investigation when the agent uses Azure's default credential chain. A Kubernetes manifest can stay unchanged while an add-on changes the identities attached to the underlying virtual machine scale set.Agent startup logs can roll away before anyone investigates. The on-disk agent logs preserved the cloud metadata error after the current console output no longer showed it. The surprising part of this incident was simple once the two paths were visible. A successful Azure lookup with no matching cluster made Elastic Agent write an empty name. The identity failure stopped the lookup earlier, so the configured name remained untouched.