Two Identical AKS Clusters, Two Different Elastic Metadata Bugs

A

Ahmed Tariq

Guest
Two AKS clusters used the equivalent Fleet processor configuration and produced different cluster metadata. One cluster produced an empty orchestrator.cluster.name. The other retained the configured name and looked correct in Kibana. I observed this behaviour with Elastic Agent 9.5.3.

The difference came from Azure identity. The cluster that looked correct had two user-assigned identities on its node pool. Elastic Agent could not choose one, so its Azure token request failed, and the cloud metadata processor skipped the update. The cluster with one identity obtained a token, completed an Azure Resource Manager request, found no matching cluster, and wrote an empty string. The failed credential path hid the metadata defect.

The Fleet processor configuration did not explain the difference​


I began with the rendered agent configuration. Running elastic-agent inspect inside an agent pod shows the policy that the agent actually received, including the processors attached to each stream. This avoids concluding what the Fleet interface claims it sent.

Both policies resolved to the same relevant structure. Each had 24 Kubernetes streams with a add_fields processor that supplied the cluster name, plus 18 System streams without that processor. The processor structure matched. That comparison ruled out the most obvious explanation. The remaining difference was the identity configuration of the AKS node pools.

Elastic Agent treats no match and lookup failure differently​


Elastic Agent's Azure cloud metadata provider reads instance metadata and then tries to add AKS cluster metadata. The provider creates an Azure Resource Manager client, lists managed clusters in the subscription, and compares each cluster's node resource group with the resource group that the Azure Instance Metadata Service reports.

The Elastic Beats v9.5.3 source shows two outcomes that look similar to an operator but produce different documents.


Outcome

Function result

Document result

The list succeeds but finds no match

Empty name, empty ID, and no error

The caller writes empty cluster fields

Token acquisition, client creation, or the list fails

An error

The caller returns without adding the cluster fields

getAKSClusterNameID reaches the end of the list and returns two empty strings with a nil error when it finds no match. The caller checks only the error. A nil error causes it to write both values:

Code:
clusterName, clusterID, err := getAKSClusterNameID(...)
if err == nil {
    result.metadata.Put("orchestrator.cluster.id", clusterID)
    result.metadata.Put("orchestrator.cluster.name", clusterName)
}

That small distinction explained the opposite results.


Environment

Node-pool identities

Token request

Managed-cluster list

Cloud metadata contribution

Production

Two user-assigned identities

Fails because the request does not select an identity

Never runs

Cloud metadata processor adds no cluster field

Non-production

One user-assigned identity

Succeeds

Returns an empty list

Cloud metadata processor writes an empty string

Microsoft's managed identity guidance requires a client ID, object ID, or resource ID when a VM has several user-assigned identities. The agent's default credential path did not select one. Its token request therefore failed on the production node pool before the managed-cluster query could run.

On the non-production node pool, the token request succeeded. The managed-cluster list returned an empty value array. That response proves that the call found no candidate. It does not, by itself, prove why Azure returned an empty list.

Two AKS identity paths produce different Elastic Agent cluster metadata results


I also checked the role assignments on both node identities. Neither identity could read managed cluster resources. The lookup could not resolve a cluster name in either environment. One path failed before the lookup, while the other path reached the lookup and returned no match.

The documented add_fields workaround exposed the empty value​


Elastic's AKS guidance says that Elastic Agent does not populate orchestrator.cluster.name on AKS. It recommends a add_fields processor that supplies the cluster name for each Kubernetes integration component.

The configuration followed that guidance. In this deployment, the field sometimes existed with an empty value. That difference changed the behaviour of the workaround.

A stored document contained both the configured cluster name and the empty value:

Code:
{
  "orchestrator": {
    "cluster": {
      "id": "",
      "name": ["", "<cluster-name>"]
    }
  }
}

Elasticsearch supports arrays without a dedicated array type. A terms aggregation therefore saw two values in the same field and placed one stored document in two buckets. The ingestion path had created a multi-valued field. The dashboard displayed the consequence.

A configured cluster name and an empty metadata value become a two-value Elasticsearch field


A blank bucket combined empty and missing fields​


Kibana's blank label hid another distinction. Some agent documents contained an empty string. Other documents had no orchestrator.cluster.name field at all. The missing-field group included application traces that never passed through Elastic Agent and System streams that had no cluster-name processor. Those documents could not support the same conclusion as an agent document with an explicit empty value.

I separated the cases with field-existence and exact-value conditions before changing the pipeline. A reduced query for the explicit empty-value case is:

Code:
GET logs-*,metrics-*,traces-*/_search
{
  "size": 0,
  "query": {
    "bool": {
      "filter": [
        { "exists": { "field": "orchestrator.cluster.name" } },
        { "term": { "orchestrator.cluster.name": "" } }
      ]
    }
  },
  "track_total_hits": true
}

A second query was used must_not with the exists clause to count documents where the field did not exist. Splitting the two populations by data_stream.dataset showed which integrations created each one. This also explained why dashboard bucket totals could exceed the number of stored documents. A terms aggregation counts each value in a multi-valued field. One document with an empty string and a cluster name contributes to both terms.

I repaired the field in a custom ingest pipeline​


The fix needed two changes. The ingest path first had to remove empty cluster metadata without damaging valid values. Fleet then had to add the configured cluster name to the System streams as well as the Kubernetes streams.

Elastic provides a global@custom hook for custom data processing. Its data-stream documentation says that default integration pipelines call this hook and that custom pipelines persist across version upgrades. Elastic also warns users to test custom processing carefully.

The pipeline normalized orchestrator.cluster.name and orchestrator.cluster.id in four cases:

  1. It left a valid scalar unchanged.
  2. It removed empty members from a list.
  3. It collapsed a one-item list to a scalar.
  4. It removed the key when no valid value remained.

The script below is a generalized version of that pattern for Kibana Dev Tools:

Code:
PUT _ingest/pipeline/global@custom
{
  "processors": [
    {
      "script": {
        "lang": "painless",
        "source": """
          if (ctx.orchestrator == null || ctx.orchestrator.cluster == null) {
            return;
          }

          def cluster = ctx.orchestrator.cluster;
          for (def key : ['name', 'id']) {
            if (!cluster.containsKey(key)) {
              continue;
            }

            def original = cluster[key];
            def cleaned = new ArrayList();

            if (original instanceof List) {
              for (def item : original) {
                if (item != null && item.toString().length() > 0) {
                  cleaned.add(item);
                }
              }
            } else if (original != null && original.toString().length() > 0) {
              cleaned.add(original);
            }

            if (cleaned.isEmpty()) {
              cluster.remove(key);
            } else if (cleaned.size() == 1) {
              cluster[key] = cleaned.get(0);
            } else {
              cluster[key] = cleaned;
            }
          }
        """
      }
    }
  ]
}

The simulate API can test the behaviour before the pipeline processes live data:

Code:
POST _ingest/pipeline/global@custom/_simulate
{
  "docs": [
    {
      "_source": {
        "orchestrator": {
          "cluster": {
            "name": ["", "example-cluster"],
            "id": ""
          }
        }
      }
    },
    {
      "_source": {
        "message": "document without cluster metadata"
      }
    }
  ]
}

The first simulated document kept example-cluster as a scalar and lost the empty ID. The second document passed through unchanged. The ingest pipeline corrected new events. Existing indices kept their previous field shape unless they were reprocessed through a pipeline.

I activated the cleanup pipeline, then added the existing cluster-name processor to the System streams. Applying the cleanup first matters because extending add_fields alone can turn a bare empty string into the same two-value array seen on the Kubernetes streams.

Repairs I rejected​


I evaluated the following changes against the stored documents, the rendered policy, or the identity model.


Candidate

What the check showed

Drop the field in each Fleet stream

The resolved agent policy contained the drop processor, but indexed documents still gained the empty value later in the agent path

Add global data tags in Fleet

Another field-writing processor would leave both the configured value and the empty value in the document

Disable Azure in BEATS_ADD_CLOUD_METADATA_PROVIDERS

This removes useful cloud.* enrichment along with the defective cluster fields

Grant managed-cluster read to the node identity

This broadens node-identity permissions and does not resolve the ambiguous-identity failure on a node pool with several identities

Add another identity to the working token path

This copies the authentication failure, so the processor aborts and the field disappears

The node identity did not need broader access for this repair. The configured add_fields processor already knew the cluster name. The ingest pipeline only removed the invalid value that the cloud metadata path added.

The checks that now catch this failure​


elastic-agent inspect provide the resolved Fleet policy. I use it to compare the processors for each stream. An empty string and an absent field need separate queries. A dashboard can render both as blank, even though they describe different execution paths. The node pool's managed identities belong in an observability investigation when the agent uses Azure's default credential chain. A Kubernetes manifest can stay unchanged while an add-on changes the identities attached to the underlying virtual machine scale set.

Agent startup logs can roll away before anyone investigates. The on-disk agent logs preserved the cloud metadata error after the current console output no longer showed it. The surprising part of this incident was simple once the two paths were visible. A successful Azure lookup with no matching cluster made Elastic Agent write an empty name. The identity failure stopped the lookup earlier, so the configured name remained untouched.
 

Thread statistics

Created
Ahmed Tariq,
Replies
0
Views
2
Back
Top