A
Ahmed Tariq
Guest
Three .NET services on an AKS cluster each pinned their CPU limit while doing no useful work. They targeted a TLS-only RabbitMQ listener with a plaintext client. Their configuration omitted the AMQP TLS flag, and the older framework package did not wire that flag into the connection factory at all. The application never logged the failure, so the proof came from seven bytes on the wire.
The investigation worked only after I separated CPU behaviour, network reachability, protocol negotiation and application startup into different tests.
Three backend services ran at roughly 100% of their CPU limit. The kernel throttled them for 100.0%, 97.6% and 95.5% of CFS periods against a 500m limit. Two comparable .NET backends on the same chart, node, and framework idled at 1m to 2m, and the kernel throttled them for 9.1% and 7.4% of periods.
The busy three had no inbound requests and nothing in flight. At the same time, the event bus reported a successful connection but created no queue. Every pod reported Ready 1/1.
I first raised the CPU limit from 200m to 500m to test whether the services were underprovisioned. All three climbed straight to the new ceiling, and the kernel kept throttling them for roughly 100% of the time. Later, a separate change raised one service to 1500m, and it stopped at 999m. The change converted around 600m of wasted CPU into around 1500m of wasted CPU without fixing the service.
The 999m figure was the useful one. A single busy thread can consume about one core but cannot use more than one core at a time. At 500m the kernel held it at half a core. At 1500m it took the full core it wanted and went no further. That is the shape of a workload spinning rather than working, and it showed that raising the limit would not solve the problem.
top -H inside the containers found one .NET ThreadPool thread sitting in state R across five consecutive samples, with wchan reading 0. On the service running against the 1500m limit, that single thread ID had consumed 16,845 of the process's 16,848 CPU-seconds, measured at 98.1% of a core over a controlled 45-second window.
wchan at zero means the sampled thread was runnable rather than sleeping in a kernel wait channel. top -H identified it as a .NET ThreadPool worker, which made a dedicated GC thread less likely. These samples still did not identify the managed frame or independently eliminate runtime activity. They established a narrower result: one pool thread accounted for almost all process CPU and remained runnable across repeated samples.
Some workloads on the cluster called the framework's
All three .NET backends that called AddEventBus sat at their ceiling. The two .NET backends that did not call it idled at 1m to 2m. The split held across build versions, so one bad image did not explain it.
Two other workloads on the cluster idled as well: an nginx SPA and a Node app. I left them out of the comparison because neither can call a .NET Framework method.
One case looked like a counterexample and became the strongest confirmation. Two services shared a repository, and only one burned CPU. Only that service called AddEventBus, from a module-specific dependency configurator that the other never loaded.
I decoded /proc/1/net/tcp and /proc/1/net/tcp6 from inside the pods. On the subscriber service, that meant 166 samples across two windows, including while driving traffic at it. Across the other two services, four additional samples over several hours also showed no broker socket, including after about twenty requests to one service. None of them held a socket to 5671, 5672 or 15672 in any state, including
PostgreSQL connections from the same pods reached
The framework exposes an endpoint listing the event definitions it knows about. On the subscriber service, it returned 23 events, every one of them Publish, and zero Subscribe. That service registers 21 subscriptions in code.
The ordering inside startup is what makes that decisive. In simplified pseudocode:
The connection step runs before the service registers any subscription. An empty subscription table therefore proves execution never got past that step. This tied the missing queues to the same startup path as the CPU burn.
The application logged nothing useful. One service produced 291 log lines with zero mentions of the broker. Another produced two lines in total. Neither carried a connection error, a retry warning or a stack trace.
So I stopped reading logs and probed the port directly. I sent a plaintext AMQP 0-9-1 protocol header to port 5671 and read what came back.
Seven bytes, and they decode cleanly:
One caveat on the second row.
That is enough. The listener on 5671 speaks TLS, it received something that was not a ClientHello, and it closed the connection with a fatal alert. An
The probe took seconds and produced a direct answer without changing or redeploying the application.
The framework's connection-string options carried two separate boolean properties. The following is simplified pseudocode rather than the proprietary implementation:
All affected connection strings set UseHttps=true and omitted UseSsl, so UseSsl remained false.
Setting UseSsl=true would not have fixed that package version. Outside the connection-string builder, two runtime consumers used UseSsl only to choose the management API scheme. The builder serialised both flags. UseHttps did not control a runtime path in that version.
The AMQP
It set the host, port, virtual host and recovery behaviour, but never set Ssl. Meanwhile, the connection string specified Port=5671, which the managed broker exposes as TLS only. The client therefore sent a plaintext AMQP header into a TLS-only listener. That observed combination could not complete an AMQP connection.
The framework registered an async lambda through Action<IServiceProvider>. Because the delegate returns void, the lambda compiled as async void. The startup loop received no Task to await or inspect, so application startup completed without waiting for the event bus.
The pod could therefore report Ready before the dependency connected. This explains the startup observability gap, but it does not identify the managed frame that consumed the CPU.
The connect path wraps the attempt in a Polly policy:
All affected services set RetryLimit to 3. That permits one initial attempt and three retries, with waits of 2, 4 and 8 seconds. The waits total fourteen seconds, excluding the time spent inside each connection attempt. This policy cannot by itself explain a thread that remained pinned the next day.
The policy should also have produced logs. Its OnRetry callback writes a warning after every failed attempt and runs inside the policy, below the async void boundary. The boundary cannot suppress it. Three warnings per connection cycle should have appeared, yet the logs contained no broker entries.
The method has a try and a finally that releases a semaphore, but no catch. When the retries are exhausted, the exception escapes through the async void boundary. An unobserved async void exception normally stops the process. These pods never restarted.
Three observations did not reconcile: the retry budget was too short for a persistent spin, warning logs should have existed, and an escaping exception should have stopped the process. Without a managed stack trace or dump from the active fault, I could not identify the exact hot frame.
The wire probe proves that the connection could never succeed. The CPU burn tracked AddEventBus across three affected services and two controls. The async void binding explains why startup could not observe the outcome. It does not identify the CPU-burning frame or explain the missing retry warnings.
The framework's AMQP connection factory was updated to apply the existing TLS flag. In simplified pseudocode:
The change also limited connection establishment with a linked CancellationTokenSource and a 15-second CancelAfter. It replaced an empty catch {} around the infrastructure-declaration step with logged warnings, then separated the two overloaded flags so UseSsl controls AMQP TLS, and UseHttps controls the management API scheme.
The second part was configuration. I added UseSsl=true to seven connection-string segments across four pipeline variable groups. An older service ignored UseSsl=true because its package never passed the flag to ConnectionFactory.Ssl. After the framework update, the same configuration enabled AMQP TLS.
The subscriber went from 501m CPU with zero sockets to 3m CPU with an established AMQP socket on 5671, and it declared 26 queues and 2 exchanges at boot. Another service confirmed it later on a subsequent framework version, moving from zero sockets at 500m to two sockets at 11m. The throttling and the missing queues cleared together.
In this incident, saturation up to one core suggested one hot thread. Raising the limit exposed the shape, and top -H confirmed that one ThreadPool worker accounted for almost all process CPU.
A Ready pod proves nothing about the dependency underneath it. This framework opens its broker connection eagerly during ApplicationStarted, but its readiness probe did not check that connection. The failure produced no useful log and did not restart the pod.
A TCP connection confirms only that a listener accepted the connection. It does not prove that AMQP negotiation succeeded. The plaintext AMQP header gave the stronger test because it forced the listener to answer at the TLS layer.
The obvious TCP-only check opens a socket and stops there:
A TCP-only check would have passed during the failure. The TLS listener accepted the connection and rejected the plaintext AMQP bytes afterwards. A useful dependency check must speak enough of the real protocol to distinguish an open port from a usable service.
A publisher with no subscriptions declares no queues even when it is healthy, so an empty broker view proves nothing about it. One of the three affected services had 18 publish calls and zero subscriptions. For that service, the valid evidence was CPU usage, 500m against 1m to 2m for its neighbours.
When the logs say nothing, probe the port from inside the pod before changing the application.
The change ships with
An alternative derives TLS from the conventional port:
The client would enable TLS automatically when configured for a TLS-only port, which would prevent this exact configuration mistake.
The rule depends on a port convention. It forces TLS whenever an endpoint uses 5671, even if a non-standard broker exposes plaintext there, and it misses TLS listeners on every other port. The standard convention makes the first case unusual, but the inference still encodes two assumptions that an explicit flag avoids. The implementation kept the explicit flag.
Moving to plaintext AMQP on 5672 would have restored connectivity immediately while removing transport encryption. Plaintext AMQP exposes the broker username and password on every network hop between the cluster and the externally hosted broker.
Keeping UseSsl=true in configuration during the package rollout avoided a second failure. The older package ignored it, and the new package started using it as soon as the deployment completed.
The investigation worked only after I separated CPU behaviour, network reachability, protocol negotiation and application startup into different tests.
What the cluster looked like
Three backend services ran at roughly 100% of their CPU limit. The kernel throttled them for 100.0%, 97.6% and 95.5% of CFS periods against a 500m limit. Two comparable .NET backends on the same chart, node, and framework idled at 1m to 2m, and the kernel throttled them for 9.1% and 7.4% of periods.
The busy three had no inbound requests and nothing in flight. At the same time, the event bus reported a successful connection but created no queue. Every pod reported Ready 1/1.
Raising the limit converted wasted CPU into more wasted CPU
I first raised the CPU limit from 200m to 500m to test whether the services were underprovisioned. All three climbed straight to the new ceiling, and the kernel kept throttling them for roughly 100% of the time. Later, a separate change raised one service to 1500m, and it stopped at 999m. The change converted around 600m of wasted CPU into around 1500m of wasted CPU without fixing the service.
The 999m figure was the useful one. A single busy thread can consume about one core but cannot use more than one core at a time. At 500m the kernel held it at half a core. At 1500m it took the full core it wanted and went no further. That is the shape of a workload spinning rather than working, and it showed that raising the limit would not solve the problem.
One thread, in state R, with wchan zero
top -H inside the containers found one .NET ThreadPool thread sitting in state R across five consecutive samples, with wchan reading 0. On the service running against the 1500m limit, that single thread ID had consumed 16,845 of the process's 16,848 CPU-seconds, measured at 98.1% of a core over a controlled 45-second window.
wchan at zero means the sampled thread was runnable rather than sleeping in a kernel wait channel. top -H identified it as a .NET ThreadPool worker, which made a dedicated GC thread less likely. These samples still did not identify the managed frame or independently eliminate runtime activity. They established a narrower result: one pool thread accounted for almost all process CPU and remained runnable across repeated samples.
Only the services that called AddEventBus burned CPU
Some workloads on the cluster called the framework's
AddEventBus and some did not, which gave me a controlled comparison.
All three .NET backends that called AddEventBus sat at their ceiling. The two .NET backends that did not call it idled at 1m to 2m. The split held across build versions, so one bad image did not explain it.
Two other workloads on the cluster idled as well: an nginx SPA and a Node app. I left them out of the comparison because neither can call a .NET Framework method.
One case looked like a counterexample and became the strongest confirmation. Two services shared a repository, and only one burned CPU. Only that service called AddEventBus, from a module-specific dependency configurator that the other never loaded.
No sampled socket to the broker, in any state
I decoded /proc/1/net/tcp and /proc/1/net/tcp6 from inside the pods. On the subscriber service, that meant 166 samples across two windows, including while driving traffic at it. Across the other two services, four additional samples over several hours also showed no broker socket, including after about twenty requests to one service. None of them held a socket to 5671, 5672 or 15672 in any state, including
SYN_SENT.SYN_SENT matters. Its absence means the client was not sitting in a half-open connection waiting on a broker that never answered. Either the application was not reaching the connect call at all, or it was connecting and tearing down faster than the sampling could catch.PostgreSQL connections from the same pods reached
ESTABLISHED and stayed there. That rules out a general egress, DNS or network policy problem, though it says nothing about the broker specifically. Proving the broker itself was reachable took a separate test, below.The application's own API showed startup never got past the connection
The framework exposes an endpoint listing the event definitions it knows about. On the subscriber service, it returned 23 events, every one of them Publish, and zero Subscribe. That service registers 21 subscriptions in code.
The ordering inside startup is what makes that decisive. In simplified pseudocode:
Code:
public async Task StartBusAsync(CancellationToken cancellationToken)
{
StartBackgroundMaintenance();
await ConnectToBrokerAsync(cancellationToken); // connection first
RegisterSubscriptions(); // only runs after the connection step returns
}
The connection step runs before the service registers any subscription. An empty subscription table therefore proves execution never got past that step. This tied the missing queues to the same startup path as the CPU burn.
Seven bytes from the broker
The application logged nothing useful. One service produced 291 log lines with zero mentions of the broker. Another produced two lines in total. Neither carried a connection error, a retry warning or a stack trace.
So I stopped reading logs and probed the port directly. I sent a plaintext AMQP 0-9-1 protocol header to port 5671 and read what came back.
Code:
15 03 03 00 02 02 0a
Seven bytes, and they decode cleanly:
| Bytes | Meaning |
|---|---|
15 | TLS record content type 21, alert |
03 03 | the record layer's legacy_record_version field |
00 02 | payload length 2 |
02 | alert level 2, fatal |
0a | alert description 10, unexpected_message |
One caveat on the second row.
0x0303 in a record header is not proof that the peer negotiated TLS 1.2. RFC 8446 keeps that field as legacy_record_version for middlebox compatibility, and a TLS 1.3 server sends the same value. It tells you the peer is speaking TLS. It does not tell you which version it would have agreed on.That is enough. The listener on 5671 speaks TLS, it received something that was not a ClientHello, and it closed the connection with a fatal alert. An
openssl s_client handshake against the same host and port completed normally, which proves the broker itself was reachable from that pod.
The probe took seconds and produced a direct answer without changing or redeploying the application.
UseSsl never reached the AMQP connection in the older package
The framework's connection-string options carried two separate boolean properties. The following is simplified pseudocode rather than the proprietary implementation:
Code:
public bool UseSsl { get; set; } // AMQP transport
public bool UseHttps { get; set; } // management API
All affected connection strings set UseHttps=true and omitted UseSsl, so UseSsl remained false.
Setting UseSsl=true would not have fixed that package version. Outside the connection-string builder, two runtime consumers used UseSsl only to choose the management API scheme. The builder serialised both flags. UseHttps did not control a runtime path in that version.
The AMQP
ConnectionFactory, the object that actually opens the connection, behaved like this in simplified pseudocode:
Code:
var client = new ConnectionFactory
{
HostName = options.Endpoint,
Port = options.Port,
VirtualHost = options.VirtualHost,
// The older package did not configure client.Ssl.
};
It set the host, port, virtual host and recovery behaviour, but never set Ssl. Meanwhile, the connection string specified Port=5671, which the managed broker exposes as TLS only. The client therefore sent a plaintext AMQP header into a TLS-only listener. That observed combination could not complete an AMQP connection.
What async void explains, and what it does not
The framework registered an async lambda through Action<IServiceProvider>. Because the delegate returns void, the lambda compiled as async void. The startup loop received no Task to await or inspect, so application startup completed without waiting for the event bus.
The pod could therefore report Ready before the dependency connected. This explains the startup observability gap, but it does not identify the managed frame that consumed the CPU.
The connect path wraps the attempt in a Polly policy:
Code:
var policy = Policy
.Handle<Exception>()
.WaitAndRetryAsync(
retryCount: settings.RetryLimit,
sleepDurationProvider: attempt => Backoff(attempt),
onRetry: (error, delay) => logger.LogWarning(error, "Broker connection failed"));
All affected services set RetryLimit to 3. That permits one initial attempt and three retries, with waits of 2, 4 and 8 seconds. The waits total fourteen seconds, excluding the time spent inside each connection attempt. This policy cannot by itself explain a thread that remained pinned the next day.
The policy should also have produced logs. Its OnRetry callback writes a warning after every failed attempt and runs inside the policy, below the async void boundary. The boundary cannot suppress it. Three warnings per connection cycle should have appeared, yet the logs contained no broker entries.
The method has a try and a finally that releases a semaphore, but no catch. When the retries are exhausted, the exception escapes through the async void boundary. An unobserved async void exception normally stops the process. These pods never restarted.
Three observations did not reconcile: the retry budget was too short for a persistent spin, warning logs should have existed, and an escaping exception should have stopped the process. Without a managed stack trace or dump from the active fault, I could not identify the exact hot frame.
The wire probe proves that the connection could never succeed. The CPU burn tracked AddEventBus across three affected services and two controls. The async void binding explains why startup could not observe the outcome. It does not identify the CPU-burning frame or explain the missing retry warnings.
The fix came in two parts
The framework's AMQP connection factory was updated to apply the existing TLS flag. In simplified pseudocode:
Code:
Ssl = new SslOption
{
Enabled = options.UseSsl,
ServerName = options.Endpoint
}
The change also limited connection establishment with a linked CancellationTokenSource and a 15-second CancelAfter. It replaced an empty catch {} around the infrastructure-declaration step with logged warnings, then separated the two overloaded flags so UseSsl controls AMQP TLS, and UseHttps controls the management API scheme.
The second part was configuration. I added UseSsl=true to seven connection-string segments across four pipeline variable groups. An older service ignored UseSsl=true because its package never passed the flag to ConnectionFactory.Ssl. After the framework update, the same configuration enabled AMQP TLS.
The subscriber went from 501m CPU with zero sockets to 3m CPU with an established AMQP socket on 5671, and it declared 26 queues and 2 exchanges at boot. Another service confirmed it later on a subsequent framework version, moving from zero sockets at 500m to two sockets at 11m. The throttling and the missing queues cleared together.
CPU saturation and pod readiness required separate tests
In this incident, saturation up to one core suggested one hot thread. Raising the limit exposed the shape, and top -H confirmed that one ThreadPool worker accounted for almost all process CPU.
A Ready pod proves nothing about the dependency underneath it. This framework opens its broker connection eagerly during ApplicationStarted, but its readiness probe did not check that connection. The failure produced no useful log and did not restart the pod.
Protocol checks that now come first
A TCP connection confirms only that a listener accepted the connection. It does not prove that AMQP negotiation succeeded. The plaintext AMQP header gave the stronger test because it forced the listener to answer at the TLS layer.
The obvious TCP-only check opens a socket and stops there:
Code:
kubectl exec <pod> -- bash -c 'cat < /dev/null > /dev/tcp/<broker-host>/5671'
A TCP-only check would have passed during the failure. The TLS listener accepted the connection and rejected the plaintext AMQP bytes afterwards. A useful dependency check must speak enough of the real protocol to distinguish an open port from a usable service.
A publisher with no subscriptions declares no queues even when it is healthy, so an empty broker view proves nothing about it. One of the three affected services had 18 publish calls and zero subscriptions. For that service, the valid evidence was CPU usage, 500m against 1m to 2m for its neighbours.
When the logs say nothing, probe the port from inside the pod before changing the application.
Why TLS remained an explicit setting
The change ships with
UseSsl defaulting to false, so every connection string has to gain UseSsl=true or the symptom comes straight back on a new service.An alternative derives TLS from the conventional port:
Code:
Enabled = EventBusOptions.UseSsl || EventBusOptions.Port == 5671
The client would enable TLS automatically when configured for a TLS-only port, which would prevent this exact configuration mistake.
The rule depends on a port convention. It forces TLS whenever an endpoint uses 5671, even if a non-standard broker exposes plaintext there, and it misses TLS listeners on every other port. The standard convention makes the first case unusual, but the inference still encodes two assumptions that an explicit flag avoids. The implementation kept the explicit flag.
Moving to plaintext AMQP on 5672 would have restored connectivity immediately while removing transport encryption. Plaintext AMQP exposes the broker username and password on every network hop between the cluster and the externally hosted broker.
Keeping UseSsl=true in configuration during the package rollout avoided a second failure. The older package ignored it, and the new package started using it as soon as the deployment completed.
References
- RabbitMQ TLS support, which documents 5671 as the TLS port
- RFC 8446 section 5.1 on
legacy_record_version - Async return types in C# on what
async voidgives up - RabbitMQ .NET client guide for connection and TLS client behaviour