N
Nurullah Acar
Guest
The quickest way to stall a network incident is to begin with a verdict: “The provider is dropping packets.” The second quickest is to paste one traceroute into a ticket and call it proof.
A Linux virtual dedicated server sits inside several systems you cannot fully observe: the guest network stack, a virtual NIC, a hypervisor, a host interface, an edge router, and one or more upstream networks. An intermittent pause can originate in any of them. It can also originate in the application while the network is behaving correctly.
The goal of a good investigation is not to collect the most command output. It is to produce a small, time-aligned evidence bundle that separates four possibilities:
This workflow applies equally to a general-purpose VPS, a conventional VDS, or a Ryzen VDS. CPU branding does not change the method. Faster cores may help an application, but they do not rule out vCPU scheduling contention, queueing, or an asymmetric network path.
Before generating traffic, write down what failed: source network, destination IP, protocol, port, UTC start and end times, and user-visible effect. “The server was slow” is not reproducible. “HTTPS time to first byte exceeded two seconds for clients on network X from 18:40 to 18:47 UTC” is.
Capture the starting state in a private directory:
Also record deployments, firewall changes, backups, traffic spikes, and kernel updates. Intermittent incidents are correlation problems; without a common clock, every later result is weaker.
Step 1: Ask
Start with a low-impact report from the VDS to a destination you control:
Then test the protocol users actually depend on. For HTTPS:
Do not accuse an intermediate router because its
HackerNoon’s traceroute troubleshooting explainer provides a useful account of TTL expiry and missing intermediate replies. The operational lesson is simple: a traceroute describes replies to probes, not a complete map of forwarding behavior.
One direction is not enough. Ask a remote probe to run the reverse test, and repeat from another access network if complaints are regional. Asymmetric routing means the return path may have little in common with the forward path.
Step 2: Let
During the incident, inspect socket state:
A persistent, large
Compare affected and healthy connections. One peer backing up can be a remote-host or path-specific issue. Many unrelated peers accumulating queues at the same time points toward a common resource: the guest, host, or upstream.
Check CPU pressure alongside the sockets:
High steal time means the guest wanted CPU but the hypervisor did not schedule it. That can delay packet processing and application work without producing a clean “network error” counter.
Find the active interface rather than assuming it is
Repeat the first two counter commands after the symptom or controlled test. Deltas matter; lifetime totals do not establish when an error occurred. The Linux kernel’s interface-statistics guide also warns that counter meanings depend on the device and driver.
Inside a VDS,
Step 4: Ask
Inspect queueing disciplines and their counters:
Run the same commands during degradation. Increasing
Traffic control is also powerful enough to disconnect a remote machine. Inspect freely, but do not replace a production qdisc without console access and a rollback plan. For a controlled introduction to
Step 5: Use
Run
The official
If one TCP stream is limited by round-trip time or CPU, compare a modest four-stream run:
This is an experiment, not a universal score. A public test server introduces unknown CPU, transit, and competing users. Watch production latency and guest CPU while testing, and stop if the test harms service.
Packet loss itself needs precise language. RFC 2680 formalizes a one-way loss metric, which is a reminder that “loss” has a source, destination, observation interval, and threshold. Round-trip tools combine two directions and cannot tell you which direction failed.
A provider-ready escalation should fit on one screen before the attachments. Include:
If the issue varies by geography, independent vantage points are stronger than repeated tests from one server. RIPE Atlas provides distributed ping and traceroute measurements; its results still require cautious interpretation, but they help separate one access network from a broader reachability event. Write observations, not accusations. The following is an illustrative escalation example, not a measured incident: “From 18:42 to 18:47 UTC, three probes showed 8–11% destination loss, beginning after hop X and persisting to the endpoint” is actionable. “Carrier X is broken” is not.
Finally, sanitize public artifacts. Remove customer addresses, credentials, private topology, and full process command lines. Retain unredacted originals privately with checksums. The result is not merely a support ticket: it is a reproducible incident record that another operator can test against host, switch, and flow telemetry unavailable inside your guest.
A Linux virtual dedicated server sits inside several systems you cannot fully observe: the guest network stack, a virtual NIC, a hypervisor, a host interface, an edge router, and one or more upstream networks. An intermittent pause can originate in any of them. It can also originate in the application while the network is behaving correctly.
The goal of a good investigation is not to collect the most command output. It is to produce a small, time-aligned evidence bundle that separates four possibilities:
- an application or socket queue problem;
- a guest or virtual-interface problem;
- a local shaper or capacity limit;
- a path problem beyond the VDS.
This workflow applies equally to a general-purpose VPS, a conventional VDS, or a Ryzen VDS. CPU branding does not change the method. Faster cores may help an application, but they do not rule out vCPU scheduling contention, queueing, or an asymmetric network path.
Start With an Incident Definition, Not a Speed Test
Before generating traffic, write down what failed: source network, destination IP, protocol, port, UTC start and end times, and user-visible effect. “The server was slow” is not reproducible. “HTTPS time to first byte exceeded two seconds for clients on network X from 18:40 to 18:47 UTC” is.
Capture the starting state in a private directory:
Code:
incident="net-$(date -u +%Y%m%dT%H%M%SZ)"
mkdir -m 700 "$incident"
date -u --iso-8601=seconds | tee "$incident/start.txt"
ip -br address | tee "$incident/ip-address.txt"
ip route show table all | tee "$incident/routes.txt"
uname -a | tee "$incident/kernel.txt"
Also record deployments, firewall changes, backups, traffic spikes, and kernel updates. Intermittent incidents are correlation problems; without a common clock, every later result is weaker.
Step 1: Ask mtr Whether Loss Reaches the Destination
Start with a low-impact report from the VDS to a destination you control:
Code:
mtr --report --report-cycles 100 --interval 0.2 \
--show-ips probe.example.net \
| tee "$incident/mtr-icmp.txt"
Then test the protocol users actually depend on. For HTTPS:
Code:
mtr --tcp --port 443 --report --report-cycles 100 \
--show-ips app.example.net \
| tee "$incident/mtr-tcp-443.txt"
Do not accuse an intermediate router because its
Loss% column is high. Routers often give control-plane replies lower priority than forwarded traffic. If hop 6 reports loss while later hops and the destination remain clean, the pattern is consistent with missing or rate-limited replies, not evidence of equivalent end-to-end forwarding loss. Different TTLs use different probes; they do not trace the fate of the same packet. Degradation that also appears at the destination warrants further investigation, but traceroute alone cannot identify the faulty link because probes and replies may follow different paths. APNIC's guide to interpreting traceroute and MTR explains this distinction in operator-focused detail.HackerNoon’s traceroute troubleshooting explainer provides a useful account of TTL expiry and missing intermediate replies. The operational lesson is simple: a traceroute describes replies to probes, not a complete map of forwarding behavior.
One direction is not enough. Ask a remote probe to run the reverse test, and repeat from another access network if complaints are regional. Asymmetric routing means the return path may have little in common with the forward path.
Step 2: Let ss Separate the Application From the Path
During the incident, inspect socket state:
Code:
ss -s | tee "$incident/ss-summary.txt"
ss -tinp state established | tee "$incident/ss-established.txt"
ss -lntup | tee "$incident/ss-listeners.txt"
A persistent, large
Recv-Q suggests the application is not consuming data promptly. A growing Send-Q, rising retransmission count, and expanding retransmission timeout suggest data is not being acknowledged quickly enough. Neither observation identifies the failing physical link, but it narrows the layer.Compare affected and healthy connections. One peer backing up can be a remote-host or path-specific issue. Many unrelated peers accumulating queues at the same time points toward a common resource: the guest, host, or upstream.
Check CPU pressure alongside the sockets:
Code:
vmstat 1 10 | tee "$incident/vmstat.txt"
High steal time means the guest wanted CPU but the hypervisor did not schedule it. That can delay packet processing and application work without producing a clean “network error” counter.
Step 3: Measure Interface Counter Deltas
Find the active interface rather than assuming it is
eth0:
Code:
iface=$(ip route show default | awk 'NR==1 {print $5}')
printf 'interface=%s\n' "$iface" | tee "$incident/interface.txt"
ip -s -s link show dev "$iface" | tee "$incident/ip-link-before.txt"
ethtool -S "$iface" 2>&1 | tee "$incident/ethtool-before.txt"
ethtool -k "$iface" 2>&1 | tee "$incident/offloads.txt"
Repeat the first two counter commands after the symptom or controlled test. Deltas matter; lifetime totals do not establish when an error occurred. The Linux kernel’s interface-statistics guide also warns that counter meanings depend on the device and driver.
Inside a VDS,
ethtool usually sees a paravirtualized device, not the physical port. Guest counters can reveal local drops or queue behavior, but they cannot exclude host-NIC, switch, or provider-edge faults. Record TSO, GSO, and GRO settings because offloads can make packet captures look counterintuitive. Do not disable them in production as a speculative fix; change one setting during a maintenance window and compare before and after.Step 4: Ask tc Whether the Guest Is Shaping Itself
Inspect queueing disciplines and their counters:
Code:
tc -s qdisc show dev "$iface" | tee "$incident/tc-qdisc-before.txt"
tc -s class show dev "$iface" | tee "$incident/tc-class-before.txt"
Run the same commands during degradation. Increasing
dropped values or backlog can identify a saturated local queue. overlimits is not automatically a fault: it is expected when an intentional shaper enforces its rate. Compare the configured policy, contracted capacity, and incident timestamp.Traffic control is also powerful enough to disconnect a remote machine. Inspect freely, but do not replace a production qdisc without console access and a rollback plan. For a controlled introduction to
tc-based impairment, HackerNoon has a practical netem overview; impairment belongs in a lab or scheduled test, not an active incident.Step 5: Use iperf3 as a Directional Experiment
Run
iperf3 only against an endpoint you control, with port 5201 restricted at the firewall. On the probe:
Code:
iperf3 --server --one-off
# The --one-off server exits after one client session.
# Restart this server before EACH client command below, including --parallel 4.
# Run client commands separately on the VDS, after the probe is listening.
Code:
iperf3 --client probe.example.net --time 20 --omit 3 \
--json > "$incident/iperf3-outbound.json"
iperf3 --client probe.example.net --reverse --time 20 --omit 3 \
--json > "$incident/iperf3-inbound.json"
iperf3 --client probe.example.net --udp --bitrate 20M --time 20 \
--json > "$incident/iperf3-udp-20m.json"
The official
iperf3 documentation defines reverse mode, JSON output, and UDP target bitrate. It is worth reading because defaults change the meaning of a result. A UDP test should begin well below known capacity and increase deliberately; flooding the path proves only that queues can fill.If one TCP stream is limited by round-trip time or CPU, compare a modest four-stream run:
Code:
iperf3 --client probe.example.net --parallel 4 --time 20 --omit 3 \
--json > "$incident/iperf3-parallel-4.json"
This is an experiment, not a universal score. A public test server introduces unknown CPU, transit, and competing users. Watch production latency and guest CPU while testing, and stop if the test harms service.
A Decision Tree That Prevents Premature Escalation
Code:
Does loss or latency reach the destination?
├── No
│ ├── Only one intermediate MTR hop looks bad
│ │ └── Do not infer forwarding loss from this hop alone
│ └── Only one client network is affected
│ └── Collect reverse and multi-vantage path evidence
└── Yes
├── Recv-Q grows and application CPU is high
│ └── Investigate application consumption
├── Send-Q/retransmissions rise across unrelated peers
│ └── Inspect interface and queue counters
├── Guest drops/errors increase
│ └── Investigate virtual NIC, driver, and host path
├── tc drops/backlog increase
│ └── Verify local shaping and service rate
├── iperf3 fails in one direction only
│ └── Investigate asymmetric routing or directional policing
└── Multiple external probes show the same failing segment
└── Escalate with a bounded UTC evidence window
Packet loss itself needs precise language. RFC 2680 formalizes a one-way loss metric, which is a reminder that “loss” has a source, destination, observation interval, and threshold. Round-trip tools combine two directions and cannot tell you which direction failed.
Turn Measurements Into Provider-Path Evidence
A provider-ready escalation should fit on one screen before the attachments. Include:
- UTC incident start and end;
- source and destination IPs, protocol, and port;
- exact user impact;
- forward and reverse
mtr, including TCP on the affected service port; - before/after
ip -s link,ethtool -S, andtc -sdeltas; - relevant
ssoutput with unrelated processes redacted; iperf3JSON, direction, duration, and offered UDP rate;- correlation with CPU steal, backups, deployments, or traffic peaks.
If the issue varies by geography, independent vantage points are stronger than repeated tests from one server. RIPE Atlas provides distributed ping and traceroute measurements; its results still require cautious interpretation, but they help separate one access network from a broader reachability event. Write observations, not accusations. The following is an illustrative escalation example, not a measured incident: “From 18:42 to 18:47 UTC, three probes showed 8–11% destination loss, beginning after hop X and persisting to the endpoint” is actionable. “Carrier X is broken” is not.
Finally, sanitize public artifacts. Remove customer addresses, credentials, private topology, and full process command lines. Retain unredacted originals privately with checksums. The result is not merely a support ticket: it is a reproducible incident record that another operator can test against host, switch, and flow telemetry unavailable inside your guest.