Skip to main content

Linux Debugging Reference

You should deeply understand: /proc, strace, tcpdump, vmstat, iostat, systemctl, journalctl, top, ss.

Senior engineers think in: resource bottlenecks, kernel states, queues, backpressure, timeouts, failure domains.

Table of Contents

  1. Practical Debugging Playbook
  2. Process Management
  3. Systemd
  4. CPU & Load Debugging
  5. Memory Management
  6. Disk & I/O Debugging
  7. Networking
  8. Logs & Observability
  9. Performance Debugging (USE / RED)
  10. Bash & Shell Skills
  11. Security & Permissions
  12. Containers
  13. Crash & Deep Debugging
  14. Boot Process

🧭 Mental modelThe practical debugging playbook, asked in orderBefore anything else, work down a fixed checklist: is the process running, is the port listening, is CPU saturated, is memory exhausted, is disk full, and are connections stuck.Process running?ps / systemctl statusPort listening?ss -lntpCPU saturated?top / vmstatMemory exhausted?free / vmstatDisk full?df / duConnections stuck?ss / netstat

When production breaks, senior engineers don't guess — they run down this fixed order: is the process running, is the port listening, is CPU saturated, is memory exhausted, is disk full, and are connections stuck, before diving into any subsystem-specific tooling.

Practical Debugging Playbook

When prod breaks, ask, in order:

  1. Is the process running?
  2. Is the port listening?
  3. Is CPU saturated?
  4. Is memory exhausted?
  5. Is disk full?
  6. Are connections stuck?
  7. Are there kernel errors?
  8. Is DNS resolving?
  9. Is the service reachable from another host?
  10. What changed?

1. Process Management

Covered in depth in Process Management & /proc: PID/PPID, zombie/orphan processes, daemons, ps/top/htop/atop/pstree, nice/renice, kill/killall/pgrep, and the /proc/<pid>/ filesystem (status, limits, fd).


2. Systemd (Most Modern Distros)

You MUST understand:

  • systemctl status
  • systemctl start/stop/restart
  • journalctl -xe
  • journalctl -u <service>
  • Unit files
  • Service dependencies
  • Restart policies
  • Targets

3. CPU & Load Debugging

🔹 Load Average

What it actually means: runnable + uninterruptible processes.

🔹 Tools

uptime, top, mpstat, vmstat, sar, pidstat

🔹 CPU Concepts

  • User vs system CPU
  • iowait
  • Context switching
  • CPU stealing (in VMs)
  • CPU throttling

4. Memory Management

🔹 Concepts

Virtual memory, paging, swapping, page cache, buffers, OOM killer

🔹 Tools

free -m, vmstat, top, htop, /proc/meminfo, smem, pmap

🔹 Debugging Memory Issues

  • Detect memory leaks
  • OOM logs (dmesg)
  • Overcommit behavior

5. Disk & I/O Debugging

🔹 Concepts

IOPS, throughput, latency, queue depth, block devices

🔹 Tools

iostat, iotop, vmstat, dstat, blktrace, lsblk

🔹 Detect

  • Disk saturation
  • High iowait
  • Filesystem corruption

Full storage & filesystem deep dive: Filesystem & Storage Playbook


6. Networking (Extremely Important for SRE)

🔹 Basics

TCP/IP model, 3-way handshake, DNS resolution, subnetting, routing tables

🔹 Tools

ip a, ip route, ss -tulpn, netstat, tcpdump, ping, traceroute, dig, nslookup, curl, nc

🔹 Debug Skills

  • Port binding conflicts
  • SYN backlog issues
  • Connection resets
  • Packet drops
  • MTU issues

7. Logs & Observability

🔹 Log Locations

  • /var/log/syslog
  • /var/log/messages
  • /var/log/auth.log
  • Application logs

🔹 Commands

tail -f, less, grep, awk, sed, journalctl

🔹 dmesg

Kernel logs: OOM killer, disk failures, driver errors.


8. Performance Debugging (The SRE Superpower)

The USE Method

  • Utilization
  • Saturation
  • Errors

The RED Method

  • Rate
  • Errors
  • Duration

🔹 Advanced Tools

strace, lsof, perf, sar, bpftrace, eBPF basics


9. Bash & Shell Skills

🔹 Must Know

Pipes, redirection, subshells, environment variables, command substitution, exit codes, set -euo pipefail

🔹 Scripting Basics

Loops, conditionals, functions, debugging scripts (set -x)


10. Security & Permissions

🔹 SSH

Key-based auth, sshd_config, port forwarding

🔹 Firewalls

iptables, nftables, ufw, firewalld

🔹 SELinux / AppArmor

Modes, troubleshooting denials

Permissions model (rwx, SUID/SGID, ACLs) is covered in Filesystem & Storage Playbook.


11. Containers (Mandatory for Modern SRE)

  • Namespaces
  • cgroups
  • PID namespace
  • Network namespace
  • OverlayFS
  • docker ps, docker logs, docker inspect, docker exec
  • Container resource limits

12. Crash & Deep Debugging

  • Core dumps
  • ulimit -c
  • gdb
  • Kernel panic basics
  • kdump

13. Boot Process

  • BIOS vs UEFI
  • GRUB
  • Init systems
  • Single-user mode
  • Emergency mode

Summary

  • 💡 The 10-question Practical Debugging Playbook is the fastest triage sequence — run it top to bottom before diving deep anywhere.
  • 🔥 USE (resource-first) and RED (service-first) are complementary, not competing — USE finds what's saturated, RED finds what users feel.
  • ⚠️ Containers add a second layer of saturation (cgroup limits) that's invisible from host-level tools — always check docker stats / cgroup files separately.
  • /proc, strace, tcpdump, vmstat, iostat, systemctl, journalctl, top, ss — if you're fluent in these nine, you can debug almost anything.

See Also