Linux Debugging Reference
You should deeply understand: /proc, strace, tcpdump, vmstat, iostat, systemctl, journalctl, top, ss.
Senior engineers think in: resource bottlenecks, kernel states, queues, backpressure, timeouts, failure domains.
Table of Contents
- Practical Debugging Playbook
- Process Management
- Systemd
- CPU & Load Debugging
- Memory Management
- Disk & I/O Debugging
- Networking
- Logs & Observability
- Performance Debugging (USE / RED)
- Bash & Shell Skills
- Security & Permissions
- Containers
- Crash & Deep Debugging
- Boot Process
When production breaks, senior engineers don't guess — they run down this fixed order: is the process running, is the port listening, is CPU saturated, is memory exhausted, is disk full, and are connections stuck, before diving into any subsystem-specific tooling.
Practical Debugging Playbook
When prod breaks, ask, in order:
- Is the process running?
- Is the port listening?
- Is CPU saturated?
- Is memory exhausted?
- Is disk full?
- Are connections stuck?
- Are there kernel errors?
- Is DNS resolving?
- Is the service reachable from another host?
- What changed?
1. Process Management
Covered in depth in Process Management & /proc: PID/PPID, zombie/orphan processes, daemons, ps/top/htop/atop/pstree, nice/renice, kill/killall/pgrep, and the /proc/<pid>/ filesystem (status, limits, fd).
2. Systemd (Most Modern Distros)
You MUST understand:
systemctl statussystemctl start/stop/restartjournalctl -xejournalctl -u <service>- Unit files
- Service dependencies
- Restart policies
- Targets
3. CPU & Load Debugging
🔹 Load Average
What it actually means: runnable + uninterruptible processes.
🔹 Tools
uptime, top, mpstat, vmstat, sar, pidstat
🔹 CPU Concepts
- User vs system CPU
- iowait
- Context switching
- CPU stealing (in VMs)
- CPU throttling
4. Memory Management
🔹 Concepts
Virtual memory, paging, swapping, page cache, buffers, OOM killer
🔹 Tools
free -m, vmstat, top, htop, /proc/meminfo, smem, pmap
🔹 Debugging Memory Issues
- Detect memory leaks
- OOM logs (
dmesg) - Overcommit behavior
5. Disk & I/O Debugging
🔹 Concepts
IOPS, throughput, latency, queue depth, block devices
🔹 Tools
iostat, iotop, vmstat, dstat, blktrace, lsblk
🔹 Detect
- Disk saturation
- High iowait
- Filesystem corruption
Full storage & filesystem deep dive: Filesystem & Storage Playbook
6. Networking (Extremely Important for SRE)
🔹 Basics
TCP/IP model, 3-way handshake, DNS resolution, subnetting, routing tables
🔹 Tools
ip a, ip route, ss -tulpn, netstat, tcpdump, ping, traceroute, dig, nslookup, curl, nc
🔹 Debug Skills
- Port binding conflicts
- SYN backlog issues
- Connection resets
- Packet drops
- MTU issues
7. Logs & Observability
🔹 Log Locations
/var/log/syslog/var/log/messages/var/log/auth.log- Application logs
🔹 Commands
tail -f, less, grep, awk, sed, journalctl
🔹 dmesg
Kernel logs: OOM killer, disk failures, driver errors.
8. Performance Debugging (The SRE Superpower)
The USE Method
- Utilization
- Saturation
- Errors
The RED Method
- Rate
- Errors
- Duration
🔹 Advanced Tools
strace, lsof, perf, sar, bpftrace, eBPF basics
9. Bash & Shell Skills
🔹 Must Know
Pipes, redirection, subshells, environment variables, command substitution, exit codes, set -euo pipefail
🔹 Scripting Basics
Loops, conditionals, functions, debugging scripts (set -x)
10. Security & Permissions
🔹 SSH
Key-based auth, sshd_config, port forwarding
🔹 Firewalls
iptables, nftables, ufw, firewalld
🔹 SELinux / AppArmor
Modes, troubleshooting denials
Permissions model (rwx, SUID/SGID, ACLs) is covered in Filesystem & Storage Playbook.
11. Containers (Mandatory for Modern SRE)
- Namespaces
- cgroups
- PID namespace
- Network namespace
- OverlayFS
docker ps,docker logs,docker inspect,docker exec- Container resource limits
12. Crash & Deep Debugging
- Core dumps
ulimit -cgdb- Kernel panic basics
kdump
13. Boot Process
- BIOS vs UEFI
- GRUB
- Init systems
- Single-user mode
- Emergency mode
Summary
- 💡 The 10-question Practical Debugging Playbook is the fastest triage sequence — run it top to bottom before diving deep anywhere.
- 🔥 USE (resource-first) and RED (service-first) are complementary, not competing — USE finds what's saturated, RED finds what users feel.
- ⚠️ Containers add a second layer of saturation (cgroup limits) that's invisible from host-level tools — always check
docker stats/ cgroup files separately. - ✅
/proc,strace,tcpdump,vmstat,iostat,systemctl,journalctl,top,ss— if you're fluent in these nine, you can debug almost anything.