Linux du Command Cheat Sheet: Find Disk Space Hogs and Suspicious Data Staging Fast
It's 2:47 a.m. and your monitoring platform fires a "disk 98% full" alert on a production Linux web server. Nobody changed anything. No deployment went out. Is it a runaway log, a forgotten backup, or something worse, like an intruder staging a compressed archive of stolen data in a temp directory?
In that moment you don't need a full SIEM query. You need one fast, reliable command. That command is du. For SOC analysts, sysadmins, and anyone doing Linux incident response, knowing how to read disk usage quickly is a core triage skill. This guide covers the du command the way it gets used in real investigations, including a few corrections to common cheat-sheet advice.
Table of Contents
- Why du Matters in Security Operations
- Basic Usage and Human-Readable Output
- Finding the Largest Directories and Files
- Investigating /var, Logs, and Docker Storage
- Exclusions, Filesystem Boundaries, and Symlinks
- Apparent Size vs. Disk Usage
- du vs. df: When the Numbers Don't Match
- Using du for Threat Detection and Log Retention
- Quick Reference Table
- Expert Tips
- FAQ
Why du Matters in Security Operations
Disk usage is a security signal, not just a housekeeping metric. A disk filling up unexpectedly can point to several problems:
- Log flooding: a brute-force attempt or misconfigured service generating gigabytes of authentication or application logs.
- Data staging: large archives appearing in writable locations such as /tmp, /var/tmp, or /dev/shm before exfiltration.
- Availability attacks: filling a volume until services crash or logging silently stops.
- Container sprawl: abandoned images and volumes quietly consuming space under Docker's data directory.
The du command (part of GNU coreutils on most distributions) estimates file space usage by walking the directory tree. It's available on virtually every Linux system, which makes it dependable during incident response when you can't install anything new.
Basic Usage and Human-Readable Output
Run plain du and it lists disk usage for the current directory and every subdirectory, in 1 KiB blocks by default on GNU systems. That output is hard to read at a glance.
du
Add -h for human-readable sizes (K, M, G), and -s for a summary of just the total:
du -h
du -sh
du -sh .
du -sh /path/to/dir
When to use it: du -sh is your first move when you only need one number, such as "how big is this directory?" Expected output is a single line like 4.2G .
To list every file as well as every directory, use -a:
du -ah /path/to/dir
Expect a long listing. Always pipe it into sort and head (shown below) rather than reading it raw.
You can check several locations in one pass:
du -sh /var /home /opt
du -sh /home/user1 /home/user2
Use the second form to compare two user home directories side by side, for example when one account's usage looks abnormal.
For your own home directory:
du -sh ~
One clarification about hidden files: du always includes hidden files and directories (those starting with a dot) when it walks a tree. The -a flag adds individual files to the listing, but it isn't what makes hidden items appear. The trap is the shell glob, because du -sh * skips dotfiles. If you're hunting for something hidden in a home directory, use:
du -sh .[!.]* * 2>/dev/null | sort -hr | head -n 15
This matters because attackers and malware commonly hide data in dot-prefixed directories. The 2>/dev/null suppresses "Permission denied" noise.
Finding the Largest Directories and Files
The most useful du pattern in any investigation is controlling depth, then sorting. The --max-depth option limits how many levels deep du reports.
du -h --max-depth=1 /path/to/dir
du -h --max-depth=2 /path/to/dir
du -h --max-depth=3 /path/to/dir
Start at depth 1 to see which immediate subdirectory is the culprit, then drill into it with a deeper level or a fresh command pointed at that folder. Note that du still adds up everything below the depth limit. It just doesn't print those lines.
Now pipe it into sort -h, which understands human-readable suffixes:
du -h --max-depth=1 /path/to/dir | sort -h
du -h --max-depth=1 /path/to/dir | sort -hr
The first command sorts smallest to largest, and the second reverses it so the biggest consumers appear first. Because the directory's own total prints last in du output, the largest-first version puts the overall total at the top.
To list the top entries across a whole tree:
du -ah /path/to/dir | sort -hr | head -n 10
The head command alone shows 10 lines by default, so head -n 10 just makes the intent explicit. Change the number for a longer or shorter list.
Correction: Showing Files Only
Many cheat sheets suggest du -ah /path | grep -v '/$' to show files only. In practice this rarely works, because du doesn't normally print a trailing slash on directory names, so the filter removes almost nothing. A more reliable way to find the largest individual files is find:
find /path/to/dir -type f -exec du -h {} + 2>/dev/null | sort -hr | head -n 10
This lists the ten largest regular files and ignores directories entirely.
Investigating /var, Logs, and Docker Storage
On servers, the usual suspects live under /var. Use sudo so du can read directories you don't own, and be aware that without it you'll get an incomplete picture.
sudo du -sh /var
sudo du -sh /var/log
sudo du -h --max-depth=1 /var | sort -hr
The last command is the standard "where did my space go under /var?" triage. Typical output ranks /var/log, /var/lib, and /var/cache at the top.
Before any cleanup, record what you have. For incident response, deleting logs before preserving them can destroy evidence.
du -sh /var/log/*
du -ah /var/log > disk-usage.txt
The first command shows each log file or directory's size. The second saves a full report to a file you can attach to a ticket or hand to a colleague.
Warning: Do not delete, truncate, or rotate logs on a system under investigation until you've confirmed with your incident response lead that the evidence is preserved. Authentication logs, audit logs, and web server logs are often the best record of what happened. Also keep in mind that retention requirements under frameworks like PCI DSS, HIPAA, or internal compliance policies may dictate how long logs must be kept, so check your policy before cleaning anything up.
Other common locations to check:
sudo du -sh /var/cache/*
du -sh /opt/*
sudo du -sh /home/*
sudo du -sh /var/lib/docker
These show package and application caches, third-party application directories, per-user home directories, and Docker's data directory. For Docker specifically, docker system df gives a cleaner breakdown of images, containers, and volumes than du alone.
Exclusions, Filesystem Boundaries, and Symlinks
Sometimes you want to ignore noise. The --exclude option skips entries matching a pattern:
du -h --exclude='node_modules' .
du -h --exclude='*.log' .
du -h --exclude='cache' /var
Quote the pattern so your shell doesn't expand it before du sees it.
The -x option deserves a clarification, because cheat sheets often describe it confusingly. By default, du crosses into mounted filesystems. With -x, du stays on the same filesystem as the starting path. It never crosses boundaries.
du -h -x /
du -h -x /path/to/dir
This is essential when you scan / on a server with NFS mounts, external volumes, or container mounts. Without it, you may waste time measuring other disks, or misattribute usage to the root volume.
For symbolic links, du doesn't follow them by default. To dereference them:
du -h -L /path/to/dir
Use this with care. Following links can double-count data or wander into directories you didn't intend to scan.
Apparent Size vs. Disk Usage
By default, du reports the space actually allocated on disk in filesystem blocks. The --apparent-size option instead reports the logical size of the file contents:
du -h --apparent-size file.iso
du -sh --apparent-size /path/to/dir
du -b file.txt
The two numbers can differ. Sparse files, such as some VM disk images and database files, may show a large apparent size but use far less real disk. Small files can also use more disk than their content size because of block rounding. The -b flag prints apparent size in bytes.
Other unit options include:
du -k /path/to/dir
du -m /path/to/dir
du -a /path/to/dir
These report in 1024-byte blocks, in 1 MiB blocks, and (for the last one) every file and directory in raw block units.
To get a grand total at the end of the listing, use -c:
du -ch /path/to/dir/*
du -csh /var/log /tmp
Remember that /path/to/dir/* skips hidden entries, so the grand total may be lower than the true directory size.
du vs. df: When the Numbers Don't Match
Here's a real-world puzzle that trips up even experienced admins. The disk reports 100% full according to df -h, yet du adds up to far less. Where did the space go?
A common cause is a deleted file that a running process still holds open. The filename is gone, so du can't see it, but the filesystem can't release the blocks until the process closes the file or exits. This frequently happens when someone deletes a large log file without restarting the service writing to it.
df -h
sudo lsof +L1
The lsof +L1 command lists open files with a link count below one, meaning deleted files still in use. Restarting the owning service typically releases the space. Another cause is du not having permission to read some directories, which is why running with sudo matters. From a security perspective, a large deleted-but-open file is also worth a second look, since malicious processes sometimes delete their own working files while continuing to use them.
Using du for Threat Detection and Log Retention
On its own, du is not a detection tool. But as part of Linux incident response it helps you answer the first triage question quickly: is something using space that shouldn't be?
Check Common Staging Locations
Writable directories are favorites for temporary data staging. Check them first:
sudo du -sh /tmp /var/tmp /dev/shm 2>/dev/null
sudo find /tmp /var/tmp /dev/shm -type f -size +100M -mtime -2 -exec ls -lh {} + 2>/dev/null
The second command lists files over 100 MB modified in the last two days. A large, recently created archive in a temp directory is not proof of compromise, but it deserves an explanation. Compare against what the system normally runs (backup jobs, application builds) before drawing conclusions.
Build a Baseline
The most effective approach is knowing what normal looks like. Capture a snapshot on a schedule and compare it over time:
du -h --max-depth=1 /var | sort -hr > /root/var-usage-baseline.txt
Many teams feed disk usage metrics into their log management or SIEM tooling so sudden jumps trigger alerts automatically. Endpoint detection and response (EDR) platforms may also flag unusual archive creation, but a simple baseline catches plenty without extra software.
Watch Usage Repeatedly
To see whether a directory is actively growing:
watch -n 5 'du -sh /path/to/dir'
This refreshes every five seconds. If the number climbs steadily, something is writing to it right now. Pair that with lsof or your audit logs to identify the process.
Prevention
- Configure log rotation and size limits so a single noisy service can't fill a volume.
- Put logs and application data on separate partitions from the root filesystem.
- Set disk usage alerts at sensible thresholds (for example, warning well before 90%).
- Ship logs to a central collector so local cleanup never destroys the only copy.
- Restrict write access to world-writable directories and consider mounting /tmp with the noexec option where it fits your environment.
Quick Reference Table
| Goal | Command |
| Total size of a directory | du -sh /path |
| One level of subdirectories | du -h --max-depth=1 /path |
| Biggest directories first | du -h --max-depth=1 /path | sort -hr |
| Top 10 largest entries | du -ah /path | sort -hr | head -n 10 |
| Stay on one filesystem | du -h -x / |
| Skip a pattern | du -h --exclude='*.log' . |
| Grand total of paths | du -csh /var/log /tmp |
| Logical (apparent) size | du -sh --apparent-size /path |
| Save a report | du -ah /path > disk-usage.txt |
Expert Tips
- Drill down in layers. Run depth 1 on /, find the biggest folder, then repeat inside it. It's faster than scanning everything deeply.
- Always use -x on /. It avoids wasted time on network and external mounts.
- Suppress permission noise. Add
2>/dev/null, but only after you've confirmed you're running with the privileges you need. - Preserve before you purge. If a system may be compromised, copy or snapshot logs first.
- Try ncdu for interactive exploration. If you're allowed to install it, ncdu is a handy interactive disk usage browser, though du remains the tool you can count on being present.
- Use thresholds. GNU du supports
--threshold=100Mto hide anything smaller, which cuts clutter on large trees.
Related Cybersecurity Topics You Should Explore
- Linux df Command Cheat Sheet: Fix Full Disk & Missing Logs
- Linux Disk Commands Cheat Sheet: df, du, mount, fsck (2026)
- Linux umask Cheat Sheet: 022 vs 027 vs 077 Explained
- Linux chgrp Cheat Sheet (2026): Commands, Examples & Audit Tips
- Linux chown Command Cheat Sheet: 40+ Examples (2026)
- Linux chmod Cheat Sheet: Stop Using 777 (Safer Fixes)
- OnePlus 15 Root Exploit: Zero-Permission Apps Can Take Full Control
- One Misconfigured chmod Command Gave Attackers Root Access
- How SOC Analysts Use sed to Catch Attacks Before the SIEM Does
- Linux tr Command Tutorial: Fix Messy SOC Logs in Seconds
FAQ
What does du stand for in Linux?
It stands for "disk usage." The command estimates how much space files and directories consume.
What is the difference between du and df?
The du command adds up the space used by files and directories you point it at. The df command reports free and used space for whole filesystems. They can disagree when deleted files are still held open by processes or when du lacks permission to read some paths.
How do I find the largest folders on my Linux server?
Run du -h --max-depth=1 /path | sort -hr | head. Repeat inside the largest folder to narrow it down.
Why does du show a different size than ls -l?
The ls command shows logical file size, while du shows allocated disk blocks by default. Use --apparent-size to make du report logical size.
Is it safe to run du on /?
Yes, it's read-only, but it can be slow and generate heavy disk activity on large systems. Use -x to stay on one filesystem and run it during quieter periods on production servers.
Can du help detect a security incident?
It can provide a useful early clue, such as unexpected large archives in temp directories or sudden log growth. It should be combined with audit logs, process inspection, and your monitoring tools before drawing conclusions.
How do I exclude certain files from du results?
Use --exclude with a quoted pattern, for example du -h --exclude='*.log' ..
Conclusion
The du command looks simple, but in the hands of a practitioner it's a fast, dependable triage tool. Whether you're clearing a full disk at 2 a.m., checking a suspicious temp directory, or building a baseline for monitoring, a handful of patterns cover most real situations: -sh for totals, --max-depth for scope, sort -hr for ranking, and -x to stay on one filesystem. Learn the corrections too, especially around files-only filtering and hidden files, because small misunderstandings can send an investigation in the wrong direction.
Practice these commands on a test system before you need them under pressure, and pair disk usage checks with solid logging and retention practices.
Analysis based on SOC monitoring experience and public technical documentation review.
