Text power tools
On Linux almost everything is text: settings, logs, lists of users, command output. Learn a handful of small tools that each do one job, snap them together with pipes, and you can answer questions in one line that would take an hour in a spreadsheet.
You will learn
grepto find lines,wcto count themsort,uniq,cutandtrto slice and tallysedto find-and-replace andawkto work with columnsxargs,teeanddiff, and how to build a pipeline step by step
The pipe |: the big idea
The pipe sends the output of one command into the next one as its input. Each tool does one small job, and together they do big ones. Build pipelines one step at a time: run the first command, look at the output, add the next |, look again.
cat names.txt | sort | uniq -c | sort -rn | head -3
# read → group → count → biggest first → top 3
grep: find lines
You met grep in lesson 4. Here are the options you'll use most:
grep Maya names.txt # lines containing Maya grep -i friday party.txt # -i: ignore upper/lower case grep -n Maya names.txt # -n: show line numbers grep -c Maya names.txt # -c: just count matching lines grep -v '^#' site.conf # -v: lines that DON'T match (here: skip comments) grep -r port ~/data # -r: search every file in a folder grep -E '404|500' access.log # -E: extended patterns, | means "or"
The ^ means “start of the line” and $ means “end of the line.” These patterns are called regular expressions. They're a whole skill of their own, and AI chatbots are great at explaining them.
Counting and sorting
wc -l scores.csv # count lines (-w words, -c bytes) sort names.txt # A to Z sort -r names.txt # Z to A sort -n numbers.txt # by number: 9 before 10 (plain sort puts 10 first!) sort names.txt | uniq # remove repeated lines sort names.txt | uniq -c # count how many times each appears sort names.txt | uniq -d # only the ones that repeat
uniq needs sortuniq only removes duplicates that sit next to each other. Always sort first, or you'll still see repeats.
Columns: cut, tr and awk
scores.csv looks like this. Each line is a row and the commas separate the columns (fields):
name,class,score Maya,10A,91 Leo,10B,78
cut -d, -f1 scores.csv # -d = delimiter (separator), -f = field number cut -d, -f1,3 scores.csv # columns 1 and 3 sort -t, -k3 -n scores.csv # sort by column 3, as numbers tr a-z A-Z < names.txt # swap characters: lowercase to UPPERCASE awk -F, '{print $1}' scores.csv # print column 1 awk -F, '$3 > 85 {print $1}' scores.csv # names of everyone who scored over 85 awk '{print $1}' access.log # no -F = split on spaces
cut is simple and fast. awk is a tiny programming language: it can compare, add up and format. In awk, $1 is column 1, $0 is the whole line, and NR is the line number. Use single quotes around awk programs so bash doesn't touch the $ signs.
Both families call it awk, but it's a different program underneath. Rocky ships gawk (GNU awk, with lots of extras) and Ubuntu ships mawk (small and fast). Everyday one-liners like the ones above work on both. If a tutorial uses fancy gawk-only features, Ubuntu users can sudo apt install gawk. Check which one you have with awk --version on Rocky or awk -W version on Ubuntu.
sed: find and replace
sed 's/Friday/Saturday/' party.txt # replace the first match on each line (prints, doesn't save) sed 's/Friday/Saturday/g' party.txt # g = every match on the line sed -i 's/Friday/Saturday/g' party.txt # -i = edit the file in place (saves!) sed -n '2p' scores.csv # print only line 2 sed '/^#/d' site.conf # delete comment lines from the output
-iWithout -i, sed only shows the result, which makes it a free preview. Check that the output looks right, then press ↑ and add -i. Admins use sed -i to change settings on hundreds of servers at once, so getting it right matters.
tee, xargs and diff
grep -v '^#' site.conf | tee clean.conf # show it AND save it (like a T-shaped pipe) ls *.conf | xargs wc -l # turn a list into arguments for another command find ~ -name '*.log' | xargs ls -lh # details for every file find found diff site-old.conf site-new.conf # what changed between two files?
2c2 < port=80 --- > port=8080 5a6 > theme=dark
In diff output, < lines come from the first file and > lines from the second. 2c2 means “line 2 was changed” and 5a6 means “a line was added after line 5.” It's the fastest way to answer “what did someone change in this config?”
Try it: data detective 🕵️
Your home folder has a data folder with class scores and settings, and logs/access.log from a small website. Something weird has been hitting that website…
Quick check
1. uniq names.txt still shows “Maya” three times. Why?
✓ Always sort before uniq.
2. You ran sed 's/80/8080/' site.conf and the output looks right, but the file didn't change. Why?
✓ That's a feature: a free preview before you commit.
3. Which prints only the second column of a file whose columns are separated by colons, like /etc/passwd?
✓ -d: sets the separator. Without it, cut splits on Tab characters. awk -F: '{print $2}' works too.
4. A tutorial's awk command works on Rocky but gives an error on Ubuntu. What's the most likely reason?
✓ Same name, different program. The same thing happens with nc in lesson 2.