Linux Sysadmin · Lesson 3 · 25 min

Text power tools

On Linux almost everything is text: settings, logs, lists of users, command output. Learn a handful of small tools that each do one job, snap them together with pipes, and you can answer questions in one line that would take an hour in a spreadsheet.

You will learn

  • grep to find lines, wc to count them
  • sort, uniq, cut and tr to slice and tally
  • sed to find-and-replace and awk to work with columns
  • xargs, tee and diff, and how to build a pipeline step by step

The pipe |: the big idea

The pipe sends the output of one command into the next one as its input. Each tool does one small job, and together they do big ones. Build pipelines one step at a time: run the first command, look at the output, add the next |, look again.

cat names.txt | sort | uniq -c | sort -rn | head -3
#  read      →  group  →  count  →  biggest first → top 3

grep: find lines

You met grep in lesson 4. Here are the options you'll use most:

Same on both
grep Maya names.txt            # lines containing Maya
grep -i friday party.txt       # -i: ignore upper/lower case
grep -n Maya names.txt         # -n: show line numbers
grep -c Maya names.txt         # -c: just count matching lines
grep -v '^#' site.conf         # -v: lines that DON'T match (here: skip comments)
grep -r port ~/data            # -r: search every file in a folder
grep -E '404|500' access.log   # -E: extended patterns, | means "or"

The ^ means “start of the line” and $ means “end of the line.” These patterns are called regular expressions. They're a whole skill of their own, and AI chatbots are great at explaining them.

Counting and sorting

Same on both
wc -l scores.csv            # count lines (-w words, -c bytes)
sort names.txt              # A to Z
sort -r names.txt           # Z to A
sort -n numbers.txt         # by number: 9 before 10 (plain sort puts 10 first!)
sort names.txt | uniq       # remove repeated lines
sort names.txt | uniq -c    # count how many times each appears
sort names.txt | uniq -d    # only the ones that repeat
uniq needs sort

uniq only removes duplicates that sit next to each other. Always sort first, or you'll still see repeats.

Columns: cut, tr and awk

scores.csv looks like this. Each line is a row and the commas separate the columns (fields):

name,class,score
Maya,10A,91
Leo,10B,78
Same on both
cut -d, -f1 scores.csv                   # -d = delimiter (separator), -f = field number
cut -d, -f1,3 scores.csv                 # columns 1 and 3
sort -t, -k3 -n scores.csv               # sort by column 3, as numbers
tr a-z A-Z < names.txt                   # swap characters: lowercase to UPPERCASE
awk -F, '{print $1}' scores.csv          # print column 1
awk -F, '$3 > 85 {print $1}' scores.csv  # names of everyone who scored over 85
awk '{print $1}' access.log              # no -F = split on spaces

cut is simple and fast. awk is a tiny programming language: it can compare, add up and format. In awk, $1 is column 1, $0 is the whole line, and NR is the line number. Use single quotes around awk programs so bash doesn't touch the $ signs.

Spot the difference

Both families call it awk, but it's a different program underneath. Rocky ships gawk (GNU awk, with lots of extras) and Ubuntu ships mawk (small and fast). Everyday one-liners like the ones above work on both. If a tutorial uses fancy gawk-only features, Ubuntu users can sudo apt install gawk. Check which one you have with awk --version on Rocky or awk -W version on Ubuntu.

sed: find and replace

Same on both
sed 's/Friday/Saturday/' party.txt     # replace the first match on each line (prints, doesn't save)
sed 's/Friday/Saturday/g' party.txt    # g = every match on the line
sed -i 's/Friday/Saturday/g' party.txt # -i = edit the file in place (saves!)
sed -n '2p' scores.csv                 # print only line 2
sed '/^#/d' site.conf                  # delete comment lines from the output
Test first, then -i

Without -i, sed only shows the result, which makes it a free preview. Check that the output looks right, then press ↑ and add -i. Admins use sed -i to change settings on hundreds of servers at once, so getting it right matters.

tee, xargs and diff

Same on both
grep -v '^#' site.conf | tee clean.conf   # show it AND save it (like a T-shaped pipe)
ls *.conf | xargs wc -l                    # turn a list into arguments for another command
find ~ -name '*.log' | xargs ls -lh        # details for every file find found
diff site-old.conf site-new.conf           # what changed between two files?
2c2
< port=80
---
> port=8080
5a6
> theme=dark

In diff output, < lines come from the first file and > lines from the second. 2c2 means “line 2 was changed” and 5a6 means “a line was added after line 5.” It's the fastest way to answer “what did someone change in this config?”

Try it: data detective 🕵️

Your home folder has a data folder with class scores and settings, and logs/access.log from a small website. Something weird has been hitting that website…

Quick check

1. uniq names.txt still shows “Maya” three times. Why?

2. You ran sed 's/80/8080/' site.conf and the output looks right, but the file didn't change. Why?

3. Which prints only the second column of a file whose columns are separated by colons, like /etc/passwd?

4. A tutorial's awk command works on Rocky but gives an error on Ubuntu. What's the most likely reason?

Finished the missions and the quiz? Mark it done to track your progress.