Run your own AI on Linux
Chatbots usually run on someone else's giant computers. But smaller AI models can run on your own Linux machine: no account, no internet once they're downloaded, and nothing leaves your computer. Let's set one up.
You will learn
- What “running a model locally” means, and what hardware you need
- How to install Ollama and chat with a model
- How to call the model from the command line and from a Python script
- The Python setup difference between Rocky and Ubuntu (it trips up almost everyone)
AI models need several gigabytes of memory, so this lesson can't run in the practice terminal. Use a Linux computer, a virtual machine, or WSL on Windows. The practice terminal at the bottom lets you rehearse the Python setup part.
What you need
- 64-bit Linux: Rocky 9, Ubuntu 22.04/24.04, or similar
- Memory (RAM): 8 GB lets you run small models comfortably. 16 GB or more lets you try bigger ones.
- Disk: a few GB free per model
- Graphics card (GPU): optional. An NVIDIA or AMD GPU makes answers much faster, but the CPU works too, just slower.
Check what you have (same on both):
free -h # memory: look at the "total" column df -h ~ # free disk space in your home folder nproc # number of CPU cores
Step 1: install Ollama
Ollama is a free tool that downloads AI models and runs them for you. Its installer is the same on both families:
curl -fsSL https://ollama.com/install.sh | sh
Wait, remember the red flag from the last lesson? curl … | sh runs a script from the internet without anyone reading it. Ollama is a well-known project, but the habit of reading first is what keeps you safe. Here's the careful version:
curl -fsSL https://ollama.com/install.sh -o install-ollama.sh
less install-ollama.sh # skim it: what does it download, where does it put things? (q to quit)
sh install-ollama.shWhat the flags mean: -f fail on errors, -s silent, -S still show errors, -L follow redirects, -o save to a file.
The installer asks for your sudo password, and it notices whether you're on Rocky or Ubuntu by itself. It also sets up Ollama as a service. That's Linux Basics, lesson 14 knowledge!
systemctl status ollama
If the installer complains that a tool is missing, install it (sudo dnf install NAME on Rocky, sudo apt install NAME on Ubuntu) and run it again.
Step 2: chat with a model
ollama run llama3.2
The first time, this downloads the model (about 2 GB for the 3-billion-parameter llama3.2), then gives you a >>> prompt. Ask it anything, for example “explain what cd .. does to a 10-year-old.” Type /bye to leave.
| Command | What it does |
|---|---|
ollama list | Show the models you've downloaded |
ollama pull NAME | Download a model without starting a chat |
ollama ps | Show which models are loaded in memory right now |
ollama rm NAME | Delete a model to free disk space |
Only 8 GB of RAM? Try the smaller llama3.2:1b (about 1.3 GB). It's faster, but less clever. Bigger models are smarter and need more memory. Picking one is always that trade-off.
Small local models make mistakes more often than big online ones. Everything from the last lesson applies double: check commands with man before you run them.
Step 3: talk to it from the command line
Ollama runs a small web service on your own machine, on port 11434. Programs talk to it with plain web requests, and you can too, using curl:
curl http://localhost:11434/api/generate -d '{
"model": "llama3.2",
"prompt": "In one sentence, what does the Linux ls command do?",
"stream": false
}'You get back JSON (structured text), with the answer in the "response" field. This is exactly how apps talk to AI models, including the big cloud ones.
By default Ollama only listens on localhost, so other computers can't reach it. Keep it that way. There's no password on port 11434, so don't open it in your firewall.
Step 4: script it with Python
Now the fun part: your own AI program. First, set up Python. This is where the families differ:
sudo dnf install -y python3-pip
Rocky's python3 already includes venv (virtual environments), so you only add pip.
sudo apt update sudo apt install -y python3-venv python3-pip
Ubuntu splits venv into its own package. Without it, python3 -m venv fails with an “ensurepip is not available” error.
Next, make a virtual environment. It's a private folder for the Python packages of one project, so they can't clash with the system's own Python. On Ubuntu it's required: if you try pip install outside one, you get an externally-managed-environment error. On Rocky it's just good practice.
python3 -m venv ~/ai-env # create it (once) source ~/ai-env/bin/activate # switch it on; your prompt now starts with (ai-env) pip install ollama # the official Ollama Python library
Now create a file called ask.py (with nano ask.py) and paste this in:
import sys
import ollama
question = " ".join(sys.argv[1:]) or "What does the ls command do?"
reply = ollama.chat(
model="llama3.2",
messages=[
{"role": "system", "content": "You are a friendly Linux tutor for high school students. Keep answers short. Always mention if a command differs between Rocky Linux and Ubuntu."},
{"role": "user", "content": question},
],
)
print(reply["message"]["content"])
python3 ask.py "how do I install tree?"
That system message is your AI's personality and rules. Change it and see how the answers change. That's the core idea behind every AI app.
Bonus project: the error explainer
Remember pipes from Linux Basics, lesson 12? Make explain.py. It reads whatever is piped into it and asks the AI to explain it:
import sys
import ollama
text = sys.stdin.read()
reply = ollama.chat(
model="llama3.2",
messages=[{"role": "user", "content": "Explain this Linux output to a beginner, and say if anything looks wrong:\n\n" + text}],
)
print(reply["message"]["content"])
ls /root 2>&1 | python3 explain.py systemctl status sshd 2>&1 | python3 explain.py
2>&1 means “send error messages down the pipe too.” Normally errors skip the pipe and go straight to the screen.
When you're done, type deactivate to leave the virtual environment.
Rehearse the Python setup
The practice terminal can't run AI models, but it does copy the Python differences, error messages included. Try it on both families.
Where to go next
- Try other models from the Ollama library and compare their answers to the same question.
- Give your
ask.pya memory by keeping themessageslist and adding each question and answer to it. - The big cloud models (like Anthropic's Claude API) work the same way, but they need an API key. Treat that key like a password: never paste it into chats or commit it to GitHub.
Quick check
1. On Ubuntu, python3 -m venv ai-env fails with “ensurepip is not available.” The fix?
✓ Ubuntu keeps venv in its own package. Rocky includes it with python3.
2. Why should you keep Ollama's port 11434 closed in your firewall?
✓ Only open what you need. You've heard that before!
3. What's the benefit of running a model locally instead of using an online chatbot?
✓ Privacy and control. The trade-off is that small local models are less capable than giant cloud ones.
Next up: AWS basics · Linux in the cloud, starting with “The cloud & AWS: getting started”.