Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Hardening a cloud server

Spin up a fresh VPS, wait an hour, and the auth log already has thousands of brute-force attempts from across the internet. Every server-hardening guide says roughly the same things — here's what each one actually stops, and where the rules are theater.

Security intro May 25, 2026 · updated Aug 25, 2026 · 15 min read

On this page

The picture version

Five pictures for a reader who has just rented their first server, following one hour: the fresh VPS whose auth log filled up on its own.

1 · The problem

Nobody knew this server existed an hour ago.

Failed password for root from 45.x.x.x Failed password for admin from 103.x.x.x Failed password for ubuntu from 61.x.x.x Failed password for oracle from 192.x.x.x Failed password for pi from 87.x.x.x Failed password for git from 209.x.x.x … and thousands more one hour old nobody knew this server existed this morning nobody is targeting you the address answered on port 22, so it joined a queue The scanning is continuous, automated, and not about you.
That is the world a hardening guide is written for. Every line of the usual advice is aimed at a specific failure mode that happens often enough to be the boring default attack — which is also why you can tell the useful rules from the cargo-culted ones by asking what each one actually stops.

2 · The three buckets

Every concrete rule slots into one of these.

shrink the surface fewer ports, fewer services, fewer accounts that could be attacked at all remove the easy wins what is exposed shouldn’t fall to a guessed password or a known CVE make it loud when the first two fail — and one will — you find out in minutes, not from the bill Three goals, in that order. and most guides stop after two
The order matters. Shrinking is cheapest, because a service the internet can’t reach can’t be exploited from the internet regardless of what it carries. The third bucket is the one the plan is built around, because it assumes the first two eventually fail.

3 · The front door

One change closes it. The rest just quieten the logs.

what actually closes it PasswordAuthentication no PermitRootLogin no removes online guessing as an attack path altogether what only quietens the logs moving SSH off port 22 fail2ban, once you are keys-only both are real operational wins and neither is a defensive layer The failure mode moves from guessing to key theft — a smaller problem. the thousands of failed-password lines go to zero, because the server stops offering that method at all
Being able to say which column a rule belongs in is most of the skill. A control that reduces noise is worth having and worth not counting — the danger is a layer that feels protective sitting on top of something that was never patched.

4 · The bucket everyone skips

Assume you lose, and arrange to find out.

the bucket most guides under-cover ship logs off the box a root shell can truncate anything local; logs elsewhere are what you’ll actually use watch the drop spots your own apps have no reason to write executables into world-writable directories alert on egress the miner phones a pool, the payload phones home — a baseline plus an alert catches it The median story isn’t nation-state malware. it’s a miner running off a binary dropped after a web-app exploit — easy to spot if you ever look This is the only bucket still working on the day the first two didn’t.
Note what these have in common: none of them prevent anything. They shorten the gap between “compromised” and “knows it is compromised”, which is the variable that decides whether an incident is an afternoon or a quarter.

5 · Keep this card

The whole thing on one index card.

server hardening = shrink the attack surface + remove the easy wins from what’s left + make compromise loud ∴ and the three are not equally weighted the first two answer the scanning that filled your auth log; the third is the one you’ll be grateful for
Picture to keep: a building with far fewer doors than it was shipped with, good locks on the ones that are left, and a camera pointed at the inside — because the plan assumes somebody eventually gets in. Where it breaks: unlike a building, most of your doors were opened by software you installed rather than by an architect, so the shrinking is something you do continuously rather than once.

Why it exists

Spin up a fresh VPS on any cloud, give it a public IP, leave the default SSH port open, walk away for an hour, then come back and read /var/log/auth.log. You’ll find thousands of failed login attempts — root, admin, ubuntu, oracle, pi, git — from IP addresses all over the world. Nobody knew your server existed an hour ago. Nobody is targeting you. The entire public IPv4 space is being continuously port-scanned by botnets, and the moment your address answered on port 22, it joined a work queue.

That’s the world a hardening guide is written for. The advice every guide repeats — keys-only SSH, no root login, close ports you don’t need, patch automatically, ship logs off-box — isn’t ritual; each line is aimed at a specific failure mode that has happened often enough to be the boring default attack. The point of this post is to walk through those rules and, for each one, say what it actually stops so you can tell the useful ones from the cargo-culted ones.

Why it matters now

Cloud servers are a commodity. A single-digit-dollar VM runs side-project APIs, small-business sites, Mastodon instances, home-lab tunnels, and increasingly self-hosted LLM frontends. The default cloud image is built to boot, not to be exposed to the open internet. The specifics vary by distro and provider, and any given image’s posture is worth checking rather than assuming — but the general pattern holds: the host firewall is yours to configure, access arrives pre-seeded by cloud-init rather than hardened, the application stack you install is yours to keep patched, and the public IP you were given will start attracting scans. The provider hands you a host; the rest is your problem.

Meanwhile the attacker side has gotten cheaper, not more sophisticated. A great deal of what compromises servers is opportunistic rather than targeted — credential stuffing, exposed Redis/MongoDB/Elasticsearch instances, mass exploitation of unpatched CVEs weeks or months old, leaked AWS keys committed to public GitHub repos. The “MongoDB ransomware” waves of the mid-2010s worked not because the attackers were clever but because so many instances were reachable from the public internet with access control switched off. Today MongoDB’s packaged installs bind to localhost and ship with authorization disabled until you enable it, so the modern version of the mistake is an operator who binds the service to a public interface and doesn’t turn auth on. The same shape of mistake still ships, just on different services.

The short answer

server hardening = shrink the attack surface + remove the easy wins + make compromise loud

Picture to keep: a building with far fewer doors than it was shipped with, good locks on the ones that are left, and a camera pointed at the inside — because the plan assumes somebody eventually gets in. (Where the picture breaks: unlike a building, most of your doors were opened by software you installed, not by an architect, so the shrinking is a thing you do continuously rather than once.)

Three goals, in that order. Shrink the attack surface means fewer ports, fewer services, fewer accounts that could be attacked at all. Remove the easy wins means the things that are exposed shouldn’t fall to a guessed password or a known CVE. Make compromise loud means when one of the first two fails — and eventually one will — you find out within minutes instead of finding out from your hosting bill.

Every concrete rule below slots into one of those three buckets.

How it works

Before you SSH in: the cloud account is the real perimeter

The first hardening step is one most server-focused guides skip, because they assume you’re already inside the VM.

Turn on MFA for the cloud console, and don’t keep long-lived root API keys on your laptop. Why: if someone phishes your AWS/GCP/Hetzner login, none of the in-VM hardening matters — they can snapshot your disk, spin up a new VM with your snapshot mounted, and read whatever they want, or just rotate the SSH key and walk in the front door. The host operating system is downstream of the cloud account; protect the upstream first. MFA with a hardware key beats TOTP beats SMS, in that order, and the gap between the first and the other two is not a matter of degree: a code you can read out is a code a convincing page can ask you for. See passkeys for why phishing-resistant, hardware-backed MFA is the rung that actually closes credential phishing.

Use the cloud firewall (security group / network ACL), not just in-VM ufw/iptables. Why: a cloud firewall filters packets before they ever reach the VM’s kernel. If a service inside your VM gets compromised and disables ufw, the cloud rules still apply. And if you misconfigure ufw itself — forgetting ufw enable, opening a port you only meant to allow from your office IP, loosening DEFAULT_FORWARD_POLICY when you meant to tighten it — the cloud firewall is a second fence. The in-VM firewall is fine to keep, but as defense in depth, not as the only layer.

Have a backup that isn’t writeable from the server. Why: the hardening rules below are all about preventing compromise; backups are the only thing that helps after compromise. The constraint is “not writeable from the server” — a backup that the compromised root user can rm -rf is no backup. Provider-side snapshots with a separate account, or restic/borg to an append-only / immutable bucket, are the usual answers.

Lock the front door: SSH

The single port every hardening guide spends the most ink on, because it’s also the one every botnet spends the most ink on.

Disable password authentication; allow public-key only. Set PasswordAuthentication no and KbdInteractiveAuthentication no in /etc/ssh/sshd_config (older configs still set the deprecated alias ChallengeResponseAuthentication no — same effect, current OpenSSH prefers the new name). Why: the entire botnet-scanning industry is built on guessing passwords. Public-key authentication removes online guessing as an attack path altogether — the realistic failure mode becomes key theft or careless key handling, which is a different and much smaller problem. The instant you turn off passwords, the thousands of Failed password for root lines in auth.log become zero, because the server stops even offering that authentication method.

Disable root login over SSH. Set PermitRootLogin no. Why: if root login is allowed, the attacker knows the username for free — root exists on every Linux box. They only need to guess one thing (the credential), not two (the username and the credential). Forcing a named user plus sudo also gives you per-person attribution: sudo logs each elevated command via syslog or the journal with the original user attached (on Debian and Ubuntu that usually lands in auth.log), so a post-incident review can tell which admin ran what. SSH still logs direct root logins too, but if multiple people share that root credential the trail stops at “root did it.”

Don’t bother changing the SSH port for “security.” Why: moving SSH from 22 to 2222 doesn’t stop a determined attacker — port scans take seconds. It does dramatically cut the volume of script-kiddie noise in your logs, which makes real anomalies easier to see. So move it if you want a quieter auth.log, not because it’s a security control. Security through obscurity isn’t worthless, but don’t count it as a layer.

Be skeptical of fail2ban as a security control on keys-only SSH. fail2ban watches log files for failed logins and temporarily blocks the source IP. Why the skepticism: with password auth disabled, those failed attempts can’t succeed anyway, so banning the IP doesn’t change the security posture much. It does cut log volume and CPU spent on doomed scans, which is a real operational win — just don’t count it as a defensive layer it isn’t. Where fail2ban earns its keep as security is in front of services that still authenticate with passwords (a database you exposed by mistake, a web admin panel, an SMTP relay).

Shrink the attack surface: close, don’t just guard

Default-deny inbound on the cloud firewall; explicitly open only what you need. Why: a service the internet can’t reach can’t be exploited from the internet, regardless of whatever CVEs it carries. Every port you open is a service that has to stay patched for as long as the server exists. The cheapest security work is the work you don’t have to do because the attack never reaches a listener. Start from “nothing open except SSH from my IPs,” then open exactly what the app needs.

Don’t bind databases or caches to a public interface. Bind Postgres, Redis, MongoDB, Elasticsearch, Memcached, etc. to 127.0.0.1 or to a private network, not to 0.0.0.0. Why: this is the failure mode behind almost every “tens of thousands of MongoDB/Elasticsearch/Redis instances ransomed” headline of the last decade. Defaults vary by product — some bind to localhost out of the box, some require auth, some don’t — but the common mistake is the same shape: an operator binds the service to a public interface because they need remote access, and either skips auth or sets a weak one because “it’s only used internally.” The robust fix isn’t “set a strong password,” it’s “don’t let the internet talk to it at all” — SSH-tunnel, VPN, or private-network access only.

Remove packages you aren’t using. Cloud images ship general-purpose — inventory yours and you’ll usually find things a single-purpose server doesn’t need. Why: every installed package is in your patch surface even if it isn’t running. A CVE in a binary you never invoke can still be a privilege-escalation primitive for an attacker who got in through a different door. “Smaller image” isn’t aesthetic; it’s fewer things to track.

Run application services as non-root, with systemd sandboxing. Set User=, Group=, and the Protect* / Private* directives in the unit file (ProtectSystem=strict, ProtectHome=true, PrivateTmp=true, NoNewPrivileges=true, etc.). Why: if your web app gets popped via SQL injection or a deserialization bug, the blast radius is whatever that process can touch. Least privilege turns “RCE in my app” from “root on the host” into “a sandboxed process that can read its own working directory and not much else.”

Make patching automatic

Enable unattended security upgrades. On Debian/Ubuntu, install and configure unattended-upgrades to apply the ${distro_id}:${distro_codename}-security channel automatically. Why: a large share of the CVEs that actually get exploited in the wild — see CISA’s “Known Exploited Vulnerabilities” catalog for the running list — have had public patches for weeks, months, or years. The bottleneck usually isn’t the patch; it’s the human who has to remember to apply it. Automating security backports (specifically security, not arbitrary upgrades) is one of the highest return-on-effort hardening steps available, because distro maintainers already do the careful work of backporting fixes without changing behaviour.

Reboot when the kernel updates. Set Unattended-Upgrade::Automatic-Reboot "true" with a sensible time window, or run an explicit reboot workflow of your own. (needrestart is a useful companion, but it restarts services whose libraries changed — it is not a substitute for booting a new kernel.) Why: a new kernel image on disk that the system never boots into isn’t actually running. Kernel CVEs (privilege escalations, especially) are common, and for most fixes the way you start running the patched kernel is a reboot. Live-patching services (Canonical Livepatch, kpatch, Ksplice) can apply a subset of kernel fixes without a reboot, with varying licensing and coverage, but none of them cover everything — plan for reboots regardless.

Pin major versions; auto-apply only security backports. Why: distro security updates aim to fix the bug without changing behaviour — that’s the whole point of a backport — so they rarely cause regressions in practice. Automatic major version bumps are a different story and regularly do break things; that’s where “the server updated overnight and now my app won’t start” stories come from. Keep the two separate.

Make compromise loud

This is the bucket most guides under-cover, and it’s the one that matters most once the first two layers eventually fail.

Ship logs off the box. Forward journald/auth.log/syslog to a separate account or service (a managed log service, an S3 bucket the VM can write but not read, another VM in another account). Why: a common move after getting root is to clobber the local logs — echo > /var/log/auth.log, journalctl --rotate && journalctl --vacuum-time=1s, or just rm. Logs that live on the compromised box can be tampered with by whoever compromised it. Logs in a separate trust boundary are the artifact you’ll actually use in incident response. The cost is small; the value the day you need it is enormous.

Watch for new binaries in the usual drop spots. World-writable directories like /tmp, /var/tmp, and /dev/shm, plus /usr/local/bin, are the obvious places for a payload that arrived via a web-app exploit to land — there’s no reason your own apps should be writing executables there. A systemd timer that periodically diffs a manifest (or a small tool like osquery, aide, or wazuh) catches the boring case. Why: the median “my server got hacked” story isn’t nation-state malware; it’s an XMRig miner running off a binary the attacker dropped after exploiting a public web app. The signature is “a process you didn’t install showed up.” That’s easy to detect if you ever look — and almost no one does, until the CPU bill arrives.

Alert on outbound traffic to unusual destinations. Why: egress is often where a compromise becomes easiest to notice, rather than ingress. The attacker’s payload phones home to a command-and-control server, or the miner connects to a mining pool, or your AWS keys start being used from an IP you’ve never seen. A baseline of “this server talks to these three IPs” plus an alert when it doesn’t is a surprisingly effective last line of detection. Cloud-native flow logs (VPC Flow Logs, equivalents elsewhere) are one practical way to get that signal.

Show the seams

A hardening guide that doesn’t tell you where it lies isn’t being honest. The honest version:

You started with server hardening = shrink the attack surface + remove the easy wins + make compromise loud. What did this post add? — + the three are not equally weighted, and most guides stop after two. Shrinking and locking are the parts everyone does, and they’re aimed at exactly the opportunistic scanning that filled your auth.log in the first hour. The third bucket is the one you’ll actually be grateful for, because it’s the only one still working on the day the first two didn’t.

Going deeper