Removing and Replacing Ceph OSDs

Removing and Replacing Ceph OSDs

Removing and Replacing Ceph OSDs: orch osd rm, rm, destroy and purge Explained

Ceph has five commands that sound like they remove an OSD:

  • ceph orch osd rm
  • ceph orch osd rm --replace
  • ceph osd rm
  • ceph osd destroy
  • ceph osd purge

They do different things, and picking the wrong one can leave orphaned daemons, stale keys, an OSD ID you can’t reuse, or a second round of rebalancing.

TL;DR

Goal cephadm cluster Classic cluster (packages + systemd)
Remove an OSD permanently ceph orch osd rm <id> --zap ceph osd purge
Replace a disk, keep the ID ceph orch osd rm <id> --replace --zap ceph osd destroy + ceph-volume lvm create --osd-id
Replace a disk without draining (degraded) ceph orch daemon stop + ceph osd destroy --force + ceph orch daemon rm --force systemctl stop + ceph osd destroy --force + ceph-volume lvm create --osd-id

On cephadm, always run ceph orch set-unmanaged <osd-spec> first, or the zapped disk is redeployed within seconds.

✅ = does it ❌ = doesn’t ➖ = not applicable

What happens orch osd rm orch osd rm --replace osd rm osd destroy osd purge
Checks safe-to-destroy first ✅ ² ✅ ² ❌ ✅ ¹ ✅ ¹
Drains data and waits ✅ (CRUSH weight 0) ✅ (marks out) ❌ ❌ ❌
Requires OSD already down ➖ ➖ ✅ ✅ ✅
Stops and removes the daemon/unit ✅ ✅ ❌ ❌ ❌
Removes cephx + dm-crypt keys ✅ ✅ ❌ ✅ ✅
Removes CRUSH entry ✅ ❌ ❌ ❌ ✅
Removes OSD map entry / frees ID ✅ ❌ ✅ ❌ ✅
Keeps ID as destroyed for reuse ❌ ✅ ❌ ✅ ❌
Wipes the disk with --zap with --zap ³ ❌ ❌ ❌

¹ Unless --force is given, or its older synonym --yes-i-really-mean-it.

² From Umbrella (v21), --force skips this check. In Tentacle and earlier it doesn’t.

³ In Squid (v19), Tentacle (v20) and Umbrella (v21), --replace always zaps the disk, even without --zap. This has been fixed on main (ceph/ceph#71938).

ceph orch osd rm does the whole job, from draining the data to wiping the disk. The ceph osd ... commands only change cluster state on the monitors. They never stop a daemon or touch a disk.

The commands

ceph orch osd rm <id> is for cephadm clusters. It puts the OSD in a removal queue, drains its data, waits until it is safe to destroy, stops and removes the container, then runs the equivalent of purge (or destroy with --replace). With --zap it also wipes the disk. In Tentacle (v20) and earlier, --force does not skip the safe-to-destroy check. That changes with Umbrella (v21): see ceph/ceph#71586 and the backport ceph/ceph#71729. Watch the progress with ceph orch osd rm status, and cancel with ceph orch osd rm stop <id>.

ceph osd rm <id> removes the OSD from the OSD map and nothing else. Its CRUSH entry remains as a DNE (“does not exist”) entry in ceph osd tree, and its cephx key stays as well. Before Luminous, removing an OSD meant running ceph osd crush remove, ceph auth del and ceph osd rm by hand. Since Luminous, ceph osd purge does all three steps in one command, so you rarely need ceph osd rm on its own.

ceph osd destroy <id> is for replacing a disk. It deletes the OSD’s cephx and dm-crypt keys and marks the OSD destroyed, but keeps the ID and its CRUSH position. A new disk created with --osd-id <id> takes over the same slot.

ceph osd purge <id> is for removing an OSD permanently. It does the same as crush remove + destroy + rm in one step, and frees the ID.

Both destroy and purge refuse to run if the OSD isn’t safe to destroy. --yes-i-really-mean-it is only a synonym for --force: it skips this safety check. Leave it out unless you really mean to skip the check.

Before you start

ceph -s                      # HEALTH_OK, no ongoing recovery
ceph osd df tree             # is there room for the data on the remaining OSDs?
ceph osd safe-to-destroy 12

Don’t move data twice. ceph orch osd rm already handles this correctly. On a classic cluster, it matters how you drain. An OSD has two weights:

  • Reweight: ceph osd out sets it to 0, and the OSD’s data moves away. But the host still counts the OSD’s CRUSH weight, so CRUSH keeps sending the host its full share of data. Part of it ends up on the host’s remaining OSDs.
  • CRUSH weight: ceph osd purge removes it. The host’s weight drops, so its share of data shrinks. PGs then move from its remaining OSDs to other hosts. That’s the second round of data movement.

For a permanent removal, drain with ceph osd crush reweight osd.N 0. The host’s weight drops right away, data moves once to its final place, and the purge afterwards moves nothing. For a replacement, out is the right choice, because the new disk takes over the CRUSH weight.

Quick disk swaps: noout and norebalance. If you want to swap a disk without draining it first, stop the cluster from rebalancing while the OSD is gone:

ceph osd set noout          # down OSD isn't marked out after 10 min (mon_osd_down_out_interval)
ceph osd set norebalance    # misplaced PGs aren't moved around
# ... stop, destroy --force, swap and recreate the OSD with the same ID ...
ceph osd unset norebalance
ceph osd unset noout

Degraded PGs still recover and backfill onto the new OSD. What stops is the shuffling of data to other OSDs and back again. Don’t set these flags during a drain: norebalance blocks exactly the data movement that a drain relies on, so safe-to-destroy never succeeds. Remember to unset them afterwards, because ceph -s only shows them as a warning.

cephadm

cephadm runs every OSD as a container and keeps track of it as a daemon. Always remove OSDs through the orchestrator, so the container is removed together with the OSD. If you run ceph osd purge directly, the container stays behind on the host.

Always set your OSD spec to unmanaged first. --zap makes the disk available again. A managed spec (for example --all-available-devices) picks it up and deploys a new OSD on it within seconds, before you’ve had a chance to pull the disk. This applies to removals and replacements alike:

ceph orch ls osd --export                          # find the spec name
ceph orch set-unmanaged osd.all-available-devices

Remove permanently:

ceph orch set-unmanaged osd.all-available-devices
ceph orch osd rm 12 --zap
ceph orch osd rm status

Pull the disk before you set the spec to managed again, or it comes back as a new OSD.

Replace a disk, keep the ID (with drain):

The safe option: the old disk keeps its data until all of it has been copied elsewhere. Redundancy never drops, but the drain can take hours.

ceph orch set-unmanaged osd.all-available-devices
ceph orch osd rm 12 --replace --zap
# wait until osd.12 shows "destroyed" in `ceph osd tree`, then swap the disk
ceph orch device ls node3 --refresh
ceph orch set-managed osd.all-available-devices

Once the spec is managed again, cephadm deploys an OSD on the new disk and reuses ID 12. Without a matching spec, run ceph orch daemon add osd node3:/dev/sdX instead. It also reuses the destroyed ID.

Replace a disk, keep the ID (without drain):

The fast option: you remove the OSD right away and accept that its PGs are degraded until the new disk has been backfilled. Data moves only once, onto the new disk. Until then, the affected PGs have one copy fewer. Don’t do this on pools with size=2, or when other OSDs are already down.

ceph orch osd rm --force doesn’t help here. In Tentacle and earlier it still waits for safe-to-destroy, which never passes while PGs are degraded. Umbrella (v21) skips that wait with --force, but usually still marks the OSD out first. Data then moves to other OSDs and back to the new disk later. Use the individual steps instead:

ceph orch set-unmanaged osd.all-available-devices
ceph osd set noout
ceph osd set norebalance
ceph orch daemon stop osd.12
ceph osd destroy 12 --force                  # PGs are degraded, so --force is required
ceph orch daemon rm osd.12 --force           # removes the container and unit
# swap the disk
ceph orch device ls node3 --refresh
ceph orch set-managed osd.all-available-devices   # or: ceph orch daemon add osd node3:/dev/sdX
# once osd.12 is up again:
ceph osd unset norebalance
ceph osd unset noout

See Quick disk swaps for what the two flags do.

Classic deployment

Remove permanently:

ceph osd crush reweight osd.12 0
while ! ceph osd safe-to-destroy osd.12; do sleep 60; done
systemctl disable --now ceph-osd@12            # on the OSD host
ceph osd purge 12
ceph-volume lvm zap --osd-id 12 --destroy      # on the OSD host

Replace a disk, keep the ID (with drain):

The safe option: redundancy never drops, but the drain can take hours.

ceph osd out 12
while ! ceph osd safe-to-destroy osd.12; do sleep 60; done
systemctl stop ceph-osd@12
ceph osd destroy 12
# swap the disk (zap it first if you reuse the same device)
ceph-volume lvm create --osd-id 12 --data /dev/sdX

If the disk is already dead, Ceph marks the OSD out after 10 minutes (mon_osd_down_out_interval). Wait for recovery to finish, then continue from destroy.

Replace a disk, keep the ID (without drain):

The fast option: the affected PGs stay degraded until the new disk has been backfilled, and data moves only once. The same caveats apply as for cephadm: don’t do this on pools with size=2, or when other OSDs are already down.

ceph osd set noout
ceph osd set norebalance
systemctl stop ceph-osd@12                     # on the OSD host
ceph osd destroy 12 --force                    # PGs are degraded, so --force is required
# swap the disk
ceph-volume lvm create --osd-id 12 --data /dev/sdX   # on the OSD host
# once osd.12 is up again:
ceph osd unset norebalance
ceph osd unset noout

If the disk has already failed, set noout within 10 minutes (mon_osd_down_out_interval). Otherwise the OSD is marked out, and recovery onto other OSDs starts anyway.

Common pitfalls

  • Orphaned osd.N in ceph orch ps: the OSD was purged or destroyed with ceph osd ..., but its daemon was never removed. Fix it with ceph orch daemon rm osd.N --force.
  • The removed OSD comes back on cephadm: a managed spec picked up the zapped disk. Always set the spec to unmanaged before running ceph orch osd rm.
  • DNE in ceph osd tree: osd rm was run without crush remove. Fix it with ceph osd crush remove osd.N.
  • “Key exists” when creating the new OSD: a stale key is left over. Fix it with ceph auth del osd.N.
Daniel Vogelbacher's Picture

About Daniel Vogelbacher

Hi, I'm Daniel, a software developer, Linux administrator and landscape photographer.

Germany https://chaospixel.com