8 February 2026

Day-two health checks every log collector needs

A short checklist administrators can run weekly to keep log shippers and collectors trustworthy.

Collection that works on launch day can drift: agents stop after package updates, disks fill, or TLS certificates expire. Without a light routine, gaps appear only during an incident.

Each week, confirm last-seen timestamps for every critical source, check free space on collector nodes, and verify that scheduled rotations completed. Keep the checklist short enough that someone will actually run it.

Watch for silent sources: hosts that still ping but send no events. Page Spruceway includes this pattern in setup handoff runbooks so administrators have a named owner for the weekly pass.

After OS patch cycles, re-check shipper service status on a sample of hosts before closing the change ticket. Many “mystery gaps” start there.

Store the checklist beside the monitoring runbook, not in a forgotten shared drive. Visibility matters as much as the steps themselves.

Back to field notes