Fleet Observability & Alerting (Zabbix)

Zabbix Server Monitor
Zabbix Server Monitor
Zabbix Server Monitor
Zabbix Server Monitor

Project information

  • Category: SysOps
  • Client: Waletech (INSTAR Germany) Shenzhen, China
  • Project date: 01 March, 2018
  • Project URL: www.zabbix.com

The problem: on a growing fleet, we were learning about outages from customers, not from our dashboards. What I did: Integrated Zabbix across the full fleet — CPU, memory, disk, network saturation, plus custom health-check API probes for every customer-facing service. The Zabbix server itself is containerized on Nomad against a PostgreSQL backend — server, frontend, and database in one job with canary-safe rollout and a health-gated deploy, so the monitoring stack is as fault-tolerant as the things it watches. I sized the Zabbix proxy/server caches (history, trend, value caches) for our event volume and tuned ulimits for the connection counts. System-critical variables feed per-service dashboards, with email and Slack alerts on threshold breach. The payoff: the cloud team gets paged on degradation, not on a support ticket — and the Zabbix dashboards are the same ones the AI pipeline project above relies on.

Designed with BootstrapMade