Monitoring and keeping an IT system operational
An outage lasting several days without any alert being raised is not an isolated technical incident: it is a monitoring design failure. This course builds the complete chain, from instrumenting the estate to structured incident management, by way of genuinely actionable alerts and an up-to-date inventory.
Training objectives
- Define a monitoring strategy aligned with service priorities.
- Deploy a monitoring, metrics and visualization platform.
- Instrument servers, hypervisors, cloud, storage and network equipment.
- Design actionable alerts and control alert fatigue.
- Centralize and correlate logs across a heterogeneous estate.
- Structure incident management and the patch lifecycle.
Target audience
- IT operations and service continuity managers.
- Systems and network administrators.
- On-call and support teams.
Prerequisites
- Linux or Windows systems administration.
- TCP/IP networking understanding and virtualization basics.
Detailed program
Day 1 — Monitoring fundamentals
- Objectives, observability and the three pillars: metrics, logs, traces
- Collection architectures, protocols and solution landscape
- Deploying the platform: Zabbix, Prometheus and Grafana
Day 2 — Instrumenting the estate
- System metrics, patch levels, licences and certificates
- Monitoring virtualization, OpenStack, Ceph and containers
- Automatic resource discovery and collection validation
Day 3 — Alerting and logging
- Designing actionable alerts, correlation and noise management
- On-call dashboards and management dashboards
- Centralized logging, format normalization and correlation
Day 4 — Keeping systems operational
- Estate inventory, lifecycle and compliance indicators
- Windows and Linux patch policy, non-disruptive updating
- Monitoring the backup chain and restore testing
Day 5 — Incident management and synthesis
- Incident management process, severity and escalation chain
- Root cause analysis and writing an incident report
- Crisis exercise, operating procedures and improvement plan
Certification
At the end of this training, you will receive a certificate of participation issued by squint.
Other training courses that might interest you
Linux Administration 1 — the fundamentals of running a server
Présentiel et distancielLinux is the foundation on which cloud, distributed storage and container platforms are built. Difficulties encountered with these advanced technologies …
Linux Administration 2 — automation, security and storage
Présentiel et distancielMoving from a managed server to a genuinely operated estate means automating repetitive tasks, logging what happens, hardening what is …
Ready to develop your skills?
Join hundreds of professionals who have trusted squint for their skills.
View all our training courses