Présentiel et distanciel
5 days (35 hours)

Monitoring and keeping an IT system operational

An outage lasting several days without any alert being raised is not an isolated technical incident: it is a monitoring design failure. This course builds the complete chain, from instrumenting the estate to structured incident management, by way of genuinely actionable alerts and an up-to-date inventory.

Training objectives

  • Define a monitoring strategy aligned with service priorities.
  • Deploy a monitoring, metrics and visualization platform.
  • Instrument servers, hypervisors, cloud, storage and network equipment.
  • Design actionable alerts and control alert fatigue.
  • Centralize and correlate logs across a heterogeneous estate.
  • Structure incident management and the patch lifecycle.

Target audience

  • IT operations and service continuity managers.
  • Systems and network administrators.
  • On-call and support teams.

Prerequisites

  • Linux or Windows systems administration.
  • TCP/IP networking understanding and virtualization basics.

Detailed program

Day 1 — Monitoring fundamentals

  • Objectives, observability and the three pillars: metrics, logs, traces
  • Collection architectures, protocols and solution landscape
  • Deploying the platform: Zabbix, Prometheus and Grafana

Day 2 — Instrumenting the estate

  • System metrics, patch levels, licences and certificates
  • Monitoring virtualization, OpenStack, Ceph and containers
  • Automatic resource discovery and collection validation

Day 3 — Alerting and logging

  • Designing actionable alerts, correlation and noise management
  • On-call dashboards and management dashboards
  • Centralized logging, format normalization and correlation

Day 4 — Keeping systems operational

  • Estate inventory, lifecycle and compliance indicators
  • Windows and Linux patch policy, non-disruptive updating
  • Monitoring the backup chain and restore testing

Day 5 — Incident management and synthesis

  • Incident management process, severity and escalation chain
  • Root cause analysis and writing an incident report
  • Crisis exercise, operating procedures and improvement plan

Certification

At the end of this training, you will receive a certificate of participation issued by squint.

Price on request

Duration

5 days (35 hours)

Format

Présentiel et distanciel

Next session

On request

Request a quote

Other training courses that might interest you

Linux Administration 1 — the fundamentals of running a server

Présentiel et distanciel

Linux is the foundation on which cloud, distributed storage and container platforms are built. Difficulties encountered with these advanced technologies …

5 days (35 hours) Learn more

Linux Administration 2 — automation, security and storage

Présentiel et distanciel

Moving from a managed server to a genuinely operated estate means automating repetitive tasks, logging what happens, hardening what is …

5 days (35 hours) Learn more

Ready to develop your skills?

Join hundreds of professionals who have trusted squint for their skills.

View all our training courses