ECI-SRV-11

Monitoring, maintenance and operational support

An outage that lasts for days without a single alert is not an isolated incident: it is a monitoring failure. We build the full chain — instrumentation, actionable alerts, an inventory kept current — and we can run it alongside you.

Scope of work

  • Deployment of a centralised monitoring and logging platform.
  • Instrumentation of servers, hypervisors, cloud platforms, storage and network.
  • Design of actionable alerts and reduction of alert noise.
  • Estate management, patch cycle and backup monitoring.
  • Operational support and maintenance, with committed response times.

Deliverables

  • Monitoring platform in production, with dedicated dashboards.
  • Documented alert catalogue and escalation matrix.
  • Estate inventory and compliance dashboard.
  • Operating procedures, incident playbooks and periodic reports.

Technologies

  • Zabbix, Prometheus, Grafana, Loki, OpenTelemetry.
  • Dedicated monitoring for OpenStack, Ceph, clusters and containers.
  • Linux and Windows patch chains, backup monitoring.

Support levels

Operational maintenance is delivered at one of three levels, chosen contractually according to how critical the platform is.

LevelCommitment
EssentialRemote assistance during business hours, response within four business hours, patch management and quarterly review.
AdvancedExtended assistance, response within two business hours, delegated monitoring, on-site intervention and monthly review.
CriticalContinuous coverage, response within one hour, on-call rota, a named technical lead and annual recovery exercises.

Engagement terms

An implementation project, followed by a monthly or annual support contract with an agreed volume of hours, response times and periodic review.

How the engagement runs

Phase 1 — Scope and priorities

  • Definition of critical services and of the expected indicators.
  • Identification of blind spots in the current monitoring.

Phase 2 — Implementation

  • Platform deployment and instrumentation of the estate.
  • Centralised logging and normalisation of formats.

Phase 3 — Alerts and reporting

  • Design of alerts, dependencies and escalation paths.
  • On-call and management dashboards.

Phase 4 — Operations

  • Incident, patch and backup management.
  • Periodic review and continuous improvement of coverage.

What every engagement includes

  • A detailed quote, issued after framing
  • A named project manager
  • Operating documentation delivered
  • Transfer of skills to your teams
  • Work carried out on site or remotely
  • Support after go-live

No work in production without a written rollback plan.

Request a framing session

Let us talk about your situation

Every engagement begins with a framing exercise and a detailed quote.

Book an appointment