ECI-SRV-11
Monitoring, maintenance and operational support
An outage that lasts for days without a single alert is not an isolated incident: it is a monitoring failure. We build the full chain — instrumentation, actionable alerts, an inventory kept current — and we can run it alongside you.
Scope of work
- Deployment of a centralised monitoring and logging platform.
- Instrumentation of servers, hypervisors, cloud platforms, storage and network.
- Design of actionable alerts and reduction of alert noise.
- Estate management, patch cycle and backup monitoring.
- Operational support and maintenance, with committed response times.
Deliverables
- Monitoring platform in production, with dedicated dashboards.
- Documented alert catalogue and escalation matrix.
- Estate inventory and compliance dashboard.
- Operating procedures, incident playbooks and periodic reports.
Technologies
- Zabbix, Prometheus, Grafana, Loki, OpenTelemetry.
- Dedicated monitoring for OpenStack, Ceph, clusters and containers.
- Linux and Windows patch chains, backup monitoring.
Support levels
Operational maintenance is delivered at one of three levels, chosen contractually according to how critical the platform is.
| Level | Commitment |
|---|---|
| Essential | Remote assistance during business hours, response within four business hours, patch management and quarterly review. |
| Advanced | Extended assistance, response within two business hours, delegated monitoring, on-site intervention and monthly review. |
| Critical | Continuous coverage, response within one hour, on-call rota, a named technical lead and annual recovery exercises. |
Engagement terms
An implementation project, followed by a monthly or annual support contract with an agreed volume of hours, response times and periodic review.
How the engagement runs
Phase 1 — Scope and priorities
- Definition of critical services and of the expected indicators.
- Identification of blind spots in the current monitoring.
Phase 2 — Implementation
- Platform deployment and instrumentation of the estate.
- Centralised logging and normalisation of formats.
Phase 3 — Alerts and reporting
- Design of alerts, dependencies and escalation paths.
- On-call and management dashboards.
Phase 4 — Operations
- Incident, patch and backup management.
- Periodic review and continuous improvement of coverage.
What every engagement includes
- A detailed quote, issued after framing
- A named project manager
- Operating documentation delivered
- Transfer of skills to your teams
- Work carried out on site or remotely
- Support after go-live
No work in production without a written rollback plan.
Request a framing session
Or contact us directly:
Let us talk about your situation
Every engagement begins with a framing exercise and a detailed quote.
Book an appointment