ECI-SRV-13

AI infrastructure — GPU sizing

Running an AI model in-house starts with a sizing question: how much VRAM, how much memory bandwidth, how much power density per rack. We help choose the GPU infrastructure, on premises or hosted, and implement it in a virtualised environment.

Scope of work

  • Workload framing: size of the target models, expected request volume, target latency.
  • Sizing of the VRAM and memory bandwidth required for inference.
  • Choice between several cards in one server or several servers connected together.
  • Rack-level power density and cooling, sized for the GPU workload.
  • Accelerator virtualisation through passthrough or vGPU, depending on the target platform.

Deliverables

  • A sizing note: required VRAM, number of cards, network and power constraints.
  • A documented target architecture, including the accelerator virtualisation choice.
  • An inference platform deployed and validated against a representative set of models.
  • Operating documentation and a scale-out procedure.

Technologies

  • PCIe passthrough and vGPU on Proxmox VE, Red Hat OpenStack and OpenShift Virtualization.
  • Common inference libraries: vLLM, TensorRT, standard quantisation formats.
  • Dedicated high-throughput networks for multi-card and multi-node configurations.

Engagement terms

Five to twenty days depending on whether the engagement covers sizing alone or full implementation, fixed price after framing.

We size and deploy the infrastructure; choosing and training the models remains the responsibility of your teams or of the software vendor in use.

How the engagement runs

Phase 1 — Workload framing

  • Target models, parameter count, planned quantisation format.
  • Expected throughput and latency constraints.

Phase 2 — Sizing

  • Calculation of the required VRAM and of the number of cards or servers.
  • Choice of accelerator virtualisation mode and of the host platform.

Phase 3 — Implementation

  • Platform installation and configuration, rack cabling and cooling.
  • Deployment of the inference chain and load testing against target models.

Phase 4 — Handover

  • Documentation of the architecture and of the scale-out procedures.
  • Training of your teams on day-to-day operation of the platform.

What every engagement includes

  • A detailed quote, issued after framing
  • A named project manager
  • Operating documentation delivered
  • Transfer of skills to your teams
  • Work carried out on site or remotely
  • Support after go-live

No work in production without a written rollback plan.

Request a framing session

Let us talk about your situation

Every engagement begins with a framing exercise and a detailed quote.

Book an appointment