ECI-SRV-13
AI infrastructure — GPU sizing
Running an AI model in-house starts with a sizing question: how much VRAM, how much memory bandwidth, how much power density per rack. We help choose the GPU infrastructure, on premises or hosted, and implement it in a virtualised environment.
Scope of work
- Workload framing: size of the target models, expected request volume, target latency.
- Sizing of the VRAM and memory bandwidth required for inference.
- Choice between several cards in one server or several servers connected together.
- Rack-level power density and cooling, sized for the GPU workload.
- Accelerator virtualisation through passthrough or vGPU, depending on the target platform.
Deliverables
- A sizing note: required VRAM, number of cards, network and power constraints.
- A documented target architecture, including the accelerator virtualisation choice.
- An inference platform deployed and validated against a representative set of models.
- Operating documentation and a scale-out procedure.
Technologies
- PCIe passthrough and vGPU on Proxmox VE, Red Hat OpenStack and OpenShift Virtualization.
- Common inference libraries: vLLM, TensorRT, standard quantisation formats.
- Dedicated high-throughput networks for multi-card and multi-node configurations.
Engagement terms
Five to twenty days depending on whether the engagement covers sizing alone or full implementation, fixed price after framing.
We size and deploy the infrastructure; choosing and training the models remains the responsibility of your teams or of the software vendor in use.
How the engagement runs
Phase 1 — Workload framing
- Target models, parameter count, planned quantisation format.
- Expected throughput and latency constraints.
Phase 2 — Sizing
- Calculation of the required VRAM and of the number of cards or servers.
- Choice of accelerator virtualisation mode and of the host platform.
Phase 3 — Implementation
- Platform installation and configuration, rack cabling and cooling.
- Deployment of the inference chain and load testing against target models.
Phase 4 — Handover
- Documentation of the architecture and of the scale-out procedures.
- Training of your teams on day-to-day operation of the platform.
What every engagement includes
- A detailed quote, issued after framing
- A named project manager
- Operating documentation delivered
- Transfer of skills to your teams
- Work carried out on site or remotely
- Support after go-live
No work in production without a written rollback plan.
Request a framing session
Or contact us directly:
Let us talk about your situation
Every engagement begins with a framing exercise and a detailed quote.
Book an appointment