ECI-SRV-14
Private AI Inference Platform
Sizing the hardware is not enough: an inference platform stands on its serving engine, its access control, its document indexing and its monitoring. We deploy that whole chain on your own infrastructure, so that the data it processes never leaves it.
Scope of work
- Framing of the intended uses, the data involved and the level of confidentiality required.
- Deployment of the inference engine and a single access interface.
- Document indexing pipeline and vector store, fed from your own repositories.
- Access control, per-team quotas and request logging.
- Measurement of throughput, latency and accelerator memory usage.
Deliverables
- Platform deployed on your own infrastructure, with its scaling procedure.
- Model register: origin, checksum of the weights file, storage format and commissioning date.
- Documented, repeatable indexing pipeline, with its reindexing procedure.
- Operating documentation, monitoring dashboards and rollback procedure.
Technologies
- Open serving engines and verifiable weight formats, deployed as containers.
- Open source vector stores and indexing pipelines, backed by your existing storage.
- Monitoring with Prometheus, Grafana and Loki, wired into your alerting chain.
Where business knowledge belongs
| Situation | Mechanism |
|---|---|
| Many documents, revised often, sources to be cited | Document indexing and retrieval at question time |
| Fixed output format, stable business vocabulary | Instructions and examples carried by the platform |
| Model behaviour to be changed in depth | Trained adapters, outside our scope |
Terms
Ten to thirty days depending on whether the engagement covers the serving platform alone or the whole chain with indexing and monitoring, fixed price after framing. Deployment takes place on GPU infrastructure we size, or on a platform you already run.
We deploy and operate the platform; the choice of models, the quality of their answers and the use made of them remain with your teams.
How the engagement runs
Phase 1 — Framing
- Priority uses, document volume and expected request throughput.
- Data involved, confidentiality and location constraints.
Phase 2 — Serving platform
- Installation of the inference engine and vector store on the target platform.
- Access control, per-team quotas and request logging.
Phase 3 — Commissioning
- Indexing of the first repositories and acceptance on an agreed set of questions.
- Throughput, latency and memory usage measured under representative load.
Phase 4 — Handover
- Operating documentation, model register and version upgrade procedure.
- Training your teams on reindexing, monitoring and rollback.
What every engagement includes
- A detailed quote, issued after framing
- A named project manager
- Operating documentation delivered
- Transfer of skills to your teams
- Work carried out on site or remotely
- Support after go-live
No work in production without a written rollback plan.
Request a framing session
Or contact us directly:
Let us talk about your situation
Every engagement begins with a framing exercise and a detailed quote.
Book an appointment