TR EN RU
Book a Technical Call
Workload and model GPU and VRAM Deploy and test LLM

GPU Servers and LLM Servers: On-Premises AI Infrastructure

We size, deploy and operate GPU servers so you can run large language models (LLMs) and AI workloads in your own data center.

Scope

GPU server selection, procurement, deployment and operations for AI workloads such as LLM inference, RAG applications, fine-tuning and computer vision. We start from model size and cover VRAM, server architecture, networking, power and cooling and the software layer in one plan.

Key Areas

  • VRAM and GPU sizing by model size
  • NVIDIA L4, L40S, H100, H200 and Blackwell options
  • High-throughput inference with vLLM and TensorRT-LLM
  • Vector database and internal API for RAG
  • Power, cooling and rack planning

Why an On-Premises LLM Server?

Every prompt and document sent to a public AI service leaves your organization. For workloads involving customer data, contracts, source code or trade secrets, that may be unacceptable for privacy and compliance. With an LLM running on your own GPU server, data stays inside.

  • Data control: prompts, documents and outputs stay in your infrastructure; you set logging and retention policy.
  • Predictable cost: with sustained, high utilization, unit cost can be lower than paying per token.
  • Low latency: a model on the same network as internal systems does not depend on external API calls.
  • Customization: open-weight models can be adapted to company data with RAG or fine-tuning.

In return, hardware investment and operations are your responsibility. If usage is periodic or you are still piloting, starting with GPU Cloud may make more sense.

Model Size and VRAM Sizing

The first step in sizing an LLM server is fitting the model into GPU memory (VRAM). Rough math for model weights: parameter count × bytes per parameter. 2 bytes in FP16/BF16, 1 byte with 8-bit quantization, about 0.5 bytes with 4-bit.

Model sizeFP16 / BF168-bit4-bit
7–8B~16 GB~8 GB~4–5 GB
13–14B~28 GB~14 GB~7–8 GB
32–34B~64–68 GB~32–34 GB~16–18 GB
70B~140 GB~70 GB~35–40 GB

The table covers model weights only. More concurrent users and longer context require extra memory for the KV cache; in practice leave at least 20–30% headroom. For example, a 70B model in FP16 does not fit on a single 80 GB GPU and is split across several GPUs.

Choosing the GPU

The right GPU depends on model size, concurrent users and whether you run inference or training. Common NVIDIA data center GPUs:

  • NVIDIA L4 (24 GB): small models, computer vision and video analytics at low power.
  • NVIDIA L40S (48 GB): inference and fine-tuning for mid-size models, plus visual workloads.
  • NVIDIA H100 (80 GB): large models, high-throughput inference and training.
  • NVIDIA H200 (141 GB HBM3e): more memory for larger models and long context.
  • Blackwell generation (B200, RTX PRO 6000 Blackwell Server Edition 96 GB): next-generation training and inference.

Availability, lead times and pricing change quickly, so the GPU model is finalized together at the proposal stage.

Server Architecture, Networking and Power

  • PCIe or SXM (HGX): PCIe cards are flexible and economical; SXM-based HGX platforms provide much higher GPU-to-GPU bandwidth over NVLink, which matters for large models.
  • CPU, RAM and storage: enough CPU cores, system memory and fast NVMe storage for data preparation and model loading.
  • Networking: 25/100 GbE may be enough for a single server; multi-node training needs 200–400 Gb/s InfiniBand or RoCE.
  • Power and cooling: an 8-GPU server can draw around 10 kW. Rack power capacity, cooling and UPS must be validated before deployment.
  • Placement: the server can live in your own server room or in a colocation facility.

Software Layer: Inference, RAG and Monitoring

  • Base layer: Linux (Ubuntu or RHEL), NVIDIA drivers and CUDA; NVIDIA GPU Operator on Docker or Kubernetes.
  • Inference engine: vLLM for high concurrency, TensorRT-LLM and Triton for peak performance on NVIDIA, Ollama for small setups.
  • Internal API: an OpenAI-compatible endpoint with authentication and quotas so applications connect without changes.
  • RAG: a vector database such as pgvector, Qdrant or Milvus and permission-aware retrieval to connect your documents to the model.
  • GPU sharing: on GPUs such as the H100, MIG (Multi-Instance GPU) splits one GPU into several isolated instances.
  • Monitoring: NVIDIA DCGM for GPU utilization, temperature, memory and errors.

Who Is This Service For?

Organizations that want to use LLMs without sending data outside, teams building an internal assistant (RAG) over documents and knowledge bases, R&D units training or fine-tuning their own models, and manufacturing and security companies running computer vision projects.

Scope and Deliverables

  • Workload, model and user count analysis
  • VRAM, GPU and server sizing report
  • Power, cooling and network pre-check
  • Hardware procurement, deployment and software layer
  • Performance testing (tokens/second, concurrency)
  • Operations documentation and handover

Example Scenario

A hypothetical law firm wants to search and summarize its contract archive with an LLM, but documents must not leave the firm. A 4-bit quantized 70B model runs with vLLM on a server with two 48 GB GPUs; documents are retrieved through RAG according to user permissions.

What Drives the Price

  • GPU model and count
  • Model size and concurrent users
  • PCIe or HGX platform
  • Storage and network needs
  • Power, cooling and hosting
  • Software and support scope

Pre-Purchase Checklist

  • Which model or model family, and how many parameters?
  • How many concurrent users or requests are expected?
  • Inference only, or fine-tuning and training as well?
  • Is there enough power and cooling in the server room or rack?
  • Does the open model's license allow commercial use?
  • Is a logging and retention policy for prompts and outputs defined?

Frequently Asked Questions

Which GPU should we choose for an LLM server?

It depends on model size, concurrent users and workload type. L4 or L40S for small and mid-size models, H100 or H200 for large models and high throughput, and Blackwell platforms for next-generation and the largest workloads.

How much VRAM does a 70B model need?

For weights alone, about 140 GB in FP16, 70 GB in 8-bit and 35–40 GB in 4-bit. KV cache for concurrent users and context length comes on top, so leave at least 20–30% headroom.

Our own GPU server or GPU Cloud?

If utilization is sustained and high and data must stay inside, your own server can be more economical long term. For pilots, periodic training or uncertain demand, GPU Cloud lets you start without investment.

Does an on-premises LLM help with data protection compliance?

Because data stays in your infrastructure, it reduces cross-border transfer and third-party processing risk. Access rights, logging, retention and deletion policies still need to be designed alongside technical safeguards.

Which open models can we use?

Open-weight families such as Llama, Qwen, Mistral and Gemma are widely used. License terms differ per model; check commercial use and user limits before deployment.

RAG or fine-tuning?

For Q&A over company documents, RAG is usually enough and needs no retraining when documents change. Fine-tuning is used to change a model's style, format or behavior on a specific task.

How much power does a GPU server use?

It depends on GPU count and model; a high-performance 8-GPU server can draw around 10 kW. Plan rack power capacity, cooling and UPS before deployment.