To estimate GPU needs, first work out how many GPUs are needed to run one copy of the model service, called a replica. Then test how many requests that copy can handle within the required response time. Use the model, numerical precision, input and output lengths, and peak request volume you expect in practice. Calculate how many copies are needed at peak demand, then add spare capacity to keep the service running if a copy fails.
Step 1: How many GPUs are needed to fit the model?
The model's size and the number format used to store its values determine the minimum memory for one replica. Each replica can receive requests independently and may run on one or several GPUs. To estimate memory for the model's learned values (weights), multiply its parameter count by the bytes used for each parameter. The formats BF16 and FP16 use about 2 bytes per parameter, FP8 about 1 byte, and INT4 about 0.5 byte.
Running the model also needs memory for information reused during text generation (the KV cache), temporary calculations, data waiting to move between devices, and a safety margin. Total memory use depends on input context length, the number of items processed together, simultaneous requests, and the software and architecture used to run the model.
Step 2: How much work must be handled at peak?
Monthly token totals help estimate cost. To size the system, focus on requests arriving at peak times, how many run at once, and input and output lengths. Also set a target for the wait until the first output token (Time to First Token, or TTFT) and the output speed (Tokens Per Second, or TPS). Average simultaneous requests can be estimated by multiplying requests per second by average processing time in seconds. This relationship, called Little’s Law, gives a starting point for testing rather than a final GPU purchase quantity.
Step 3: Convert benchmark results into GPU quantity
Test the actual model, number format, input and output lengths, model-running software, GPU configuration, and simultaneous request load. Record requests and output tokens per second, success rate, time to first token, output speed, time spent waiting in the queue, and the response time that 95% of requests meet (P95 latency). Divide peak demand by one replica’s tested capacity and round up. Multiply by the GPUs needed per replica, then add spare capacity for failures.
Example
If peak demand is 6 requests per second and one two-GPU replica sustains 2 requests per second at the target latency, three replicas—or six GPUs—provide base capacity. If one replica may fail while peak demand is maintained, deploy four replicas, or eight GPUs.
Why training requires a separate estimate
Training needs extra memory for model updates (gradients and optimizer states), intermediate results (activations), and temporary data. Methods that train a small set of added parameters, such as LoRA or QLoRA, have different needs from updating all parameters or training a model from scratch. Run a small-scale test of the intended method to estimate the final configuration.
KONST GPU deployment options
KONST Group provides bare-metal GPUs and GPU clusters, while Konstra AI supports AI data center construction, infrastructure operations, and compute delivery. Enterprises can choose rented capacity, dedicated infrastructure, or a private deployment according to benchmark results, usage duration, control requirements, and growth plans.
FAQ
How many GPUs does a 70B model need?
It depends on the number format, GPU memory, input context length, simultaneous requests, and required response time. For a model with 70 billion parameters, the model weights alone theoretically need about 140 GB in BF16, 70 GB in FP8, or 35 GB in INT4. Allow extra memory for the KV cache and for running the model safely.
Can monthly token volume determine GPU count?
No. Monthly volume supports cost planning; capacity depends on peak arrival patterns and response targets.
Why add GPUs if the model fits in memory?
Fitting the model only proves it can run. More copies of the service may be needed to handle the request volume, respond quickly enough, and keep running during failures.
- GPU
- AI Infrastructure
- Capacity Planning
- Inference
KONST Editorial Team
AIDC Engineering



