GPU Specifications / NVIDIA
NVIDIA T4
Turing datacenter GPU with 16 GB of GDDR6 memory, 320 GB/s of bandwidth and up to 65 TFLOPS of FP16 tensor compute.
Key Specifications
Memory
16 GB GDDR6
Memory Bandwidth
320 GB/s
TDP
70 W
Architecture
Turing
Interconnect
PCIe 3.0 · 32 GB/s
Est. On-demand Price
~$0.50/h
Hourly rates are indicative on-demand estimates; actual pricing varies by provider and commitment.
Compute Performance
| Precision | Peak throughput |
|---|---|
| FP64 (double precision) | 0.25 TFLOPS |
| FP32 (single precision) | 8.1 TFLOPS |
| FP24 | 16.2 TFLOPS |
| FP16 (tensor) | 65 TFLOPS |
| INT8 (tensor) | 130 TFLOPS |
| INT4 (tensor) | 260 TFLOPS |
Tensor figures use the vendor's peak numbers (with structured sparsity where supported).
System Requirements
Recommended CPU
Intel Xeon Silver 4214 or AMD EPYC 7302
Max VRAM per node (8 GPUs)
128 GB
System RAM (min / recommended)
64 / 128 GB
Minimum PSU
450 W
LLMs on the T4
Number of T4 GPUs needed to serve popular models at 8-bit quantization (including 20% overhead for activations and KV cache).
Frequently Asked Questions
How much VRAM does the NVIDIA T4 have?
The NVIDIA T4 has 16 GB of GDDR6 memory with 320 GB/s of memory bandwidth.
Which LLMs can run on a single T4?
At 8-bit quantization, a single T4 (16 GB) can serve models up to roughly 14B parameters, such as Phi 4 (14B). Larger models require multiple GPUs or more aggressive quantization.
How much does it cost to rent a NVIDIA T4?
On-demand cloud pricing for the T4 is around $0.50/hour, i.e. about $365/month running 24/7. Actual prices vary by provider, region, and commitment.
What are the power and system requirements of the NVIDIA T4?
The T4 has a TDP of 70W. A power supply of at least 450W per GPU is recommended. Recommended host CPUs: Intel Xeon Silver 4214 or AMD EPYC 7302.
Deploy on a GPU cloud
Rent the NVIDIA T4 by the hour instead of buying hardware.