The Conversation That Happens in Every AI Infrastructure Meeting
IT Director: “We need GPUs for inference. Engineering wants H100s.”
Infrastructure Architect: “Do they need 80GB of HBM3 and FP8 Transformer Engine for serving a fine-tuned 7B model?”
IT Director: “…They said H100.”
Infrastructure Architect: “They need an L40S. Half the power budget, 30% of the cost, and it handles their inference workload without breaking a sweat. Let us spend the H100 budget on the training cluster where it actually matters.”
This conversation happens because the GPU market segments cleanly into training GPUs and inference GPUs — and confusing the two is the fastest way to overspend on AI infrastructure. The NVIDIA L40S is the inference workhorse: 48 ECC 付き GB GDDR6, PCIe Gen4 x16, and 300W TDP in a dual-slot form factor that fits in standard air-cooled servers.
技術仕様
| パラメータ | 仕様 |
|---|---|
| GPU アーキテクチャ | NVIDIA エイダ ラブレス |
| CUDAの色 | 18,176 |
| テンソルコア | 568 (4第 3 世代) |
| RTコア | 142 (3第 3 世代) |
| メモリー | 48 ECC 付き GB GDDR6 |
| メモリ帯域幅 | 864 GB/s |
| メモリバス | 384-少し |
| インタフェース | PCIe 4.0 x16 |
| FP32のパフォーマンス (non-Tensor) | 91.6 TFLOPS |
| TF32 Tensor Core | 183 TFLOPS (まばらさのある: 366 TFLOPS) |
| FP16 テンソルコア | 362 TFLOPS (まばらさのある: 733 TFLOPS) |
| INT8 テンソルコア | 733 トップス (まばらさのある: 1,466 トップス) |
| FP8 Tensor Core | 733 TFLOPS (まばらさのある: 1,466 TFLOPS) |
| TDP (マックス) | 300W |
| フォームファクタ | Dual-slot, フルハイト, full-length (FHFL), passive cooling |
| 電源コネクタ | 1x PCIe CEM 16 ピン (12VHPER) or 1x 8-pin CPU power |
| NVリンク / NVSwitch | サポートされていません (single GPU inference, no multi-GPU fabric) |
| MIG (Multi-Instance GPU) | サポートされていません (use A100 or H100 for MIG workloads) |
| vGPU のサポート | NVIDIA vWS, vPC, vApps — supported for VDI and virtualization |
| ECCメモリ | Full ECC on GDDR6 (not all GDDR6 GPUs offer ECC — the L40S does) |
| 物理的寸法 | 267 バツ 112 んん (デュアルスロット), 1.35 kg |
| 動作温度 | 0℃~45℃ |
| 保証 | 3-年 (メーカー) |
L40S vs H100 vs A100 — Which GPU for Your Workload?
| ワークロード | L40S | A100 80GB | H100 80GB |
|---|---|---|---|
| LLM inference (7B-13B models, FP16) | 素晴らしい. 48GB fits most 13B models. | Overkill unless running multiple models concurrently. | Overkill. FP8 inference is faster but not needed at this scale. |
| LLM inference (70B+ models) | Struggles. 48GB insufficient for 70B models without quantization. | Good with INT8 quantization. 80GB fits 70B models. | Best. FP8 + Transformer Engine for maximum throughput. |
| Full model training (70B+) | Not recommended. No NVLink, limited memory. | Acceptable for fine-tuning. Full training needs NVLink. | Best. NVSwitch + HBM3 for multi-GPU training. |
| Fine-tuning (LoRA, QLoRA, 7B-13B) | 素晴らしい. 48GB + FP16 is sufficient. | Good but more expensive than needed. | Overkill. Spend the savings on more L40S units. |
| VDI / Virtual Desktops | 素晴らしい. vGPU + 48GB supports 6-12 concurrent users. | Good but expensive per-user. | Not designed for VDI workloads. |
| Rendering / RTX workloads | 素晴らしい. RTコア + 18K CUDA cores. | RTコアなし. Not designed for rendering. | RTコアなし. |
| HPC シミュレーション (FP64) | Not recommended. No FP64 Tensor Cores. | 良い. FP64 Tensor Cores enabled. | Limited FP64 compared to A100. |
| 価格 (relative to H100) | ~25-30% | ~55-65% | 100% (ベースライン) |
互換性のあるサーバー プラットフォーム
The L40S is a PCIe 4.0 x16 dual-slot FHFL card compatible with any server that supports 300W GPU power delivery via 8-pin or 12VHPWR connectors. Confirmed compatible platforms include Dell PowerEdge R760xa, R770 (up to 6x L40S), R7715, xFusion G5500 V7, 2288H V7, and 2488H V7. Requires adequate chassis airflow — the passive cooling design depends on server fan speed profiles. Verify your server’s GPU enablement kit includes the correct power cables before ordering.
新川経由のソース
We supply L40S GPUs pre-installed in Dell PowerEdge and xFusion servers, or as standalone upgrades for existing platforms. 工場直送価格, 3-年保証, そして世界的な発送. Our AI infrastructure engineers can help you right-size your GPU fleet — training cluster vs inference cluster vs mixed workload — to maximize throughput per dollar.
新川サーバー | エンタープライズ サーバー ハードウェア サプライヤー











