Back to all GPUs

NVIDIA L40S Cloud GPU Specs

HexGrid Cloud rents the NVIDIA L40S at $1.09 per GPU hour, billed per minute. The L40S carries 48GB of GDDR6 memory on the Ada Lovelace architecture.

48GB VRAMAda LovelaceDeploys in ~1 minSecure CloudUpdated

HexGrid Cloud price

$1.09/hr

Billed per minute

VRAM

48GB

GDDR6

FP32

91.6 TFLOPS

Compute performance

Market average

$1.09/hr

Other clouds, list price

Running the L40S on HexGrid Cloud

What you get when you deploy with us, beyond the hourly rate.

Live in under a minute

Pick a region, hit deploy, and SSH in. No quota requests, no sales call.

Per-minute billing

You pay for the minutes you use. Stop the instance and billing stops with it.

Ready for training day one

CUDA, PyTorch and the usual drivers are preinstalled, or bring your own image.

No egress charges

Move checkpoints and datasets out without a surprise bandwidth bill.

L40S pricing across the market

Published on-demand rates from other clouds, so you can see where our L40S price sits.

Market average $1.09/hr across 1 cloud

  • HexGrid CloudYou're here$1.09/hr
  • RunPod$1.09/hr

Other providers' rates are their published on-demand list prices, last checked Sep 15, 2026. They exclude committed-use discounts and can change at any time. Shown for reference only.

About the NVIDIA L40S

What the L40S is built for, and where it falls short.

The NVIDIA L40S is a 48GB datacenter GPU based on NVIDIA Ada Lovelace architecture. It combines high FP32 performance with fourth-generation Tensor Cores and FP8 support, making it suitable for AI inference, model fine-tuning, image and video generation, rendering, and GPU-accelerated virtual workstations.

Best suited for

The L40S is particularly strong for inference and generative AI workloads that benefit from 48GB of VRAM and FP8 Tensor Core acceleration.

  • LLM inference
  • LLM fine-tuning
  • Image generation
  • Video generation
  • 3D rendering
  • Virtual workstations

Strengths

  • 48GB of ECC GPU memory
  • Strong FP8 and FP16 Tensor performance
  • Excellent generative image and video performance
  • High rendering performance
  • AV1 hardware encoding

Limitations

  • No NVLink support
  • No MIG support
  • Less suitable for very large distributed training jobs than H100-class accelerators

NVIDIA L40S specifications

Full technical specifications for the NVIDIA L40S.

VendorNVIDIAArchitectureAda LovelaceChipAD102CategoryDatacenter GPUReleased2023Form factorDual-slotInterfacePCIe 4.0 x16

Memory

Memory
48 GB
Type
GDDR6
Bandwidth
864 GB/s
Bus width
384-bit
ECC
Yes

Compute

CUDA cores
18,176
Tensor cores
568
Tensor generation
Gen 4
RT cores
142
Boost clock
2,520 MHz

AI & compute performance

FP32
91.6 TFLOPS
TF32 Tensor
183 TFLOPS / 366 TFLOPS sparse
BF16 Tensor
362.05 TFLOPS / 733 TFLOPS sparse
FP16 Tensor
362.05 TFLOPS / 733 TFLOPS sparse
FP8 Tensor
733 TFLOPS / 1,466 TFLOPS sparse
INT8 Tensor
733 TOPS / 1,466 TOPS sparse
RT
212 TFLOPS

Power & physical

TDP
350 W
Power connector
16-pin
Form factor
Dual-slot
Cooling
Passive
Length
267 mm

Interconnect

Interface
PCIe 4.0 x16
NVLink
No
Infinity Fabric
No

Virtualization

MIG
No
vGPU
Yes

Media engines

NVENC engines
3
NVDEC engines
3
AV1 encode
Yes
AV1 decode
Yes

Software

CUDA
Yes
ROCm
No

Precision support

FP64
No
FP32
Yes
TF32
Yes
BF16
Yes
FP16
Yes
FP8
Yes
INT8
Yes

L40S frequently asked questions

Common questions about NVIDIA L40S pricing, specifications and workloads on HexGrid Cloud.

How much VRAM does the NVIDIA L40S have?

The NVIDIA L40S has 48GB of GDDR6 ECC memory with 864 GB/s of memory bandwidth.

How much does an L40S cost per hour?

Cloud L40S pricing varies by provider and region. HexGrid compares currently available offers on this page, with the lowest available hourly price shown above.

Is the L40S good for LLM inference?

Yes. Its 48GB VRAM, FP8 support, and fourth-generation Tensor Cores make the L40S well suited to LLM inference and many fine-tuning workloads.

Does the NVIDIA L40S support NVLink?

No. The NVIDIA L40S uses PCIe 4.0 x16 and does not support NVLink.

Does the L40S support FP8?

Yes. The L40S supports FP8 Tensor Core operations, which can be particularly useful for supported AI inference and training workloads.

Other GPUs on HexGrid Cloud

If the L40S isn't the right fit, these are also available to deploy.

Browse every GPU we rent

Start training on the L40S today

$1.09 per GPU hour, billed per minute, running in about a minute.

Specification sources