The NVIDIA A40 is the data-center version of the RTX A6000: 48 GB of ECC memory in a passively cooled card designed for servers. It is widely used for AI inference, rendering and virtual workstations.

In this 2026 guide we cover the NVIDIA A40’s specs, what it can (and cannot) do today, where you can still rent it, and which GPU to choose instead if you need more power for AI.
Table of Contents
NVIDIA A40 Specifications
| Spec | NVIDIA A40 |
|---|---|
| Architecture | Ampere (GA102), 2020 |
| CUDA cores | 10,752 |
| Tensor cores | 336 (3rd gen) |
| Memory | 48 GB GDDR6 with ECC |
| Memory bandwidth | 696 GB/s |
| TDP | 300 W (passive, data center) |
Is the NVIDIA A40 Still Worth Renting in 2026?
Yes. 48 GB of VRAM at around $439–$473/month makes it one of the more affordable ways to run 70B-class LLMs in 4-bit or large rendering jobs.
What it is good for
- 70B LLMs in 4-bit
- AI inference with many concurrent users
- Rendering and VDI
- Video generation models
AI workloads on the NVIDIA A40
70B LLMs in 4-bit (Llama 3.3 70B, DeepSeek-R1-Distill 70B), 30B models in 8-bit and high-resolution image and video generation.
Best NVIDIA A40 Hosting Providers
1. GPU Mart

- Dedicated GPU servers and GPU VPS in US data centers
- Enterprise Dedicated A40 (48 GB), 256 GB RAM
- Windows or Linux with full admin access
- A40 $439/mo
Pros
- Very low monthly prices
- Dedicated hardware, no noisy neighbours
- One-click AI apps (Ollama, Stable Diffusion)
Cons
- Monthly billing only
- US locations only
GPU Mart offers a dedicated A40 server for $439/month.
2. Cherry Servers

- Bare-metal servers with data-center GPUs
- Hourly or monthly billing
- API and Terraform automation
- A40 $0.801/hr · $472.66/mo (GPU add-on)
Pros
- True bare metal
- EU, US and Asia locations
- Automation friendly
Cons
- Limited stock on some GPUs
- No consumer RTX cards
Cherry Servers adds an A40 to bare-metal servers for $0.801/hour or $472.66/month (plus the base server).
3. Vast.ai

- Marketplace with 68+ GPU types
- Per-second billing, interruptible discounts
- Docker templates for AI and gaming workloads
- Market pricing: RTX 4090 from ~$0.30/hr · A100 80GB from ~$0.43/hr · interruptible 50%+ cheaper
Pros
- Often the cheapest hourly prices
- Huge choice of consumer GPUs
- No long-term commitment
Cons
- Quality varies by host
- Availability of older cards changes daily
A40s appear on Vast.ai from data-center hosts at marketplace prices.
NVIDIA A40 vs Alternatives (Monthly Prices)
If your workload needs more speed or VRAM, these are the closest upgrades available as hosted servers today:
| GPU | VRAM | From | Notes |
|---|---|---|---|
| RTX A6000 | 48 GB | $409/mo | Same chip, workstation version |
| L40S | 48 GB | $1.80/hr (OVHcloud) | Newer Ada, FP8 |
| A100 80GB | 80 GB | $1,559/mo | More VRAM, faster memory |
How to Choose
- Match VRAM to your AI model. 8 GB for 7–8B LLMs, 16 GB for 14–24B, 24 GB for 27–32B, 48 GB for 70B in 4-bit.
- Hourly vs monthly. For tests and short jobs use RunPod or Vast.ai. For 24/7 use, compare the monthly total (hourly price × 730 hours) with a fixed-price monthly server from GPU Mart or HOSTKEY, which includes the whole machine, storage and bandwidth.
- Dedicated vs shared. Dedicated servers give stable performance; marketplace GPUs are cheaper but vary by host.
- Location. Pick EU hosting (HOSTKEY, OVHcloud) for GDPR-sensitive data.
FAQ
GPU Mart offers a dedicated A40 for $439/month; Cherry Servers adds it to bare metal for about $473/month.
They share the same GA102 chip and 48 GB. The A40 is passively cooled for data centers; performance is similar.
Conclusion
The A40 is a reliable 48 GB data-center GPU. GPU Mart offers the best monthly price at $439, while Cherry Servers suits automated bare-metal setups.
Prices were checked in September 2026 and change often. Always confirm current pricing on the provider’s website before you order.