4ROUTER / RESOURCES & PRICING

Compute. Training. Infrastructure.

From a single GPU to a complete cluster. Start with the compute your workload needs, then shape the configuration, delivery, and support around it.

Discuss your workload

PLAN A WORKLOAD

GPU compute

Start with memory requirements and runtime to plan a configuration for inference, training, or batch workloads.

4Router resource pricingUSD / GPU-hour
GPU compute
GPU configurationVRAMRateChoose
24 GB$0.74 / h
32 GB$0.99 / h
48 GB$1.09 / h
80 GB$1.59 / h
80 GB$3.49 / h
141 GB$4.59 / h
180 GB$6.79 / h

GPU rates are listed per GPU-hour. CPU, system memory, networking, and engineering services are quoted separately. Configuration and availability are confirmed before provisioning.

Billing units and included costs

GPU prices are in USD per GPU-hour. The estimate is the listed rate × GPU count × runtime hours. Storage prices are in USD per GB-month and depend on capacity and the period of use.

GPU compute and storage are priced separately. CPU, system memory, networking, cluster interconnect, and engineering are quoted to configuration and delivery scope. Available resources, billing periods, and settlement terms are confirmed before provisioning.

COMPLETE THE INSTANCE

CPU and memory. Sized to the workload.

Data loading, preprocessing, and model execution place different demands on the host. CPU, system memory, and networking are quoted separately to the GPU configuration and workload.

CPU

Size cores for preprocessing, data delivery, and scheduling.

Quoted to your configuration

System memory

Plan memory for model loading, datasets, and caches.

Quoted to your configuration

Networking & interconnect

Match data throughput and multi-node communication requirements.

Quoted to your configuration

4ROUTER / SYSTEMS RESEARCH

Capacity planning: size KV before choosing resources

Understand how operators, KV state and scheduling shape an inference path.

Read the systems notes in Resources
STATE / CAPACITY

KV cache: from tensor shapes to capacity budgets

Context length is a state budget that grows with concurrency.

TIERS / TRANSFER

LMCache: placing KV state in a storage hierarchy

Avoiding prefill requires paying for state access.

SCHEDULING / SLO

From prefill to decode: designing the scheduling experiment

When throughput rises, know which requests are waiting.

KEEP DATA CLOSE

Storage

Plan storage for datasets, checkpoints, and runtime environments separately.

Network storage

Network storage capacity under 1 TB

$0.07 / GB·mo

Container disk

Container disk space for the runtime environment

$0.10 / GB·mo

Storage capacity is priced separately from GPU compute. Network transfer is quoted separately. Storage type, capacity, and location are confirmed before provisioning.

FROM RESOURCES TO DELIVERY

The infrastructure. And the engineering behind it.

Scope the services that bring models, systems, and operating environments together.

MODEL / DATA / EVALUATION

Training & tuning

From fine-tuning and alignment to kernel and inference optimization. Scope work around the model, data, and evaluation requirements.

Quoted to project scopeExplore the scope

COMPUTE / FABRIC / SCHEDULING

GPU clusters & HPC

Hardware selection, networking, deployment, scheduling, and monitoring. Plan around cluster size, interconnect, and operational responsibilities.

Quoted to project scopeExplore the scope

ENVIRONMENT / ACCESS / DELIVERY

Private deployment

Deploy models and services in your environment. Define data, access, and delivery requirements from bare metal to a callable endpoint.

Quoted to project scopeExplore the scope

MODEL GATEWAY

Using a hosted model API?

Gateway access is billed per request, separately from GPU resource planning.

Explore gateway integration
Explore gateway billing

Actual upstream cost × account multiplier → per-request charge

The account multiplier is agreed during onboarding and retained with each usage record. Decimal arithmetic uses HALF_UP rounding to eight places; a non-zero upstream cost has a minimum charge of USD 0.00000001.

Usage, routing decisions, and ledger entries are linked by request_id. Balance and quota checks run before execution.

START WITH YOUR WORKLOAD

Tell us what you want to run.

Model and dataset size, expected runtime, deployment environment, and delivery goals are the starting point. We will confirm the resource configuration and service scope with you.

ops@4router.cnSend your requirements