CPU
Size cores for preprocessing, data delivery, and scheduling.
Quoted to your configuration4ROUTER / RESOURCES & PRICING
From a single GPU to a complete cluster. Start with the compute your workload needs, then shape the configuration, delivery, and support around it.
Discuss your workloadPLAN A WORKLOAD
Start with memory requirements and runtime to plan a configuration for inference, training, or batch workloads.
| GPU configuration | VRAM | Rate | Choose |
|---|---|---|---|
| 24 GB | $0.74 / h | ||
| 32 GB | $0.99 / h | ||
| 48 GB | $1.09 / h | ||
| 80 GB | $1.59 / h | ||
| 80 GB | $3.49 / h | ||
| 141 GB | $4.59 / h | ||
| 180 GB | $6.79 / h |
GPU rates are listed per GPU-hour. CPU, system memory, networking, and engineering services are quoted separately. Configuration and availability are confirmed before provisioning.
GPU prices are in USD per GPU-hour. The estimate is the listed rate × GPU count × runtime hours. Storage prices are in USD per GB-month and depend on capacity and the period of use.
GPU compute and storage are priced separately. CPU, system memory, networking, cluster interconnect, and engineering are quoted to configuration and delivery scope. Available resources, billing periods, and settlement terms are confirmed before provisioning.
COMPLETE THE INSTANCE
Data loading, preprocessing, and model execution place different demands on the host. CPU, system memory, and networking are quoted separately to the GPU configuration and workload.
Size cores for preprocessing, data delivery, and scheduling.
Quoted to your configurationPlan memory for model loading, datasets, and caches.
Quoted to your configurationMatch data throughput and multi-node communication requirements.
Quoted to your configuration4ROUTER / SYSTEMS RESEARCH
Understand how operators, KV state and scheduling shape an inference path.
Read the systems notes in Resources ↗Context length is a state budget that grows with concurrency.
↗TIERS / TRANSFERAvoiding prefill requires paying for state access.
↗SCHEDULING / SLOWhen throughput rises, know which requests are waiting.
↗KEEP DATA CLOSE
Plan storage for datasets, checkpoints, and runtime environments separately.
Network storage capacity under 1 TB
Container disk space for the runtime environment
Storage capacity is priced separately from GPU compute. Network transfer is quoted separately. Storage type, capacity, and location are confirmed before provisioning.
FROM RESOURCES TO DELIVERY
Scope the services that bring models, systems, and operating environments together.
MODEL / DATA / EVALUATION
From fine-tuning and alignment to kernel and inference optimization. Scope work around the model, data, and evaluation requirements.
Quoted to project scopeExplore the scopeCOMPUTE / FABRIC / SCHEDULING
Hardware selection, networking, deployment, scheduling, and monitoring. Plan around cluster size, interconnect, and operational responsibilities.
Quoted to project scopeExplore the scopeENVIRONMENT / ACCESS / DELIVERY
Deploy models and services in your environment. Define data, access, and delivery requirements from bare metal to a callable endpoint.
Quoted to project scopeExplore the scopeMODEL GATEWAY
Gateway access is billed per request, separately from GPU resource planning.
Explore gateway integrationActual upstream cost × account multiplier → per-request charge
The account multiplier is agreed during onboarding and retained with each usage record. Decimal arithmetic uses HALF_UP rounding to eight places; a non-zero upstream cost has a minimum charge of USD 0.00000001.
Usage, routing decisions, and ledger entries are linked by request_id. Balance and quota checks run before execution.
START WITH YOUR WORKLOAD
Model and dataset size, expected runtime, deployment environment, and delivery goals are the starting point. We will confirm the resource configuration and service scope with you.
ops@4router.cnSend your requirements