All systems operational
Support
EN
Language

More languages are on the way.

GPU servers arrive in Singapore

H100 and RTX inventory in our APAC hub, the same per-hour-to-monthly pricing, and why latency to Tokyo, Sydney and Mumbai matters for inference.

CPCheapServ ProductWritten by 3 min read
A GPU card above a map of South-East Asia
On this page5
  1. What is available
  2. Why latency matters for inference
  3. Pricing
  4. The hub around it
  5. What comes next

Singapore has been our APAC hub since 2022 for VPS and dedicated servers. From today it also carries GPU inventory: H100 SXM nodes and RTX 4090 servers, on the same monthly pricing as Frankfurt and New York, with the same templates. The reason is latency, and the reason latency matters for GPUs is inference.

What is available

  • H100 SXM 80 GB, single-card and 2× configurations, with 256 GB and 512 GB of host RAM.
  • RTX 4090 24 GB, single-card, for image generation, fine-tuning small models and rendering.
  • Every template in the order form: vLLM, TGI, Ollama, PyTorch, ComfyUI, JupyterLab and the rest, on Ubuntu 24.04 with CUDA 12 and the current driver.

Stock is honest as always: the counts on the pricing page are the free pool, and “sold out” means sold out rather than “ask us”.

Why latency matters for inference

A chat completion streams tokens as they are generated, and the user experiences two delays: time to first token and the network round trip on every chunk. From Tokyo, Sydney or Mumbai, a model served in Frankfurt adds 200 to 280 ms of round trip before the first token appears, and every subsequent chunk carries the same delay. The same model in Singapore is 50 to 90 ms away from those cities. For an interactive product, that is the difference between a response that feels live and one that feels like a page load.

Data residency points the same way. Customers serving Singapore, Indonesia, India or Australia often need the model and the prompts to stay in the region, and a GPU in Frankfurt does not satisfy that whatever its speed.

Pricing

Identical to the other sites. The pricing rule compares each plan against the cheapest credible provider for the same hardware, and in APAC that comparison lands at a higher number than in Europe; we charge the European price anyway, because a customer should pick a location for their users, not for our cost structure.

The hub around it

Singapore peers at the Equinix exchange and has three transit providers, the same design as Frankfurt and New York. DDoS scrubbing capacity is local. The looking glass shows the paths, and the location guide has the round-trip numbers to the cities that matter for APAC products.

What comes next

RTX 5090 inventory follows as soon as the cards clear the same burn-in the Frankfurt batch went through. Larger H100 configurations, 4× and 8×, arrive with the next expansion of the hall; if you need one before then, a ticket with the timeline gets you a straight answer.

CP
CheapServ Product

Pricing, plans, locations and everything customers see first.

Deploy your first server in under a minute.

Top up from $25 in BTC, ETH, XMR or USDT. Your balance never expires and unused funds are refundable.

Sign up now