Singapore has been our APAC hub since 2022 for VPS and dedicated servers. From today it also carries GPU inventory: H100 SXM nodes and RTX 4090 servers, on the same monthly pricing as Frankfurt and New York, with the same templates. The reason is latency, and the reason latency matters for GPUs is inference.
What is available
- H100 SXM 80 GB, single-card and 2× configurations, with 256 GB and 512 GB of host RAM.
- RTX 4090 24 GB, single-card, for image generation, fine-tuning small models and rendering.
- Every template in the order form: vLLM, TGI, Ollama, PyTorch, ComfyUI, JupyterLab and the rest, on Ubuntu 24.04 with CUDA 12 and the current driver.
Stock is honest as always: the counts on the pricing page are the free pool, and “sold out” means sold out rather than “ask us”.
Why latency matters for inference
A chat completion streams tokens as they are generated, and the user experiences two delays: time to first token and the network round trip on every chunk. From Tokyo, Sydney or Mumbai, a model served in Frankfurt adds 200 to 280 ms of round trip before the first token appears, and every subsequent chunk carries the same delay. The same model in Singapore is 50 to 90 ms away from those cities. For an interactive product, that is the difference between a response that feels live and one that feels like a page load.
Data residency points the same way. Customers serving Singapore, Indonesia, India or Australia often need the model and the prompts to stay in the region, and a GPU in Frankfurt does not satisfy that whatever its speed.
Pricing
Identical to the other sites. The pricing rule compares each plan against the cheapest credible provider for the same hardware, and in APAC that comparison lands at a higher number than in Europe; we charge the European price anyway, because a customer should pick a location for their users, not for our cost structure.
The hub around it
Singapore peers at the Equinix exchange and has three transit providers, the same design as Frankfurt and New York. DDoS scrubbing capacity is local. The looking glass shows the paths, and the location guide has the round-trip numbers to the cities that matter for APAC products.
What comes next
RTX 5090 inventory follows as soon as the cards clear the same burn-in the Frankfurt batch went through. Larger H100 configurations, 4× and 8×, arrive with the next expansion of the hall; if you need one before then, a ticket with the timeline gets you a straight answer.
