Malaysia’s first openly-priced RTX PRO 6000 Blackwell (96GB) GPU compute. Dedicated cards, real vCPUs, RM billing — no USD surprises, no setup fees, no lock-in. Built for DeepSeek, Qwen, GLM, Llama and every LLM you want to run yourself.
Seven reasons Malaysian teams pick us over foreign clouds and quote-only local providers.
Not “contact us for a quote”. Real RM prices on the page, today.
Published RM pricing belowEntry AI inference card at a price your budget can feel.
From RM1,811/moFPX, local cards, Maybank billing. No foreign-exchange shock on your invoice.
Billed & paid in MYRSimple monthly billing. Cancel when you want.
Others charge RM500–1,599 setupWe install Ollama / vLLM / DeepSeek / Qwen and hand you an OpenAI-compatible API the same day.
Free managed deploy add-onYour AI server, domain, SSL and backups under a single MYR invoice.
Bundled servicesMalaysian engineers on WhatsApp, not a ticket queue in another timezone.
wa.link/wn5vxtWhole-card GPU passthrough with dedicated vCPUs — predictable performance for production AI, no noisy neighbours.
Every model below is open-weight — you own the weights, we provide the silicon. No per-token API bills, no rate limits; your data stays yours.
96GB Blackwell runs Qwen2.5-72B at FP8/INT4 on a single GPU.
Per card on 49B-class models (FP4, 100 concurrent) — real serving, not marketing.
Blackwell generation vs previous-gen RTX 4000 Ada-class cards.
Compared with comparable GPU plans — the LLM-serving cost king.
Dedicated GPU + dedicated vCPUs. Predictable performance for production APIs.
Outbound transfer costs far less than hyperscaler per-GB fees.
Transparent Ringgit pricing. No setup fees, no lock-in — simple monthly billing.
| Plan | Specs | RM/mo |
|---|---|---|
| RTX PRO 6000 Blackwell x1 | 176GB RAM · 16 vCPU · 1TB NVMe | RM8,616 |
| RTX PRO 6000 Blackwell x2 | 352GB RAM · 32 vCPU · 2TB NVMe | RM17,233 |
| RTX PRO 6000 Blackwell x4 | 704GB RAM · 64 vCPU · 4TB NVMe | RM34,466 |
| Plan | Specs | RM/mo |
|---|---|---|
| RTX 4000 Ada x1 Small | 16GB RAM · 4 vCPU · 500GB | RM1,811 |
| RTX 4000 Ada x1 Medium | 32GB RAM · 8 vCPU · 500GB | RM2,308 |
| RTX 4000 Ada x1 Large | 64GB RAM · 16 vCPU · 500GB | RM3,302 |
| RTX 4000 Ada x1 X-Large | 128GB RAM · 32 vCPU · 500GB | RM5,289 |
| RTX 4000 Ada x2 Small | 32GB RAM · 2x GPU · 1TB | RM3,622 |
| RTX 4000 Ada x2 Medium | 64GB RAM · 2x GPU · 1TB | RM4,616 |
| RTX 4000 Ada x4 Small | 128GB RAM · 4x GPU · 2TB | RM10,226 |
| RTX 4000 Ada x4 Medium | 196GB RAM · 4x GPU · 2TB | RM12,337 |
| Plan | Specs | RM/mo |
|---|---|---|
| RTX 6000 Quadro x1 | 32GB RAM · 8 vCPU · 640GB · 16TB transfer | RM5,175 |
| RTX 6000 Quadro x2 | 64GB RAM · 2x GPU · 1.28TB · 20TB transfer | RM10,350 |
| RTX 6000 Quadro x3 | 96GB RAM · 3x GPU · 1.92TB · 20TB transfer | RM15,525 |
| RTX 6000 Quadro x4 | 128GB RAM · 4x GPU · 2.56TB · 20TB transfer | RM20,700 |
DeepSeek, Qwen, GLM, Kimi, Llama — a quick-fit guide for Malaysia’s AI builders.
Up to 14B models comfortably:
From RM1,811/mo — usually enough for most startups.
30B-class at Q4 — the reasoning sweet spot:
Flagship reasoning quality on a single card — from RM5,175/mo.
Flagship-class on ONE card:
From RM8,616/mo. Multi-GPU clusters (2–8x) for 284B–1.6T models.
For the giants — talk to us, we’ll size it:
Free sizing consultation — WhatsApp us your model & workload.
| Open-weight model | Params | VRAM at Q4 | Best Big Domain tier | Notes |
|---|---|---|---|---|
| DeepSeek-R1-Distill-Qwen-7B | 7B | ~5GB | RTX 4000 Ada (20GB) | Full FP16 even; fast chat |
| DeepSeek-R1-Distill-Qwen-14B | 14B | ~10GB | RTX 4000 Ada (20GB) | Q4 comfortable |
| DeepSeek-R1-Distill-Qwen-32B | 32B | ~19GB | RTX 6000 Quadro (24GB) | Q4 sweet spot |
| DeepSeek-V4-Flash | 284B MoE (13B active) | ~120–150GB | 2x Blackwell | MIT · 1M ctx · FP4+FP8 native |
| DeepSeek-V4-Pro | 1.6T MoE (49B active) | ~700–870GB | 8x Blackwell cluster / API | MIT · 1M ctx · frontier-class |
| DeepSeek-R1-0528 (full) | 671B MoE | ~400GB | 4–8x Blackwell | Reasoning; updated May 2026 |
| Qwen3.5-9B | 9B | ~5GB | RTX 4000 Ada (20GB) | Apache 2.0 · multimodal |
| Qwen3.5-27B | 27B dense | ~16GB | RTX 6000 Quadro (24GB) | Q4 |
| Qwen3.5-35B-A3B | 35B MoE | ~19GB | RTX 6000 Quadro (24GB) | Q4 |
| Qwen3.5-122B-A10B | 122B MoE | ~65GB | RTX PRO 6000 Blackwell (96GB) | Flagship on ONE card (Q4) |
| Qwen3.5-397B-A17B | 397B MoE | ~210–400GB | 3–5x Blackwell / API | 256K ctx · Apache 2.0 |
| Qwen2.5-72B | 72B | ~45GB | RTX PRO 6000 Blackwell (96GB) | FP8 / INT4 single card |
| GLM-4-9B | 9B | ~6GB | RTX 4000 Ada (20GB) | INT4 |
| GLM-5-Air | ~67B active | ~85GB | RTX PRO 6000 Blackwell (96GB) | Q4 single card |
| GLM-5 | 744B MoE (40B active) | ~380–456GB | 4x Blackwell / API | 200K ctx · MIT |
| Kimi K2.7 | ~1T-class MoE | ~325GB (2-bit) | 4x Blackwell (Q2) / API | Practical local Kimi |
| Kimi K3 | 2.8T MoE (16/896 experts) | ~350–700GB (Q4) | 8x Blackwell (Q4) / API | 1M ctx · native MXFP4 ~1.4TB |
| Llama 3.1 / 3.3 8B | 8B | ~5GB | RTX 4000 Ada (20GB) | — |
| Llama 3.3 70B | 70B | ~41GB | RTX PRO 6000 Blackwell (96GB) | Q4/Q8 |
From enquiry to a working OpenAI-compatible API — same day.
WhatsApp us or send the enquiry — your workload, your model, your budget.
Free consultation picks the right card — or cluster — for your LLM.
Free managed setup: Ollama / vLLM / DeepSeek / Qwen, OpenAI-compatible API, same day.
SSH/API access. Upgrade from 1 to 8 cards without re-provisioning from scratch.
Domain, SSL, backups and GPU on a single Malaysian invoice. Cancel anytime.
Get a free sizing consultation — we’ll tell you exactly which card (or cluster) your model needs.