Servers that arrive ready.
Done-for-you high-performance VPS and compute provisioning. Our engineers hand-configure your kernel, firewall, Docker runtime, and DGX Spark unified memory, then deliver working credentials to your dashboard in 6–8 hours.

Why spend days configuring servers when we deliver them ready to run?
Standard cloud providers give you an empty Linux prompt and charge you for every support ticket. We take responsibility for installation, hardening, Docker environment, and initial configuration so your team can deploy code immediately.

Infrastructure Standard Inclusions
EVERY INSTANCE INCLUDESDone-for-You Setup (Our Moat)
Skip unmanaged raw images, broken driver packages, and hours of debugging network routing. Every server is verified running and hardened before delivery.
DGX Spark 128 GB @ 273 GB/s
Massive unified memory pool for prototyping, fine-tuning large MoE models (gpt-oss-120B, Qwen3-Next-80B), and high-concurrency batch inference pipelines.
99.9% Uptime with Pro-Rata Credits
Contractually backed availability guarantee. If monthly uptime drops below 99.9%, you receive proportional service credits automatically applied to your account.
Backups Managed by Our Team
Full system state and configuration backups are maintained by our operations team, ensuring rapid recovery without managing complex backup scripts.
In-Dashboard Ticket Support
Direct support tickets handled by systems engineers inside your customer portal. No lost emails, no WhatsApp chaos, and full issue audit history.
Audited Metric Snapshots
Verifiable daily telemetry snapshots stamped with captured_at timestamps. Clean, truthful historical sparklines without fake live gauge animations.
DGX Spark: 128 GB unified LPDDR5X at 273 GB/s
We believe in complete architectural transparency. DGX Spark delivers a massive 128 GB unified memory pool at 273 GB/s bandwidth. It is bandwidth-bound by design — making it an exceptional fit for specific ML workloads, and a poor fit for others.

Where DGX Spark Wins
MoE Architecture Inferences
Mixture-of-Experts models like gpt-oss-120B and Qwen3-Next-80B run at ~45–60 tok/s because only a fraction of sparse parameters activate per token.
Massive Batch & Agent Pipelines
Aggregate throughput scales to ~863 tok/s across 256 concurrent streams (~26x single-stream), ideal for automated background agents, scraping extractors, and bulk evaluation.
Prototyping & LoRA / QLoRA Fine-Tuning
Fit large model weights, KV caches, and adapter checkpoints into 128 GB @ 273 GB/s of unified memory that would otherwise trigger Out-Of-Memory errors on standard 24GB or 32GB cards.
What NOT to Use It For
Not for Interactive Dense 70B
A dense 70B model fits into memory, but decodes at ~2.7 tok/s due to the 273 GB/s memory bandwidth limit. We will not sell it for real-time single-user dense 70B chat.
Not Foundational Training Hardware
DGX Spark is for fine-tuning adapters (LoRA/QLoRA) and batch inference. It is not designed or sold as full pre-training hardware.
Single-Session Chat Latency
For single-stream interactive chat, an RTX 5090 with high memory bandwidth beats it ~4:1. We recommend DGX Spark strictly for batch, MoE, and fine-tuning.
Geographic Transparency
USA GPUs · European Compute Servers
Our DGX Spark GPU clusters operate in US Tier-3 datacenters while primary host VPS compute nodes are provisioned in Europe (~80–100 ms transit apart).
No "Low-Latency GPU Attach" Claims
We never falsely claim microsecond GPU interconnects. This topology is engineered specifically for bulk batch pipelines, async fine-tuning jobs, and distributed agent workloads.
Dedicated High-Throughput Routing
Optimized for multi-gigabit bulk transfer of datasets, model checkpoints, and batch inference requests with automatic error recovery and retry queues.
Hardware & Throughput Benchmark Summary
Verified performance metrics across supported model architectures
| Model / Workload | Architecture | Throughput | Usability Verdict |
|---|---|---|---|
| gpt-oss-120B | MoE (Sparse Activation) | ~45–60 tok/s | Optimal for Production |
| Qwen3-Next-80B | MoE (Sparse Activation) | ~50–65 tok/s | Optimal for Production |
| Concurrent Batch (256 streams) | Parallel Inferences | ~863 tok/s (aggregate) | 26x Single-Stream Gain |
| Dense 70B (Single-Stream) | Dense Architecture | ~2.7 tok/s | Too slow for interactive chat |
| LoRA / QLoRA Fine-Tuning | 128 GB @ 273 GB/s Unified Checkpoints | Fits 70B+ weights | Full Memory Headroom |
Ready-to-work servers. Zero surprise invoices.
Every server is fully hand-configured by engineers and delivered within 6–8 hours with working credentials. All tiers include our binding 99.9% pro-rata uptime SLA.
Scale
Top TierFlagship compute power with DGX Spark 128 GB @ 273 GB/s unified memory for MoE and batch pipelines.
- 32 vCPU Dedicated Cores
- 128 GB ECC RAM
- 2 TB NVMe High-Speed Array
- 100 TB Monthly Bandwidth
- DGX Spark 128 GB @ 273 GB/s
- Priority 6–8 hour hand setup
- Managed by our team backups
- In-dashboard ticket support
- 99.9% Uptime SLA (Pro-Rata)
Growth
High-throughput compute and expanded NVMe storage for scaling backends, microservices, and databases.
- 16 vCPU Dedicated Cores
- 64 GB ECC RAM
- 1 TB NVMe Storage
- 50 TB Monthly Bandwidth
- Hand-provisioned in 6–8 hours
- Managed by our team backups
- In-dashboard ticket support
- 99.9% Uptime SLA (Pro-Rata)
- DGX Spark 128 GB @ 273 GB/s
Starter
Base ComputeDedicated high-performance compute node for lightweight workloads, background services, and proxies.
- 8 vCPU Dedicated Cores
- 32 GB ECC RAM
- 500 GB NVMe Storage
- 20 TB Monthly Bandwidth
- Hand-provisioned in 6–8 hours
- In-dashboard ticket support
- 99.9% Uptime SLA (Pro-Rata)
- DGX Spark 128 GB @ 273 GB/s
- Managed by our team backups
How hostnetvps compares to raw VPS and hyperscale clouds
See the practical differences between our done-for-you model, bare unmanaged servers, and complex cloud platforms.
| Feature / Capability | hostnetvps.com (Done-For-You) | Raw Unmanaged VPS | Hyperscale Clouds (AWS / GCP) |
|---|---|---|---|
| Provisioning & Setup | Hand-provisioned in 6–8h (Ready to deploy) | Unmanaged empty OS (Hours of manual setup) | Complex IAM, VPC, and Cloud-init scripts |
| Large Model Memory | 128 GB @ 273 GB/s unified memory (Scale) | Limited to host RAM; no unified GPU pool | Hourly GPU instances ($3–$8/hr dynamic billing) |
| Backup Management | Managed by our team | Customer writes and tests own backup scripts | Billed per GB snapshot + transfer fees |
| Support Channel | Direct in-dashboard ticket system | Community forums or unmanaged / zero support | $100–$1,000/mo enterprise support contracts |
| Uptime SLA | 99.9% Uptime with pro-rata service credits | Best-effort with no financial SLA | Complex multi-page credit claim bureaucracy |
| Pricing Structure | Predictable flat monthly billing | Flat compute, but hidden engineer time cost | Surprise bandwidth, IOPS, and egress fees |
| Credential Security | AES-256-GCM encrypted vault with re-auth & audit logs | Plain text root passwords sent via email | Separate Secret Manager billable service |
Clear answers about our hardware & operations
Everything you need to know about our hand-provisioning guarantee, DGX Spark specifications, and SLA terms.
How does the 6–8 hour done-for-you provisioning work?
Once you complete checkout, your order enters our engineering queue. An infrastructure engineer personally provisions your dedicated hardware, configures Linux kernel parameters, applies SSH and firewall security hardening, installs the container runtime, and attaches your DGX Spark memory module (on Scale tiers). Within 6–8 hours, your working credentials appear encrypted in your dashboard.
How are backups handled?
Backups are managed by our team. Our operations staff maintains configuration and system state safeguards so your servers can be recovered swiftly in the event of hardware failure.
What is your 99.9% Uptime SLA and credit policy?
We guarantee 99.9% monthly network and server availability. If availability falls below this threshold during any calendar month, you are eligible for pro-rata service credits calculated against the outage duration and applied to your subsequent billing invoice.
What are the exact hardware capabilities of DGX Spark 128 GB @ 273 GB/s?
DGX Spark features 128 GB of unified LPDDR5X memory running at 273 GB/s bandwidth. It is optimized for large Mixture-of-Experts (MoE) models (such as gpt-oss-120B and Qwen3-Next-80B running at ~45–60 tok/s) and high-concurrency batch inference pipelines (~863 tok/s aggregate across 256 streams). Because bandwidth is 273 GB/s, running single-session dense 70B models decodes at ~2.7 tok/s (not suitable for real-time interactive chat). It is designed for fine-tuning, prototyping, and batch agents — not full foundational training.
Where are the servers and GPUs physically located?
Our compute nodes are hosted in European datacenters, while DGX Spark GPU clusters operate in US Tier-3 facilities (~80–100 ms network transit apart). This architecture is engineered for bulk batch pipelines, async fine-tuning, and background agent workloads, not synchronous low-latency GPU attach.
How do I securely access my server credentials?
Your server IP, root username, temporary password, and SSH keys are encrypted at rest using AES-256-GCM. To reveal them in the customer portal, you must re-authenticate with your account password. Every reveal event is logged in your account audit trail.
How do I get technical support?
Technical support is handled directly through the in-dashboard ticket form in your customer portal. This ensures every inquiry is tracked, auditable, and answered by infrastructure engineers.
What payment methods and regions do you accept?
We process payments securely via Stripe. We serve customers in the US, UK, Canada, Australia, New Zealand, and Europe. Billing is a transparent, flat monthly recurring charge with no hidden setup fees.