Apple discontinued · rare configuration

Mac Studio M3 Ultra, 512 GB unified memory
— full specs & remote rental

The highest-memory Apple Silicon machine ever sold, and the only consumer hardware that holds 405B–1T-class open-weight models entirely in memory. Apple no longer sells it. Here, you can rent one.

Measured on this exact machine: DeepSeek R1 at 20.3 tok/s, a 1T-parameter MoE at 24.6 tok/s — see the full benchmarks.

Apple discontinued the 512 GB unified-memory configuration of the Mac Studio M3 Ultra. We secured our machines before it was pulled, and rent out dedicated remote access — you connect over Tailscale, nothing is shipped.

Can you still buy one?

Can I still buy a new Mac Studio with 512 GB of memory?
No. Apple discontinued the 512 GB unified-memory configuration of the Mac Studio M3 Ultra, and it can no longer be ordered new from Apple or authorized resellers. The current lineup tops out below it.
Why was the 512 GB configuration discontinued?
Apple pulled the configuration amid the DRAM supply shortage. The 512 GB build used an unusually large amount of high-bandwidth memory per unit, and it was dropped from the lineup rather than repriced.
What are my options if I need 512 GB of unified memory now?
Realistically, renting is the only immediate option. New sales have ended — Apple could not keep sourcing the memory to build it — and used units almost never surface. We secured machines before the configuration was pulled and rent dedicated remote access from $499 for 7 days.

Hardware specifications

Source: Apple — Mac Studio (2025) technical specifications

Chip
Apple M3 Ultra
CPU
32 cores
GPU
80 cores
Neural Engine
32 cores
Unified memory
512 GB
Memory bandwidth
819 GB/s
Storage
4 TB SSD
Thunderbolt
Thunderbolt 5 × 6 (4 rear + 2 front)
USB-A
USB 3 × 2 (up to 5 Gb/s)
HDMI
HDMI 2.1 × 1
Ethernet
10 Gb Ethernet
Wi-Fi
Wi-Fi 6E (802.11ax)
Bluetooth
Bluetooth 5.3

What 512 GB actually runs

With 512 GB of unified memory at 819 GB/s, the weights of frontier open-weight models stay fully resident — no sharding across GPUs, no offloading to disk. Ollama, llama.cpp, and MLX work out of the box on macOS.

GLM-5.2 (744B MoE)
4-bit quant (~390 GB) fits entirely in memory — MIT license, 1M context
DeepSeek-V3 / R1 (671B MoE)
4-bit quant (~380 GB) fits with no offloading to disk
Kimi K2.7 Code (1T MoE)
Runs at 2-bit quantization (~325 GB) with room for context
DeepSeek V4-Flash (284B MoE)
4-bit (~145 GB) — huge headroom for context or a second model
Qwen 3.6 / 70B-class
Several models resident at once — hot-swap without reloading

Throughput depends on the model and quantization. Rent it and run your own benchmark — that is exactly what the short plans are for.

Measured throughput — on this exact machine

We rented our own machine through the normal checkout and benchmarked it over Tailscale like any customer (July 2026, single-run snapshots, 300 generated tokens, temperature 0):

DeepSeek R1-0528 (4-bit MLX, 351.7 GiB)
Generation 20.3 tok/s · peak 380.7 GB RAM — no swap, no OOM
Kimi K2.7 Code (2-bit GGUF, 316.2 GiB)
Generation 24.6 tok/s on llama.cpp — a 1T-parameter MoE on one Mac
gpt-oss-120B (MXFP4 MLX)
Generation 79.2 tok/s — and 62.6 tok/s even at a 32K-token prompt
Qwen3.6 35B-A3B (4-bit MLX)
Generation 94.6 tok/s, prompt processing 2,892 tok/s — fully interactive
Qwen3-Coder-Next (4-bit MLX)
77.2 tok/s + a 29/29 local Codex agent soak on the same machine

Prompt processing, load times, peak memory, power draw, the failures, and every caveat are in the full benchmark write-up. Speed measurements only — not a model-quality comparison.

The rental environment

Connection
Tailscale (WireGuard) — SSH / VNC / macOS Screen Sharing
Machine type
Dedicated bare metal — not shared, not virtualized
Uptime
Runs 24/7 for your whole rental period
Your account
Your own macOS administrator user
Security
No open inbound ports; end-to-end encrypted tunnel only
Network
1 Gbps fiber, hosted in Kobe, Japan
Start
Access page is provisioned automatically right after checkout

Pricing — one-time payment, no subscription

7 days
$499 (~$71/day)
14 days
$799 (~$57/day)

Pay once with Stripe Checkout (cards, Apple Pay, Google Pay). Access starts right after payment and simply ends when the period is over — nothing renews, nothing to cancel.

What people rent it for

Frontier open-weight models

Run the current open-weight frontier (GLM-5.2, Kimi K2.7, DeepSeek R1 / V4-Flash) locally over Ollama, llama.cpp, or MLX — no API rate limits, no per-token billing.

AI agents, running 24/7

Keep Claude Code, Codex, or autonomous agents running around the clock on a machine that is not your laptop. Multiple models stay resident.

Try before you buy

The 512 GB configuration is discontinued and used units almost never surface. Verify your workload actually fits and performs for $499 — before hunting for hardware you may never find.

Research & prototyping

Benchmark quantizations, test fine-tuned weights, or run evaluations on hardware you could not otherwise access — for exactly as long as you need.

Rent the 512 GB machine

7 days from $499. Checkout to SSH in about 10 minutes.

See plans & availabilityAsk a question