Apple discontinued the 512 GB unified-memory configuration of the Mac Studio M3 Ultra. We secured our machines before it was pulled, and rent out dedicated remote access — you connect over Tailscale, nothing is shipped.
Can you still buy one?
- Can I still buy a new Mac Studio with 512 GB of memory?
- No. Apple discontinued the 512 GB unified-memory configuration of the Mac Studio M3 Ultra, and it can no longer be ordered new from Apple or authorized resellers. The current lineup tops out below it.
- Why was the 512 GB configuration discontinued?
- Apple pulled the configuration amid the DRAM supply shortage. The 512 GB build used an unusually large amount of high-bandwidth memory per unit, and it was dropped from the lineup rather than repriced.
- What are my options if I need 512 GB of unified memory now?
- Realistically, renting is the only immediate option. New sales have ended — Apple could not keep sourcing the memory to build it — and used units almost never surface. We secured machines before the configuration was pulled and rent dedicated remote access from $499 for 7 days.
Hardware specifications
Source: Apple — Mac Studio (2025) technical specifications
- Chip
- Apple M3 Ultra
- CPU
- 32 cores
- GPU
- 80 cores
- Neural Engine
- 32 cores
- Unified memory
- 512 GB
- Memory bandwidth
- 819 GB/s
- Storage
- 4 TB SSD
- Thunderbolt
- Thunderbolt 5 × 6 (4 rear + 2 front)
- USB-A
- USB 3 × 2 (up to 5 Gb/s)
- HDMI
- HDMI 2.1 × 1
- Ethernet
- 10 Gb Ethernet
- Wi-Fi
- Wi-Fi 6E (802.11ax)
- Bluetooth
- Bluetooth 5.3
What 512 GB actually runs
With 512 GB of unified memory at 819 GB/s, the weights of frontier open-weight models stay fully resident — no sharding across GPUs, no offloading to disk. Ollama, llama.cpp, and MLX work out of the box on macOS.
- GLM-5.2 (744B MoE)
- 4-bit quant (~390 GB) fits entirely in memory — MIT license, 1M context
- DeepSeek-V3 / R1 (671B MoE)
- 4-bit quant (~380 GB) fits with no offloading to disk
- Kimi K2.7 Code (1T MoE)
- Runs at 2-bit quantization (~325 GB) with room for context
- DeepSeek V4-Flash (284B MoE)
- 4-bit (~145 GB) — huge headroom for context or a second model
- Qwen 3.6 / 70B-class
- Several models resident at once — hot-swap without reloading
Throughput depends on the model and quantization. Rent it and run your own benchmark — that is exactly what the short plans are for.
Measured throughput — on this exact machine
We rented our own machine through the normal checkout and benchmarked it over Tailscale like any customer (July 2026, single-run snapshots, 300 generated tokens, temperature 0):
- DeepSeek R1-0528 (4-bit MLX, 351.7 GiB)
- Generation 20.3 tok/s · peak 380.7 GB RAM — no swap, no OOM
- Kimi K2.7 Code (2-bit GGUF, 316.2 GiB)
- Generation 24.6 tok/s on llama.cpp — a 1T-parameter MoE on one Mac
- gpt-oss-120B (MXFP4 MLX)
- Generation 79.2 tok/s — and 62.6 tok/s even at a 32K-token prompt
- Qwen3.6 35B-A3B (4-bit MLX)
- Generation 94.6 tok/s, prompt processing 2,892 tok/s — fully interactive
- Qwen3-Coder-Next (4-bit MLX)
- 77.2 tok/s + a 29/29 local Codex agent soak on the same machine
Prompt processing, load times, peak memory, power draw, the failures, and every caveat are in the full benchmark write-up. Speed measurements only — not a model-quality comparison.
The rental environment
- Connection
- Tailscale (WireGuard) — SSH / VNC / macOS Screen Sharing
- Machine type
- Dedicated bare metal — not shared, not virtualized
- Uptime
- Runs 24/7 for your whole rental period
- Your account
- Your own macOS administrator user
- Security
- No open inbound ports; end-to-end encrypted tunnel only
- Network
- 1 Gbps fiber, hosted in Kobe, Japan
- Start
- Access page is provisioned automatically right after checkout
Pricing — one-time payment, no subscription
- 7 days
- $499 (~$71/day)
- 14 days
- $799 (~$57/day)
Pay once with Stripe Checkout (cards, Apple Pay, Google Pay). Access starts right after payment and simply ends when the period is over — nothing renews, nothing to cancel.
What people rent it for
Frontier open-weight models
Run the current open-weight frontier (GLM-5.2, Kimi K2.7, DeepSeek R1 / V4-Flash) locally over Ollama, llama.cpp, or MLX — no API rate limits, no per-token billing.
AI agents, running 24/7
Keep Claude Code, Codex, or autonomous agents running around the clock on a machine that is not your laptop. Multiple models stay resident.
Try before you buy
The 512 GB configuration is discontinued and used units almost never surface. Verify your workload actually fits and performs for $499 — before hunting for hardware you may never find.
Research & prototyping
Benchmark quantizations, test fine-tuned weights, or run evaluations on hardware you could not otherwise access — for exactly as long as you need.
Rent the 512 GB machine
7 days from $499. Checkout to SSH in about 10 minutes.
Weighing this against renting GPUs? See Mac Studio vs cloud GPU. 日本語版は ai-kizai.jp へ。