Two-Phase Cooling Squeezes Dual 600W GPUs Into 1U Edge AI Server at AMD EPYC 9006 Venice Launch
2 hour ago / Read about 39 minute
Source:TechTimes

gettyimages.com

Taiwanese server specialist AEWIN Technologies — a subsidiary of the BenQ Qisda Group — announced a four-platform server portfolio on July 28, 2026, timed to the same week AMD formally launched its 6th-generation EPYC 9006 "Venice" processors. The move gives enterprise and edge AI operators an immediately available production path to Venice, which AMD unveiled at its Advancing AI 2026 conference in San Francisco on July 22 and 23. That timing matters: EPYC 9006 SP7 systems will not begin commercial shipping to customers until Q4 2026, meaning AEWIN's announcement lands at the front edge of the procurement window rather than six months after it.

The headline platform is the BG18-A10710 — a 1U AI server that integrates a full two-phase direct liquid cooling system covering both the CPU and two single-width 600W GPUs. That combination in a 1U chassis is the specific technical claim worth scrutinizing, because it addresses a thermal constraint that has confined high-density GPU compute to taller chassis or more complex facility infrastructure.

What AMD EPYC 9006 Venice Actually Is

Venice is the first server CPU to reach production on TSMC's 2nm process node — which uses gate-all-around nanosheet transistors rather than the FinFET architecture that defined the prior decade of chip manufacturing. That process shift delivers roughly 10 to 15 percent more performance per watt at the transistor level before any architectural improvements layer on top.

Read more: AMD Advancing AI 2026 Opens With Zen 6 Venice, Helios, and Open AI Rack Bet

The architectural improvements are substantial. Venice introduces the Zen 6 microarchitecture with two distinct core configurations: standard Zen 6 chips reach up to 96 cores and 192 threads per socket, while compact Zen 6c chips reach up to 256 cores and 512 threads per socket. The only named SKU at launch is the EPYC 9996 — the 256-core, 512-thread dense model packing approximately 1 GB (1,024 MB) of L3 cache and 203 billion transistors in a single socket. AMD has not yet released the full SKU table, pricing, or power/frequency details for the broader lineup.

Three platform-level changes accompany the core count increase. First, Venice adopts PCIe Gen 6, which doubles per-lane bandwidth to 64 GT/s — a shift from PCIe Gen 5's 32 GT/s — using PAM-4 (Pulse Amplitude Modulation with 4 signal levels) and FLIT-based encoding to maintain signal integrity at the higher speed. Second, Venice adds a new SP7 socket that is not backward-compatible with the SP5 infrastructure Turin currently runs on, meaning a full platform replacement is required to adopt it. Third, the memory subsystem upgrades to 16-channel DDR5 with MRDIMM (Multiplexed Registered DIMM) support at 12,800 MT/s — doubling per-channel bandwidth compared to standard DDR5-6400 — and delivering an aggregate 1.6 TB/s of memory bandwidth.

AMD is positioning Venice specifically for "agentic AI" workloads — systems where autonomous AI agents plan, call external tools, execute multi-step reasoning, and loop through complex tasks rather than answering a single prompt. As compute shifts from prompt-and-response inference to agentic pipelines, the CPU becomes the orchestration layer feeding data to GPU accelerators — making per-core performance, memory bandwidth, and interconnect speed more strategically important than they were in the pure-training era.

AMD's own benchmark data — run on AMD test systems and not yet independently verified — claims the 256-core Venice configuration delivers roughly 3.3 times the rack throughput of NVIDIA Vera in SPEC CPU 2017 integer workloads within the same 100 kW rack power budget, and approximately twice the throughput of the Intel Xeon 6980P. Independent third-party results are expected when SP7 systems reach reviewers closer to Q4 2026 availability.

Intel's next-generation P-core Xeon, Diamond Rapids, is not expected until 2027 at the earliest, leaving AMD with an uncontested process-node lead through at least the end of next year.

Why Boiling Coolant Beats Pumped Coolant at the Edge

The BG18-A10710's engineering claim requires some unpacking, because the difference between single-phase and two-phase direct-to-chip cooling is not simply "better" versus "worse" — it is a specific trade-off that becomes advantageous under specific density conditions.

In single-phase direct-to-chip cooling — currently the dominant approach, with an estimated 55% of new liquid cooling deployments as of 2026, according to Greg Alexander, senior thermal engineer at Motivair by Schneider Electric — a coolant (typically a 75% water / 25% glycol mixture) circulates continuously through cold plates mounted directly over CPU and GPU surfaces. The coolant stays liquid throughout, absorbing heat through convection and carrying it to a heat exchanger. Industry standard flow rate: 1.5 liters per minute per kilowatt of thermal load.

In two-phase direct-to-chip cooling, a dielectric fluid or refrigerant actually boils at the chip surface. The key physical advantage is the latent heat of vaporization: the energy required to change a substance from liquid to vapor is far larger than the energy required to raise the same liquid's temperature by a given amount. Research from Opteon, a Chemours division specializing in thermal fluids, estimates two-phase fluids can provide 10 to 100 times greater heat transfer capacity per unit of coolant compared to single-phase alternatives. Because the coolant boils at a constant temperature, the chip surface temperature also stays more uniform — a specific advantage for sustained GPU workloads where thermal gradients accelerate component fatigue.

The engineering implication for a 1U chassis is concrete. Two 600W GPUs and a high-power EPYC Venice CPU in a 1U enclosure represent a total thermal load that could approach 1.5 kW or more. At the 1.5 L/min/kW industry flow standard, a single-phase system would require pushing roughly 2.25 liters of coolant per minute through the tight plumbing of a chassis that is only 44.45 mm (1.75 inches) tall. Two-phase DLC dramatically reduces the required flow rate by extracting heat through phase change rather than temperature rise — making the compact plumbing physically viable where single-phase systems would require inconveniently large pipes or unacceptably high pump speeds.

As Greg Alexander, senior thermal engineer at Motivair by Schneider Electric, noted in May 2026, two-phase DTC "is seldom used" at scale today, with operators favoring single-phase for its lower complexity and faster deployment. Josh Claman, CEO of Accelsius, a two-phase cooling company, put the long-term case plainly: "Why would anyone expect data centers with some of the highest heat fluxes over the next five to 10 years not to go to two-phase? Every other industrial sector has gone to two-phase."

AEWIN is not new to this technology. The company and its cooling subsidiary, Arivor Technologies, demonstrated a rack-scale 2P DLC system at Computex 2026 in June, designed for infrastructure above 100 kW per rack and capable of handling chips above 3,000 W. Today's 1U BG18-A10710 announcement scales that expertise down to the edge form factor.

Four Platforms, One Modular Standard

All four AEWIN platforms share a DC-MHS (Data Center Modular Hardware System) form factor, an Open Compute Project standard that defines interoperable mechanical, electrical, thermal, and management interfaces. The standard, backed by AMD, Dell, Google, HPE, Microsoft, NVIDIA, and others, allows enterprise operators to mix components from different vendors and upgrade compute or networking modules independently rather than replacing entire servers.

The BG18-A10710 (1U) is the headline platform, as described above. Its 16 DDR5 MRDIMM slots support up to 12,800 MT/s, and its eight PCIe Gen 6 x4 E1.S NVMe SSD slots each deliver 32 GB/s of aggregate read bandwidth per slot — double what PCIe Gen 5 equivalents offer — feeding the CPU with high-speed storage for AI inference context and KV cache retrieval.

The BG28-A10710 (2U) targets large language model training and generative AI workloads at greater scale, accepting up to four full-height, full-length GPU cards with PCIe 6.0 connectivity and support for accelerators drawing up to 600W each. This is the platform suited to inference clusters and HPC deployments where air cooling's limits have already been hit and rack density is the primary operational constraint.

The BX28-A10710 (2U) is the general-purpose enterprise server in the portfolio, balancing compute and storage: up to 12 PCIe Gen 5 x4 U.2 NVMe drives, two PCIe 6.0 expansion slots, one OCP 3.0 network slot, and 16 DDR5 RDIMM/MRDIMM slots for high-memory-bandwidth virtualization and distributed data processing.

The BS28-A10710 (2U) closes the portfolio as a storage-focused platform: up to 16 PCIe 6.0 x4 NVMe E3.S drives, 16 DDR5 MRDIMM slots at 12,800 MT/s, and the full PCIe Gen 6 bandwidth headroom that makes it practical for data lakes, backup workloads, and enterprise storage tiers where I/O throughput is the bottleneck.

AEWIN at Scale: Where the Portfolio Fits in the Venice Ecosystem

"The launch of our new portfolio, based on 6th Gen AMD EPYC, reflects AEWIN's ability to rapidly translate next-generation processor technologies into production-ready platforms," said David Chung, VP of R&D Division II at AEWIN. "From AI acceleration and enterprise computing to high-density storage, we provide purpose-built infrastructure that helps customers shorten deployment cycles, accelerate solution development, and scale with confidence."

Related
EPYC Venice Arrives Wednesday: AMD's Zen 6 on TSMC 2nm Resets Server Race
AI Data Center Water Use Is Not Solved: Nvidia's Cooling Fix Stops at the Walls

How Does AEWIN Fit Into the Broader Venice Ecosystem?

AEWIN's announcement competes in a crowded field. More prominent Taiwanese ODMs — including Supermicro, Wiwynn, and Inventec — have also been aggressively aligning server launches with chip announcements, compressing the traditional 12- to 18-month hardware qualification cycle into same-day portfolio reveals. For commodity enterprise deployments, AEWIN's modular DC-MHS architecture and rapid platform readiness are differentiators within a competitive ODM market but not unique claims.

What may distinguish AEWIN's play specifically is the 1U 2-phase DLC integration. As AI inference pushes into edge locations — telecom facilities, industrial sites, distributed enterprise locations that cannot support the full cooling infrastructure of a hyperscale data center — a 1U chassis with integrated 2-phase DLC becomes load-bearing infrastructure rather than just a thermal checkbox. Operators in these environments cannot retrofit a facility-scale coolant distribution unit; they need the cooling system to be self-contained and rack-local. The BG18-A10710 is positioned precisely for that constraint.

What Is Not Yet Known

Pricing and specific availability timelines for all four AEWIN platforms were not disclosed at announcement. The SP7 Venice processor family — which these servers are built around — will begin shipping commercially in Q4 2026. Availability of the AEWIN platforms presumably follows that timeline, but AEWIN has not confirmed.

AMD's Venice benchmark claims remain self-reported. Independent test results from hardware reviewers are expected as SP7 systems reach labs in the approach to Q4, and will be the meaningful data point for enterprise procurement teams evaluating the EPYC 9006 against NVIDIA Vera and the Intel Xeon 6 lineup.


Frequently Asked Questions

What is two-phase liquid cooling and how does it differ from standard liquid cooling?

In standard (single-phase) direct-to-chip liquid cooling, a water-glycol mixture circulates through cold plates on CPU and GPU surfaces, absorbs heat as a liquid, and transfers that heat to a radiator. The coolant never changes state. In two-phase direct-to-chip cooling, the coolant — typically a dielectric fluid or refrigerant — actually boils at the chip surface, absorbing heat through the phase transition from liquid to vapor. The latent heat of vaporization is far larger than the sensible heat absorption of a liquid, so two-phase systems can extract more heat per unit of coolant with much lower flow rates. That reduced flow rate is the key engineering advantage in a 1U chassis: two 600W GPUs in a compact enclosure would require inconveniently high pump speeds under single-phase, but two-phase's lower flow requirement fits in tight 1U plumbing. Two-phase systems are more complex and costly to deploy, which is why single-phase still holds roughly 55% of new liquid cooling installations as of 2026.

How does PCIe Gen 6 change what an AI inference server can do?

PCIe Gen 6 doubles per-lane bandwidth to 64 GT/s (gigatransfers per second) compared to PCIe Gen 5's 32 GT/s, using a new signaling scheme called PAM-4 (Pulse Amplitude Modulation with 4 levels) plus forward error correction to maintain reliability at higher speeds. In an AI inference server, this translates directly to faster data movement between the CPU, GPUs, and NVMe storage — which matters for agentic AI workloads where the CPU must orchestrate multi-step agent tasks, retrieve context from fast storage, and continuously feed GPU accelerators. AEWIN's 1U server uses eight PCIe Gen 6 x4 NVMe slots, each delivering 32 GB/s, double what Gen 5 equivalents offer. For inference at the edge, where latency is the primary metric, that bandwidth headroom is the difference between a CPU bottleneck and a balanced pipeline.

When will AMD EPYC 9006 Venice servers actually be available?

AMD's SP7 flagship Venice server platform will begin commercial shipping in Q4 2026. AEWIN's announcement of these four platforms is production-intent, not a roadmap preview, meaning the hardware exists — but customers will wait until the underlying processor reaches production availability before taking delivery. The SP8 mainstream platform targets the first half of 2027. AMD's 3D V-Cache (Venice-X) and LPDDR5X-based Verano variants are scheduled for the second half of 2027. Pricing for AEWIN's platforms has not been disclosed; AMD has similarly not published the full EPYC 9006 SKU pricing table. Enterprise buyers should note that the SP7 socket is incompatible with the SP5 infrastructure Turin uses, so adopting Venice requires a full server replacement rather than a CPU upgrade.

Is the AEWIN 1U edge server suitable for AI deployments at industrial or telecom sites?

The BG18-A10710's two-phase direct liquid cooling makes it more viable in constrained facilities than single-phase alternatives, because the self-contained 2-phase loop does not require the facility-scale coolant distribution infrastructure that larger liquid cooling implementations need. Industrial sites, telecom edge facilities, and distributed enterprise locations where rack space is limited to a few units and facility cooling plants are not available are precisely the deployment scenarios this platform targets. The important caveat: two-phase systems require tighter sealing and more careful fluid management than single-phase, and operators without previous DLC experience will face a steeper operational learning curve. AEWIN has not yet published maintenance specifications, coolant type, or service interval information for this platform.