
Global.fujitsu
Fujitsu announced on September 14 that it will begin global commercial sales of MONAKA — a 144-core Arm server processor it describes as the world's first commercialized 2-nanometer 3D-stacked CPU — in November 2026, alongside a made-in-Japan MONAKA Server the company is positioning as sovereign AI infrastructure. The announcement matters not only because a new chip is entering the market, but because of what it is specifically built to do: run AI inference in facilities where no GPU-based alternative can currently go.
Hospitals, government agencies, regional data centers, and older enterprise server rooms share a constraint that is rarely discussed in AI hardware coverage: they run on air cooling and do not have the power density headroom to add liquid-cooled GPU racks. For those facilities, today's most capable AI inference hardware is simply off the table. Fujitsu's claim — that the MONAKA Server operates at 40°C with air cooling and reduces server cooling power consumption by as much as 80% — describes a chip that could unlock a deployment class that the current GPU-centric AI buildout has structurally bypassed.
If the performance numbers hold under independent testing, that is the story. If they do not, that is equally important. Fujitsu's 2× AI inference throughput claim — allowing the same workloads to run on roughly half the number of servers and half the power of other CPUs — remains an internal estimate; no independent benchmarks have been published as of this writing.
FUJITSU-MONAKA is a 144-core Armv9.3-A server processor organized as four 36-core compute chiplets, specifically engineered for AI inference rather than training or general-purpose compute. Its architecture combines two different manufacturing process nodes in a single package: a 2nm compute die for the processing cores, stacked face-to-face on top of a 5nm SRAM die that holds the entire last-level cache, with a third 5nm I/O die connected via silicon interposer below. The dies are joined using hybrid copper bonding for die stacking — a technique that creates direct copper-to-copper connections at sub-10-micrometer pitch, delivering much higher interconnect density and lower power than traditional solder-based packaging.
The rationale for this three-die split is economic as much as it is technical. The 2nm compute silicon accounts for less than 30% of the chip's total die area. That percentage matters because 2nm silicon on TSMC's N2P process is the most expensive real estate in semiconductor manufacturing. By confining the cutting-edge process to the compute cores — where it delivers the most benefit — and building the cache and I/O on the larger, less expensive 5nm node, Fujitsu says it can accelerate time to market and keep production costs manageable for a chip this ambitious.
The processor's core differentiating claim is its ultra-low voltage operation technology. Fujitsu runs the MONAKA cores at approximately 30% below the nominal voltage typical of comparable designs. Because power consumption scales with the square of the voltage, a 30% voltage reduction translates to roughly a 51% reduction in dynamic power at comparable frequencies — an efficiency gain Fujitsu's lead architect Ryohei Okazaki described at Hot Chips 2026 as equivalent to a full process generation advance. Achieving this required custom SRAM cells and a proprietary CAD design flow built specifically for non-standard low-voltage operation. The chip also moves low-dropout voltage regulators — analog circuits that scale poorly at 2nm — onto the 5nm SRAM die and places them directly beneath the core's floating-point units, enabling per-core dynamic voltage and frequency scaling.
The practical result is that MONAKA will ship in two SKUs: a 350W air-cooled variant at 2.1 GHz, and a 500W liquid-cooled variant at 2.9 GHz base, with a maximum turbo frequency of 3.8 GHz and a memory transfer rate of 8,800 MT/s. Fujitsu's performance estimates — 4,355 GFLOPS and 69.7 TOPS in INT8 for the 350W chip, 6,013 GFLOPS and 96.2 TOPS for the 500W chip — remain Fujitsu's own projections until independent production testing at the 2027 volume launch confirms or challenges them.
For AI inference specifically, the chip adds dedicated hardware acceleration for matrix operations — the mathematical operations at the heart of neural network inference — combined with dual 256-bit SVE2 vector units per core and a floating-point register cache that reduces power on GEMM workloads with high temporal locality. The choice to use 256-bit SVE2 rather than the 512-bit SVE of the A64FX predecessor was deliberate: Fujitsu's engineers determined the wider 512-bit units were optimal for the scientific HPC workloads that dominated Fugaku, but that 256-bit fits data center inference workloads better and keeps the core die smaller and cheaper.
The argument for MONAKA that no Nvidia, AWS, or Ampere product currently makes is precisely the 350W air-cooled SKU. The dominant GPU-based inference systems — Nvidia's H200 and GB200 clusters — require aggressive liquid cooling and power densities that older or smaller data centers cannot support. Enterprise IT organizations in hospitals, government buildings, regional financial institutions, and academic facilities frequently operate air-cooled server rooms at densities and power budgets designed for conventional workloads. For those organizations, GPU inference requires costly facility retrofits that can cost millions of dollars and take years to complete.
Fujitsu explicitly designed MONAKA's 350W SKU and the MONAKA Server's 40°C (104°F) ambient air-cooled operating envelope to reach those facilities. The company claims its cooling technology reduces server cooling power consumption by up to 80% versus conventional approaches — which means not only that existing air-cooled facilities can run the hardware without upgrading their cooling infrastructure, but that organizations with liquid cooling can substantially reduce the energy budget dedicated to keeping the hardware cool.
The CDI/CXL (Composable Disaggregated Infrastructure / Compute Express Link) technology on the 1U MONAKA Server adds another dimension relevant to inference workloads: it pools memory across server boundaries, allowing organizations serving large language models to address memory constraints without fitting the entire model into a single server's physical memory footprint.
The chip's confidential computing support — implemented via Arm's Confidential Compute Architecture, which creates hardware-isolated "Realm" execution environments that shield workloads from the hypervisor and other tenants — is especially relevant for the healthcare and government markets where MONAKA's air-cooled positioning is most compelling. An organization running patient data or classified workloads through AI inference on shared cloud infrastructure has a strong motivation to deploy on hardware where the cloud operator cannot access the computation. MONAKA's Arm CCA implementation directly addresses that concern.
Fujitsu is pitching MONAKA as sovereign AI infrastructure — a term that has become central to national AI strategies in Japan and Europe, tracking the growing political priority of reducing dependence on US and Taiwanese hardware supply chains. The CNAS Sovereign AI Index, updated through June 2026, tracked 185 sovereign AI projects across 67 government actors, with total disclosed investment reaching $83.9 billion.
The pitch has substance. The MONAKA Server is designed, integrated, and assembled at Fujitsu's Kasashima Plant in Japan, drawing on the same integrated domestic production system Fujitsu built for Fugaku manufacturing. Component traceability — the ability to verify the origin and manufacturing history of every part — is central to what Fujitsu means when it says "sovereign": a government or defense agency deploying MONAKA Server can audit its supply chain in a way that is not possible with infrastructure assembled from globally sourced components of unknown provenance.
The limitation of the sovereignty claim is the one that belongs in every MONAKA article and is frequently omitted: the compute dies are manufactured at TSMC fabs in Taiwan, not in Japan. Fujitsu's chip is designed and developed in Japan and its server is assembled in Japan, but the most advanced manufacturing step — the silicon itself — happens in Taiwan. For customers with the strictest supply-chain requirements, this is a meaningful distinction between "Japanese-designed sovereign server" and "fully domestic fabrication."
Japan is actively working to close that gap. As of 2026, Japan's government has committed approximately ¥2.9 trillion in Rapidus support (approximately $18.8 billion USD at current rates), the domestic chipmaker targeting 2nm mass production at its Hokkaido facility beginning in 2027. Separately, NEDO provided approximately ¥76 billion (approximately $493 million USD) for Fujitsu and IBM Japan's 1.4nm program to develop an AI chip to be fabricated entirely in Japan via Rapidus. That successor chip — already named Monaka-X — is slated for the FugakuNEXT supercomputer, a roughly $750 million RIKEN project targeting more than 600 FP8 exaFLOPS within a 40-megawatt envelope and approximately 100 times the application performance of Fugaku, with operation targeted around 2030.
In other words, MONAKA is the first commercial generation of a chip architecture with a clear domestic fabrication roadmap — but the domestic fabrication part is on the roadmap, not yet in production.
Fujitsu has spent the past two years building MONAKA's commercial ecosystem before the chip reached buyers. Three partnerships are worth understanding:
NVIDIA NVLink Fusion: Fujitsu has outlined plans to connect MONAKA CPUs with NVIDIA GPUs via NVLink Fusion, the interconnect NVIDIA opened to third-party CPUs in 2025. This means MONAKA is not limited to pure CPU-only inference deployments; it can serve as the host processor in heterogeneous CPU-GPU systems, expanding the chip's addressable workload significantly. The Monaka-X successor for FugakuNEXT is specifically built around this CPU-GPU configuration.
Broadcom: In February 2026, Broadcom disclosed shipment of a 2nm 3.5D SoC supporting the MONAKA program — providing early third-party confirmation that the stacked silicon architecture was manufacturable.
Arrcus and 1Finity: In March 2026, Fujitsu, Arrcus, and 1Finity outlined a distributed AI infrastructure using MONAKA compute combining Arrcus's ArcOS networking software and 1Finity's optical interconnect — positioning MONAKA for edge-to-core AI deployments that extend beyond the data center.
Fujitsu is also selling MONAKA as standalone silicon to cloud providers and third-party server vendors — a dual-track strategy that maximizes the chip's reach. Cloud operators who want to incorporate MONAKA into their own custom platforms can do so without adopting Fujitsu's server; governments and enterprises that want the full sovereignty story can take the integrated MONAKA Server.
Positioning MONAKA in the Arm server CPU landscape requires clarity about what it is and is not trying to win. At 144 cores, MONAKA sits in the middle of the current Arm field: below AWS's Graviton5 at 192 cores (Neoverse V3 cores on a single 3nm die, though AWS-exclusive) and Ampere Computing's AmpereOne Aurora (targeting 512 cores), roughly comparable to Microsoft's Cobalt 200 (132 cores).
Nvidia's Grace CPU — 72 Neoverse V2 cores tightly integrated with HBM memory — is aimed at feeding data to Nvidia GPUs, not at CPU-only inference. Grace is not a direct competitor; it is designed for a fundamentally different deployment model.
What makes MONAKA distinct within its competitive field is not core count but architectural choices: its 3D-stacked cache (the entire last-level cache on a separate die beneath the compute die, rather than in the compute die itself), its 12-channel DDR5 memory without HBM, and its ultra-low-voltage operation technique. Those choices prioritize sustained inference efficiency and air-cooled deployability over peak throughput. A customer running continuous AI inference workloads on constrained power budgets gets a different value proposition from MONAKA than from any hyperscaler Arm CPU, most of which optimize for cloud-native throughput-per-watt in well-cooled environments.
The 256-bit SVE2 vector width also deserves context. Most hyperscaler Arm cores use 128-bit SVE2. MONAKA's 256-bit width is wider than the field — matching SiPearl's Rhea1 (the European sovereign AI CPU) — but narrower than the 512-bit SVE of the A64FX that powered Fugaku. Fujitsu's architects concluded that 256-bit SVE2 fits inference workloads, which do not require the ultra-wide vectors that justified Fugaku's HPC-optimized design.
The three open questions that matter for procurement decisions:
Independent benchmark validation. Fujitsu's 2× inference throughput claim against "other CPUs" and 50% TCO reduction assertion are internal performance estimates against unspecified comparisons. No independent production testing has been published. At volume production in 2027, industry benchmarking organizations and cloud providers will produce the data that decides whether MONAKA's efficiency pitch holds on real workloads. The performance estimates of 4,355 GFLOPS DGEMM and 69.7 TOPS INT8 for the 350W SKU are Fujitsu's own projections.
Core-to-core latency. At Hot Chips 2026, Dr. Ian Cutress of More Than Moore asked Fujitsu's lead architect specifically whether MONAKA was doing anything to minimize core-to-core latency, given that the four compute chiplets sit on opposite sides of the package with traffic routing through the I/O die. Fujitsu declined to disclose latency figures in response. For workloads with tight inter-core communication requirements, this is a gap in the public technical record that needs independent measurement.
US availability timeline. Initial November 2026 server sales target Japan and Europe. The US is not in the first wave. Global CPU availability including the US is planned during Fujitsu's fiscal Q4 2026, which ends March 31, 2027. American data center operators and enterprises evaluating MONAKA for sovereign AI deployments will be purchasing hardware in 2027, with benchmarks potentially available around the same time.
Fujitsu's A64FX powered Fugaku, which held the top position on the TOP500 supercomputer list for four consecutive biannual rankings between 2020 and 2022 — a record of technical credibility that no paper announcement can replicate. That pedigree means MONAKA enters the market with institutional standing that most first-generation sovereign AI chips lack.
The broader context matters: global spending on sovereign AI infrastructure is now a documented investment phenomenon. As of June 2026, 185 sovereign AI projects had been tracked across 67 government actors globally, with total disclosed investment reaching $83.9 billion. Japan's own semiconductor investment — ¥2.9 trillion (approximately $18.8 billion USD) into Rapidus semiconductor funding alone — signals that the government views chip sovereignty as a national security priority, not just an economic preference.
MONAKA is launching into a market that has been asking for exactly this category of product. The technical question is whether it delivers on the efficiency claims under independent testing. The strategic question is whether Japan and European customers, evaluating their AI infrastructure options through the lens of supply chain auditability and export-control risk, assign enough value to MONAKA's auditable Japanese supply chain to accept a chip whose silicon is still fabricated in Taiwan — at least until Rapidus reaches production scale.
For data center operators in Japan and Europe with air-cooled facilities, November 2026 and April 2027 are dates worth tracking. For buyers anywhere with strict sovereignty requirements, the domestic fabrication gap is the decision variable that MONAKA's current generation cannot yet resolve.
Yes — that is Fujitsu's primary claim for the 350W air-cooled SKU. The MONAKA Server is rated to operate in ambient temperatures up to 40°C (104°F) with air cooling and is designed to reduce server cooling power consumption by up to 80% compared to conventional approaches. This means hospitals, government buildings, academic institutions, and other organizations with existing air-cooled server infrastructure can, in principle, deploy real AI inference workloads without a facility retrofit or specialized cooling upgrade. That claim remains vendor-stated until independent production testing at the 2027 volume launch confirms it on real workloads.
The three chips target different problems. Nvidia Grace (72 Arm Neoverse V2 cores plus HBM) is designed to feed data to Nvidia GPUs efficiently, not to run inference on its own. AWS Graviton5 (192 cores on a single 3nm die) is a cloud-native CPU exclusive to AWS infrastructure. MONAKA's 144 cores fall between those in count, but the comparison Fujitsu is making is on efficiency and deployability in power-constrained, sovereignty-sensitive environments — dimensions where Grace and Graviton are not competing. The direct competitors on workload and deployment model are Ampere Computing's AmpereOne (512-core roadmap) and, notably, SiPearl's Rhea1, which is Europe's own sovereign AI CPU with a similar 256-bit SVE2 design philosophy.
The correct characterization is that MONAKA is designed and developed in Japan, and the MONAKA Server is assembled and integrated in Japan at the Kasashima Plant. The silicon dies — both the 2nm compute chiplets and the 5nm cache and I/O dies — are fabricated at TSMC facilities in Taiwan. Fujitsu's sovereignty claim rests on supply chain traceability and design ownership, not full onshore fabrication. Japan is investing heavily in domestic 2nm fabrication via Rapidus, with mass production targeted for 2027. Fujitsu's next-generation Monaka-X chip, planned for the FugakuNEXT supercomputer around 2030, is intended to be the first generation with fully domestic Japanese fabrication.
The core mechanism is ultra-low voltage operation: Fujitsu runs MONAKA's compute cores at approximately 30% below the voltage used by comparable designs. Because power consumption scales with the square of the voltage, cutting voltage by 30% reduces dynamic power by roughly half, at comparable performance levels. This is not a standard design technique — it required Fujitsu to develop custom SRAM cells and proprietary computer-aided design tools built specifically for sub-nominal voltage operation. The additional factor is architecture: CPUs like MONAKA run inference at lower power than GPUs because inference is not as parallelism-intensive as training. For batch inference on standard models, a CPU's more flexible compute pipeline can be nearly as effective as GPU matrix hardware, at a fraction of the power and without the facility infrastructure GPU racks require.
