Apple Plots Server Market Comeback with Custom Silicon, Taps NVIDIA for Interconnect Tech

In a move that signals its most ambitious hardware pivot in over a decade, Apple is developing a line of enterprise-grade artificial Intelligence (AI) servers powered by its own custom-designed chips. The project would mark the iPhone maker’s first serious re-entry into the server market since it quietly discontinued the Xserve rack-mounted server line back in 2011 – a retreat that, at the time, reflected Apple’s decision to concentrate its engineering resources on the then-burgeoning iPhone and iPad businesses.

Two configurations, one goal: AI inference

Sources familiar with the matter describe two distinct server configurations currently in development. The smaller, more compact variant will be built around two M8 Ultra chips – Apple’s highest-end silicon, expected to succeed the M-series Ultra processors currently powering the Mac Studio and Mac Pro. The larger, higher-capacity version will pack four M8 Ultra processors into a single enclosure, targeting data-center deployments that demand maximum compute density.

Perhaps more notably, Apple has reportedly entered into discussions with NVIDIA to license and integrate NVLink Fusion, NVIDIA’s next-generation high-speed interconnect technology, into its server design. NVLink Fusion enables ultra-low-latency, high-bandwidth data transfer between multiple processors within a server – a critical capability for running large language models and other AI workloads that require chips to share memory and coordinate computations in real time. If the partnership materializes, it would represent a rare instance of Apple relying on a competitor’s core infrastructure technology, underscoring how difficult it is to build multi-chip AI systems from scratch.

Targeting the inference opportunity

Unlike the massive Graphics Processing Unit (GPU) clusters used to train frontier AI models – a market currently dominated by NVIDIA’s H100 and H200 accelerators – Apple’s servers are being designed specifically for AI inference: the phase where a trained model processes live user requests and generates outputs in real time. As AI assistants, image generators, and code-completion tools move from research labs into everyday products, inference workloads are projected to become the largest and fastest-growing segment of AI compute.

This focus aligns closely with Apple’s own product roadmap. The company has been rolling out Apple Intelligence, its on-device and cloud-based AI feature suite, across iPhone, iPad, and Mac. Running inference for hundreds of millions of users requires enormous server-side capacity – and currently, Apple relies heavily on third-party cloud providers, including Amazon Web Services and Google Cloud, to handle that load. Bringing inference hardware in-house would give Apple greater control over performance, cost, and – crucially – user privacy, a pillar of the company’s brand identity. Processing AI requests on its own servers, using its own chips, would allow Apple to minimize the amount of user data that passes through third-party infrastructure.

A long road to market

Despite the ambitious scope of the project, The Information reports that the servers are not expected to launch until as early as 2029. That timeline leaves Apple with roughly three years of engineering, validation, and supply-chain work ahead – and it means the company will be entering a market that is likely to look very different by then. NVIDIA, Advanced Micro Devices (AMD), Intel, and a raft of startups are all racing to ship increasingly powerful AI inference chips, and hyperscalers like Amazon (with its Trainium and Inferentia chips), Google (TPUs), and Microsoft (Maia) are already several generations into their own custom silicon programs.

Apple’s advantages, however, are substantial. Its M-series chips have consistently delivered industry-leading performance-per-watt, a metric that matters enormously in data centers where power and cooling costs dominate operating expenses. The company also controls the full vertical stack – from silicon to operating system to end-user applications – which could allow it to optimize inference pipelines in ways that off-the-shelf hardware cannot. And its massive installed base of over 2 billion active devices gives it a built-in, guaranteed demand for inference capacity that few competitors can match.

Markets react positively

News of Apple’s server ambitions moved both companies’ stocks. Apple shares rose approximately 0.6% in intraday trading following the report, while NVIDIA stock climbed more than 2.0% – a reflection of investor enthusiasm over the prospect of NVIDIA’s interconnect technology finding a major new customer outside its traditional GPU base. The dual gain also suggests that the market views the Apple-NVIDIA collaboration as complementary rather than competitive: Apple provides the compute silicon and the end-user demand, while NVIDIA provides the glue that makes multi-chip AI systems work at scale.

What comes next

For Apple, the server project is about more than a new product category. It is a statement of intent: that the company intends to be a first-party player in the AI infrastructure era, not merely a consumer of someone else’s cloud. Whether the 2029 timeline holds, whether the NVIDIA partnership solidifies, and whether Apple can translate its consumer-chip prowess into data-center reliability remain open questions. But one thing is clear: after 18 years away, Apple is serious about getting back into the server business – and this time, the stakes are far higher than they were in the Xserve era.

Published

17/09/2026