d-Matrix and NVIDIA: The Interconnect Race for Real-Time Inference

AI inference startup d-Matrix is adopting NVIDIA’s NVLink Fusion to connect its Raptor XPUs to NVIDIA’s infrastructure. This partnership highlights the growing need for high-bandwidth interconnects in the race for real-time AI inference.

Share
d-Matrix and NVIDIA: The Interconnect Race for Real-Time Inference

In the world of semiconductors, the bottleneck is no longer just how fast a single chip can compute, but how quickly data can move between chips. d-Matrix, a rising star in the AI inference space, has announced it will adopt NVIDIA’s NVLink Fusion technology for its next-generation Raptor XPUs. This move allows d-Matrix to integrate its specialized inference hardware directly into the broader NVIDIA ecosystem.

The Raptor XPU is designed specifically for the high-demand, low-latency requirements of Large Language Model (LLM) inference. By using NVLink Fusion, d-Matrix can achieve rack-scale connectivity, essentially treating a massive cluster of chips as a single, unified processor. This is crucial for "Physical AI" applications, where a robot or autonomous vehicle needs to process vast amounts of sensor data and generate motor commands in milliseconds.

This partnership is a strategic win for both companies. For NVIDIA, it solidifies NVLink as the industry-standard interconnect. For d-Matrix, it provides a clear path to market by ensuring their chips are compatible with the most widely used AI infrastructure. As AI models grow in complexity, the "interconnect war" will become just as important as the "transistor war" in determining who leads the next generation of silicon.


Source: NVIDIA Blog