Part 7 — The NVIDIA AI Stack: From Silicon to Hyperscale Fabric

Machine learning means teaching a machine to learn from data — for example, teaching a computer to play chess by showing it millions of games rather than coding every rule by hand. Researchers use software frameworks to build and train AI models. The main ones are:

ToolFocus Area
PyTorchFlexible and research-friendly; very popular in academia and modern AI labs.
TensorFlowGoogle’s framework, built for large-scale production
scikit-learnThe “go-to” for traditional machine learning (regression, clustering, etc.).
MATLABMaths and engineering simulations
Apache SparkProcessing massive datasets across many machines
NVIDIA Isaac LabSimulating robots so they can learn in virtual environments

NVIDIA NGC Catalog

Think of NGC as an App Store for AI. Instead of spending days installing and configuring software, you download a ready-made container a pre-packaged software environment that is already optimised to run at full speed on NVIDIA GPUs.

It also includes pre-trained models, meaning the AI has already been taught the basics and you just fine-tune it for your specific use case.

Website: catalog.ngc.nvidia.com

Nvidia Workflows

These are ready-made solutions for common business problems. Instead of building from scratch, a developer downloads a Reference Application which includes the trained model, the data pipeline, and the deployment configuration all in one package.

Example of some workflows available in NCC catalog

WorkflowWhat it doesReal-world example
Intelligent Virtual Assistant24/7 automated customer support.A banking bot that helps you reset a password via voice or text.
Audio TranscriptionHigh-speed, accurate speech-to-text.Generating automated captions for a live video meeting.
Digital FingerprintingUses AI to detect weird patterns in network traffic.Identifying a hacker trying to steal company data in real-time.
Next Item PredictionSuggests products based on user history.Customers who bought this also liked… on an e-commerce site.
Route OptimizationFinds the most efficient path for vehicles/robots.A delivery company saving fuel by optimizing 1,000 truck routes at once.
Generative AIUses a knowledge base to create or summarise contentAn internal tool that summarizes thousands of legal documents for a firm.

The Individual Chips

A CPU handles general-purpose tasks running the operating system, managing memory etc

A GPU does the heavy mathematical lifting behind AI training — running billions of calculations in parallel. Think of it as a factory floor with thousands of workers doing simple jobs at once.

The Superchips

Normally a CPU and GPU are on separate chips, connected by PCIe a reasonably fast but still limited bus. NVIDIA fuses them together into a single Superchip using NVLink-C2C, a much faster internal connection. The result is less data movement, lower latency, and higher overall performance.

Agentic AI represents the next major shift in artificial intelligence, enabling systems that can think, plan, and act independently.

NVIDIA AI networking — Photonics

NVIDIA uses two distinct highways to keep the data moving at the speed of light launched in GTCS

Quantum-X800 InfiniBand (east-west fabric)

This is the east-west fabric — purely for GPUs talking to each other during training so the equivalent of a dedicated storage backend network but for collective operations (AllReduce, AllGather). SHARP v4 doing the reduction inside the switch itself is the same concept as offloading RAID parity to a storage controller

Spectrum-X800 Photonics -(Hyperscale Ethernet Switch)

512 ports at 800 Gb/s, totalling 409.6 Tb/s

This is the scale-out fabric connecting servers, storage, and everything else. The switch has no SFP or QSFP slot so the lasers are built in here which means we plug the fibre cable straight into the box eliminating the sfp’s completely..This is the world’s first Ethernet platform purpose-built for AI that uses Co-Packaged Optics. The advantage is Hyperscale AI fabric switches at 800G/1.6T density so now they can do lot more.

Agentic AI and Physical AI

Traditional AI answers a question when asked. Agentic AI can plan a sequence of tasks, take actions, and improve over time with minimal human involvement. The Vera CPU is purpose-built for this.

Physical AI

Physical AI brings intelligence into the real world – robots, autonomous vehicles, and industrial systems that use sensors to perceive their environment and make decisions without a human in the loop.

AI Factories

An AI factory is the full infrastructure stack – compute, networking, storage, and software needed to run scalable, automated AI workloads across an enterprise. It is the data centre re-imagined around AI as the primary workload rather than traditional IT.

(Visited 6 times, 1 visits today)

By Ash Thomas

Ash Thomas is a seasoned IT professional with extensive experience as a technical expert, complemented by a keen interest in blockchain technology.