NVIDIA is turning to Palantir’s Foundry platform and its cuOpt optimisation engine to automate allocation decisions across its global hardware supply chain, in a move designed to reduce delays in delivering the components that power its AI systems. According to NVIDIA, the effort is focused on shortening the route from wafer-out to first token, a measure that spans the journey from chip fabrication through rack assembly and, finally, the point at which a system is ready for live use...
Continue Reading This Article
Enjoy this article as well as all of our content, including reports, news, tips and more.
By registering or signing into your SRM Today account, you agree to SRM Today's Terms of Use and consent to the processing of your personal information as described in our Privacy Policy.
The company’s challenge has grown with the scale of its latest infrastructure. A single Grace Blackwell NVL72 rack contains 18 compute trays, each dependent on two Grace CPUs, four Blackwell GPUs and 32 HBM3e memory packages, with parts sourced through a sprawling network of suppliers, original equipment manufacturers and design partners. NVIDIA has said the next-generation supply chain for its Vera Rubin architecture will be roughly twice as large again, underscoring how quickly logistics complexity is rising as its systems become more ambitious.
At the heart of the new approach is a digital operations command centre built in Palantir Foundry. Palantir’s Ontology layer maps factories, supplier commitments, stock levels and production targets into linked objects, giving planners a more structured view of the chain. NVIDIA says cuOpt then pulls directly from that operational model and formulates allocation as a mixed-integer linear programme aimed at minimising what it calls time of ownership, the period between a site receiving materials and finished sub-assemblies leaving the factory.
The system is intended to do more than generate weekly shipping schedules. NVIDIA says it can also expose bottlenecks such as assembly capacity limits in particular regions or shortages of raw memory, helping planners rebalance factory allocations over rolling two-quarter horizons. That matters because parts can arrive through direct inventory, consignment stock or external suppliers, and early shipments may have to wait for missing components before assembly can continue.
NVIDIA and Palantir have both presented the collaboration as a broader effort to bring so-called sovereign intelligence into critical supply chains. StorageReview reported that the Vera Rubin rack alone involves about 1.3 million parts, illustrating why the company has opted for a system that combines optimisation with a live operational data model. NVIDIA has said the aim is to surface constraints earlier and codify the judgement normally applied by experienced planners.
The company also found that pure mathematical optimisation was not enough to reflect the messy realities of supply management. Human planners were factoring in information such as supplier call transcripts, weather forecasts, email exchanges with partners and geopolitical developments. To account for that, NVIDIA said it post-trained Nemotron 3.5 Lightning, an open-weight mixture-of-experts model with 30 billion total parameters and about three billion active parameters per inference pass.
The training pipeline makes use of NVIDIA’s own software stack. NeMo Anonymizer removes sensitive fields, NeMo Data Designer helps balance the examples with synthetic disruption scenarios, and NeMo AutoModel applies low-rank adaptation while leaving the base model frozen. Palantir Autopilot, according to the company, handles lineage, tracking and the delivery of recommendations back into operations.
NVIDIA says the resulting model delivered 86.7 per cent decision accuracy when tested against historical allocation records, comfortably ahead of the larger Nemotron 3 Ultra model and the untuned Lightning base model. It also recorded stronger balanced accuracy and macro-F1 scores, suggesting the domain-specific tuning improved its ability to handle uneven operational scenarios. Fine-tuning took only minutes on two NVIDIA B200 GPUs, though the company acknowledged that forecasting risk further into the future remains difficult.
The system is not static. Operational choices, planner edits, overrides and factory outcomes are continually fed back into the Ontology, creating a live record that NVIDIA says will later be used to build preference pairs for reinforcement learning. Those routines will judge recommendations on allocation precision, policy compliance and evidence grounding, while production models remain isolated from any live, unmonitored retraining.
For NVIDIA, the project is as much about industrial discipline as technical ambition. As the company pushes into increasingly complex AI infrastructure, it is trying to make supply decisions with the same speed and precision as the compute systems those decisions are meant to deliver.
Source: Noah Wire Services



