
Ep17. Welcome Jensen Huang | BG2 w/ Bill Gurley & Brad Gerstner
What this covers
Open Source bi-weekly convo w/ Bill Gurley and Brad Gerstner on all things tech, markets, investing & capitalism. This week, Jensen Huang, CEO of NVIDIA, makes a guest appearance. In Bill’s absence, Brad is joined by Clark Tang (Partner at Altimeter) as they discuss with Jensen scaling intelligence towards AGI, the acceleration of machine learning, NVIDIA's competitive advantages, the significance of inference alongside training, future market dynamics in the AI landscape, the impact of AI on various industries, the future of work, inference time reasoning, AI’s potential to enhance productivity, the balance between open source and closed source, Elon’s Memphis Supercluster, X.ai, OpenAI, the safe development of AI, & more. Enjoy another episode of BG2.
Chapters
(00:00) Introduction (1:50) The Evolution of AGI and Personal Assistants (06:03) NVIDIA's Competitive Moat (15:51 ) The Future of Inference and Training in AI (19:01) Building the AI Infrastructure (31:35) Inventing a New Market in an AI Future (38:40) The Impact of OpenAI (43:25) The Future of AI Models (46.44) X.ai and Memphis Supercluster (51:21) Distributed Computing and Inference Scaling (55:54) Inference Time Reasoning and Its Importance (01:00:46) AI's Role in Growing Business and Improving Productivity (01:08:00) Ensuring Safe AI Development (01:12:31)The Balance of Open Source and Closed Source AI
Produced by Benny Beausoleil Music by Yung Spielberg
#jensenhuang #nvidia #bradgerstner #billgurley #clarktang #xai #memphiscluster #elonmusk #noambrown #openai #gptstrawberry
Source description (no synthesized summary yet).
Jensen Huang argues that Nvidia's moat has grown stronger over the past 3-4 years because the company optimizes the entire AI computing stack rather than just chips, enabling 100-1000x computing acceleration through combinatorial improvements across hardware, software, algorithms, and system integration that competitors cannot replicate.
- Nvidia achieved 100,000x marginal cost reduction in computing over 10 years through accelerated computing, tensor cores, HBM, NVLink, new numerical precisions, and algorithms—far exceeding Moore's Law at 100x
- Unlike Intel's single-threaded focus, Nvidia's moat is built on domain-specific libraries (cuDNN, Triton, RAPIDS) that sit below PyTorch and enable parallel computing architecture that requires custom compilers for each ISA
- The AI training flywheel extends beyond model training to data curation, synthetic generation, post-training, and inference scaling—optimizing each step compounds competitive advantage more than optimizing single components
This asset isn't compiled yet
You're seeing its claims, ranked. Compile it to build the argument threads, weight them, and check each claim against your library — the full view.
Building a new supercomputer cluster with 100,000 GPUs ready for training in 19 days is singular—never been done before—and required superhuman engineering from Elon and his team; normally it takes 3 years to plan a supercomputer, 1 year for equipment delivery and setup, but Nvidia's hardened platform made this possible.
“what they achieved is is singular never been done before just to put in perspective 100,000 gpus that's you know easily the fastest supercomputer on the planet as one cluster um a supercomputer uh that you would build would take normally three years to plan right and then they deliver the equipment and it takes one year to get it all working yes we're talking about 19 days”
Pre-training scaling laws are now being supplemented by post-training scaling, multimodality scaling, synthetic data generation, and inference-time scaling, with inference-time scaling now scaling up to achieve intelligence through reasoning, simulation, and reflection rather than just model size.
“we used to talk about pre-trained models and scaling at that level...but now we're seeing scaling with post trining and we're seeing scaling at inference...people used to think that pre-training was hard and inference was easy now everything is hard right...there's a concept of fast thinking and slow thinking and reasoning and reflection and iteration and simulation”
In parallel computing, unlike CPUs with multiple ISAs that can share C compilers, each accelerated computing architecture requires its own compiler, and Nvidia's domain-specific libraries like cuDNN sit below PyTorch and are fundamental to why the company dominates accelerated computing rather than just GPU manufacturing.
“there are a whole lot of people who build gpus...but they're not accelerated Computing companies...you can have uh three different Isa CPU Isis they all have their own C compilers you could take you could take software and compile down to the ISA that's not possible in accelerated Computing that's not possible in parallel Computing the company who comes up with the architecture has to come up with their own openg GL...we revolutionize deep learning because of our domain specific Library called cunn without cudn nobody talks about cunn because it's one layer underneath p”
Nvidia achieved a marginal cost reduction in computing of 100,000x over the course of 10 years, with Moore's Law alone accounting for about 100x, through accelerated computing, tensor cores, HBM, NVLink, new numerical precisions, and new architectures.
“we drove the marginal cost of computing down by 100,000x over the course of 10 years. Moe's law would have been about 100x”
Machine learning can learn faster than static software because software is no longer pre-compiled shrink-wrapped static code but continuously learning systems, enabling the entire stack including hardware, algorithms, and training methods to compound and innovate together.
“back in the old days if you look at the way Moe's law was working the software was static right it was pre it was pre-compiled as shrink wrapped put into a store it was static and the hardware underneath was growing at Moe's law rate right now we've got the whole stack growing right innovating across the whole stack”
Nvidia uses AI agents internally in cyber security, chip design, software engineering, and verification—Hopper and Blackwell wouldn't be possible without AI designers and AI engineers—demonstrating the company's dogfooding of its own technology.
“our cyber security system today can't run without without um uh our own agents uh we have we have uh agents helping design chips Hopper wouldn't be possible black wall would be possible Ruben don't even think about it U we have digital we have we have ai chip designers AI software Engineers AI verification Engineers”
The software was static in the old Moore's Law era, but today the entire computing stack is innovating simultaneously—both hardware and software are improving at extraordinary rates—which compounds into exponential scaling rather than linear Moore's Law improvement.
“back in the old days if you look at the way Moore's law was working the software was static...the hardware underneath was growing at Moore's law rate right now we've got the whole stack growing right innovating across the whole stack...we're seeing scaling that is that is extraordinary”
AI is already revolutionizing work in practice across multiple fields—digital biologists, climate researchers, materials researchers, physical scientists, astrophysicists, quantum chemists, video game designers, manufacturing engineers, and roboticists are all using AI right now to advance their work.
“go go talk to uh digital biologists uh climate Tech researchers material researchers um uh physical sciences astrophysicists Quantum chemists um you go ask uh uh video game designers um uh manufacturing uh um engineers uh roboticist pick your favorite whatever industry you want to go pick right and you go deep in there and you talk to the people that matter and you ask them has AI revolutionized the way you work right”
Grace Blackwell with NVLink was architected specifically to optimize for inference with low time-to-first-token, requiring infinite bandwidth and infinite compute simultaneously to achieve millisecond response times with rich context.
“how do we create such a thing and what came out of it was MV link right you know MV link so that we could take uh these systems that are excellent for training U but when you're done with it the the inference performance is ex ex exceptional...optimize for this time to First token right...insanely hard to do actually because time to First token requires a lot of bandwidth but if your context is also Rich then um you need a lot of flops yeah and so you need an infinite amount of bandwidth infinite amount of flops at the same time...Grace Blackwell MV link for that right”
Analysts in January 2023 forecasted Nvidia would do $26 billion in revenue for 2023, but the company actually did $60 billion—a $34 billion miss—representing the single greatest failure of forecasting the world has ever seen due to analysts' focus on crypto winners and inability to imagine the generative AI transformation.
“at that dinner in January of 23 when we were having dinner the forecast for NVIDIA at that dinner in January of 23 was that you would do 26 billion of Revenue uh for the year 2023 you did 60 billion right the the 25 people let's just let let the truth be known that is the single greatest failure of forecasting the world has ever seen right right”
Nvidia will never be worth more than a billion dollars was a prediction someone made about the company, and it is now worth over $1 trillion, demonstrating the misjudgment of fixed TAM assumptions.
“you made this clear you've got a trillion of old stuff you got to modernize you at least have a trillion of new AI workloads coming on you're give or take you'll do 125 billion in Revenue this year you know there was at one point somebody told you the company would never be worth more than a billion”
Accelerated computing requires different architecture than serial processing: parallel processing benefits from many slower transistors (10x more transistors 20% slower) rather than fewer faster transistors, which is the opposite of what serial/CPU optimization requires.
“parallel processing doesn't require every transistor to be excellent serial processing requires every transistor to be excellent parallel processing requires lots and lots of transistors to be more cost effective I rather have 10 times more transistors 20% slower right than 10 times less transistor 20 % faster does that make sense they would like the opposite”
The Cambrian moment in generative AI occurred with ChatGPT's launch in November 2022, and analysts' failure to forecast Nvidia's 2023 growth was attributable to their focus on cryptocurrency winners rather than imagination of the generative AI transformation.
“of course the Cambrian moment occurred with chat GPT...these 25 analysts were so focused on the crypto winner that they couldn't get their head around an imagination of what was what was happening in the world”
Nemotron 340B is the best model in the world for reward systems and critiquing other models, making it valuable for enhancing any other model including Llama and other competitors regardless of how great those models are.
“our model neotron 350b is 340b is is H the best in the world for reward systems and so it is the best critique okay interesting yeah and so so um a fantastic model for enhancing everybody else El's model so irrespective of how how great somebody else's model is I'd recommend using neotron 340b to enhance and make it better and we've already seen made llama better made all the other models better”
U.S. productivity growth has slowed from 2.5-3% annually in the 1990s to 1.8% in the 2000s to the slowest rate on record in the past 10 years, suggesting the world may be on the verge of a dramatic productivity expansion if AI manufacturing of intelligence is as transformative as described.
“if you look at the 90s our productivity growth in the United States was about 2 and a half to 3% a year okay and then in the 2000s it slowed down to about 1.8% and then the last 10 years has been the slowest productivity growth so that's the amount of Labor and capital or the amount of output we have for a fixed amount of Labor and capital the slowest we've had had on record”
AI will not eliminate jobs but will dramatically change every job and enable more productivity, leading to more job creation rather than elimination, because the production of intelligence increases human capacity to solve problems and generate new economic opportunities.
“AI is not...AI will change every job right AI will have a a a seismic impact on how people think about work...when companies become more productive using artificial intelligence it is likely that it manifests itself into either better earnings or better growth or both right and when that happens the next email from the CEO is likely not a layoff announcement of course because you're growing...we have more ideas than we can explore and we need people to help us think through it...as a result we're going to hire more people as we become more productive”
Nvidia produces approximately 4 million dollars of revenue per employee and 2 million dollars of profits/free cash flow per employee, representing extraordinary operational efficiency and culture of execution.
“you look at the business you know we talk a lot about Fitness and efficiency flat organizations that can execute quickly smaller teams um you know Nvidia is in a league of its own really um you know at about 4 million of Revenue per employee about 2 million of profits are free cash flow per employee”
Nvidia's moat in inference will be greater than in training because Nvidia-trained models are architecturally built for Nvidia and will run on Nvidia, older Nvidia generations left behind from new training purchases become free inference infrastructure that remains Cuda compatible, and Nvidia continuously reinvents algorithms making older architectures better over time.
“if you INF if you train well it is very likely you'll inference well if you built it on this architecture without any consideration it will run on this architecture...you could still go and optimize it for other architectures but at the very minimum since it's already been architect you know built on Nvidia it will run on Nvidia...when you're training new models you want your best new gear to be used for training right which leaves behind gear that you used yesterday well that gear is perfect for inference...there's a trail of free gear there's a trail of free infrastructure behind the new INF structure that's Cuda compatible”
Nvidia is not a market-share-focused company but a market-maker focused on solving new problems and accelerating the machine learning flywheel, and the company does not discuss market share in internal meetings, only new opportunities and making the flywheel faster.
“Nvidia is a market maker not share taker if you look at our company slides we don't we don't show not one day does this company talk about market share not inside all we're talking about is how do we create the next thing what's the next problem we can solve in that flywheel how can we do a better job for people how do we take that that flywheel that used to take about a year how do we crank crank it down to about a month”
Nvidia's fusion of mathematics (algorithms), architecture (GPU), and software (CUDA ecosystem) is the real source of moat, not any single component; this mathematico-architectural-software synthesis cannot be easily replicated piecemeal.
“the mathematics is really what Nvidia is really good at is algorithm right that in that Fusion between the the science above the architecture on the bottom that's what we're really good at”
Inference-time reasoning (OpenAI's o1 model) is a huge deal because it decouples compute from instant response—some inference can take a second, some can take 5 minutes, some can take a night. The quality of output depends on the time allowed and the consequentiality of the question, not on having instant answers.
“it's a huge deal I think the um uh a lot of intelligence can't be done a priori right...and so so whether you think about it from a computer science perspective or you think about it from from a from a intelligence P perspective uh too much of it requires context right um the circumstance right uh the quality the type of answer you're looking for uh sometimes just a quick answer is good enough depends on depends on the the uh um the consequential you know impact of the answer”
Safety in AI should be understood as a system of AIs and engineered systems well-engineered from first principles and well-tested, not a single monolithic AI regulator; regulation should be at the application layer through existing agencies like FAA, NHTSA, FDA rather than creating universal galactic AI councils.
“somebody needs to to to um well everybody needs to start talking about AI as a system of AIS and system of of engineered systems engineered systems that are that are well engineered built from first principles well tested...regulation uh...it's also don't don't don't overreach to the point where some of the regulation ought to be um done most of the regulation ought to be done at the applications right the FAA Nitsa FDA you name it...don't overlook the overwhelming amount of regulation in the world that are going to have to be activated for AI”
Nvidia's relationships with decades-old supply chain partners in Taiwan, Korea, and Japan are critical to the company's competitive moat because they enable the hardened APIs, design rules, and coordination methodologies that ensure components manufactured worldwide come together seamlessly when integrated into cloud data centers.
“the entire U ecosystem of electronics today is dedicated in working with us to build ultimately this cube of a computer integrated into all of these different ecosystems and the coordination is so seamless so there's obviously apis and and methodologies and business processes and design rules that we've propagated back backwards and methodologies and architectures and a apis that we propagated forward that have been hardened for decades Harden for decades”
Future work will be conducted by humans prompting AI agents rather than humans programming computers in C++, similar to how leaders today prompt their teams with context, constraints, and missions while leaving creative space, and will result in inbox full of AI icons representing specialized AI agents that work together and with humans.
“I'm no longer going to program computers with C++ I'm going to program AIS with prompting isn't that right now this is no different than me talking to my you know this morning I I wrote a bunch of emails before I came here I was prompting my team course right yeah and I I would describe the context I would describe the the the fundamental constraints that I I know of and I would describe the mission for them...I'm going to be sending them...in your inbox you have all these little dots and these little faces in the future there's going be little icons of AIS”
Open source models are necessary for enabling industries to develop domain-specific AIs and should coexist with closed-source models that sustain commercial innovation, rather than being viewed as open versus closed but rather open and closed.
“it is it is I believe uh wrong-minded to be uh closed versus open right it should be closed and open open yeah right because open is necessary for many Industries to be activated right now if we didn't have open source how would all these different fields of science be able to activate be activated on AI right right because they have to develop their own domain specific AIS”
A difference exists between a GPU (Graphics chip) and accelerated computing, and a difference between accelerated computing and AI infrastructure, where each layer requires different skills and each success at one layer doesn't guarantee success at the next layer.
“there's a fundamental difference between a model yes and artificial intelligence yes right yeah a model is an essential ingredient correct for artificial intelligence it's necessary but not sufficient correct...and so and the and artificial intelligence is a capability but for what right then what's the application right...you have to understand the taxonomy yes of Stack yeah of the stack yeah and at every layer of the stack there will be opportunities but not infinite opportunities for everybody at every single layer dis sack”
There is a fundamental difference between a model and artificial intelligence: a model is necessary but not sufficient for AI. Models are tools, but artificial intelligence is applied to specific use cases (self-driving cars, robots, chatbots) which are domain-specific and not interchangeable.
“there's a different fundamental difference between a model yes and artificial intelligence yes right...a model is an essential ingredient correct for artificial intelligence it's necessary but not sufficient correct...artificial intelligence is a capability but for what right then what's the application”
Nvidia hopes to become a 50,000-employee company with 100 million AI assistants, where AIs recruit other AIs to solve problems, AI agents work in Slack channels with humans and each other, and the company operates as one large employee base with some biological and some digital members.
“I'm hoping that Nvidia Sunday will be a 50,000 employee company with a 100 million you know AI assistants wow and they're in every single group all right um uh we'll have a whole directory of uh AIS that are just generally good at doing things we'll also have our inbox is going to full of directories of AIS that we work with that we know are really good specialized at our skill and so so um AIS will recruit other AIS to solve problems right AIS will be in you know slack Channels with each other and with humans right and with humans and so so we'll just be one large you know employee base if you will uh some of them are digital in AI some some of them are biological”
Huang believes AI as a tutor, assistant, and brainstorming partner is completely revolutionary for information workers, enabling him to stay relevant and continue learning daily by using AI to research topics, double-check answers, and reveal new knowledge through follow-up questions.
“AI as a tutor AI as an assistant um AI as a you know a a partner to brainstorm with um double check my work um you know boy you guys it's completely revolutionary...there's not one piece of research that I don't involve AI with there's not one question that I even if I know the answer I double check on it with AI and surprisingly you know the next two or three questions I ask it reveals something I didn't know”
The demand for Nvidia Blackwell is insane and will remain insane for the foreseeable future because the computing stack is being reinvented for machine learning and almost everything that was hand-engineered (Excel, PowerPoint, Photoshop, AutoCAD) will become machine-learned, requiring massive compute infrastructure modernization.
“the demand for Blackwell is insane...you said in very plain English the demand is insane for Blackwell that it's going to be that way for as far as you can you know for as far as you can see...the first thing that we are doing is we are Reinventing Computing...almost everything that we do almost every single application word excel PowerPoint uh Photoshop Premiere You AutoCAD you you give me your favorite application that was all hand hand engineered I promise you it will be highly machine learned in the future”
Nvidia is not worried about custom ASICs from cloud providers building in-house because Nvidia's singular mission is to build a universal computing platform for AI that works everywhere, and custom ASICs serve specific applications while Nvidia enables the entire ecosystem.
“as a company we want to be sit situationally aware and I'm very situationally aware of everything around our company and our ecosystem right I'm aware of all the people doing alternative things and and what they're doing and and and and sometimes sometimes it's adversarial to us sometimes it's not I'm I'm super aware of it but that doesn't change what the purpose of the company is yeah the singular purpose of the company is to build an architecture that a platform that could be everywhere right that is our goal we're not trying to take any share from anybody”
A company with $50 billion capex should allocate 100% to generative AI rather than maintaining legacy CPUs, since legacy infrastructure won't improve much with Moore's Law largely ended, and generative AI represents new opportunity growth.
“you have $50 billion of capex you like to spend option a option b bill capex for the future right or build capex like the past right know um you already have the capex of the past corre it's sitting right there it's not getting much better anyways Mor's law has largely ended and so why rebuild that let's just take $50 billion put it into generative AI isn't that right...I would put in 100% of the 50 billion”
Compounding a 4x per-year growth in model and compute size with continuously growing usage demand means demand for millions of GPUs is not speculative—it is a near-certainty based on the mathematics of scaling.
“if you just did that math um and you compound it with you add you compound that with 4X per year on model size and Computing size and then on the other hand demand continues to grow in usage uh do we think that we need millions of gpus No Doubt yeah yeah that is that is a for certainty now”
Distributed training and distributed computing across multiple clusters will eventually scale beyond single clusters; asynchronous distributed computing and federated learning will enable scaling to millions of GPUs, but this doesn't mean demand diminishes—it means the architecture changes while compute demand continues to grow.
“the last part is no um my sense is that uh distributed training will have to work right and my sense is that that uh distributed computing will be invented right and and some form of Federated learning and and distributed you know um asynchronous distributed computing uh is going to is going to um uh be discovered”
The industry is building safety infrastructure at pace and coordinating on best practices without government mandate. This self-coordination among AI developers is under-celebrated and demonstrates responsibility.
“there are many things that that are being done that I think are excellent...the methodologies the red teaming the process um the the the the model cards the you know the evaluation systems the benchmarking systems...all of the harnesses that are being built at the velocity that's been built is incredible and one thing under celebrated...the actors in the space today who are building these AIS are taking seriously and coordinating around best practices”
Open AI is one of the most consequential companies of this era because it's pursuing the AGI vision and building an economic engine that can finance the next frontier of AI models, achieving escape velocity while many other model companies struggle to fund their next generation.
“this this is one of the one of the most consequential companies of our time uh the um uh a uh a pure play um AI company uh pursuing the the uh uh the vision of uh AGI right...they build an economic engine that can Finance the next you know Frontier of models right...open AI clearly has hit that escape velocity they can fund their own future it's not clear to me that many of these other companies can”
Given that Nvidia currently does $125B in revenue out of a several-trillion-dollar total addressable market (modernizing $1T of existing infrastructure plus building new AI factories), there is no mathematical reason why the company's revenue could not 2x or 3x in the future.
“you give or take you'll do 125 billion in Revenue this year you know there was at one point somebody told you the company would never be worth more than a billion as you sit here today is there any reason right if you're only 125 billion out of a multi- trillion Tam that you're not going to have 2X the revenue 3x the revenue in the future that you have today is there any reason your Revenue doesn't”
Jensen loves his work and cannot imagine missing this moment. The work is important enough to pursue even when not always fun. He takes the work very seriously and his quality of life is incredible.
“I couldn't imagine anything else I'd rather be doing...my job isn't fun all the time nor nor do I expect it to be fun all the time...I take the work very seriously I take our responsibility very seriously...my quality of life is incredible”