a16z Podcast
Summary & Insights
Energy capacity is becoming the definitive ceiling for economic capacity in the AI era. As the conversation reveals, the ability to design a sophisticated chip is no longer the primary hurdle; rather, the bottleneck has shifted to the physical realities of powering, cooling, and connecting those chips at scale. We have entered a paradoxical phase where AI can compress a chip’s design cycle from years to months, yet the time it takes to fabricate the silicon, package it in 3D structures, and integrate it into a rack remains a rigid, multi-month process. This disconnect means that by the time a specialized chip is deployed, the AI workloads it was designed for may have already evolved, rendering the hardware prematurely obsolete.
The dialogue between Pat Gelsinger and the hosts highlights a tension between extreme hardware specialization and the necessity of generalizability. While a proliferation of AI startups are creating “niche” chips for specific tasks like pre-fill or decoding, Gelsinger argues that this level of heterogeneity is unsustainable. He posits that the industry will inevitably consolidate around a few winning architectures—not necessarily because of the hardware itself, but because of the software ecosystems and capital required to scale. The “winning” chips will be those that can adapt to evolving algorithmic domains, such as the shift toward reasoning models or the integration of high-precision HPC (High Performance Computing) for chemical and biological analysis.
A significant portion of the discussion centers on the “hideous” state of current memory. Gelsinger asserts that for the first time in thirty years, we are on the precipice of a true memory innovation breakthrough. Current HBM (High Bandwidth Memory) is limited by shoreline bandwidth and thermal constraints, creating a massive gap between compute speed and data retrieval. The path forward involves bringing memory and compute closer together through new materials—such as ferroelectrics—and non-capacitive high-density structures. However, the physical limits of “stacking” provide a hard ceiling; while the industry pushes for taller memory stacks, yield rates drop exponentially, suggesting that a “mid-rise” approach of three to five layers is the likely engineering sweet spot.
The conversation then pivots to the “plumbing” of AI: networking and power. Gelsinger predicts the “death of copper,” arguing that as clusters grow, the cost and energy required to push signals through copper wires become prohibitive. The industry is moving toward optical connectivity, with a predicted tipping point between 2028 and 2029. This shift will likely lead to Optical Circuit Switching (OCS) and a convergence of “scale-up” and “scale-out” architectures. Furthermore, the energy crisis is described as a “head-butt” for the industry. The lack of new nuclear capacity and the inefficiency of current power grids mean that many data center projects may default simply because the energy cannot be delivered to the site.
Finally, the speakers reflect on the evolution of virtualization, drawing from their shared history at VMware. They explore the concept of “virtualization for agents,” suggesting that the next great abstraction layer won’t be designed for human users, but for AI agents. This requires a complete reimagining of security profiles, migration (a “vMotion for agents”), and resource management. In this future, the virtual machine serves as a performant, secure container for an agent’s operations, governed by human-set “constitutions” or policies but optimized for the near-instantaneous speed and scale at which agents operate.
Major Discussion Themes
The Shift in Hardware Bottlenecks
The conversation emphasizes that the “chip” is no longer the unit of innovation; the “rack” is. While AI-driven EDA (Electronic Design Automation) tools have made logic design easier, the physical constraints of fabrication, 3D packaging, and thermal dissipation have become the new primary bottlenecks. The “time-to-silicon” lag creates a dangerous gap where hardware development cannot keep pace with the weekly evolution of AI software.
The Memory Wall and Material Science
Gelsinger argues that the industry has been stagnant in memory innovation for three decades due to violent commoditization cycles. The current AI boom, however, provides the capital incentive to move beyond DRAM and HBM. The focus is shifting toward new physics—non-capacitive memory and 3D integration—to solve the “shoreline bandwidth” problem, where the physical edge of the chip limits how much data can enter the compute core.
Optical Networking and the End of Copper
There is a strong consensus that copper is becoming an inefficient waveguide for large-scale AI clusters. The transition to Co-Packaged Optics (CPO) and optical switching is viewed as inevitable to reduce power consumption and latency. This transition is expected to materialize fully by 2028-2029, transforming how clusters are architected and reducing the distinction between internal (scale-up) and external (scale-out) networking.
The Energy-Economic Link
The discussion frames energy not as a utility, but as a hard limit on economic growth. With national energy capacity growing at a fraction of the rate of AI demand, the “revenge of Edison” is seen in the push for 800V DC data centers to reduce conversion losses. The survival of the AI build-out depends on a “Renaissance” in power generation (specifically nuclear) and more efficient power delivery networks.
Agent-Centric Virtualization
The speakers propose a shift from human-centric computing to agent-centric computing. This involves rebuilding the virtualization stack to manage the lifecycle, security, and migration of AI agents. The design constraint shifts from “how does a human interact with this OS?” to “how does an agent execute this task securely and performantly across a distributed hardware fabric?”
Surprising Insights
- Design vs. Embodiment Gap: AI can now design a chip in 3 months, but it still takes 9 months to get that design into a usable rack-scale solution.
- The “Mid-Rise” Stack: While there is a push for massive 16-32 layer memory stacks, yield math suggests the industry will settle on “mid-rise” stacks of 3-5 layers to avoid exponential failure rates.
- The Death of Specialized Chips: Contrary to the current trend of “one chip for every task,” Gelsinger predicts a return to generalizability because AI workloads evolve too fast for specialized silicon to remain relevant.
- Memory’s 30-Year Drought: The realization that almost no major new memory architectures have successfully launched in 30 years, making the current movement toward new materials a historical anomaly.
- Copper as a Waveguide: The counterintuitive point that copper is currently being forced to act as an optical waveguide, making it more expensive and less efficient than actual optics at distances as short as five meters.
- Energy-Driven Defaults: The prediction that we will see a wave of data center project defaults not because of a lack of capital or GPUs, but because the power grid cannot support the sites.
- Agent Patience: The insight that humans are “patient” compared to AI agents, meaning the virtualization layers for agents must have drastically lower startup and latency times than current VMs.
- The Return of High Precision: While the trend has been toward lower precision (INT8, FP8), reasoning models and scientific AI are bringing a need for 64-bit precision back into the fold.
Practical Takeaways
Infrastructure Planning Framework
- Audit Energy First: Before committing to GPU clusters, validate the 10-year power trajectory of the site. Energy capacity is now the primary leading indicator of project viability.
- Design for Modularity: Given the lag in silicon fabrication, utilize chiplet architectures and modular racks to allow for partial hardware updates without replacing the entire system.
Hardware Strategy
- Prioritize Memory Bandwidth over Raw TFLOPS: Focus on the “memory-to-compute” ratio; raw compute power is useless if the shoreline bandwidth cannot feed the cores.
- Prepare for the Optical Transition: Begin planning for a transition to optical networking (CPO) by 2028 to avoid the “copper wall” in cluster scaling.
Software and Abstraction
- Build for Agent-to-Agent Interaction: When designing new software layers, assume the primary user is an agent. Optimize for API-driven orchestration rather than GUI-driven human interaction.
- Implement “Constitutional” Policy Layers: Create a separate policy and governance layer (the “constitution”) that sits above the virtualization layer to manage agent behavior and security.
Engineering Focus
- Focus on “Plumbing”: In the current era, the most significant gains are found in power conversion (e.g., 800V DC) and thermal dissipation (e.g., liquid cooling and new materials) rather than just logic gates.
- Balance Precision: Ensure hardware fleets can handle a mix of low-precision (for inference) and high-precision (for reasoning/HPC) workloads to remain flexible as models evolve.
🛍️ Products & Resources Mentioned
- ⚡ DeviceNVIDIA Blackwell — Mentioned as a potential benchmark for the last generation of chips designed before AI fully took over the design process.View on Amazon →
- ⚡ DeviceD-Matrix AI Chips — Discussed as a company working on bringing memory and compute closer together to solve bandwidth limitations.View on Amazon →
Năng lực năng lượng đang trở thành rào cản định hình cho năng lực kinh tế trong kỷ nguyên AI. Như cuộc thảo luận đã tiết lộ, khả năng thiết kế một con chip tinh vi không còn là trở ngại chính; thay vào đó, nút thắt cổ chai đã chuyển sang những thực tế vật lý về việc cung cấp năng lượng, làm mát và kết nối các con chip đó ở quy mô lớn. Chúng ta đã bước vào một giai đoạn nghịch lý khi AI có thể nén chu kỳ thiết kế chip từ nhiều năm xuống còn vài tháng, nhưng thời gian để chế tạo silicon, đóng gói trong các cấu trúc 3D và tích hợp vào một tủ rack vẫn là một quy trình cứng nhắc kéo dài nhiều tháng. Sự mất kết nối này đồng nghĩa với việc vào thời điểm một con chip chuyên dụng được triển khai, các khối lượng công việc AI mà nó được thiết kế để xử lý có thể đã tiến hóa, khiến phần cứng trở nên lỗi thời một cách sớm sủa.
Cuộc đối thoại giữa Pat Gelsinger và các người dẫn chương trình làm nổi bật sự căng thẳng giữa tính chuyên biệt hóa cực độ của phần cứng và sự cần thiết của tính tổng quát. Trong khi hàng loạt các startup AI đang tạo ra các con chip “ngách” cho các tác vụ cụ thể như tiền nạp (pre-fill) hoặc giải mã (decoding), Gelsinger lập luận rằng mức độ không đồng nhất này là không bền vững. Ông cho rằng ngành công nghiệp chắc chắn sẽ hợp nhất xung quanh một vài kiến trúc chiến thắng—không nhất thiết vì bản thân phần cứng, mà vì hệ sinh thái phần mềm và nguồn vốn cần thiết để mở rộng quy mô. Những con chip “chiến thắng” sẽ là những con chip có thể thích nghi với các lĩnh vực thuật toán đang phát triển, chẳng hạn như sự chuyển dịch sang các mô hình suy luận hoặc việc tích hợp HPC (Điện toán hiệu năng cao) độ chính xác cao cho phân tích hóa học và sinh học.
Một phần đáng kể của cuộc thảo luận tập trung vào tình trạng “tệ hại” của bộ nhớ hiện nay. Gelsinger khẳng định rằng lần đầu tiên trong ba mươi năm qua, chúng ta đang đứng trước ngưỡng cửa của một bước đột phá thực sự về đổi mới bộ nhớ. HBM (Bộ nhớ băng thông cao) hiện nay bị hạn chế bởi băng thông đường biên (shoreline bandwidth) và các ràng buộc về nhiệt, tạo ra một khoảng cách khổng lồ giữa tốc độ tính toán và truy xuất dữ liệu. Hướng đi tiếp theo là đưa bộ nhớ và tính toán lại gần nhau hơn thông qua các vật liệu mới—như vật liệu sắt điện (ferroelectrics)—và các cấu trúc mật độ cao không dựa trên tụ điện. Tuy nhiên, các giới hạn vật lý của việc “xếp chồng” tạo ra một mức trần cứng; trong khi ngành công nghiệp thúc đẩy các chồng bộ nhớ cao hơn, tỷ lệ thành phẩm (yield rates) lại giảm theo cấp số nhân, cho thấy phương pháp “tầng trung” từ ba đến năm lớp có khả năng là điểm tối ưu về kỹ thuật.
Cuộc hội thoại sau đó chuyển sang “hệ thống dẫn” của AI: mạng kết nối và năng lượng. Gelsinger dự báo về “cái chết của đồng”, lập luận rằng khi các cụm máy chủ (clusters) phát triển, chi phí và năng lượng cần thiết để đẩy tín hiệu qua dây đồng trở nên quá đắt đỏ. Ngành công nghiệp đang chuyển sang kết nối quang học, với điểm xoay dự kiến nằm trong khoảng từ năm 2028 đến 2029. Sự chuyển dịch này có khả năng dẫn đến Chuyển mạch mạch quang (OCS) và sự hội tụ của kiến trúc “mở rộng theo chiều dọc” (scale-up) và “mở rộng theo chiều ngang” (scale-out). Hơn nữa, cuộc khủng hoảng năng lượng được mô tả như một “cú húc đầu” đối với ngành công nghiệp. Việc thiếu hụt năng lực hạt nhân mới và sự kém hiệu quả của lưới điện hiện tại có nghĩa là nhiều dự án trung tâm dữ liệu có thể thất bại đơn giản vì năng lượng không thể được cung cấp đến địa điểm xây dựng.
Cuối cùng, các diễn giả suy ngẫm về sự phát triển của ảo hóa, rút ra từ lịch sử chung của họ tại VMware. Họ khám phá khái niệm “ảo hóa cho các tác nhân” (virtualization for agents), gợi ý rằng lớp trừu tượng lớn tiếp theo sẽ không được thiết kế cho người dùng là con người, mà dành cho các tác nhân AI. Điều này đòi hỏi một sự tái tư duy hoàn toàn về hồ sơ bảo mật, di trú (một dạng “vMotion cho các tác nhân”) và quản lý tài nguyên. Trong tương lai này, máy ảo đóng vai trò là một container hiệu suất cao, bảo mật cho các hoạt động của tác nhân, được điều phối bởi các “hiến pháp” hoặc chính sách do con người thiết lập nhưng được tối ưu hóa cho tốc độ và quy mô gần như tức thời mà các tác nhân vận hành.
Các Chủ đề Thảo luận Chính
Sự dịch chuyển của các nút thắt cổ chai phần cứng
Cuộc hội thoại nhấn mạnh rằng “con chip” không còn là đơn vị đổi mới; mà là “tủ rack”. Trong khi các công cụ EDA (Tự động hóa thiết kế điện tử) do AI thúc đẩy giúp việc thiết kế logic trở nên dễ dàng hơn, thì các ràng buộc vật lý của chế tạo, đóng gói 3D và tản nhiệt đã trở thành những nút thắt cổ chai chính mới. Độ trễ “từ thiết kế đến silicon” tạo ra một khoảng cách nguy hiểm, nơi phát triển phần cứng không thể theo kịp sự tiến hóa hàng tuần của phần mềm AI.
Bức tường bộ nhớ và Khoa học vật liệu
Gelsinger lập luận rằng ngành công nghiệp đã trì trệ trong đổi mới bộ nhớ suốt ba thập kỷ do các chu kỳ hàng hóa hóa (commoditization) khốc liệt. Tuy nhiên, sự bùng nổ AI hiện nay cung cấp động lực về vốn để tiến xa hơn DRAM và HBM. Trọng tâm đang chuyển sang vật lý mới—bộ nhớ không tụ điện và tích hợp 3D—để giải quyết vấn đề “băng thông đường biên”, nơi cạnh vật lý của chip giới hạn lượng dữ liệu có thể đi vào lõi tính toán.
Mạng quang học và sự kết thúc của dây đồng
Có một sự đồng thuận mạnh mẽ rằng đồng đang trở thành một ống dẫn tín hiệu kém hiệu quả cho các cụm AI quy mô lớn. Việc chuyển sang Quang học đóng gói chung (CPO) và chuyển mạch quang học được xem là tất yếu để giảm tiêu thụ điện năng và độ trễ. Quá trình chuyển đổi này dự kiến sẽ hoàn tất vào năm 2028-2029, thay đổi cách kiến trúc các cụm máy chủ và giảm bớt sự phân biệt giữa mạng nội bộ (scale-up) và mạng bên ngoài (scale-out).
Mối liên kết giữa Năng lượng và Kinh tế
Cuộc thảo luận định nghĩa năng lượng không phải là một tiện ích, mà là một giới hạn cứng đối với tăng trưởng kinh tế. Với năng lực năng lượng quốc gia tăng trưởng chỉ bằng một phần nhỏ so với nhu cầu AI, “sự trả thù của Edison” được nhìn thấy trong nỗ lực thúc đẩy các trung tâm dữ liệu DC 800V để giảm tổn thất chuyển đổi. Sự sinh tồn của quá trình xây dựng hạ tầng AI phụ thuộc vào một cuộc “Phục hưng” trong phát điện (đặc biệt là hạt nhân) và các mạng phân phối điện hiệu quả hơn.
Ảo hóa lấy tác nhân làm trung tâm
Các diễn giả đề xuất một sự chuyển dịch từ điện toán lấy con người làm trung tâm sang điện toán lấy tác nhân làm trung tâm. Điều này bao gồm việc xây dựng lại ngăn xếp ảo hóa để quản lý vòng đời, bảo mật và di trú của các tác nhân AI.
Hệ điều hành không còn tập trung vào câu hỏi “con người tương tác với OS này như thế nào?” mà chuyển sang “làm thế nào để một tác nhân (agent) thực thi tác vụ này một cách bảo mật và hiệu quả trên một cấu trúc phần cứng phân tán?”
Những hiểu biết bất ngờ
- Khoảng cách giữa Thiết kế và Hiện thực hóa: AI hiện có thể thiết kế một con chip trong 3 tháng, nhưng vẫn mất đến 9 tháng để đưa thiết kế đó vào một giải pháp thực tế ở quy mô tủ rack.
- Ngăn xếp “Tầm trung” (Mid-Rise Stack): Trong khi có xu hướng đẩy mạnh các ngăn xếp bộ nhớ khổng lồ từ 16-32 lớp, các bài toán về tỷ lệ thành phẩm (yield math) cho thấy ngành công nghiệp sẽ dừng lại ở các ngăn xếp “tầm trung” từ 3-5 lớp để tránh tỷ lệ lỗi tăng theo cấp số nhân.
- Sự kết thúc của các chip chuyên dụng: Trái ngược với xu hướng hiện nay là “mỗi tác vụ một con chip”, Gelsinger dự báo sự trở lại của tính tổng quát, bởi vì khối lượng công việc của AI tiến hóa quá nhanh khiến các chip chuyên dụng không còn phù hợp.
- Cơn hạn hán 30 năm của Bộ nhớ: Sự nhận ra rằng gần như không có kiến trúc bộ nhớ lớn mới nào được triển khai thành công trong 30 năm qua, khiến phong trào chuyển sang các vật liệu mới hiện nay trở thành một hiện tượng hiếm hoi trong lịch sử.
- Đồng làm ống dẫn sóng: Một điểm ngược với trực giác là đồng hiện đang bị ép đóng vai trò như một ống dẫn sóng quang học, khiến nó trở nên đắt đỏ và kém hiệu quả hơn so với quang học thực sự ngay cả ở khoảng cách ngắn chỉ 5 mét.
- Vỡ nợ do năng lượng: Dự báo về một làn sóng vỡ nợ của các dự án trung tâm dữ liệu, không phải vì thiếu vốn hay thiếu GPU, mà vì lưới điện không thể đáp ứng được nhu cầu của các địa điểm này.
- Sự kiên nhẫn của Agent: Nhận định rằng con người “kiên nhẫn” hơn so với các agent AI, điều này có nghĩa là các lớp ảo hóa cho agent phải có thời gian khởi động và độ trễ thấp hơn đáng kể so với các máy ảo (VM) hiện nay.
- Sự trở lại của độ chính xác cao: Trong khi xu hướng hiện nay là giảm độ chính xác (INT8, FP8), các mô hình lập luận và AI khoa học đang đưa nhu cầu về độ chính xác 64-bit quay trở lại.
Bài học thực tiễn
Khung lập kế hoạch hạ tầng
- Kiểm tra năng lượng trước tiên: Trước khi cam kết đầu tư vào các cụm GPU, hãy xác thực quỹ đạo cung cấp điện trong 10 năm của địa điểm đó. Công suất năng lượng hiện là chỉ số dẫn dắt chính cho tính khả thi của dự án.
- Thiết kế theo hướng mô-đun: Do độ trễ trong chế tạo silicon, hãy sử dụng kiến trúc chiplet và tủ rack mô-đun để cho phép cập nhật một phần phần cứng mà không cần thay thế toàn bộ hệ thống.
Chiến lược phần cứng
- Ưu tiên băng thông bộ nhớ hơn TFLOPS thuần túy: Tập trung vào tỷ lệ “bộ nhớ trên tính toán” (memory-to-compute); sức mạnh tính toán thuần túy sẽ trở nên vô dụng nếu băng thông truyền dẫn không thể cung cấp đủ dữ liệu cho các lõi xử lý.
- Chuẩn bị cho chuyển đổi quang học: Bắt đầu lập kế hoạch chuyển sang mạng quang học (CPO) vào năm 2028 để tránh “bức tường đồng” khi mở rộng quy mô cụm máy chủ.
Phần mềm và Trừu tượng hóa
- Xây dựng cho tương tác Agent-với-Agent: Khi thiết kế các lớp phần mềm mới, hãy giả định người dùng chính là một agent. Tối ưu hóa cho điều phối dựa trên API thay vì tương tác của con người dựa trên giao diện đồ họa (GUI).
- Triển khai các lớp chính sách “Hiến pháp”: Tạo một lớp quản trị và chính sách riêng biệt (gọi là “hiến pháp”) nằm trên lớp ảo hóa để quản lý hành vi và bảo mật của agent.
Trọng tâm kỹ thuật
- Tập trung vào “Hệ thống dẫn”: Trong kỷ nguyên hiện nay, những cải tiến đáng kể nhất nằm ở chuyển đổi năng lượng (ví dụ: 800V DC) và tản nhiệt (ví dụ: làm mát bằng chất lỏng và vật liệu mới) chứ không chỉ đơn thuần là các cổng logic.
- Cân bằng độ chính xác: Đảm bảo các đội hình phần cứng có thể xử lý hỗn hợp các khối lượng công việc độ chính xác thấp (cho suy luận) và độ chính xác cao (cho lập luận/HPC) để duy trì sự linh hoạt khi các mô hình tiến hóa.
在 AI 時代,能源容量正成為經濟能力的決定性天花板。正如對話中所揭示的,設計複雜晶片的能力已不再是主要障礙;相反,瓶頸已轉移到大規模供電、冷卻以及連接這些晶片的物理現實上。我們進入了一個矛盾的階段:AI 能將晶片的設計週期從數年縮短至數月,然而製造矽片、將其封裝在 3D 結構中並整合到機架中的時間,依然是一個僵化的、需耗時數月的過程。這種脫節意味著,當一款專用晶片部署時,其設計之初所針對的 AI 工作負載可能已經演進,導致硬體過早過時。
Pat Gelsinger 與主持人之間的對話凸顯了極端硬體專業化與通用性必要性之間的緊張關係。儘管大量 AI 初創公司正在為預填(pre-fill)或解碼(decoding)等特定任務開發「利基」晶片,但 Gelsinger 認為這種程度的異質性是不可持續的。他斷言,產業將不可避免地圍繞少數獲勝的架構進行整合——這未必是因為硬體本身,而是因為擴展所需的軟體生態系統和資本。那些能適應演進中算法領域(例如向推理模型的轉移,或整合用於化學與生物分析的高精度高性能計算 HPC)的晶片,將成為最終的「贏家」。
討論的一個重要部分集中在目前記憶體「糟糕」的狀態。Gelsinger 主張,三十年來我們首次處於真正的記憶體創新突破邊緣。目前的高頻寬記憶體(HBM)受限於岸線頻寬(shoreline bandwidth)和熱限制,導致計算速度與數據檢索之間存在巨大的鴻溝。未來的方向是通過新材料(如鐵電體)和非電容式高密度結構,讓記憶體與計算單元更加接近。然而,「堆疊」的物理極限提供了一個硬天花板;雖然業界在追求更高的記憶體堆疊,但良率呈指數級下降,這表明三到五層的「中層」方案可能是工程上的最佳平衡點。
對話隨後轉向 AI 的「管線」:網絡與電力。Gelsinger 預言「銅線的死亡」,他認為隨著集群規模的增長,將信號推過銅線所需的成本和能量將變得高不可攀。產業正向光互連(optical connectivity)邁進,預計臨界點將在 2028 年至 2029 年之間。這一轉變可能會導致光迴路切換(OCS)的普及,以及「向上擴展」(scale-up)與「向外擴展」(scale-out)架構的融合。此外,能源危機被描述為產業的一次「正面碰撞」。由於缺乏新的核電容量以及現有電網的低效,許多數據中心項目可能會僅因無法將電力輸送至現場而宣告失敗。
最後,講者們結合在 VMware 的共同經歷,反思了虛擬化的演進。他們探討了「針對 Agent(智能體)的虛擬化」概念,認為下一個偉大的抽象層將不再是為人類用戶設計,而是為 AI Agent 設計。這需要對安全配置文件、遷移(類似於「Agent 版 vMotion」)和資源管理進行完全的重新構思。在這樣的未來中,虛擬機將成為 Agent 運作的高性能、安全容器,由人類設定的「憲法」或政策管轄,但針對 Agent 運作的近乎即時的速度與規模進行優化。
主要討論主題
硬體瓶頸的轉移
對話強調「晶片」不再是創新的單位,「機架」才是。雖然 AI 驅動的 EDA(電子設計自動化)工具使邏輯設計變得更容易,但製造、3D 封裝和散熱的物理限制已成為新的主要瓶頸。「從設計到矽片」(time-to-silicon)的滯後創造了一個危險的差距,使得硬體開發無法跟上 AI 軟體每週一次的演進速度。
記憶體牆與材料科學
Gelsinger 認為,由於劇烈的商品化週期,記憶體創新在過去三十年一直停滯不前。然而,目前的 AI 熱潮提供了超越 DRAM 和 HBM 的資本動力。焦點正轉向新物理學——非電容式記憶體和 3D 整合——以解決「岸線頻寬」問題(即晶片的物理邊緣限制了進入計算核心的數據量)。
光網絡與銅線的終結
各方達成強烈共識,認為銅線對於大規模 AI 集群而言已成為低效的波導。轉向共封裝光學(CPO)和光切換被視為降低功耗和延遲的必然選擇。這一轉型預計在 2028-2029 年全面實現,從而改變集群的構建方式,並縮小內部(scale-up)與外部(scale-out)網絡之間的區別。
能源與經濟的鏈接
討論將能源定義為經濟增長的硬限制,而非僅僅是一種公共事業。由於國家能源容量的增長速度遠低於 AI 需求的增長速度,推動 800V 直流電數據中心以減少轉換損失被視為「愛迪生的反擊」。AI 基礎設施建設的生存取決於發電(特別是核電)的「文藝復興」以及更高效的電力輸送網絡。
以 Agent 為中心的虛擬化
講者們提出從「以人類為中心的計算」轉向「以 Agent 為中心的計算」。這涉及重建虛擬化棧,以管理 AI Agent 的生命週期、安全性和遷移。
設計限制已從「人類如何與此作業系統(OS)互動?」轉向「代理人(Agent)如何在分佈式硬件構架中,安全且高效地執行此項任務?」。
令人驚訝的洞察
- 設計與實現的差距: AI 現在可以在 3 個月內設計出晶片,但將該設計轉化為可用的機架規模(rack-scale)解決方案仍需 9 個月。
- 「中層」堆疊: 雖然業界正追求 16-32 層的大規模記憶體堆疊,但良率計算顯示,為了避免故障率呈指數級增長,業界最終將定格在 3-5 層的「中層」堆疊。
- 專用晶片的消亡: 與目前「每項任務使用一種晶片」的趨勢相反,Gelsinger 預測將回歸通用性,因為 AI 工作負載的演進速度太快,導致專用矽晶片難以保持時效性。
- 記憶體 30 年的乾涸期: 意識到過去 30 年幾乎沒有任何重大的新記憶體架構成功推出,這使得目前向新材料的轉型成為一個歷史性的反常現象。
- 將銅作為波導: 一個反直覺的觀點是,目前的銅線被強行要求扮演光波導的角色,這使得在短至五公尺的距離上,銅線比真正的光學元件更昂貴且效率更低。
- 能源驅動的違約: 預測將出現一波數據中心項目的違約潮,原因並非缺乏資金或 GPU,而是電網無法支持這些場地。
- 代理人的耐心: 洞察到人類相對於 AI 代理人來說是「有耐心的」,這意味著為代理人設計的虛擬化層,其啟動和延遲時間必須比目前的虛擬機(VM)低得多。
- 高精度的回歸: 儘管目前的趨勢是追求低精度(INT8, FP8),但推理模型和科學 AI 正在將對 64 位元精度的需求重新帶回主流。
實踐要點
基礎設施規劃框架
- 優先審核能源: 在投入 GPU 集群之前,請驗證場地 10 年的電力軌跡。能源容量現在是衡量項目可行性的首要領先指標。
- 為模組化而設計: 考慮到矽晶片製造的滯後性,應利用小晶片(chiplet)架構和模組化機架,以便在不更換整個系統的情況下進行部分硬件更新。
硬件策略
- 記憶體頻寬優先於原始 TFLOPS: 關注「記憶體與計算」的比率;如果邊界頻寬(shoreline bandwidth)無法為核心提供足夠數據,原始計算能力將毫無用處。
- 準備光學轉型: 開始規劃在 2028 年前過渡到光學網絡(CPO),以避免在集群擴展時遇到「銅牆」瓶頸。
軟體與抽象化
- 為「代理人對代理人」的交互而構建: 在設計新軟體層時,假設主要用戶是代理人。針對 API 驅動的編排進行優化,而非針對 GUI 驅動的人機交互。
- 實施「憲法級」策略層: 在虛擬化層之上創建一個獨立的策略與治理層(即「憲法」),用以管理代理人的行為與安全性。
工程重心
- 關注「管線工程」: 在當前時代,最顯著的提升來自於電源轉換(例如 800V DC)和散熱(例如液冷和新材料),而非僅僅是邏輯門。
- 平衡精度: 確保硬件集群能夠處理低精度(用於推理)和高精度(用於推理/高效能計算 HPC)的混合工作負載,以便在模型演進時保持靈活性。
La capacité énergétique devient le plafond définitif de la capacité économique à l’ère de l’IA. Comme le révèle la discussion, la capacité à concevoir une puce sophistiquée n’est plus l’obstacle principal ; le goulot d’étranglement s’est déplacé vers les réalités physiques de l’alimentation, du refroidissement et de l’interconnexion de ces puces à grande échelle. Nous sommes entrés dans une phase paradoxale où l’IA peut compresser le cycle de conception d’une puce de plusieurs années à quelques mois, alors que le temps nécessaire pour fabriquer le silicium, l’encapsuler dans des structures 3D et l’intégrer dans une baie reste un processus rigide de plusieurs mois. Ce décalage signifie qu’au moment où une puce spécialisée est déployée, les charges de travail d’IA pour lesquelles elle a été conçue ont peut-être déjà évolué, rendant le matériel prématurément obsolète.
Le dialogue entre Pat Gelsinger et les animateurs met en lumière une tension entre l’hyper-spécialisation du matériel et la nécessité de généralisation. Alors qu’une prolifération de startups en IA créent des puces « de niche » pour des tâches spécifiques comme le pré-remplissage (*pre-fill*) ou le décodage, Gelsinger soutient que ce niveau d’hétérogénéité n’est pas viable. Il avance que l’industrie convergera inévitablement vers quelques architectures gagnantes — non pas nécessairement en raison du matériel lui-même, mais à cause des écosystèmes logiciels et des capitaux nécessaires pour passer à l’échelle. Les puces « gagnantes » seront celles capables de s’adapter à l’évolution des domaines algorithmiques, comme le passage vers des modèles de raisonnement ou l’intégration du HPC (*High Performance Computing*) de haute précision pour les analyses chimiques et biologiques.
Une partie importante de la discussion se concentre sur l’état « hideux » de la mémoire actuelle. Gelsinger affirme que pour la première fois en trente ans, nous sommes au seuil d’une véritable percée en matière d’innovation mémorielle. La HBM (*High Bandwidth Memory*) actuelle est limitée par la bande passante du littoral (*shoreline bandwidth*) et les contraintes thermiques, créant un fossé massif entre la vitesse de calcul et la récupération des données. La voie à suivre consiste à rapprocher la mémoire et le calcul grâce à de nouveaux matériaux — tels que les ferroélectriques — et des structures haute densité non capacitives. Cependant, les limites physiques de l’empilement constituent un plafond rigide ; alors que l’industrie pousse pour des piles de mémoire plus hautes, les taux de rendement chutent exponentiellement, suggérant qu’une approche « moyenne » de trois à cinq couches est le point d’équilibre technique probable.
La conversation pivote ensuite vers la « plomberie » de l’IA : le réseau et l’énergie. Gelsinger prédit la « mort du cuivre », arguant qu’à mesure que les clusters s’agrandissent, le coût et l’énergie nécessaires pour propulser les signaux à travers des fils de cuivre deviennent prohibitifs. L’industrie se tourne vers la connectivité optique, avec un point de basculement prévu entre 2028 et 2029. Ce changement mènera probablement à la commutation de circuits optiques (OCS) et à une convergence des architectures de montée en échelle verticale (*scale-up*) et horizontale (*scale-out*). Par ailleurs, la crise énergétique est décrite comme un « choc frontal » pour l’industrie. Le manque de nouvelles capacités nucléaires et l’inefficacité des réseaux électriques actuels signifient que de nombreux projets de centres de données pourraient échouer simplement parce que l’énergie ne peut être acheminée jusqu’au site.
Enfin, les intervenants réfléchissent à l’évolution de la virtualisation, s’appuyant sur leur histoire commune chez VMware. Ils explorent le concept de « virtualisation pour agents », suggérant que la prochaine grande couche d’abstraction ne sera pas conçue pour des utilisateurs humains, mais pour des agents d’IA. Cela nécessite une refonte complète des profils de sécurité, de la migration (un « vMotion pour agents ») et de la gestion des ressources. Dans ce futur, la machine virtuelle sert de conteneur performant et sécurisé pour les opérations d’un agent, gouvernée par des « constitutions » ou des politiques définies par l’homme, mais optimisée pour la vitesse et l’échelle quasi instantanées auxquelles opèrent les agents.
Thèmes principaux de la discussion
Le déplacement des goulots d’étranglement matériels
La conversation souligne que la « puce » n’est plus l’unité d’innovation ; c’est la « baie » (*rack*) qui l’est. Alors que les outils EDA (*Electronic Design Automation*) pilotés par l’IA ont facilité la conception logique, les contraintes physiques de fabrication, d’encapsulation 3D et de dissipation thermique sont devenues les nouveaux goulots d’étranglement primaires. Le délai de « mise en silicium » crée un écart dangereux où le développement matériel ne peut suivre l’évolution hebdomadaire des logiciels d’IA.
Le mur de la mémoire et la science des matériaux
Gelsinger soutient que l’industrie a stagné en matière d’innovation mémorielle pendant trois décennies en raison de cycles de commoditisation violents. Le boom actuel de l’IA offre cependant l’incitation financière nécessaire pour dépasser la DRAM et la HBM. L’accent est mis sur une nouvelle physique — mémoire non capacitive et intégration 3D — pour résoudre le problème de la « bande passante du littoral », où le bord physique de la puce limite la quantité de données pouvant entrer dans le cœur de calcul.
Réseautage optique et fin du cuivre
Il existe un fort consensus sur le fait que le cuivre devient un guide d’ondes inefficace pour les clusters d’IA à grande échelle. La transition vers l’optique co-encapsulée (CPO) et la commutation optique est jugée inévitable pour réduire la consommation d’énergie et la latence. Cette transition devrait se concrétiser pleinement d’ici 2028-2029, transformant l’architecture des clusters et réduisant la distinction entre le réseautage interne (*scale-up*) et externe (*scale-out*).
Le lien Énergie-Économie
La discussion présente l’énergie non pas comme un service public, mais comme une limite stricte à la croissance économique. Alors que la capacité énergétique nationale croît à une fraction du rythme de la demande en IA, la « vengeance d’Edison » se manifeste par la promotion de centres de données en courant continu 800V pour réduire les pertes de conversion. La survie du déploiement de l’IA dépend d’une « Renaissance » de la production d’énergie (spécifiquement nucléaire) et de réseaux de distribution d’électricité plus efficaces.
Virtualisation centrée sur l’agent
Les intervenants proposent de passer d’une informatique centrée sur l’humain à une informatique centrée sur l’agent. Cela implique de reconstruire la pile de virtualisation pour gérer le cycle de vie, la sécurité et la migration des agents d’IA.
La contrainte de conception passe de « comment un humain interagit-il avec ce système d’exploitation ? » à « comment un agent exécute-t-il cette tâche de manière sécurisée et performante à travers une infrastructure matérielle distribuée ? »
Perspectives Surprenantes
- Écart entre Conception et Incarnation : L’IA peut désormais concevoir une puce en 3 mois, mais il faut toujours 9 mois pour transformer cette conception en une solution exploitable à l’échelle d’une baie (rack).
- La Pile « Mid-Rise » : Alors qu’une tendance pousse vers des piles de mémoire massives de 16 à 32 couches, les calculs de rendement suggèrent que l’industrie s’orientera vers des piles « mid-rise » de 3 à 5 couches pour éviter des taux de défaillance exponentiels.
- La Mort des Puces Spécialisées : Contrairement à la tendance actuelle d’« une puce pour chaque tâche », Gelsinger prévoit un retour à la généralisation, car les charges de travail de l’IA évoluent trop rapidement pour que le silicium spécialisé reste pertinent.
- La Sécheresse de 30 Ans de la Mémoire : Le constat que presque aucune nouvelle architecture majeure de mémoire n’a été lancée avec succès depuis 30 ans, faisant du mouvement actuel vers de nouveaux matériaux une anomalie historique.
- Le Cuivre comme Guide d’Ondes : Le point contre-intuitif selon lequel le cuivre est actuellement contraint d’agir comme un guide d’ondes optiques, ce qui le rend plus coûteux et moins efficace que l’optique réelle, même sur des distances aussi courtes que cinq mètres.
- Les Défauts de Paiement Dictés par l’Énergie : La prédiction que nous verrons une vague de défauts de projets de centres de données, non pas par manque de capital ou de GPU, mais parce que le réseau électrique ne peut pas supporter les sites.
- La Patience des Agents : L’idée que les humains sont « patients » comparés aux agents d’IA, ce qui signifie que les couches de virtualisation pour les agents doivent avoir des temps de démarrage et de latence radicalement inférieurs à ceux des machines virtuelles (VM) actuelles.
- Le Retour de la Haute Précision : Alors que la tendance était à la basse précision (INT8, FP8), les modèles de raisonnement et l’IA scientifique réintroduisent le besoin d’une précision 64 bits.
enseignements Pratiques
Cadre de Planification des Infrastructures
- Auditer l’Énergie en Priorité : Avant de s’engager dans des clusters de GPU, validez la trajectoire énergétique du site sur 10 ans. La capacité énergétique est désormais le principal indicateur avancé de la viabilité d’un projet.
- Concevoir pour la Modularité : Compte tenu du délai de fabrication du silicium, utilisez des architectures de chiplets et des baies modulaires pour permettre des mises à jour matérielles partielles sans remplacer l’intégralité du système.
Stratégie Matérielle
- Prioriser la Bande Passante Mémoire sur les TFLOPS Bruts : Concentrez-vous sur le ratio « mémoire-calcul » ; la puissance de calcul brute est inutile si la bande passante de bordure (shoreline bandwidth) ne peut pas alimenter les cœurs.
- Préparer la Transition Optique : Commencez à planifier la transition vers le réseau optique (CPO) d’ici 2028 pour éviter le « mur du cuivre » lors de la mise à l’échelle des clusters.
Logiciel et Abstraction
- Construire pour l’Interaction Agent-à-Agent : Lors de la conception de nouvelles couches logicielles, partez du principe que l’utilisateur principal est un agent. Optimisez l’orchestration pilotée par API plutôt que l’interaction humaine pilotée par interface graphique (GUI).
- Implémenter des Couches de Politiques « Constitutionnelles » : Créez une couche de politique et de gouvernance distincte (la « constitution ») située au-dessus de la couche de virtualisation pour gérer le comportement et la sécurité des agents.
Focus Ingénierie
- Se Concentrer sur la « Plomberie » : À l’ère actuelle, les gains les plus significatifs se trouvent dans la conversion de puissance (ex: 800V DC) et la dissipation thermique (ex: refroidissement liquide et nouveaux matériaux) plutôt que simplement dans les portes logiques.
- Équilibrer la Précision : Assurez-vous que les parcs matériels peuvent gérer un mélange de charges de travail à basse précision (pour l’inférence) et à haute précision (pour le raisonnement/HPC) afin de rester flexibles face à l’évolution des modèles.
Die Energiekapazität entwickelt sich in der Ära der KI zur entscheidenden Obergrenze für die wirtschaftliche Leistungsfähigkeit. Wie das Gespräch verdeutlicht, ist die Fähigkeit, einen hochentwickelten Chip zu entwerfen, nicht mehr das primäre Hindernis; vielmehr hat sich der Engpass auf die physischen Realitäten der Stromversorgung, Kühlung und Vernetzung dieser Chips in großem Maßstab verschoben. Wir sind in eine paradoxe Phase eingetreten, in der die KI den Designzyklus eines Chips von Jahren auf Monate verkürzen kann, während die Zeit für die Herstellung des Siliziums, die Verpackung in 3D-Strukturen und die Integration in ein Rack ein starrer, mehrmonatiger Prozess bleibt. Diese Diskrepanz bedeutet, dass zu dem Zeitpunkt, an dem ein spezialisierter Chip eingesetzt wird, die KI-Workloads, für die er entwickelt wurde, möglicherweise bereits weiterentwickelt wurden, wodurch die Hardware vorzeitig veraltet.
Der Dialog zwischen Pat Gelsinger und den Moderatoren beleuchtet das Spannungsfeld zwischen extremer Hardware-Spezialisierung und der Notwendigkeit der Generalisierbarkeit. Während eine Vielzahl von KI-Startups „Nischen-Chips“ für spezifische Aufgaben wie Pre-fill oder Decoding entwickeln, argumentiert Gelsinger, dass dieses Maß an Heterogenität nicht nachhaltig sei. Er vertritt die Ansicht, dass sich die Branche unweigerlich auf einige wenige siegreiche Architekturen konsolidieren wird – nicht notwendigerweise aufgrund der Hardware selbst, sondern aufgrund der Software-Ökosysteme und des Kapitals, das für die Skalierung erforderlich ist. Die „Gewinner-Chips“ werden diejenigen sein, die sich an evolvierende algorithmische Bereiche anpassen können, wie etwa den Trend hin zu Reasoning-Modellen oder die Integration von hochpräzisem HPC (High Performance Computing) für chemische und biologische Analysen.
Ein erheblicher Teil der Diskussion konzentriert sich auf den „entsetzlichen“ Zustand des aktuellen Speichers. Gelsinger behauptet, dass wir uns zum ersten Mal seit dreißig Jahren an der Schwelle zu einem echten Durchbruch bei der Speicherinnovation befinden. Der aktuelle HBM (High Bandwidth Memory) ist durch die Shoreline-Bandbreite und thermische Einschränkungen limitiert, was eine massive Lücke zwischen Rechengeschwindigkeit und Datenabruf schafft. Der Weg nach vorne besteht darin, Speicher und Recheneinheiten durch neue Materialien – wie Ferroelektrika – und nicht-kapazitive Hochdichtestrukturen näher zusammenzubringen. Die physischen Grenzen des „Stacking“ (Stapelns) bilden jedoch eine harte Obergrenze; während die Industrie auf höhere Speicherstapel drängt, sinken die Ausbeuteraten exponentiell, was darauf hindeutet, dass ein „Mid-Rise“-Ansatz von drei bis fünf Schichten der wahrscheinlich optimale Punkt in der Technik ist.
Das Gespräch wendet sich dann der „Installation“ der KI zu: Vernetzung und Stromversorgung. Gelsinger prognostiziert den „Tod des Kupfers“ und argumentiert, dass mit wachsenden Clustern die Kosten und die Energie, die benötigt werden, um Signale durch Kupferleitungen zu drücken, prohibitiv werden. Die Branche bewegt sich in Richtung optischer Konnektivität, wobei ein Wendepunkt zwischen 2028 und 2029 erwartet wird. Dieser Wechsel wird wahrscheinlich zu Optical Circuit Switching (OCS) und einer Konvergenz von „Scale-up“- und „Scale-out“-Architekturen führen. Darüber hinaus wird die Energiekrise als ein „Frontalzusammenstoß“ für die Branche beschrieben. Der Mangel an neuen Kernkraftkapazitäten und die Ineffizienz der aktuellen Stromnetze bedeuten, dass viele Rechenzentrumsprojekte scheitern könnten, schlichtweg weil die Energie nicht an den Standort geliefert werden kann.
Abschließend reflektieren die Sprecher über die Evolution der Virtualisierung, basierend auf ihrer gemeinsamen Geschichte bei VMware. Sie untersuchen das Konzept der „Virtualisierung für Agenten“ und legen nahe, dass die nächste große Abstraktionsschicht nicht für menschliche Nutzer, sondern für KI-Agenten konzipiert wird. Dies erfordert eine komplette Neugestaltung von Sicherheitsprofilen, Migration (ein „vMotion für Agenten“) und Ressourcenmanagement. In dieser Zukunft dient die virtuelle Maschine als performanter, sicherer Container für die Operationen eines Agenten, gesteuert durch vom Menschen festgelegte „Konstitutionen“ oder Richtlinien, aber optimiert für die nahezu augenblickliche Geschwindigkeit und Skalierung, mit der Agenten operieren.
Zentrale Diskussionsthemen
Die Verschiebung der Hardware-Engpässe
Das Gespräch betont, dass nicht mehr der „Chip“ die Einheit der Innovation ist, sondern das „Rack“. Während KI-gestützte EDA-Tools (Electronic Design Automation) das Logik-Design vereinfacht haben, sind die physischen Einschränkungen der Fertigung, des 3D-Packaging und der Wärmeableitung zu den neuen primären Engpässen geworden. Die Verzögerung bei der „Time-to-Silicon“ schafft eine gefährliche Lücke, in der die Hardwareentwicklung nicht mit der wöchentlichen Evolution der KI-Software Schritt halten kann.
Die Speicherwand und Materialwissenschaft
Gelsinger argumentiert, dass die Speicherinnovation aufgrund gewaltsamer Kommoditätszyklen drei Jahrzehnte lang stagniert hat. Der aktuelle KI-Boom bietet jedoch den kapitalen Anreiz, über DRAM und HBM hinauszugehen. Der Fokus verschiebt sich hin zu neuer Physik – nicht-kapazitivem Speicher und 3D-Integration –, um das Problem der „Shoreline-Bandbreite“ zu lösen, bei dem die physische Kante des Chips begrenzt, wie viele Daten in den Rechenkern gelangen können.
Optische Vernetzung und das Ende des Kupfers
Es besteht ein starker Konsens darüber, dass Kupfer ein ineffizienter Wellenleiter für groß angelegte KI-Cluster wird. Der Übergang zu Co-Packaged Optics (CPO) und optischem Switching wird als unvermeidlich angesehen, um den Stromverbrauch und die Latenz zu senken. Es wird erwartet, dass sich dieser Übergang bis 2028–2029 vollständig vollzieht, was die Architektur von Clustern transformiert und die Unterscheidung zwischen interner (Scale-up) und externer (Scale-out) Vernetzung aufhebt.
Die Energie-Ökonomie-Verknüpfung
Die Diskussion betrachtet Energie nicht als bloße Versorgungsleistung, sondern als harte Grenze für das Wirtschaftswachstum. Da die nationale Energiekapazität nur mit einem Bruchteil der Geschwindigkeit der KI-Nachfrage wächst, zeigt sich die „Rache Edisons“ im Bestreben nach 800V-DC-Rechenzentren, um Wandlungsverluste zu reduzieren. Das Überleben des KI-Ausbaus hängt von einer „Renaissance“ der Stromerzeugung (insbesondere Kernenergie) und effizienteren Stromliefernetzwerken ab.
Agentenzentrierte Virtualisierung
Die Sprecher schlagen einen Wechsel vom menschenzentrierten Computing zum agentenzentrierten Computing vor. Dies beinhaltet den Neuaufbau des Virtualisierungs-Stacks, um den Lebenszyklus, die Sicherheit und die Migration von KI-Agenten zu verwalten.
Die Design-Beschränkung verschiebt sich von „Wie interagiert ein Mensch mit diesem Betriebssystem?“ hin zu „Wie führt ein Agent diese Aufgabe sicher und performant über ein verteiltes Hardware-Fabric aus?“
Überraschende Erkenntnisse
- Lücke zwischen Design und Verkörperung: KI kann mittlerweile in 3 Monaten einen Chip entwerfen, aber es dauert immer noch 9 Monate, bis dieses Design in eine nutzbare Lösung auf Rack-Ebene überführt wird.
- Der „Mid-Rise“-Stack: Während es einen Trend zu massiven Speicher-Stacks mit 16–32 Layern gibt, legt die Ausbeutemathematik nahe, dass sich die Branche auf „Mid-Rise“-Stacks von 3–5 Layern einigen wird, um exponentielle Ausfallraten zu vermeiden.
- Das Ende spezialisierter Chips: Entgegen dem aktuellen Trend „ein Chip für jede Aufgabe“ prognostiziert Gelsinger eine Rückkehr zur Generalisierbarkeit, da sich KI-Workloads zu schnell entwickeln, als dass spezialisiertes Silizium relevant bleiben könnte.
- Die 30-jährige Speicher-Dürre: Die Erkenntnis, dass in 30 Jahren fast keine bedeutenden neuen Speicherarchitekturen erfolgreich eingeführt wurden, was die aktuelle Bewegung hin zu neuen Materialien zu einer historischen Anomalie macht.
- Kupfer als Wellenleiter: Der kontraintuitive Punkt, dass Kupfer derzeit gezwungen wird, als optischer Wellenleiter zu fungieren, was es bereits bei Distanzen von nur fünf Metern teurer und weniger effizient macht als echte Optik.
- Energiegetriebene Ausfälle: Die Vorhersage, dass wir eine Welle von Ausfällen bei Rechenzentrumsprojekten sehen werden – nicht wegen eines Mangels an Kapital oder GPUs, sondern weil das Stromnetz die Standorte nicht unterstützen kann.
- Agenten-Geduld: Die Erkenntnis, dass Menschen im Vergleich zu KI-Agenten „geduldig“ sind, was bedeutet, dass die Virtualisierungsschichten für Agenten drastisch geringere Start- und Latenzzeiten haben müssen als aktuelle VMs.
- Die Rückkehr der hohen Präzision: Während der Trend zu geringerer Präzision ging (INT8, FP8), bringen Reasoning-Modelle und wissenschaftliche KI den Bedarf an 64-Bit-Präzision zurück ins Spiel.
Praktische Ableitungen
Rahmenwerk für die Infrastrukturplanung
- Zuerst Energie prüfen: Bevor man sich auf GPU-Cluster festlegt, sollte die 10-jährige Energieprognose des Standorts validiert werden. Die Energiekapazität ist mittlerweile der primäre Frühindikator für die Durchführbarkeit eines Projekts.
- Auf Modularität setzen: Angesichts der Verzögerungen in der Siliziumfertigung sollten Chiplet-Architekturen und modulare Racks genutzt werden, um teilweise Hardware-Updates zu ermöglichen, ohne das gesamte System ersetzen zu müssen.
Hardware-Strategie
- Speicherbandbreite vor rohen TFLOPS priorisieren: Konzentration auf das Verhältnis von „Speicher zu Rechenleistung“; rohe Rechenleistung ist nutzlos, wenn die Shoreline-Bandbreite die Kerne nicht versorgen kann.
- Auf den optischen Übergang vorbereiten: Beginn der Planung für den Übergang zu optischen Netzwerken (CPO) bis 2028, um die „Kupferwand“ bei der Cluster-Skalierung zu vermeiden.
Software und Abstraktion
- Interaktion von Agent zu Agent planen: Gehen Sie bei der Entwicklung neuer Softwareschichten davon aus, dass der Hauptnutzer ein Agent ist. Optimieren Sie auf API-gesteuerte Orchestrierung statt auf GUI-gesteuerte menschliche Interaktion.
- „Konstitutionelle“ Richtlinienschichten implementieren: Erstellen Sie eine separate Policy- und Governance-Schicht (die „Verfassung“), die über der Virtualisierungsschicht sitzt, um das Verhalten und die Sicherheit der Agenten zu steuern.
Engineering-Fokus
- Fokus auf die „Leitungen“: In der aktuellen Ära liegen die signifikantesten Gewinne in der Stromumwandlung (z. B. 800V DC) und der Wärmeabfuhr (z. B. Flüssigkeitskühlung und neue Materialien) und nicht nur in Logikgattern.
- Präzision ausbalancieren: Stellen Sie sicher, dass Hardware-Flotten eine Mischung aus Low-Precision (für Inferenz) und High-Precision (für Reasoning/HPC) Workloads bewältigen können, um flexibel auf die Modellentwicklung zu reagieren.
a16z’s Raghu Raghuram and Guido Appenzeller sit down with Playground Global General Partner and former Intel CEO Pat Gelsinger to discuss the next wave of semiconductor innovation and the physical constraints shaping the AI buildout.
Drawing on his experience designing Intel’s 386 and 486 processors, Pat explains how AI could transform chip design, but also why faster design alone won’t solve the industry’s biggest problems. They examine the bottlenecks in manufacturing, memory bandwidth, advanced packaging, and power, and why today’s explosion of specialized AI chips may eventually consolidate around a smaller number of architectures.
They also discuss the potential for new memory technologies, the shift from copper to optical networking, and why energy capacity could become a major constraint on AI growth. Finally, they revisit Pat’s VMware years to ask what virtualization might look like when infrastructure is built for agents rather than humans.
Resources:
Follow Pat Gelsinger on X: https://x.com/PGelsinger
Follow Raghu Raghuram on X:https://x.com/RaghuRaghuram
Follow Guido Appenzeller: https://x.com/appenz
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
-
Three Startups Reinventing Critical Infrastructure
This week, a16z American Dynamism Films premiered three short documentaries highlighting companies tackling some of America’s biggest industrial challenges: Ulysses, Mariana Materials, and Radiant. Before watching those films, we’re revisiting conversations with the founders behind…
-
OpenAI’s Joshua Achiam: Did We Already Reach AGI?
Theo Jaffee is joined by Joshua Achiam, Chief Futurist at OpenAI, for a conversation on AI cybersecurity, frontier model capabilities, and why he believes society may have already crossed the threshold into an AGI-era without…
-
Ruby Thelot on Internet Culture, AI, and the Future of Taste
Sophia Dew and Sofia Puccini are joined by Ruby Thelot, designer, artist, cyberethnographer, professor at NYU, and founder of 13101401, for a wide-ranging conversation about internet culture, AI, digital communities, and how technology is reshaping…
-
Marc Andreessen and Chris Dixon: What’s at Stake in Crypto Regulation
Marc Andreessen, Chris Dixon, and Robert Hackett discuss one of the most consequential policy debates facing the crypto industry: the push for comprehensive U.S. market structure legislation and what regulatory clarity could mean for innovation,…
-
How Enterprise AI Really Gets Deployed
Sarah Wang and Kimberly Tan are joined by Jesse Zhang and Ashwin Sreenivas, co-founders of Decagon, to discuss the evolution of enterprise AI agents, why the company increasingly relies on open-source models, and how it…
-
AI for America’s Small Businesses | Lassie
Alex Rampell and Olivia Moore speak with Lassie cofounders Steijn Pelle and Frédéric Renken about bringing AI to one of the most overlooked parts of the economy: small businesses. Inspired by time spent working inside…
-
AI Micro Dramas, Generative Media, and the Future of Creativity
Justine Moore, partner at Andreessen Horowitz, joins New Economies to explore the rapid evolution of generative media and why AI-native content is reaching an inflection point. They discuss the rise of AI micro-dramas, how creators…
-
Fei-Fei Li on Spatial Intelligence and Robotics
Last week, World Labs announced its acquisition of SceniX, bringing together two teams working on one of AI’s biggest unsolved problems: how to give machines a true understanding of the physical world. Martin Casado sits…
-
Steven Sinofsky: AI Doesn’t Need New Rules Yet
Steven Sinofsky joins Theo Jaffee and Sofia Puccini for a conversation on AI regulation, open-source models, and what history can teach us about technological revolutions. Drawing on decades of experience leading products at Microsoft, Sinofsky…
-
Ben Horowitz: The Fight Over Open Source AI
Ben Horowitz joins Theo Jaffee and Sofia Puccini to discuss one of the biggest debates in AI today: the future of open-source models. They examine the growing push to restrict open models, why Ben believes…
