a16z Podcast
Summary & Insights
The scale of artificial intelligence infrastructure has transitioned from a software problem to a massive industrial and geopolitical coordination challenge. AWS is currently operating at a magnitude where capital expenditures—projected at $220 billion for 2026—are no longer just about buying servers, but about fundamentally altering energy grids and supply chains. The conversation centers on the transition from traditional cloud computing to “agentic workflows,” where the primary users of cloud resources are shifting from human developers to autonomous AI agents that require different permissions, lower latencies, and more transient infrastructure.
The shift toward agentic workflows is forcing a re-architecture of legacy cloud services. Speaker 2 notes that while humans might tolerate certain latencies, AI agents are frequently blocked by “tail latencies” (P999), making performance optimization critical for agent efficiency. Furthermore, the nature of resource consumption is changing; where a human developer might spin up a database for a project, an agent might create a database, perform a task, and destroy it in seconds. This “transient” nature of infrastructure challenges the traditional “five nines” of durability and availability that AWS was built upon, necessitating a new balance between extreme durability and rapid, disposable compute.
Startups remain the “lifeblood” of AWS, but their profile has evolved. Speaker 2 highlights that modern AI startups launch with far more capital and higher valuations than the “two people in a garage” of the early 2000s, yet they still rely on AWS for the security and scalability frameworks that “Neo-clouds” often lack. To lower the barrier to entry, AWS is stripping away the complexity of initial setup—such as VPC and IAM configuration—allowing users to launch accounts in under 30 seconds. This creates a frictionless “on-ramp” that leads into the deeper, more complex enterprise features required as these startups scale into the enterprises of tomorrow.
The “GPU war” is managed through a deliberate allocation strategy. While frontier labs (like Anthropic and OpenAI) consume the vast majority of high-end compute, AWS intentionally reserves capacity for smaller startups to ensure ecosystem health and diversification. This is coupled with a massive vertical integration strategy. By developing their own silicon—Graviton for general compute and Trainium for AI—AWS is reducing the “virtualization tax” and offering a cost-performance alternative to NVIDIA, with Bedrock’s inference traffic largely running on Trainium.
Enterprise adoption of AI is currently stalled not by a lack of capability, but by a “trust gap.” CEOs are primarily concerned with safety: preventing an agent from accidentally deleting a production database or leaking proprietary data. AWS is addressing this through “Bedrock,” which ensures data never leaves the customer’s VPC, and through specialized “Foundational Design Engineering” (FDE) teams. These teams operate on a 45-day sprint model designed to teach customers how to build their own evaluation (eval) loops and guardrails, rather than creating a permanent dependency on external consultants.
The internal transformation at Amazon mirrors the external offering. From HR teams using agents for resource management to finance teams automating tax compliance via “Amazon Quick,” the company is moving toward “agent-first” development. This is fundamentally altering organizational structure; teams that once required ten people to maintain a capability can now function with three or four, allowing the company to be more agile and rotate talent across projects more rapidly.
Major Discussion Themes
- The Industrialization of AI Compute: The conversation reveals that the bottleneck for AI is no longer just chips, but “physical world” constraints: power grids, data center construction, and the availability of specialized labor. AWS is now investing directly in nuclear and solar projects and managing supply chains four to five tiers deep to prevent the kind of shortages seen during previous global crises (e.g., the Thailand floods).
- Agent-Centric Infrastructure: A pivotal shift is occurring from “people-facing” to “agent-facing” APIs. This includes “AWS Context,” a layer that allows agents to find data across fragmented data lakes (S3, Aurora) more efficiently than a human could, and a rethink of IAM permissions to provide time-boxed, granular access specifically for autonomous agents.
- The Vertical Integration of Silicon: AWS is aggressively pursuing a “silicon-first” strategy to escape the limitations of generalized hardware. The progression from Nitro (offloading virtualization) to Graviton (ARM-based efficiency) to Trainium (AI-specific accelerators) is designed to offer a 20% better performance at a 20% lower cost, creating a powerful economic moat.
- The Enterprise Trust Framework: The transition from “Proof of Concept” to “Production” in the enterprise depends on “Eval loops”—the ability to constantly test and back-test agent behavior to prevent “drift” and catastrophic errors. AWS is positioning itself as the “safe” harbor by guaranteeing data residency within the customer’s own virtual private cloud.
Surprising Insights
- The “Durability Paradox”: Agents often want “non-durable” databases that can be created and destroyed instantly, which contradicts the traditional cloud goal of extreme durability (five nines).
- The Invisible Tax Benefit: Data centers provide massive, often invisible, economic boons to local communities; in one case, AWS’s presence reduced the annual tax bill for every resident in a county by $5,000.
- The “Virtualization Tax”: A significant portion of compute power was historically wasted on the “tax” of virtualization; moving these functions to dedicated hardware (Nitro) was the catalyst for AWS’s current performance leads.
- Agent-First Development: AWS’s “frontier teams” are no longer using AI for code completion; they are managing “swarms” of agents that write the bulk of the code.
- Diversification vs. Concentration: Unlike some Neo-clouds that have 30–60% revenue concentration in one or two customers, AWS’s risk is spread across a vast array of clients, making them less vulnerable to the failure of a single “unicorn.”
- The 30-Second Account: AWS is removing the “barrier of entry” (credit cards, VPC setup) for new users to compete with simpler platforms, knowing that users will eventually need the complex tools they’ve built.
Practical Takeaways
- Adopt a “Greenfield” Approach to AI: Do not simply replicate a human’s step-by-step workflow (e.g., “Bob does steps 1-5”). Instead, redesign the process for how a computer would solve it, leveraging massive parallelism.
- Prioritize “Tail Latency” for Agents: When building agentic workflows, focus on P999 latency rather than average latency, as agents are more sensitive to “stalls” in the pipeline than humans are.
- Implement Time-Boxed Permissions: Move away from permanent service roles for agents. Use fine-grained, short-term permissions that expire immediately after a task is completed.
- Build an “Eval-First” Pipeline: Before deploying agents to production, establish a constant loop of testing with labeled data to ensure the agent’s output doesn’t drift over time.
- Leverage ARM Architecture for Cost Reduction: For general-purpose workloads, migrating to ARM-based instances (like Graviton) is the fastest way to reduce cloud spend while maintaining or increasing performance.
- Focus on “Security at Machine Speed”: Shift from human-led alarm monitoring to AI-powered security (like “Continuum”) that can identify and prioritize vulnerabilities in real-time.
🛍️ Products & Resources Mentioned
- 📚 BookThe Goal by Eliyahu M. Goldratt — Mentioned by Matt Garman as a book he read in undergrad that taught him about the nature of constraints in business and systems.View on Amazon →
- ⚡ DeviceAWS Graviton — An ARM-based server chip designed by AWS that offers better price-performance and lower costs compared to traditional architectures.View on Amazon →
- ⚡ DeviceAWS Trainium — A high-performance AI chip designed specifically for training and inference of large-scale machine learning models.View on Amazon →
- ⚡ DeviceAWS Nitro — A system of hardware offload cards that removes virtualization overhead to provide bare-metal performance and improved security.View on Amazon →
Quy mô hạ tầng trí tuệ nhân tạo đã chuyển dịch từ một bài toán phần mềm sang một thách thức điều phối công nghiệp và địa chính trị khổng lồ. AWS hiện đang vận hành ở một tầm vóc mà chi phí đầu tư vốn — dự kiến đạt 220 tỷ USD vào năm 2026 — không còn đơn thuần là việc mua máy chủ, mà là việc thay đổi căn bản các lưới điện và chuỗi cung ứng. Trọng tâm của cuộc thảo luận xoay quanh sự chuyển đổi từ điện toán đám mây truyền thống sang “luồng công việc tác tử” (agentic workflows), nơi người dùng chính của tài nguyên đám mây đang chuyển từ các lập trình viên con người sang các tác tử AI tự trị — những đối tượng yêu cầu các quyền hạn khác biệt, độ trễ thấp hơn và hạ tầng mang tính tạm thời hơn.
Sự chuyển dịch sang luồng công việc tác tử đang buộc các dịch vụ đám mây kế thừa phải tái cấu trúc. Diễn giả 2 lưu ý rằng trong khi con người có thể chấp nhận một mức độ trễ nhất định, các tác tử AI thường xuyên bị cản trở bởi “độ trễ đuôi” (tail latencies – P999), khiến việc tối ưu hóa hiệu suất trở thành yếu tố then chốt cho hiệu quả của tác tử. Hơn nữa, bản chất của việc tiêu thụ tài nguyên cũng đang thay đổi; nếu một lập trình viên con người có thể khởi tạo một cơ sở dữ liệu cho một dự án, thì một tác tử có thể tạo ra cơ sở dữ liệu, thực hiện tác vụ và xóa bỏ nó chỉ trong vài giây. Bản chất “tạm thời” này của hạ tầng thách thức tiêu chuẩn “năm số chín” (99,999%) về độ bền và tính sẵn sàng mà AWS từng xây dựng, đòi hỏi một sự cân bằng mới giữa độ bền cực cao và khả năng tính toán nhanh, dùng một lần.
Các startup vẫn là “mạch máu” của AWS, nhưng đặc điểm của họ đã tiến hóa. Diễn giả 2 nhấn mạnh rằng các startup AI hiện đại khởi nghiệp với số vốn lớn hơn và định giá cao hơn nhiều so với hình ảnh “hai người trong gara” hồi đầu những năm 2000, tuy nhiên họ vẫn dựa vào AWS để có được các khung bảo mật và khả năng mở rộng mà các “Neo-clouds” (đám mây thế hệ mới) thường thiếu. Để hạ thấp rào cản gia nhập, AWS đang lược bỏ sự phức tạp của việc thiết lập ban đầu — chẳng hạn như cấu hình VPC và IAM — cho phép người dùng khởi tạo tài khoản trong chưa đầy 30 giây. Điều này tạo ra một “lối vào” không ma sát, dẫn dắt người dùng đến các tính năng doanh nghiệp sâu hơn và phức tạp hơn khi các startup này phát triển thành những doanh nghiệp của tương lai.
“Cuộc chiến GPU” được quản lý thông qua một chiến lược phân bổ có tính toán. Trong khi các phòng lab tiên phong (như Anthropic và OpenAI) tiêu thụ phần lớn năng lực tính toán cao cấp, AWS chủ động dự phòng công suất cho các startup nhỏ hơn để đảm bảo sức khỏe và sự đa dạng của hệ sinh thái. Điều này đi đôi với chiến lược tích hợp dọc mạnh mẽ. Bằng cách tự phát triển chip silicon của riêng mình — Graviton cho tính toán tổng quát và Trainium cho AI — AWS đang giảm thiểu “thuế ảo hóa” (virtualization tax) và cung cấp một lựa chọn thay thế về chi phí-hiệu năng cho NVIDIA, với lưu lượng suy luận của Bedrock chạy phần lớn trên Trainium.
Việc áp dụng AI trong doanh nghiệp hiện đang bị đình trệ không phải do thiếu năng lực, mà do “khoảng cách niềm tin”. Các CEO chủ yếu lo ngại về an toàn: ngăn chặn một tác tử vô tình xóa cơ sở dữ liệu vận hành hoặc làm rò rỉ dữ liệu độc quyền. AWS đang giải quyết vấn đề này thông qua “Bedrock”, đảm bảo dữ liệu không bao giờ rời khỏi VPC của khách hàng, và thông qua các đội ngũ “Kỹ thuật Thiết kế Nền tảng” (FDE) chuyên biệt. Các đội ngũ này vận hành theo mô hình sprint 45 ngày, được thiết kế để hướng dẫn khách hàng cách xây dựng các vòng lặp đánh giá (eval loops) và rào chắn (guardrails) của riêng họ, thay vì tạo ra sự phụ thuộc vĩnh viễn vào các tư vấn viên bên ngoài.
Sự chuyển đổi nội bộ tại Amazon phản chiếu những gì họ cung cấp ra bên ngoài. Từ các đội ngũ nhân sự sử dụng tác tử để quản lý nguồn lực đến các đội ngũ tài chính tự động hóa tuân thủ thuế thông qua “Amazon Quick”, công ty đang tiến tới phát triển “ưu tiên tác tử” (agent-first). Điều này đang thay đổi căn bản cấu trúc tổ chức; những đội ngũ từng cần mười người để duy trì một năng lực hiện có thể vận hành chỉ với ba hoặc bốn người, cho phép công ty linh hoạt hơn và luân chuyển nhân tài giữa các dự án nhanh chóng hơn.
Các chủ đề thảo luận chính
- Công nghiệp hóa tính toán AI: Cuộc thảo luận tiết lộ rằng điểm nghẽn của AI không còn chỉ là chip, mà là những hạn chế của “thế giới vật lý”: lưới điện, xây dựng trung tâm dữ liệu và sự sẵn có của lao động chuyên môn. AWS hiện đang đầu tư trực tiếp vào các dự án điện hạt nhân và điện mặt trời, đồng thời quản lý chuỗi cung ứng sâu tới bốn hoặc năm cấp để ngăn chặn tình trạng thiếu hụt như trong các cuộc khủng hoảng toàn cầu trước đây (ví dụ: lũ lụt ở Thái Lan).
- Hạ tầng lấy tác tử làm trung tâm: Một sự chuyển dịch then chốt đang diễn ra từ các API “hướng con người” sang API “hướng tác tử”. Điều này bao gồm “AWS Context”, một lớp cho phép các tác tử tìm kiếm dữ liệu xuyên suốt các hồ dữ liệu phân mảnh (S3, Aurora) hiệu quả hơn con người, và việc tư duy lại các quyền IAM để cung cấp quyền truy cập chi tiết, giới hạn thời gian dành riêng cho các tác tử tự trị.
- Tích hợp dọc Silicon: AWS đang quyết liệt theo đuổi chiến lược “ưu tiên silicon” để thoát khỏi những hạn chế của phần cứng tổng quát. Tiến trình từ Nitro (tách biệt ảo hóa) đến Graviton (hiệu quả dựa trên ARM) và Trainium (bộ tăng tốc chuyên dụng cho AI) được thiết kế để mang lại hiệu suất tốt hơn 20% với chi phí thấp hơn 20%, tạo ra một “con hào kinh tế” mạnh mẽ.
- Khung niềm tin doanh nghiệp: Việc chuyển đổi từ “Chứng minh khái niệm” (PoC) sang “Triển khai thực tế” trong doanh nghiệp phụ thuộc vào các “Vòng lặp đánh giá” (Eval loops) — khả năng liên tục kiểm tra và kiểm tra ngược hành vi của tác tử để ngăn chặn sự “sai lệch” (drift) và các lỗi thảm khốc. AWS đang định vị mình là một “bến cảng an toàn” bằng cách đảm bảo dữ liệu lưu trú trong chính đám mây riêng ảo của khách hàng.
Những hiểu biết bất ngờ
- “Nghịch lý về độ bền”: Các tác tử thường yêu cầu các cơ sở dữ liệu “không bền” (non-durable) có thể được tạo ra và hủy bỏ tức thì, điều này trái ngược với mục tiêu truyền thống của đám mây là độ bền cực cao (năm số chín).
- Lợi ích thuế vô hình: Các trung tâm dữ liệu mang lại những lợi ích kinh tế khổng lồ và thường vô hình cho cộng đồng địa phương; trong một trường hợp, sự hiện diện của AWS đã giúp giảm hóa đơn thuế hàng năm cho mỗi cư dân trong một quận xuống 5.000 USD.
Here is the professional translation into Vietnamese, maintaining the HTML structure:
Bài học thực tiễn
- Áp dụng phương pháp “Greenfield” cho AI: Đừng chỉ đơn thuần mô phỏng quy trình làm việc từng bước của con người (ví dụ: “Bob thực hiện các bước 1-5”). Thay vào đó, hãy thiết kế lại quy trình theo cách máy tính sẽ giải quyết, tận dụng khả năng xử lý song song quy mô lớn.
- Ưu tiên “Độ trễ đuôi” (Tail Latency) cho các Agent: Khi xây dựng quy trình làm việc cho agent, hãy tập trung vào độ trễ P999 thay vì độ trễ trung bình, vì các agent nhạy cảm với tình trạng “tắc nghẽn” trong đường ống xử lý hơn con người.
- Triển khai phân quyền giới hạn thời gian (Time-Boxed Permissions): Loại bỏ các vai trò dịch vụ vĩnh viễn cho agent. Sử dụng các quyền hạn chi tiết, ngắn hạn và hết hạn ngay sau khi nhiệm vụ được hoàn thành.
- Xây dựng đường ống “Ưu tiên đánh giá” (Eval-First Pipeline): Trước khi triển khai agent vào môi trường thực tế, hãy thiết lập một vòng lặp kiểm thử liên tục với dữ liệu đã được gán nhãn để đảm bảo đầu ra của agent không bị sai lệch (drift) theo thời gian.
- Tận dụng kiến trúc ARM để giảm chi phí: Đối với các khối lượng công việc đa mục đích, việc chuyển sang các instance dựa trên ARM (như Graviton) là cách nhanh nhất để giảm chi phí đám mây trong khi vẫn duy trì hoặc tăng cường hiệu suất.
- Tập trung vào “Bảo mật với tốc độ máy tính”: Chuyển dịch từ việc giám sát cảnh báo do con người dẫn dắt sang bảo mật hỗ trợ bởi AI (như “Continuum”), có khả năng nhận diện và ưu tiên xử lý các lỗ hổng trong thời gian thực.
人工智慧基礎設施的規模已從單純的軟體問題,轉變為一場巨大的工業與地緣政治協調挑戰。AWS 目前的營運規模已達到一個臨界點,預計 2026 年的資本支出將達 2,200 億美元——這不再僅僅是購買伺服器,而是從根本上改變能源電網與供應鏈。目前的討論核心在於從傳統雲端運算向「代理工作流」(agentic workflows)的轉型;雲端資源的主要使用者正從人類開發者轉向自主 AI 代理(autonomous AI agents),而後者需要不同的權限設定、更低的延遲以及更具瞬時性的基礎設施。
向代理工作流的轉型正迫使舊有的雲端服務進行架構重組。發言者 2 指出,雖然人類可以容忍一定的延遲,但 AI 代理經常被「尾端延遲」(tail latencies, P999)所阻礙,這使得性能優化對於代理效率至關重要。此外,資源消耗的性質也在改變:人類開發者可能會為一個專案啟動一個資料庫,而 AI 代理則可能在數秒內創建資料庫、執行任務並將其銷毀。這種基礎設施的「瞬時性」挑戰了 AWS 建立之上的傳統「五個九」(99.999%)耐用性與可用性標準,因此需要在極端耐用性與快速、可拋棄式運算之間尋找新的平衡。
新創公司仍是 AWS 的「生命線」,但其特徵已演變。發言者 2 強調,現代 AI 新創公司的啟動資金與估值遠高於 2000 年代初期「兩個人在車庫裡」的模式,但他們依然依賴 AWS 提供的安全性與可擴展性框架,而這些正是「新興雲端」(Neo-clouds)通常缺乏的。為了降低進入門檻,AWS 正在簡化初始設定的複雜度(例如 VPC 和 IAM 配置),讓使用者能在 30 秒內啟動帳戶。這創造了一個無摩擦的「進入匝道」,引導使用者在公司規模擴大並轉型為未來的企業時,進而使用更深層、更複雜的企業級功能。
「GPU 戰爭」是透過深思熟慮的分配策略來管理的。雖然前沿實驗室(如 Anthropic 和 OpenAI)消耗了絕大部分的高階運算資源,但 AWS 有意為小型新創公司保留容量,以確保生態系統的健康與多元化。與此同時,AWS 採取了大規模的垂直整合策略。透過開發自有晶片——用於通用運算的 Graviton 和用於 AI 的 Trainium——AWS 正在降低「虛擬化稅」(virtualization tax),並提供 NVIDIA 之外的成本性能替代方案,目前 Bedrock 的推論流量很大一部分運行在 Trainium 上。
企業對 AI 的採納目前停滯不前,並非因為缺乏能力,而是因為存在「信任差距」。執行長們主要擔心安全性:防止代理不小心刪除生產環境的資料庫或洩露專有數據。AWS 透過「Bedrock」解決此問題,確保數據絕不離開客戶的 VPC,並透過專業的「基礎設計工程」(Foundational Design Engineering, FDE)團隊提供支援。這些團隊採用 45 天的衝刺(sprint)模式,旨在教導客戶如何建立自己的評估(eval)迴圈與護欄(guardrails),而非讓客戶對外部顧問產生永久依賴。
亞馬遜內部的轉型與其外部產品相呼應。從使用代理進行資源管理的人力資源團隊,到透過「Amazon Quick」自動化稅務合規的財務團隊,公司正邁向「代理優先」的開發模式。這從根本上改變了組織結構;過去需要十個人維護的功能,現在只需三四個人即可運作,使公司能夠更加靈活,並在不同專案之間更快速地輪調人才。
主要討論主題
- AI 運算的工業化: 討論揭示了 AI 的瓶頸已不再僅僅是晶片,而是「物理世界」的限制:電網、數據中心建設以及專業勞動力。AWS 目前正直接投資核能與太陽能項目,並管理深達四到五層的供應鏈,以防止像先前全球危機(如泰國水災)時出現的那類短缺。
- 以代理為中心的基礎設施: API 正在發生從「面向人類」到「面向代理」的關鍵轉移。這包括「AWS Context」——一個讓代理能比人類更高效地在碎片化數據湖(S3, Aurora)中尋找數據的層級,以及對 IAM 權限的重新思考,以針對自主代理提供限時且精細的訪問權限。
- 晶片的垂直整合: AWS 正積極追求「晶片優先」策略,以擺脫通用硬體的限制。從 Nitro(卸載虛擬化)到 Graviton(基於 ARM 的效率),再到 Trainium(AI 專用加速器)的演進,旨在提供性能提升 20% 且成本降低 20% 的方案,從而建立強大的經濟護城河。
- 企業信任框架: 企業從「概念驗證」(PoC)轉向「正式生產」取決於「評估迴圈」(Eval loops)——即能夠不斷測試和回測代理行為,以防止「漂移」和災難性錯誤。AWS 通過保證數據保留在客戶自己的虛擬私有雲(VPC)中,將自己定位為「安全」的避風港。
驚人洞察
- 「耐用性悖論」: 代理程式通常需要可以瞬間創建和銷毀的「非耐用性」資料庫,這與傳統雲端追求極端耐用性(五個九)的目標背道而馳。
- 隱形的稅收紅利: 數據中心為當地社區提供了巨大且往往被忽略的經濟利益;在一個案例中,AWS 的進駐使該郡每位居民的年度稅單降低了 5,000 美元。
實務啟示
- 對 AI 採取「綠地」(Greenfield) 方法: 不要簡單地複製人類的逐步工作流(例如:「Bob 執行步驟 1-5」)。相反,應針對電腦如何解決問題重新設計流程,利用大規模的平行運算。
- 優先考慮代理的「尾端延遲」(Tail Latency): 在構建代理工作流時,應關注 P999 延遲而非平均延遲,因為代理對管線中的「停頓」比人類更敏感。
- 實施時限權限管理: 捨棄為代理設定永久性的服務角色。改用細粒度且短期的權限,在任務完成後立即失效。
- 構建「評估優先」(Eval-First) 的管線: 在將代理部署到生產環境之前,建立一個使用標記數據的持續測試循環,以確保代理的輸出不會隨著時間而偏移。
- 利用 ARM 架構降低成本: 對於通用型工作負載,遷移到基於 ARM 的實例(如 Graviton)是在維持或提升性能的同時,降低雲端支出最快的方式。
- 專注於「機器速度的安全防護」: 從由人類主導的警報監控,轉向由 AI 驅動的安全系統(如 “Continuum”),以便即時識別並優先處理漏洞。
L’échelle des infrastructures d’intelligence artificielle est passée d’un problème logiciel à un défi massif de coordination industrielle et géopolitique. AWS opère actuellement à une magnitude telle que les dépenses d’investissement — projetées à 220 milliards de dollars pour 2026 — ne consistent plus seulement à acheter des serveurs, mais à modifier fondamentalement les réseaux énergétiques et les chaînes d’approvisionnement. Le débat se concentre sur la transition du cloud computing traditionnel vers des « flux de travail agentiques » (agentic workflows), où les principaux utilisateurs des ressources cloud ne sont plus des développeurs humains, mais des agents d’IA autonomes nécessitant des permissions différentes, des latences plus faibles et une infrastructure plus éphémère.
Le passage aux flux de travail agentiques impose une réarchitecture des services cloud hérités. L’intervenant 2 note que si les humains peuvent tolérer certaines latences, les agents d’IA sont fréquemment bloqués par les « latences de queue » (P999), rendant l’optimisation des performances critique pour l’efficacité des agents. De plus, la nature de la consommation des ressources évolue ; là où un développeur humain pourrait déployer une base de données pour un projet, un agent pourrait créer une base de données, accomplir une tâche et la détruire en quelques secondes. Cette nature « transitoire » de l’infrastructure remet en question les traditionnels « cinq neuf » (99,999 %) de durabilité et de disponibilité sur lesquels AWS a été bâti, nécessitant un nouvel équilibre entre durabilité extrême et calcul rapide et jetable.
Les startups restent le « moteur » d’AWS, mais leur profil a évolué. L’intervenant 2 souligne que les startups d’IA modernes se lancent avec beaucoup plus de capitaux et des valorisations bien plus élevées que les « deux personnes dans un garage » du début des années 2000, tout en continuant de s’appuyer sur AWS pour les cadres de sécurité et d’évolutivité qui font souvent défaut aux « Néo-clouds ». Pour abaisser la barrière à l’entrée, AWS simplifie la complexité de la configuration initiale — comme le VPC et l’IAM — permettant aux utilisateurs d’ouvrir des comptes en moins de 30 secondes. Cela crée une « rampe d’accès » sans friction qui mène vers les fonctionnalités d’entreprise plus profondes et complexes nécessaires à mesure que ces startups deviennent les entreprises de demain.
La « guerre des GPU » est gérée par une stratégie d’allocation délibérée. Alors que les laboratoires de pointe (comme Anthropic et OpenAI) consomment la vaste majorité du calcul haut de gamme, AWS réserve intentionnellement de la capacité pour les plus petites startups afin de garantir la santé et la diversification de l’écosystème. Ceci s’accompagne d’une stratégie massive d’intégration verticale. En développant ses propres puces — Graviton pour le calcul général et Trainium pour l’IA — AWS réduit la « taxe de virtualisation » et propose une alternative coût-performance face à NVIDIA, le trafic d’inférence de Bedrock tournant largement sur Trainium.
L’adoption de l’IA par les entreprises est actuellement freinée non pas par un manque de capacités, mais par un « déficit de confiance ». Les PDG s’inquiètent principalement de la sécurité : empêcher un agent de supprimer accidentellement une base de données de production ou de divulguer des données propriétaires. AWS répond à cela via « Bedrock », qui garantit que les données ne quittent jamais le VPC du client, et via des équipes spécialisées de « Foundational Design Engineering » (FDE). Ces équipes opèrent selon un modèle de sprint de 45 jours conçu pour apprendre aux clients à construire leurs propres boucles d’évaluation (eval loops) et garde-fous, plutôt que de créer une dépendance permanente envers des consultants externes.
La transformation interne chez Amazon reflète son offre externe. Des équipes RH utilisant des agents pour la gestion des ressources aux équipes financières automatisant la conformité fiscale via « Amazon Quick », l’entreprise s’oriente vers un développement « agent-first ». Cela modifie fondamentalement la structure organisationnelle ; des équipes qui nécessitaient autrefois dix personnes pour maintenir une capacité peuvent désormais fonctionner avec trois ou quatre, permettant à l’entreprise d’être plus agile et de faire pivoter les talents entre les projets plus rapidement.
Thèmes majeurs de discussion
- L’industrialisation du calcul IA : La conversation révèle que le goulot d’étranglement de l’IA n’est plus seulement les puces, mais des contraintes du « monde physique » : les réseaux électriques, la construction de centres de données et la disponibilité d’une main-d’œuvre spécialisée. AWS investit désormais directement dans des projets nucléaires et solaires et gère des chaînes d’approvisionnement sur quatre ou cinq niveaux pour éviter le genre de pénuries observées lors de crises mondiales précédentes (ex: les inondations en Thaïlande).
- Infrastructure centrée sur l’agent : Un basculement pivot s’opère des API « orientées humain » vers des API « orientées agent ». Cela inclut « AWS Context », une couche permettant aux agents de trouver des données à travers des lacs de données fragmentés (S3, Aurora) plus efficacement qu’un humain, ainsi qu’une refonte des permissions IAM pour fournir des accès granulaires et limités dans le temps, spécifiquement pour les agents autonomes.
- L’intégration verticale du silicium : AWS poursuit agressivement une stratégie « silicon-first » pour s’affranchir des limites du matériel généraliste. La progression de Nitro (déchargement de la virtualisation) vers Graviton (efficacité basée sur ARM) puis vers Trainium (accélérateurs spécifiques à l’IA) est conçue pour offrir une performance 20 % supérieure pour un coût 20 % inférieur, créant ainsi un fossé économique puissant.
- Le cadre de confiance pour l’entreprise : Le passage de la « preuve de concept » à la « production » en entreprise dépend des « boucles d’évaluation » — la capacité de tester et de rétro-tester constamment le comportement des agents pour prévenir la « dérive » et les erreurs catastrophiques. AWS se positionne comme un refuge sûr en garantissant la résidence des données au sein du cloud privé virtuel du client.
Perspectives surprenantes
- Le « paradoxe de la durabilité » : Les agents demandent souvent des bases de données « non durables » pouvant être créées et détruites instantanément, ce qui contredit l’objectif traditionnel du cloud visant une durabilité extrême (cinq neuf).
- L’avantage fiscal invisible : Les centres de données apportent des retombées économiques massives et souvent invisibles aux communautés locales ; dans un cas précis, la présence d’AWS a réduit la facture fiscale annuelle de chaque résident d’un comté de 5 000 $.
Here is the professional translation into French, maintaining the HTML structure:
Enseignements Pratiques
- Adopter une approche « Greenfield » pour l’IA : Ne vous contentez pas de reproduire le flux de travail étape par étape d’un humain (ex: « Bob effectue les étapes 1 à 5 »). Repensez plutôt le processus selon la manière dont un ordinateur le résoudrait, en exploitant un parallélisme massif.
- Prioriser la « latence de queue » (Tail Latency) pour les agents : Lors de la création de flux de travail agentiques, concentrez-vous sur la latence P999 plutôt que sur la latence moyenne, car les agents sont plus sensibles aux « blocages » du pipeline que ne le sont les humains.
- Implémenter des permissions limitées dans le temps : Abandonnez les rôles de service permanents pour les agents. Utilisez des permissions granulaires et à court terme qui expirent immédiatement après l’achèvement d’une tâche.
- Bâtir un pipeline « Eval-First » : Avant de déployer des agents en production, établissez une boucle de tests constants avec des données étiquetées pour garantir que les résultats de l’agent ne dérivent pas avec le temps.
- Exploiter l’architecture ARM pour réduire les coûts : Pour les charges de travail polyvalentes, la migration vers des instances basées sur ARM (comme Graviton) est le moyen le plus rapide de réduire les dépenses cloud tout en maintenant ou en augmentant les performances.
- Se concentrer sur la « sécurité à vitesse machine » : Passez d’une surveillance des alertes dirigée par l’humain à une sécurité propulsée par l’IA (comme « Continuum ») capable d’identifier et de prioriser les vulnérabilités en temps réel.
Die Dimension der KI-Infrastruktur hat sich von einem Softwareproblem zu einer massiven industriellen und geopolitischen Koordinationsherausforderung gewandelt. AWS operiert derzeit in einer Größenordnung, in der Investitionsausgaben – für 2026 auf 220 Milliarden Dollar prognostiziert – nicht mehr nur dem Kauf von Servern dienen, sondern die Energienetze und Lieferketten grundlegend verändern. Im Zentrum der Diskussion steht der Übergang vom traditionellen Cloud-Computing zu „agentischen Workflows“, bei denen sich die primären Nutzer von Cloud-Ressourcen von menschlichen Entwicklern hin zu autonomen KI-Agenten verschieben, die andere Berechtigungen, geringere Latenzen und eine transientere Infrastruktur benötigen.
Der Trend zu agentischen Workflows erzwingt eine Neugestaltung (Re-Architecture) bestehender Cloud-Dienste. Sprecher 2 merkt an, dass Menschen gewisse Latenzen tolerieren mögen, KI-Agenten jedoch häufig durch „Tail Latencies“ (P999) blockiert werden, was die Performance-Optimierung für die Effizienz der Agenten kritisch macht. Zudem verändert sich die Art des Ressourcenverbrauchs: Während ein menschlicher Entwickler eine Datenbank für ein Projekt erstellt, könnte ein Agent eine Datenbank aufbauen, eine Aufgabe ausführen und diese innerhalb von Sekunden wieder löschen. Diese „transiente“ Natur der Infrastruktur stellt die traditionelle „Five Nines“-Garantie (99,999 %) für Haltbarkeit und Verfügbarkeit, auf der AWS aufgebaut wurde, infrage und erfordert ein neues Gleichgewicht zwischen extremer Beständigkeit und schnellen, verfügbaren Rechenkapazitäten (Disposable Compute).
Start-ups bleiben das „Lebenselixier“ von AWS, doch ihr Profil hat sich gewandelt. Sprecher 2 hebt hervor, dass moderne KI-Start-ups mit wesentlich mehr Kapital und höheren Bewertungen starten als die „zwei Personen in einer Garage“ der frühen 2000er Jahre; dennoch verlassen sie sich weiterhin auf AWS für die Sicherheits- und Skalierbarkeits-Frameworks, die es „Neo-Clouds“ oft mangelt. Um die Eintrittshürden zu senken, reduziert AWS die Komplexität der Ersteinrichtung – etwa bei VPC- und IAM-Konfigurationen –, sodass Nutzer Konten in weniger als 30 Sekunden eröffnen können. Dies schafft einen reibungslosen „On-Ramp“, der zu den tieferen, komplexeren Enterprise-Funktionen führt, die benötigt werden, wenn diese Start-ups zu den Unternehmen von morgen heranwachsen.
Der „GPU-Krieg“ wird durch eine bewusste Allokationsstrategie gesteuert. Während Frontier-Labs (wie Anthropic und OpenAI) den Großteil der High-End-Rechenkapazitäten verbrauchen, reserviert AWS gezielt Kapazitäten für kleinere Start-ups, um die Gesundheit und Diversifizierung des Ökosystems zu gewährleisten. Dies ist gekoppelt mit einer massiven Strategie der vertikalen Integration. Durch die Entwicklung eigener Chips – Graviton für allgemeine Berechnungen und Trainium für KI – reduziert AWS die „Virtualisierungssteuer“ und bietet eine Kosten-Leistungs-Alternative zu NVIDIA, wobei der Inferenz-Traffic von Bedrock größtenteils auf Trainium läuft.
Die Einführung von KI in Unternehmen stockt derzeit nicht an mangelnden Fähigkeiten, sondern an einer „Vertrauenslücke“. CEOs sind primär besorgt um die Sicherheit: sie wollen verhindern, dass ein Agent versehentlich eine Produktionsdatenbank löscht oder proprietäre Daten preisgibt. AWS begegnet dem mit „Bedrock“, wodurch sichergestellt wird, dass Daten die VPC des Kunden niemals verlassen, sowie durch spezialisierte „Foundational Design Engineering“ (FDE)-Teams. Diese Teams arbeiten in einem 45-Tage-Sprint-Modell, das darauf ausgelegt ist, Kunden beizubringen, wie sie eigene Evaluations-Loops (Eval-Loops) und Leitplanken (Guardrails) aufbauen, anstatt eine dauerhafte Abhängigkeit von externen Beratern zu schaffen.
Die interne Transformation bei Amazon spiegelt das externe Angebot wider. Von HR-Teams, die Agenten für das Ressourcenmanagement einsetzen, bis hin zu Finanzteams, die die Steuerehrlichkeit via „Amazon Quick“ automatisieren – das Unternehmen bewegt sich hin zu einer „Agent-First“-Entwicklung. Dies verändert die Organisationsstruktur grundlegend: Teams, die einst zehn Personen benötigten, um eine Funktion aufrechtzuerhalten, können nun mit drei oder vier Personen auskommen, was das Unternehmen agiler macht und es ermöglicht, Talente schneller zwischen Projekten zu rotieren.
Zentrale Diskussionsthemen
- Die Industrialisierung von KI-Rechenkapazitäten: Die Diskussion zeigt, dass der Flaschenhals für KI nicht mehr nur die Chips sind, sondern Einschränkungen der „physischen Welt“: Stromnetze, der Bau von Rechenzentren und die Verfügbarkeit von Fachkräften. AWS investiert nun direkt in Kern- und Solarprojekte und steuert Lieferketten über vier bis fünf Ebenen tief, um Engpässe wie in früheren globalen Krisen (z. B. den Thailand-Fluten) zu vermeiden.
- Agentenzentrierte Infrastruktur: Es findet ein entscheidender Wechsel von „menschenorientierten“ zu „agentenorientierten“ APIs statt. Dazu gehört „AWS Context“, eine Ebene, die es Agenten ermöglicht, Daten über fragmentierte Data Lakes (S3, Aurora) effizienter zu finden als es ein Mensch könnte, sowie ein Überdenken der IAM-Berechtigungen, um zeitlich begrenzte, granulare Zugriffe speziell für autonome Agenten bereitzustellen.
- Die vertikale Integration von Silizium: AWS verfolgt aggressiv eine „Silicon-First“-Strategie, um die Einschränkungen generalisierter Hardware zu überwinden. Die Entwicklung von Nitro (Auslagerung der Virtualisierung) über Graviton (ARM-basierte Effizienz) zu Trainium (KI-spezifische Beschleuniger) zielt darauf ab, eine um 20 % bessere Performance bei 20 % geringeren Kosten zu bieten und so einen starken wirtschaftlichen Wettbewerbsvorteil (Economic Moat) zu schaffen.
- Das Trust-Framework für Unternehmen: Der Übergang vom „Proof of Concept“ zur „Produktion“ in Unternehmen hängt von „Eval-Loops“ ab – der Fähigkeit, das Verhalten von Agenten kontinuierlich zu testen und zu prüfen, um „Drift“ und katastrophale Fehler zu vermeiden. AWS positioniert sich als „sicherer Hafen“, indem es die Datenresidenz innerhalb der eigenen Virtual Private Cloud des Kunden garantiert.
Überraschende Erkenntnisse
- Das „Haltbarkeits-Paradoxon“: Agenten benötigen oft „nicht-haltbare“ (non-durable) Datenbanken, die sofort erstellt und wieder gelöscht werden können, was dem traditionellen Cloud-Ziel extremer Haltbarkeit (Five Nines) widerspricht.
- Der unsichtbare Steuervorteil: Rechenzentren bieten lokalen Gemeinden massive, oft unsichtbare wirtschaftliche Vorteile; in einem Fall reduzierte die Präsenz von AWS die jährliche Steuerlast für jeden Einwohner eines Countys um 5.000 Dollar.
Praktische Erkenntnisse
- „Greenfield“-Ansatz bei KI verfolgen: Replizieren Sie nicht einfach den Schritt-für-Schritt-Workflow eines Menschen (z. B. „Bob erledigt die Schritte 1–5“). Gestalten Sie den Prozess stattdessen so um, wie ein Computer ihn lösen würde, und nutzen Sie dabei massives Parallelcomputing.
- „Tail Latency“ für Agenten priorisieren: Konzentrieren Sie sich beim Aufbau agentenbasierter Workflows auf die P999-Latenz anstatt auf die durchschnittliche Latenz, da Agenten empfindlicher auf „Stalls“ (Stockungen) in der Pipeline reagieren als Menschen.
- Zeitlich begrenzte Berechtigungen implementieren: Verabschieden Sie sich von permanenten Service-Rollen für Agenten. Nutzen Sie feingranulare, kurzfristige Berechtigungen, die sofort nach Abschluss einer Aufgabe ablaufen.
- Eine „Eval-First“-Pipeline aufbauen: Bevor Sie Agenten in die Produktion überführen, etablieren Sie einen kontinuierlichen Testzyklus mit gelabelten Daten, um sicherzustellen, dass der Output des Agenten im Laufe der Zeit nicht driftet.
- ARM-Architektur zur Kostensenkung nutzen: Für allgemeine Workloads ist die Migration auf ARM-basierte Instanzen (wie Graviton) der schnellste Weg, die Cloud-Ausgaben zu senken und gleichzeitig die Performance beizubehalten oder zu steigern.
- Fokus auf „Security at Machine Speed“: Wechseln Sie von einer menschlich geführten Alarmüberwachung zu KI-gestützter Sicherheit (wie „Continuum“), die Schwachstellen in Echtzeit identifizieren und priorisieren kann.
a16z’s Raghu Raghuram sits down with AWS CEO Matt Garman to discuss how AI is reshaping the cloud, from the needs of AI-native startups to infrastructure increasingly designed for agents.
Matt explains how AWS is adapting as agents write code and manage infrastructure, why it’s reserving scarce GPU capacity for startups, and where custom chips like Trainium and Graviton fit into the AI stack.
They also discuss Amazon’s $220 billion capital investment, the shifting bottlenecks in the infrastructure buildout, what enterprises need to trust autonomous agents, and how AWS’s own teams are building with agents.
Resources:
Follow Matt Garman on X: https://x.com/mattsgarman
Follow Matt Garman on LinkedIn: https://www.linkedin.com/in/mattgarman
Follow Raghu Raghuram on X:https://x.com/RaghuRaghuram
Learn more about AWS:https://aws.amazon.com/
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
-
What Comes After Mobile? Meta’s Andrew Bosworth on AI and Consumer Tech
Are we nearing the end of the smartphone era? In this episode, a16z Growth General Partner David George talks with Meta CTO Andrew “Boz” Bosworth about what comes after apps and touchscreens. From smart glasses…
-
The Dual-Use Founder: Vets Now Building For America
In today’s world, the battlefield extends far beyond war zones—it’s embedded in our tech stacks, supply chains, and airspace security systems. So who better to solve these modern challenges than those who’ve served on the…
-
The Future of Drone Warfare
War has always been shaped by technology—from steel and gunpowder to GPS and nuclear weapons. But the decisive technologies of tomorrow aren’t coming—they’re already here. In this episode, recorded live at our third annual American…
-
How to Build with the Department of Defense
When people think about startups working with the government, the phrase “black box” often comes up. But what if that box is finally being pried open? In this episode—recorded live at the American Dynamism Summit…
-
The Top 100 GenAI Products, Ranked and Explained
This month, a16z’s Consumer team released the fourth edition of the GenAI 100 — a data-driven ranking of the top 50 AI-first web products and mobile apps, based on unique monthly visits and active users.…
-
Jensen Huang and Arthur Mensch on Winning the Global AI Race
The global race for AI leadership is no longer just about companies—it’s about nations. AI isn’t just computing infrastructure; it’s cultural infrastructure, economic strategy, and national security all rolled into one. In this episode, Jensen…
-
Why AI Voice Feels More Human Than Ever
AI voice technology has been around for years — think Siri or Alexa — but the magic has been missing. That’s changing, and quickly! In this episode, Anish Acharya, General Partner at a16z, and Olivia…
-
From Thesis to Meme to Fund: Building American Dynamism
For over a century, the United States has been the birthplace of world-changing innovation – from the Wright brothers’ first flight to the invention of the transistor, the moon landing, and the birth of the…
-
What Founders Get Wrong About Scaling – With Carta’s CEO
Henry Ward, cofounder and CEO of Carta, has spent over a decade scaling his company from an early-stage startup to a 2,000-person industry leader. In this Speedrun conversation with a16z Games partner Josh Lu, Henry…
-
Creativity vs Control: Where AI Fits in the Creative Toolbox
Being a creator in 2025 is tough—but building for creators is even harder. Perhaps no one understands this conundrum better than Scott Belsky. As the founder of Behance, a longtime executive at Adobe, and an…
