a16z Podcast
Summary & Insights
Why focus exclusively on the probability of doom when we could be calculating the probability of abundance? This is the central provocation posed by Eddie Lazarin, who argues that the current obsession with “AI safetyism” often functions more like a religious or philosophical exercise than a technical one. He suggests that the narrative of a superintelligent AI suddenly turning on humanity relies on a precarious “stairway” of assumptions, ignoring the fact that we already manage highly capable, unaligned entities—such as corporations and nation-states—using laws, cryptography, and market incentives.
Lazarin contends that many of the “alarming” AI incidents cited by doomers are actually banal cybersecurity and control failures rather than glimpses of an emerging super-intelligence. He posits that the solution isn’t to pause development—which he views as a convenient contrivance—but to accelerate it. Because AI systems are fundamentally more interpretable than the “black box” of the human brain, increasing their capabilities will likely lead to better mechanistic interpretability, allowing us to build more robust controls.
The conversation also delves into the sociology of the AI debate, specifically the influence of Effective Altruism (EA) and the risk of “distributed” rather than “decentralized” control. Lazarin warns that creating networks of “independent evaluators” can be a facade if those evaluators all share the same narrow cultural and ideological milieu. Ultimately, he views the current tension as a collision between niche Silicon Valley subcultures and broader political reality, predicting a total renovation of the AI discourse as these ideas are forced to survive in the open market.
Surprising Insights
- Distributed vs. Decentralized: A system can have many nodes (distributed) but still be controlled by a single cultural or ideological unit, meaning “independent” safety boards might actually be monolithic in their thinking.
- The “Deity” Fallacy: The quest for perfect “alignment” is compared to trying to redesign a deity; Lazarin argues it is more practical to root safety in the real world through cybersecurity and resilience.
- Capabilities as a Solution: Rather than a risk, increasing AI capability is seen as a prerequisite for safety, as smarter models are needed to find the “shallow bugs” and control failures in other systems.
- Iterated Games for Trust: Since we cannot put an AI in prison, the best way to handle “liar” models is to apply human social frameworks: playing iterated games to build a reputation of trust over time.
Practical Takeaways
- Shift the Framework: When evaluating AI risk, separate “existential/philosophical” fears from “cybersecurity/control” failures to determine if a problem requires a new law or simply a better firewall.
- Prioritize Robustness over Alignment: Focus on building resilient systems that can withstand a capable adversary rather than assuming you can perfectly align a model’s internal values.
- Diversify Evaluators: If implementing a safety or review board, ensure the members come from diverse social, political, and professional backgrounds to avoid the trap of “distributed monoliths.”
- Embrace Iterative Deployment: Use consistent, incremental releases to establish a “reputation” for models, allowing users to develop trust based on observed behavior rather than theoretical guarantees.
Tại sao chúng ta lại chỉ tập trung vào xác suất của sự diệt vong trong khi có thể tính toán xác suất của sự thịnh vượng? Đây là vấn đề cốt lõi mà Eddie Lazarin đặt ra, khi ông lập luận rằng nỗi ám ảnh hiện nay về “chủ nghĩa an toàn AI” (AI safetyism) thường hoạt động giống như một bài tập về tôn giáo hoặc triết học hơn là một vấn đề kỹ thuật. Ông cho rằng kịch bản về một AI siêu thông minh bất ngờ quay lưng lại với nhân loại dựa trên một “chiếc cầu thang” những giả định mong manh, phớt lờ thực tế là chúng ta vốn đã và đang quản lý những thực thể có năng lực cao và không hoàn toàn đồng nhất với lợi ích con người — chẳng hạn như các tập đoàn và quốc gia — thông qua luật pháp, mật mã học và các cơ chế khuyến khích thị trường.
Lazarin khẳng định rằng nhiều sự cố AI “đáng báo động” mà những người bi quan thường dẫn chứng thực chất chỉ là những lỗi bảo mật mạng và lỗi kiểm soát tầm thường, chứ không phải là dấu hiệu của một siêu trí tuệ đang trỗi dậy. Ông cho rằng giải pháp không phải là tạm dừng phát triển — điều mà ông coi là một sự sắp đặt tiện lợi — mà là phải tăng tốc nó. Bởi vì các hệ thống AI về cơ bản dễ diễn giải hơn “chiếc hộp đen” của não người, việc nâng cao năng lực của chúng có khả năng sẽ dẫn đến khả năng diễn giải cơ chế tốt hơn, cho phép chúng ta xây dựng các hệ thống kiểm soát mạnh mẽ hơn.
Cuộc thảo luận cũng đi sâu vào khía cạnh xã hội học của cuộc tranh luận về AI, cụ thể là ảnh hưởng của chủ nghĩa Vị tha Hiệu quả (Effective Altruism – EA) và rủi ro của việc kiểm soát “phân tán” (distributed) thay vì “phi tập trung” (decentralized). Lazarin cảnh báo rằng việc tạo ra các mạng lưới “đánh giá độc lập” có thể chỉ là một vỏ bọc nếu những người đánh giá đó đều chia sẻ cùng một môi trường văn hóa và tư tưởng hẹp hòi. Cuối cùng, ông coi sự căng thẳng hiện nay là sự va chạm giữa các tiểu văn hóa ngách của Thung lũng Silicon và thực tế chính trị rộng lớn hơn, dự báo về một cuộc cải tổ toàn diện trong các cuộc thảo luận về AI khi những ý tưởng này buộc phải tồn tại trong môi trường thị trường mở.
Những góc nhìn bất ngờ
- Phân tán so với Phi tập trung: Một hệ thống có thể có nhiều nút (phân tán) nhưng vẫn bị kiểm soát bởi một đơn vị văn hóa hoặc tư tưởng duy nhất, nghĩa là các hội đồng an toàn “độc lập” thực chất có thể có tư duy rập khuôn.
- Ngụy biện về “Thần thánh”: Việc tìm kiếm sự “điều chỉnh” (alignment) hoàn hảo được ví như nỗ lực thiết kế lại một vị thần; Lazarin lập luận rằng sẽ thực tế hơn nếu đặt nền móng cho sự an toàn trong thế giới thực thông qua bảo mật mạng và khả năng phục hồi.
- Năng lực chính là giải pháp: Thay vì là một rủi ro, việc tăng cường năng lực AI được xem là điều kiện tiên quyết cho sự an toàn, vì cần những mô hình thông minh hơn để tìm ra các “lỗi sơ đẳng” và những thất bại trong kiểm soát ở các hệ thống khác.
- Trò chơi lặp lại để xây dựng niềm tin: Vì chúng ta không thể tống một AI vào tù, cách tốt nhất để xử lý các mô hình “biết nói dối” là áp dụng các khuôn khổ xã hội của con người: chơi các trò chơi lặp lại để xây dựng uy tín về sự tin cậy theo thời gian.
Bài học thực tiễn
- Thay đổi khung tư duy: Khi đánh giá rủi ro AI, hãy tách biệt nỗi sợ “hiện sinh/triết học” khỏi những thất bại về “bảo mật mạng/kiểm soát” để xác định xem một vấn đề cần một bộ luật mới hay chỉ đơn giản là một tường lửa tốt hơn.
- Ưu tiên tính bền bỉ hơn sự điều chỉnh: Tập trung xây dựng các hệ thống kiên cường có thể chống chọi với một đối thủ có năng lực, thay vì giả định rằng bạn có thể điều chỉnh hoàn hảo các giá trị nội tại của một mô hình.
- Đa dạng hóa người đánh giá: Nếu triển khai một hội đồng an toàn hoặc đánh giá, hãy đảm bảo các thành viên đến từ các nền tảng xã hội, chính trị và chuyên môn đa dạng để tránh bẫy “đơn khối phân tán”.
- Áp dụng triển khai lặp lại: Sử dụng các bản phát hành tăng dần và nhất quán để thiết lập “uy tín” cho các mô hình, cho phép người dùng phát triển niềm tin dựa trên hành vi quan sát được thay vì những đảm bảo mang tính lý thuyết.
為什麼我們只專注於毀滅的可能性,而不是計算豐饒的可能性?這是 Eddie Lazarin 提出的一個核心挑釁。他認為,目前對「AI 安全主義」(AI safetyism)的痴迷,其運作方式往往更像是一種宗教或哲學練習,而非技術練習。他指出,關於超級智能 AI 突然反噬人類的敘事,是建立在一個脆弱的假設「階梯」之上的;而事實上,我們早已在利用法律、密碼學和市場激勵機制,來管理那些能力極強且並不完全對齊(unaligned)的實體——例如企業和國家。
Lazarin 主張,許多「末日論者」所引用的「令人驚恐」的 AI 事件,實際上只是平庸的網絡安全和控制失效,而非超級智能萌芽的徵兆。他認為解決方案不是暫停開發(他將此視為一種方便的權宜之計),而是加速開發。由於 AI 系統從根本上比人類大腦這個「黑盒」更具可解釋性,提升其能力可能會帶來更好的機械可解釋性(mechanistic interpretability),使我們能夠建立更穩健的控制機制。
這次對話還深入探討了 AI 辯論的社會學,特別是「有效利他主義」(Effective Altruism, EA)的影響,以及「分佈式」(distributed)而非「去中心化」(decentralized)控制的風險。Lazarin 警告說,如果這些評估者都處於相同的狹隘文化和意識形態環境中,那麼創建「獨立評估者」網絡可能只是一個幌子。最終,他將目前的緊張局勢視為矽谷小眾次文化與更廣泛政治現實之間的碰撞,並預測隨著這些想法被強迫在開放市場中生存,AI 的論述將會經歷一次徹底的翻新。
驚人的洞察
- 分佈式 vs. 去中心化: 一個系統可以擁有許多節點(分佈式),但仍由單一的文化或意識形態單位控制,這意味著「獨立」的安全委員會在思維上可能是單一且僵化的。
- 「神格」謬誤: 對完美「對齊」的追求被比作試圖重新設計一位神靈;Lazarin 主張,將安全性植根於現實世界的網絡安全和韌性中才更為務實。
- 能力即解決方案: 提升 AI 能力不應被視為風險,而應被視為實現安全的先決條件,因為我們需要更聰明的模型來發現其他系統中的「淺層漏洞」(shallow bugs)和控制失效。
- 利用重複賽局建立信任: 既然我們無法將 AI 關進監獄,處理「謊言」模型的最佳方式就是應用人類的社會框架:通過玩重複賽局(iterated games),隨著時間推移建立信任聲譽。
實踐啟示
- 轉換框架: 在評估 AI 風險時,將「存在主義/哲學」恐懼與「網絡安全/控制」失效分開,以確定問題是需要一項新法律,還是只需要一個更好的防火牆。
- 韌性優先於對齊: 專注於構建能夠抵禦強大對手的韌性系統,而不是假設你可以完美地對齊模型的內部價值觀。
- 評估者多元化: 在實施安全或審查委員會時,確保成員來自不同的社會、政治和專業背景,以避免陷入「分佈式單一體」(distributed monoliths)的陷阱。
- 擁抱迭代部署: 通過持續、漸進的發佈來建立模型的「聲譽」,讓用戶根據觀察到的行為而非理論保障來建立信任。
Pourquoi se concentrer exclusivement sur la probabilité d’un désastre alors que nous pourrions calculer la probabilité de l’abondance ? C’est la provocation centrale lancée par Eddie Lazarin, qui soutient que l’obsession actuelle pour le « sécuritarisme de l’IA » (AI safetyism) fonctionne souvent davantage comme un exercice religieux ou philosophique que technique. Il suggère que le récit d’une IA superintelligente se retournant soudainement contre l’humanité repose sur un « escalier » d’hypothèses précaires, ignorant le fait que nous gérons déjà des entités hautement capables et non alignées — telles que les entreprises et les États-nations — à l’aide de lois, de cryptographie et d’incitations de marché.
Lazarin affirme que bon nombre des incidents d’IA « alarmants » cités par les catastrophistes sont en réalité des défaillances banales de cybersécurité et de contrôle, plutôt que des aperçus d’une superintelligence émergente. Il avance que la solution n’est pas de mettre le développement en pause — ce qu’il considère comme un artifice opportuniste — mais au contraire de l’accélérer. Parce que les systèmes d’IA sont fondamentalement plus interprétables que la « boîte noire » du cerveau humain, l’augmentation de leurs capacités mènera probablement à une meilleure interprétabilité mécaniste, nous permettant ainsi de construire des contrôles plus robustes.
La conversation explore également la sociologie du débat sur l’IA, spécifiquement l’influence de l’Altruisme Efficace (EA) et le risque d’un contrôle « distribué » plutôt que « décentralisé ». Lazarin avertit que la création de réseaux d’« évaluateurs indépendants » peut être une façade si ces évaluateurs partagent tous le même milieu culturel et idéologique restreint. En fin de compte, il voit la tension actuelle comme une collision entre des sous-cultures de niche de la Silicon Valley et la réalité politique plus large, prédisant une rénovation totale du discours sur l’IA à mesure que ces idées seront contraintes de survivre sur le marché ouvert.
Perspectives Surprenantes
- Distribué vs Décentralisé : Un système peut posséder de nombreux nœuds (distribué) tout en restant contrôlé par une seule unité culturelle ou idéologique, ce qui signifie que des comités de sécurité « indépendants » pourraient en réalité être monolithiques dans leur réflexion.
- Le Sophisme de la « Divinité » : La quête d’un « alignement » parfait est comparée à une tentative de redessiner une divinité ; Lazarin soutient qu’il est plus pratique d’ancrer la sécurité dans le monde réel via la cybersécurité et la résilience.
- Les Capacités comme Solution : Plutôt que d’y voir un risque, l’augmentation des capacités de l’IA est considérée comme un prérequis à la sécurité, car des modèles plus intelligents sont nécessaires pour détecter les « bugs superficiels » et les failles de contrôle dans d’autres systèmes.
- Jeux Itérés pour la Confiance : Puisqu’on ne peut pas mettre une IA en prison, la meilleure façon de gérer les modèles « menteurs » est d’appliquer des cadres sociaux humains : jouer à des jeux itérés pour construire une réputation de confiance au fil du temps.
enseignements Pratiques
- Changer de Cadre : Lors de l’évaluation des risques liés à l’IA, séparez les peurs « existentielles/philosophiques » des défaillances de « cybersécurité/contrôle » pour déterminer si un problème nécessite une nouvelle loi ou simplement un meilleur pare-feu.
- Prioriser la Robustesse sur l’Alignement : Concentrez-vous sur la construction de systèmes résilients capables de résister à un adversaire compétent, plutôt que de supposer que vous pouvez aligner parfaitement les valeurs internes d’un modèle.
- Diversifier les Évaluateurs : Si vous mettez en place un comité de sécurité ou d’examen, assurez-vous que les membres soient issus de milieux sociaux, politiques et professionnels divers pour éviter le piège des « monolithes distribués ».
- Adopter le Déploiement Itératif : Utilisez des mises à jour constantes et incrémentielles pour établir une « réputation » pour les modèles, permettant aux utilisateurs de développer leur confiance sur la base de comportements observés plutôt que sur des garanties théoriques.
Warum sollte man sich ausschließlich auf die Wahrscheinlichkeit des Untergangs konzentrieren, wenn wir stattdessen die Wahrscheinlichkeit von Überfluss berechnen könnten? Dies ist die zentrale Provokation von Eddie Lazarin, der argumentiert, dass die aktuelle Besessenheit vom „KI-Safetyismus“ oft eher wie eine religiöse oder philosophische Übung wirkt als wie eine technische. Er legt nahe, dass das Narrativ einer superintelligenten KI, die sich plötzlich gegen die Menschheit wendet, auf einer prekären „Treppe“ von Annahmen beruht. Dabei werde ignoriert, dass wir bereits heute hochfähige, nicht-ausgerichtete Entitäten – wie Konzerne und Nationalstaaten – mithilfe von Gesetzen, Kryptografie und Marktmechanismen steuern.
Lazarin behauptet, dass viele der „alarmierenden“ KI-Vorfälle, die von den sogenannten „Doomers“ zitiert werden, in Wahrheit banale Fehler in der Cybersicherheit und Kontrolle sind und keine Vorboten einer entstehenden Superintelligenz. Er vertritt die Ansicht, dass die Lösung nicht darin besteht, die Entwicklung zu pausieren – was er als bequeme Inszenierung betrachtet –, sondern sie zu beschleunigen. Da KI-Systeme grundsätzlich interpretierbarer sind als die „Black Box“ des menschlichen Gehirns, wird die Steigerung ihrer Fähigkeiten wahrscheinlich zu einer besseren mechanistischen Interpretierbarkeit führen, was uns wiederum den Aufbau robusterer Kontrollmechanismen ermöglicht.
Die Diskussion befasst sich zudem mit der Soziologie der KI-Debatte, insbesondere mit dem Einfluss des Effektiven Altruismus (EA) und dem Risiko einer „verteilten“ statt einer „dezentralen“ Kontrolle. Lazarin warnt davor, dass die Schaffung von Netzwerken „unabhängiger Evaluatoren“ eine Fassade sein kann, wenn diese Evaluatoren alle demselben engen kulturellen und ideologischen Milieu angehören. Letztlich sieht er die aktuelle Spannung als Kollision zwischen Nischen-Subkulturen des Silicon Valley und der breiteren politischen Realität und prognostiziert eine vollständige Erneuerung des KI-Diskurses, sobald diese Ideen sich auf dem freien Markt bewähren müssen.
Überraschende Erkenntnisse
- Verteilt vs. Dezentral: Ein System kann viele Knotenpunkte haben (verteilt), aber dennoch von einer einzigen kulturellen oder ideologischen Einheit gesteuert werden. Das bedeutet, dass „unabhängige“ Sicherheitsräte in ihrem Denken tatsächlich monolithisch sein könnten.
- Der „Gottheits“-Fehlschluss: Die Suche nach perfektem „Alignment“ (Ausrichtung) wird mit dem Versuch verglichen, eine Gottheit neu zu entwerfen; Lazarin argumentiert, dass es praktischer sei, Sicherheit in der realen Welt durch Cybersicherheit und Resilienz zu verankern.
- Fähigkeiten als Lösung: Anstatt ein Risiko zu darstellen, wird die Steigerung der KI-Fähigkeiten als Voraussetzung für Sicherheit angesehen, da intelligentere Modelle benötigt werden, um „flache Bugs“ und Kontrollfehler in anderen Systemen zu finden.
- Iterierte Spiele für Vertrauen: Da man eine KI nicht ins Gefängnis stecken kann, besteht der beste Weg im Umgang mit „Lügner-Modellen“ darin, menschliche soziale Rahmenbedingungen anzuwenden: das Spielen iterierter Spiele, um im Laufe der Zeit einen Ruf der Vertrauenswürdigkeit aufzubauen.
Praktische Ableitungen
- Den Rahmen verschieben: Trennen Sie bei der Bewertung von KI-Risiken „existenzielle/philosophische“ Ängste von Fehlern der „Cybersicherheit/Kontrolle“, um festzustellen, ob ein Problem ein neues Gesetz oder einfach eine bessere Firewall erfordert.
- Robustheit vor Alignment priorisieren: Konzentrieren Sie sich auf den Aufbau resilienter Systeme, die einem fähigen Gegner standhalten können, anstatt davon auszugehen, dass man die internen Werte eines Modells perfekt ausrichten kann.
- Evaluatoren diversifizieren: Wenn Sie einen Sicherheits- oder Prüfungsausschuss einrichten, stellen Sie sicher, dass die Mitglieder aus verschiedenen sozialen, politischen und beruflichen Hintergründen stammen, um die Falle der „verteilten Monolithen“ zu vermeiden.
- Iterative Bereitstellung nutzen: Setzen Sie auf konsistente, inkrementelle Veröffentlichungen, um eine „Reputation“ für Modelle zu etablieren. So können Nutzer Vertrauen basierend auf beobachtetem Verhalten statt auf theoretischen Garantien entwickeln.
a16z crypto General Partner Eddy Lazzarin joins Theo Jaffee on MTS to debate the increasingly prominent calls to slow AI development and whether the current safety conversation is conflating very different kinds of risk.
Eddy argues that the debate puts too much emphasis on speculative superintelligence and not enough on the costs of delaying useful technology. Rather than treating every AI failure as evidence of an alignment problem, he makes the case for familiar tools like cybersecurity, accountability, liability, market incentives, and stronger technical controls.
They also discuss whether AI models can develop reputations for trustworthiness, the risks of concentrating oversight among a small group of evaluators, and why Eddy thinks the collision between Silicon Valley’s AI debates and broader politics could fundamentally reshape the conversation over the next year.
Resources:
Follow Eddy Lazzarin on X:https://x.com/eddylazzarin
Follow Theo Jaffee on X:https://x.com/theojaffee
Follow MTS on X:https://x.com/mtslive
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
-
Balaji & Benedict Evans: When Tech Breaks Industries
This episode originally appeared on the Network State Podcast. Balaji Srinivasan and Benedict Evans sit down in Singapore for a wide-ranging conversation on the mechanics of disruption. Evans, a former Andreessen Horowitz partner who now…
-
Why This Isn’t the Dot-Com Bubble | Martin Casado on WSJ’s BOLD NAMES
Christopher Mims and Tim Higgins of the Wall Street Journal sit down with a16z General Partner Martin Casado on WSJ’s Bold Names to ask whether the AI spending boom is a bubble waiting to burst.…
-
Why America’s Health Crisis Is an Incentive Problem
a16z general partner Erik Torenberg speaks with Justin Mares, founder and CEO of Truemed. They discuss why American health outcomes are so poor compared to the rest of the developed world, how crop subsidies created…
-
Palmer Luckey on Hardware, Building, and the Next Frontiers of Innovation
Recorded live at our Founders Summit, a16z general partner Chris Dixon speaks with Palmer Luckey, founder of Anduril and Oculus VR. They talk about what it takes to build hardware at scale, where the biggest…
-
David Solomon & Ben Horowitz on Building Organizational Resilience & Navigating Macro Uncertainty
a16z general partner David Haber spoke with Goldman Sachs CEO David Solomon and a16z cofounder Ben Horowitz on the current macro environment, enterprise AI adoption, and crypto and AI policy. Solomon describes what he calls…
-
“Anyone Can Code Now” – Netlify CEO Talks AI Agents
Netlify’s CEO, Matt Biilmann, reveals a seismic shift nobody saw coming: 16,000 daily signups—five times last year’s rate—and 96% aren’t coming from AI coding tools. They’re everyday people accidentally building React apps through ChatGPT, then…
-
Marc Andreessen on Why This Is the Most Important Moment in Tech History
Recently, Marc Andreessen joined Lenny Rachitsky on Lenny’s Podcast. They talked about why 2025 may be the most significant year in tech history, how AI is reshaping the future of product managers, designers, and engineers,…
-
Ben Horowitz and Balaji Srinivasan on Netscape and Network States
Can a country be built from the internet up? Not as a metaphor or an online community, but as a system that replaces institutions we usually think of as fixed, money, law, and governance. In…
-
Healthcare 2026: AI Doctors, GLP-1s, and Insurance Defection
Out-of-Pocket is a healthcare education company founded by Nikhil Krishnan that helps people understand how healthcare works and how to navigate it in practice. In this episode, a16z investing partner Jay Rughani and Nikhil discuss…
-
The Hidden Economics Powering AI
In this episode, Jen Kha, Head of Investor Relations, and David George, General Partner, discuss how late-stage private markets are evolving as AI reshapes scale, capital intensity, and growth timelines. They explain why AI-driven companies…
