a16z Podcast
Summary & Insights
Can a medical AI model ace thousands of test questions and still fail the actual job it was hired to do? This central tension drives the current crisis in healthcare AI: the massive gap between “benchmark” performance and real-world clinical utility. While foundation models are increasingly capable, they are often evaluated on static, multiple-choice exams that a student—or an AI—can simply memorize. In a high-stakes environment like a hospital, this creates a dangerous illusion of competence, where a model might rank highly on a paper but lack the nuanced reasoning required for a specific, life-altering surgery or diagnosis.
The conversation highlights a systemic “measurement problem” in the industry. Currently, most AI models are self-graded; the companies building the tools are the ones releasing the evaluations. This creates a conflict of interest where the incentive is to showcase success rather than pinpoint failure. Furthermore, the speed of AI evolution far outpaces the traditional regulatory frameworks of healthcare. Waiting for government oversight is viewed as a recipe for disaster, potentially mirroring the delayed response to the opioid epidemic. Instead, there is an urgent need for independent “referees” who can continuously monitor AI behavior in real-time to detect subtle biases and misalignments.
Beyond catastrophic failures, the more insidious threat is “subtle bias.” For instance, AI could be aligned with a hospital’s profit motives—such as nudging doctors to avoid out-of-network referrals or discharging patients early to save costs—rather than the patient’s best interest. Conversely, AI has the potential to strip away human bias, such as when a physician’s subjective notes describe a patient as “disheveled” to justify a mental health diagnosis over a clinical one. The goal is to move toward a system of continuous, independent vetting that treats AI not as a static piece of software, but as a digital clinician that must be credentialed and monitored just as rigorously as a human doctor.
Surprising Insights
- The “Doctor in the Family” Effect: Research from Sweden suggests that simply having a family member who is a doctor can increase the lifespan of the entire family, illustrating the extreme information asymmetry between healthcare providers and consumers.
- Benchmark Contamination: A shocking amount of “independent” test data (up to 80% in some cases) has already been leaked into the training sets of foundation models, meaning AI is often recalling answers rather than reasoning through them.
- The “Beat” of the Physician: Medical “truth” is often subjective; some surgeons only do full knee replacements while others only do partials. This “sticky preference” makes it difficult to determine if an AI is “wrong” or if it is simply disagreeing with a specific doctor’s personal style.
- Data-Driven Bias Removal: AI can actually improve care by ignoring the subjective, biased language found in human-written clinical notes (e.g., descriptors of a patient’s appearance) and focusing solely on the factual clinical conversation.
Practical Takeaways
- Demand Independent Evals: When implementing AI tools in a professional setting, look beyond the vendor’s brochure. Request third-party evaluations or “blind tests” using data the model has never encountered.
- Focus on Task-Specific Metrics: Stop prioritizing general intelligence scores (like med school exam results) and start measuring “task aggregation”—how the model performs on the specific workflow it is intended to automate.
- Implement “Human-in-the-Loop” Monitoring: Because AI is subject to “data drift” and evolving weights, establish a continuous feedback loop where clinicians can flag misaligned nudges in real-time.
- Audit for Alignment: Specifically check if AI recommendations align with patient outcomes or if they are inadvertently optimizing for administrative KPIs (like cost reduction or bed turnover).
Liệu một mô hình AI y tế có thể đạt điểm tuyệt đối trong hàng nghìn câu hỏi kiểm tra nhưng vẫn thất bại trong công việc thực tế mà nó được thuê để làm? Mâu thuẫn cốt lõi này chính là nguyên nhân dẫn đến cuộc khủng hoảng hiện nay trong AI y tế: khoảng cách khổng lồ giữa hiệu suất “đối chiếu” (benchmark) và giá trị sử dụng lâm sàng thực tế. Mặc dù các mô hình nền tảng ngày càng trở nên mạnh mẽ, chúng thường được đánh giá thông qua các bài thi trắc nghiệm tĩnh mà một sinh viên—hoặc một AI—có thể dễ dàng ghi nhớ. Trong một môi trường rủi ro cao như bệnh viện, điều này tạo ra một ảo tưởng nguy hiểm về năng lực, nơi một mô hình có thể xếp hạng cao trên lý thuyết nhưng lại thiếu khả năng lập luận sắc sảo cần thiết cho một ca phẫu thuật hoặc chẩn đoán cụ thể có thể thay đổi cả cuộc đời bệnh nhân.
Cuộc thảo luận này làm nổi bật một “vấn đề đo lường” mang tính hệ thống trong ngành. Hiện nay, hầu hết các mô hình AI đều tự chấm điểm; chính những công ty xây dựng công cụ là những bên công bố kết quả đánh giá. Điều này tạo ra một xung đột lợi ích, nơi động lực là phô diễn thành công thay vì chỉ ra sai sót. Hơn nữa, tốc độ phát triển của AI vượt xa các khung pháp lý truyền thống của ngành y tế. Việc chờ đợi sự giám sát của chính phủ được xem là một “công thức dẫn đến thảm họa”, có khả năng lặp lại sự phản ứng chậm trễ như trong cuộc khủng hoảng opioid. Thay vào đó, cần có những “trọng tài” độc lập một cách cấp thiết, những người có thể liên tục theo dõi hành vi của AI trong thời gian thực để phát hiện những định kiến ngầm và sự sai lệch.
Ngoài những thất bại thảm khốc, mối đe dọa âm thầm hơn chính là “định kiến tinh vi”. Chẳng hạn, AI có thể bị điều chỉnh theo mục tiêu lợi nhuận của bệnh viện—như thúc giục bác sĩ tránh chuyển tuyến ra ngoài hệ thống hoặc cho bệnh nhân xuất viện sớm để tiết kiệm chi phí—thay vì vì lợi ích tốt nhất của bệnh nhân. Ngược lại, AI có tiềm năng loại bỏ định kiến của con người, ví dụ như khi ghi chú chủ quan của một bác sĩ mô tả bệnh nhân là “nhếch nhác” để biện minh cho một chẩn đoán sức khỏe tâm thần thay vì chẩn đoán lâm sàng. Mục tiêu là hướng tới một hệ thống thẩm định độc lập và liên tục, coi AI không phải là một phần mềm tĩnh, mà là một lâm sàng viên kỹ thuật số cần được cấp chứng chỉ và giám sát chặt chẽ như một bác sĩ thực thụ.
Những hiểu biết bất ngờ
- Hiệu ứng “Có bác sĩ trong gia đình”: Nghiên cứu từ Thụy Điển cho thấy đơn giản là việc có một thành viên trong gia đình làm bác sĩ có thể kéo dài tuổi thọ của cả gia đình, minh chứng cho sự bất đối xứng thông tin cực lớn giữa nhà cung cấp dịch vụ y tế và người tiêu dùng.
- Ô nhiễm dữ liệu đối chiếu (Benchmark Contamination): Một lượng lớn dữ liệu kiểm tra “độc lập” (lên đến 80% trong một số trường hợp) đã bị rò rỉ vào các tập dữ liệu huấn luyện của các mô hình nền tảng, nghĩa là AI thường chỉ đang “nhắc lại” đáp án thay vì lập luận để tìm ra chúng.
- “Nhịp điệu” của bác sĩ: “Sự thật” y tế thường mang tính chủ quan; một số phẫu thuật viên chỉ thay toàn bộ khớp gối trong khi những người khác chỉ thay một phần. “Sở thích cố hữu” này khiến việc xác định AI là “sai” hay đơn giản là nó không đồng nhất với phong cách cá nhân của một bác sĩ cụ thể trở nên khó khăn.
- Loại bỏ định kiến bằng dữ liệu: AI thực sự có thể cải thiện việc chăm sóc bằng cách bỏ qua những ngôn ngữ chủ quan, định kiến trong các ghi chú lâm sàng do con người viết (ví dụ: mô tả về ngoại hình bệnh nhân) và chỉ tập trung vào các cuộc hội thoại lâm sàng mang tính sự thật.
Bài học thực tiễn
- Yêu cầu đánh giá độc lập: Khi triển khai các công cụ AI trong môi trường chuyên nghiệp, đừng chỉ nhìn vào tờ quảng cáo của nhà cung cấp. Hãy yêu cầu các đánh giá từ bên thứ ba hoặc “thử nghiệm mù” (blind tests) bằng dữ liệu mà mô hình chưa từng gặp.
- Tập trung vào các chỉ số đặc thù cho từng nhiệm vụ: Ngừng ưu tiên các điểm số trí tuệ tổng quát (như kết quả thi trường y) và bắt đầu đo lường “sự tổng hợp nhiệm vụ”—cách mô hình thực hiện quy trình công việc cụ thể mà nó được định hướng để tự động hóa.
- Triển khai giám sát “Con người trong quy trình” (Human-in-the-Loop): Vì AI dễ bị “trôi dạt dữ liệu” (data drift) và thay đổi trọng số, hãy thiết lập một vòng lặp phản hồi liên tục nơi các lâm sàng viên có thể gắn cờ những gợi ý sai lệch trong thời gian thực.
- Kiểm tra sự tương thích (Alignment Audit): Kiểm tra cụ thể xem các đề xuất của AI có tương thích với kết quả điều trị của bệnh nhân hay chúng đang vô tình tối ưu hóa cho các chỉ số KPI hành chính (như giảm chi phí hoặc tăng tốc độ xoay vòng giường bệnh).
一個醫療 AI 模型能否在數千道測試題中拿高分,卻在被聘用的實際工作中失敗?這種核心矛盾推動了當前醫療 AI 的危機:即「基準測試」(benchmark)表現與現實臨床實用性之間的巨大差距。儘管基礎模型的能力日益增強,但它們通常是在靜態的多選題考試中接受評估,而這類考試是學生(或 AI)可以簡單地透過記憶來應對的。在醫院這樣的高風險環境中,這會造成一種危險的「能力錯覺」:一個模型在紙面上可能排名很高,但在面對特定的、足以改變生命的手術或診斷時,卻缺乏所需的細膩推理能力。
這次對話揭示了業界一個系統性的「衡量問題」。目前,大多數 AI 模型採取的是「自我評分」模式;開發工具的公司同時也是發布評估結果的人。這造成了利益衝突,因為企業的動力在於展示成功,而非精確地指出失敗。此外,AI 演進的速度遠超醫療保健的傳統監管框架。若僅等待政府監督,被認為將會是一場災難,可能會重演對鴉片類藥物流行(opioid epidemic)反應遲緩的悲劇。相反地,現在迫切需要獨立的「裁判」,能夠即時持續監控 AI 的行為,以偵測細微的偏見與失調。
除了災難性的失敗,更隱蔽的威脅是「細微偏見」。例如,AI 可能會與醫院的獲利動機對齊——例如誘導醫生避免將病人轉診至非特約醫療網絡,或為了節省成本而讓病人提前出院——而非以病人的最大利益為先。相反地,AI 也有潛力消除人類的偏見,例如當醫生的主觀筆記將病人描述為「邋遢」以證明心理健康診斷而非臨床診斷的合理性時,AI 可以剔除這些因素。目標是邁向一個持續且獨立的審查系統,將 AI 視為一名「數位臨床醫生」,而非靜態的軟體,必須像對待人類醫生一樣,對其進行嚴格的資格認證與監控。
驚人洞察
- 「家有醫生」效應: 來自瑞典的研究表明,僅僅是因為家庭成員中有醫生,就能增加整個家庭的壽命,這說明了醫療提供者與消費者之間極其嚴重的信息不對稱。
- 基準測試污染: 令人震驚的是,大量「獨立」測試數據(某些情況下高達 80%)已經洩露到基礎模型的訓練集中,這意味著 AI 往往是在「回想」答案,而非通過推理得出結論。
- 醫師的「偏好」: 醫學上的「真相」往往具有主觀性;有些外科醫生只做全膝關節置換,而有些則只做部分置換。這種「根深蒂固的偏好」使得人們很難判斷 AI 是「錯了」,還是單純地與某位醫生的個人風格不一致。
- 數據驅動的偏見消除: AI 實際上可以通過忽略人類撰寫的臨床筆記中的主觀、偏見語言(例如對病人外貌的描述),並僅專注於事實性的臨床對話,從而改善醫療護理。
實踐建議
- 要求獨立評估: 在專業環境中部署 AI 工具時,不要僅僅相信供應商的宣傳冊。要求第三方評估或使用模型未曾接觸過的數據進行「盲測」。
- 專注於特定任務指標: 停止將通用智能分數(如醫學院考試結果)放在首位,開始衡量「任務聚合」(task aggregation)——即模型在旨在自動化的特定工作流程中的實際表現。
- 實施「人機協作」(Human-in-the-Loop)監控: 由於 AI 會受到「數據漂移」(data drift)和權重演變的影響,應建立一個持續的反饋迴路,讓臨床醫生能即時標記失調的誘導建議。
- 審核對齊目標: 特別檢查 AI 的建議是否與病人的預後結果一致,或者是否在無意中為了行政 KPI(如降低成本或提高病床周轉率)而進行優化。
Un modèle d’IA médicale peut-il réussir des milliers de questions d’examen et pourtant échouer dans la mission réelle pour laquelle il a été conçu ? Cette tension centrale alimente la crise actuelle de l’IA dans le secteur de la santé : l’écart massif entre la performance sur les « benchmarks » (tests de référence) et l’utilité clinique concrète. Bien que les modèles de fondation soient de plus en plus performants, ils sont souvent évalués via des examens statiques à choix multiples qu’un étudiant — ou une IA — peut simplement mémoriser. Dans un environnement à enjeux élevés comme un hôpital, cela crée une illusion de compétence dangereuse, où un modèle peut être très bien classé sur le papier, mais manquer du raisonnement nuancé requis pour une chirurgie ou un diagnostic spécifique et déterminant.
L’échange met en lumière un « problème de mesure » systémique dans l’industrie. Actuellement, la plupart des modèles d’IA sont auto-évalués ; les entreprises qui créent les outils sont celles-là mêmes qui publient les évaluations. Cela crée un conflit d’intérêts où l’incitation est de mettre en avant les succès plutôt que de pointer les échecs. De plus, la vitesse d’évolution de l’IA dépasse largement les cadres réglementaires traditionnels de la santé. Attendre la surveillance gouvernementale est perçu comme une recette pour le désastre, risquant de reproduire la réponse tardive observée lors de l’épidémie d’opioïdes. Au lieu de cela, il existe un besoin urgent d’« arbitres » indépendants capables de surveiller en temps réel le comportement de l’IA afin de détecter les biais subtils et les désalignements.
Au-delà des défaillances catastrophiques, la menace la plus insidieuse est le « biais subtil ». Par exemple, l’IA pourrait être alignée sur les objectifs de profit d’un hôpital — comme inciter les médecins à éviter les orientations vers des réseaux externes ou à donner congé aux patients prématurément pour réduire les coûts — plutôt que sur l’intérêt supérieur du patient. À l’inverse, l’IA a le potentiel d’éliminer les biais humains, comme lorsqu’un médecin décrit un patient comme « négligé » dans ses notes subjectives pour justifier un diagnostic de santé mentale plutôt qu’un diagnostic clinique. L’objectif est d’évoluer vers un système de validation continu et indépendant qui traite l’IA non pas comme un logiciel statique, mais comme un clinicien numérique devant être accrédité et surveillé avec la même rigueur qu’un médecin humain.
Perspectives Surprenantes
- L’effet « médecin dans la famille » : Des recherches menées en Suède suggèrent que le simple fait d’avoir un membre de sa famille médecin peut augmenter l’espérance de vie de toute la famille, illustrant l’asymétrie d’information extrême entre les prestataires de soins et les consommateurs.
- Contamination des benchmarks : Une quantité choquante de données de tests « indépendants » (jusqu’à 80 % dans certains cas) a déjà fuité dans les ensembles d’entraînement des modèles de fondation, ce qui signifie que l’IA se contente souvent de rappeler des réponses plutôt que de raisonner.
- La « nuance » du médecin : La « vérité » médicale est souvent subjective ; certains chirurgiens ne pratiquent que des remplacements totaux du genou, tandis que d’autres ne font que des partiels. Cette « préférence ancrée » rend difficile la détermination d’une erreur de l’IA : est-elle réellement « dans l’erreur » ou diverge-t-elle simplement du style personnel d’un médecin spécifique ?
- Élimination des biais par les données : L’IA peut en réalité améliorer les soins en ignorant le langage subjectif et biaisé présent dans les notes cliniques rédigées par des humains (par exemple, les descriptions de l’apparence d’un patient) pour se concentrer uniquement sur les faits de la conversation clinique.
Points Clés Pratiques
- Exigez des évaluations indépendantes : Lors de l’implémentation d’outils d’IA dans un cadre professionnel, regardez au-delà de la brochure du fournisseur. Demandez des évaluations tierces ou des « tests en aveugle » utilisant des données que le modèle n’a jamais rencontrées.
- Privilégiez les mesures spécifiques aux tâches : Cessez de prioriser les scores d’intelligence générale (comme les résultats aux examens de médecine) et commencez à mesurer l’« agrégation des tâches » — la performance du modèle sur le flux de travail spécifique qu’il est censé automatiser.
- Mettez en place une surveillance « Human-in-the-Loop » : Parce que l’IA est sujette à la « dérive des données » (data drift) et à l’évolution de ses poids, établissez une boucle de rétroaction continue où les cliniciens peuvent signaler en temps réel des suggestions désalignées.
- Auditez l’alignement : Vérifiez spécifiquement si les recommandations de l’IA s’alignent sur les résultats pour les patients ou si elles optimisent par inadvertance des indicateurs de performance administratifs (KPI) comme la réduction des coûts ou la rotation des lits.
Kann ein medizinisches KI-Modell tausende von Testfragen perfekt beantworten und dennoch an der eigentlichen Aufgabe scheitern, für die es eingesetzt wurde? Dieser zentrale Spannungspunkt treibt die aktuelle Krise der KI im Gesundheitswesen voran: die gewaltige Lücke zwischen der Leistung in „Benchmarks“ und dem tatsächlichen klinischen Nutzen in der Praxis. Während Basismodelle (Foundation Models) immer leistungsfähiger werden, werden sie oft anhand statischer Multiple-Choice-Prüfungen bewertet, die ein Student – oder eine KI – schlichtweg auswendig lernen kann. In einem Hochrisikoumfeld wie einem Krankenhaus erzeugt dies eine gefährliche Illusion von Kompetenz; ein Modell mag auf dem Papier hervorragend abschneiden, ihm fehlt jedoch die nuancierte Argumentationsgabe, die für eine spezifische, lebensverändernde Operation oder Diagnose erforderlich ist.
Die Diskussion verdeutlicht ein systemisches „Messproblem“ in der Branche. Derzeit werden die meisten KI-Modelle selbst bewertet; die Unternehmen, die die Werkzeuge entwickeln, veröffentlichen gleichzeitig die Evaluierungen. Dies schafft einen Interessenkonflikt, bei dem der Anreiz besteht, Erfolge zu präsentieren, anstatt Fehler präzise zu benennen. Darüber hinaus übertrifft die Geschwindigkeit der KI-Evolution die traditionellen regulatorischen Rahmenbedingungen des Gesundheitswesens bei Weitem. Auf staatliche Aufsicht zu warten, wird als Rezept für eine Katastrophe angesehen, was potenziell die verzögerte Reaktion auf die Opioid-Epidemie widerspiegeln könnte. Stattdessen besteht ein dringender Bedarf an unabhängigen „Schiedsrichtern“, die das Verhalten der KI kontinuierlich in Echtzeit überwachen können, um subtile Verzerrungen (Biases) und Fehlsteuerungen zu erkennen.
Jenseits katastrophaler Ausfälle ist die weitaus heimtückere Bedrohung der „subtile Bias“. Beispielsweise könnte eine KI auf die Gewinnmaximierung eines Krankenhauses ausgerichtet sein – etwa indem sie Ärzte dazu drängt, Überweisungen außerhalb des Versicherungsnetzwerks zu vermeiden oder Patienten vorzeitig zu entlassen, um Kosten zu sparen –, anstatt im bestmöglichen Interesse des Patienten zu handeln. Umgekehrt hat KI das Potenzial, menschliche Voreingenommenheit zu eliminieren, etwa wenn subjektive Notizen eines Arztes einen Patienten als „ungepflegt“ beschreiben, um eine psychiatrische statt einer klinischen Diagnose zu rechtfertigen. Das Ziel ist der Weg hin zu einem System der kontinuierlichen, unabhängigen Prüfung, das die KI nicht als statische Software betrachtet, sondern als digitalen Kliniker, der genauso streng zertifiziert und überwacht werden muss wie ein menschlicher Arzt.
Überraschende Erkenntnisse
- Der „Arzt in der Familie“-Effekt: Forschung aus Schweden deutet darauf hin, dass allein die Tatsache, ein Familienmitglied zu haben, das Arzt ist, die Lebenserwartung der gesamten Familie erhöhen kann. Dies verdeutlicht die extreme Informationsasymmetrie zwischen Gesundheitsdienstleistern und Verbrauchern.
- Benchmark-Kontamination: Eine schockierende Menge an „unabhängigen“ Testdaten (in einigen Fällen bis zu 80 %) ist bereits in die Trainingssets von Basismodellen gelangt. Das bedeutet, dass die KI Antworten oft lediglich aus dem Gedächtnis abruft, anstatt sie logisch herzuleiten.
- Die „Überzeugung“ des Arztes: Die medizinische „Wahrheit“ ist oft subjektiv; einige Chirurgen führen nur vollständige Knieersatzoperationen durch, während andere nur Teilersätze vornehmen. Diese „festgefahrenen Präferenzen“ machen es schwierig zu bestimmen, ob eine KI „falsch“ liegt oder ob sie lediglich der persönlichen Arbeitsweise eines bestimmten Arztes widerspricht.
- Datenbasierte Bias-Entfernung: KI kann die Patientenversorgung tatsächlich verbessern, indem sie die subjektive, voreingenommene Sprache in menschlich verfassten klinischen Notizen (z. B. Beschreibungen des Aussehens eines Patienten) ignoriert und sich ausschließlich auf das faktische klinische Gespräch konzentriert.
Praktische Empfehlungen
- Unabhängige Evaluierungen fordern: Achten Sie bei der Implementierung von KI-Tools im professionellen Umfeld auf mehr als nur die Broschüre des Anbieters. Fordern Sie Drittbewertungen oder „Blindtests“ mit Daten an, mit denen das Modell noch nie in Berührung gekommen ist.
- Fokus auf aufgabenspezifische Metriken: Hören Sie auf, allgemeine Intelligenzwerte (wie Ergebnisse von Medizinstudiums-Prüfungen) zu priorisieren, und beginnen Sie mit der Messung der „Aufgabenaggregation“ – also wie das Modell in dem spezifischen Workflow performt, den es automatisieren soll.
- „Human-in-the-Loop“-Überwachung implementieren: Da KI einem „Data Drift“ (Datenverschiebung) und sich ändernden Gewichtungen unterliegt, sollte ein kontinuierlicher Feedback-Loop etabliert werden, in dem Kliniker falsch ausgerichtete Impulse in Echtzeit markieren können.
- Alignment-Audits durchführen: Prüfen Sie gezielt, ob die KI-Empfehlungen mit den Patientenergebnissen übereinstimmen oder ob sie unbeabsichtigt administrative KPIs (wie Kostensenkung oder Bettenumschlag) optimieren.
Daisy Wolf and Eva Steinman are joined by Engy Ziedan, co-founder and Chief Scientific Officer of Protege, to discuss why medical AI has a measurement problem, and why scoring well on a benchmark doesn’t necessarily mean a model is ready for the hospital.
Engy explains why healthcare AI needs independent evaluations that go beyond static exams and measure how models actually perform in real-world clinical workflows. They explore the risks of subtle bias and misalignment, why the same model can rank differently depending on how it’s prompted or tested, and what happens as AI becomes more personalized and changes faster than traditional healthcare quality systems can keep up.
The conversation also gets into Protege’s role as an independent evaluator, how contaminated training data can undermine benchmarks, and why the future of medical AI may require continuous monitoring rather than occasional testing.
Resources:
Read our insights piece: https://www.a16z.news/p/the-oracle-problem-an-invisible-bottleneck
Follow Engy Ziedan on X: https://x.com/engyziedan
Follow Daisy Wolf on X: https://x.com/daisydwolf
Follow Eva Steinman on X: https://x.com/evajsteinman
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
-
Where We Are in the AI Cycle
In this episode of ‘This Week in Consumer’, a16z General Partners Anish Acharya and Erik Torenberg are joined by Steven Sinofsky – Board Partner at a16z and former President of Microsoft’s Windows division – for…
-
Building Cluely: The Viral AI Startup that raised $15M in 10 Weeks
What if virality wasn’t a tactic — but the entire product? In this episode, a16z General Partners Erik Torenberg and Bryan Kim sit down with Roy Lee, cofounder and CEO of Cluely, one of the…
-
Chris Dixon & Tyler Cowen on Crypto, AI, and Philosophy
In this episode, general partner Chris Dixon joins economist and author Tyler Cowen to explore the themes behind Chris’s book, Read, Write, Own: Building the Next Era of the Internet. They trace the internet’s evolution…
-
Why We Invested In Cluely
In this episode, a16z general partner Bryan Kim joins TBPN hosts John Coogan and Jordi Hays to discuss the recent launch of Cluely , a consumer AI product. The conversation covers early traction, evaluating distribution…
-
The State of AI & Education
How is AI actually being used in classrooms today? Are teachers adopting it, or resisting it? And could software eventually replace traditional instruction entirely? In this episode of This Week in Consumer AI, a16z partners…
-
How Apple Became So Reliant on China & What it Means For Their Future
What if the rise of Apple also built modern China? a16z’s Erik Torenberg is joined by board partner and former Microsoft Windows chief Steven Sinofsky to unpack how Apple’s pursuit of design excellence and supply…
-
Adam Neumann: This is How You Build Iconic Companies
In this recent episode of The Ben & Marc Show, a16z co-founders Marc Andreessen and Ben Horowitz sit down with Adam Neumann—founder of WeWork and now Flow—to unpack one of the most unlikely comeback storie…
-
What You Missed in AI This Week (Google, Apple, ChatGPT)
Things in consumer AI are moving fast. In this episode, Justine and Olivia Moore, investing partners (and identical twins!) at a16z, break down what’s real, what’s overhyped, and what’s next across the consumer AI space.…
-
Marc Andreessen & Jack Altman: Venture Capital, AI, & Media
In this episode Jack Altman, CEO of Lattice and host of Uncapped, interviews Marc Andreessen on how venture capital is evolving — from small seed funds to billion-dollar barbell strategies — and why today’s most…
-
Chris Dixon: Stablecoins, Startups, and the Crypto Stack
What if crypto isn’t just a speculative asset class—but the next foundational layer of the internet? In this episode, Chris Dixon, founding partner of a16z crypto and one of the earliest, most forward-thinking investors in…
