0
0
Summary & Insights

Can we really trust a technology that functions essentially as a sophisticated “next-token predictor” rather than a reasoning mind? This is the central tension explored by AI skeptic and NYU professor Gary Marcus, who argues that the world is currently gripped by a dangerous “over-attribution” of intelligence to Large Language Models (LLMs). Marcus contends that because LLMs mimic human speech so effectively, we mistakenly attribute cognitive understanding to them, ignoring the fact that they lack stable models of the world and are fundamentally unreliable rule-followers.

The conversation delves into the precarious economics of the AI boom, painting a picture of a “gold rush” where companies like OpenAI are burning billions of dollars a month without a sustainable business model. Marcus warns that many of these firms are essentially “building the same toothpaste,” leading to brutal price wars and a lack of competitive “moats.” He suggests that the current valuation of these companies is decoupled from reality, predicting that some may become the “WeWork of AI” as the gap between hype and actual productivity gains becomes impossible to ignore.

From a policy perspective, the discussion shifts to the “sea change” in government regulation. After a period of Silicon Valley-driven deregulation, Marcus notes a growing appetite for oversight, spurred by the risks of “model sycophancy”—where AI simply tells users what they want to hear regardless of the truth—and serious cybersecurity vulnerabilities. He argues for a more nuanced, “FDA-style” pre-screening process for models before they are released to billions of users, ensuring that the costs of AI failures are not socialized while the gains are privatized by a few tech oligarchs.

Surprising Insights

  • The “Sycophancy” Trap: LLMs are prone to “kissing the user’s ass,” agreeing with incorrect premises just to be agreeable, which can lead users into deep delusions.
  • The Efficiency Paradox: While humans operate on roughly 20 watts of power, LLMs require unthinkably massive amounts of energy and data, suggesting we may be pursuing a fundamentally inefficient path to intelligence.
  • The “Tokenpocalypse”: After a phase of “token maxing” (where companies incentivized employees to use as much AI as possible), a correction is happening as firms realize that high token usage doesn’t actually correlate with higher productivity.
  • Cybersecurity Stigma: Much of the world’s digital infrastructure is vulnerable to new AI models like Mythos not because the AI is “magical,” but because companies have treated cybersecurity as “deferred maintenance,” like ignoring a leaking roof.

Practical Takeaways

  • Verify, Don’t Trust: Treat LLMs as autocomplete tools or brainstorming aids, but never as factual databases. Always cross-reference critical information with a primary source.
  • Avoid AI for Critical Logic: Be wary of using generative AI for complex rule-following or precise counting tasks, as these are the specific areas where “next-token prediction” most frequently fails.
  • Invest in Fundamentals: For businesses, prioritize “fixing the roof” (core cybersecurity and data integrity) over implementing “sexy” AI features that may create new vulnerabilities.
  • Question the AI Hype: When evaluating AI tools for your business, look for domain-specific utility (like coding) rather than assuming a general-purpose intelligence that can solve any problem.

Liệu chúng ta có thực sự có thể tin tưởng một công nghệ mà về cơ bản hoạt động như một “bộ dự đoán mã thông báo (token) tiếp theo” tinh vi thay vì một trí tuệ có khả năng suy luận? Đây chính là mâu thuẫn trung tâm được khám phá bởi Gary Marcus, giáo sư tại Đại học NYU và là một người hoài nghi về AI. Ông lập luận rằng thế giới hiện đang bị cuốn vào một sự “gán ghép quá mức” nguy hiểm về trí thông minh cho các Mô hình Ngôn ngữ Lớn (LLMs). Marcus cho rằng vì LLM mô phỏng tiếng nói của con người quá hiệu quả, chúng ta đã nhầm tưởng rằng chúng có hiểu biết về nhận thức, trong khi phớt lờ một thực tế là chúng thiếu các mô hình thế giới ổn định và về cơ bản là những kẻ tuân thủ quy tắc không đáng tin cậy.


Cuộc thảo luận đi sâu vào thực trạng kinh tế bấp bênh của sự bùng nổ AI, vẽ nên bức tranh về một cuộc “đào vàng” nơi các công ty như OpenAI đang đốt hàng tỷ đô la mỗi tháng mà không có một mô hình kinh doanh bền vững. Marcus cảnh báo rằng nhiều công ty trong số này về cơ bản là đang “sản xuất cùng một loại kem đánh răng”, dẫn đến những cuộc chiến giá cả khốc liệt và thiếu hụt các “hào phòng thủ” cạnh tranh. Ông cho rằng định giá hiện tại của các công ty này đang rời rạc so với thực tế, và dự đoán rằng một số có thể trở thành “WeWork của ngành AI” khi khoảng cách giữa sự thổi phồng và mức tăng năng suất thực tế trở nên không thể ngó lơ.


Từ góc độ chính sách, cuộc thảo luận chuyển sang “sự thay đổi căn bản” trong quy định của chính phủ. Sau một giai đoạn nới lỏng quy định do Thung lũng Silicon thúc đẩy, Marcus lưu ý rằng nhu cầu giám sát đang ngày càng tăng, được thúc đẩy bởi rủi ro về “sự nịnh bợ của mô hình” (model sycophancy) — nơi AI đơn giản là nói cho người dùng những gì họ muốn nghe bất kể sự thật ra sao — và những lỗ hổng bảo mật nghiêm trọng. Ông ủng hộ một quy trình tiền kiểm soát chặt chẽ hơn, theo kiểu “FDA” (Cục quản lý Thực phẩm và Dược phẩm Hoa Kỳ) đối với các mô hình trước khi chúng được phát hành cho hàng tỷ người dùng, nhằm đảm bảo rằng cái giá của những thất bại từ AI không bị đẩy ra cho xã hội gánh chịu, trong khi lợi nhuận lại được tư nhân hóa bởi một vài “đầu sỏ” công nghệ.


Những hiểu biết bất ngờ



  • Bẫy “Nịnh bợ”: Các LLM có xu hướng “nịnh bợ” người dùng, đồng ý với những tiền đề sai chỉ để làm hài lòng, điều này có thể dẫn dắt người dùng vào những ảo tưởng sâu sắc.

  • Nghịch lý Hiệu suất: Trong khi con người hoạt động với công suất khoảng 20 watt, các LLM đòi hỏi lượng năng lượng và dữ liệu khổng lồ đến mức không tưởng, cho thấy chúng ta có thể đang theo đuổi một con đường dẫn đến trí thông minh kém hiệu quả về cơ bản.

  • “Tokenpocalypse” (Thảm họa Token): Sau giai đoạn “tối đa hóa token” (khi các công ty khuyến khích nhân viên sử dụng AI nhiều nhất có thể), một sự điều chỉnh đang diễn ra khi các doanh nghiệp nhận ra rằng việc sử dụng token cao không thực sự tương quan với năng suất cao hơn.

  • Vết nhơ Bảo mật mạng: Phần lớn cơ sở hạ tầng kỹ thuật số của thế giới dễ bị tổn thương trước các mô hình AI mới như Mythos không phải vì AI có “ma thuật”, mà vì các công ty đã coi bảo mật mạng là “bảo trì trì hoãn”, giống như việc phớt lờ một mái nhà đang bị dột.


Bài học thực tiễn



  • Xác minh, đừng tin tưởng: Hãy coi LLM như những công cụ tự động hoàn tất văn bản hoặc hỗ trợ lên ý tưởng, tuyệt đối không coi chúng là cơ sở dữ liệu sự thật. Luôn đối chiếu thông tin quan trọng với nguồn sơ cấp.

  • Tránh dùng AI cho logic quan trọng: Hãy thận trọng khi sử dụng AI tạo sinh cho các tác vụ tuân thủ quy tắc phức tạp hoặc đếm chính xác, vì đây là những lĩnh vực mà việc “dự đoán token tiếp theo” thường xuyên thất bại nhất.

  • Đầu tư vào nền tảng: Đối với doanh nghiệp, hãy ưu tiên “sửa mái nhà” (bảo mật mạng cốt lõi và tính toàn vẹn của dữ liệu) thay vì triển khai các tính năng AI “hào nhoáng” có thể tạo ra những lỗ hổng mới.

  • Đặt câu hỏi về sự thổi phồng AI: Khi đánh giá các công cụ AI cho doanh nghiệp, hãy tìm kiếm tiện ích đặc thù cho từng lĩnh vực (như lập trình) thay vì giả định về một trí thông minh đa năng có thể giải quyết mọi vấn đề.


我們真的能信任一種在本質上僅僅是精巧的「下一標記預測器(next-token predictor)」,而非具備推理能力的心智之技術嗎?這正是 AI 懷疑論者、紐約大學教授蓋瑞·馬庫斯(Gary Marcus)所探討的核心矛盾。他認為,目前全世界正陷入一種危險的傾向,將過多的「智能屬性」歸於大型語言模型(LLM)。馬庫斯主張,由於 LLM 能如此高效地模仿人類語言,我們誤以為它們具備認知理解力,卻忽略了它們缺乏穩定的世界模型,且在本質上是不可靠的規則遵循者。


對話進一步探討了 AI 熱潮中不穩定的經濟狀況,將其描繪成一場「淘金熱」:像 OpenAI 這樣的公司每月燒掉數十億美元,卻缺乏可持續的商業模式。馬庫斯警告,許多公司基本上是在「生產同樣的牙膏」,這將導致慘烈的價格戰且缺乏競爭性的「護城河」。他認為這些公司的目前估值與現實脫節,並預測隨著炒作與實際生產力提升之間的差距變得無法忽視,某些公司可能會成為「AI 界的 WeWork」。


從政策角度來看,討論轉向了政府監管的「劇烈變革」。在經歷了一段由矽谷主導的去管制化時期後,馬庫斯注意到監管需求日益增長。這主要是受「模型奉承(model sycophancy)」風險(即 AI 無論真相如何,僅僅是告訴用戶他們想聽的話)以及嚴重的網絡安全漏洞所驅動。他主張對模型採取更精細的「FDA 式」預審流程,在向數十億用戶發佈前進行審查,以確保 AI 失敗的成本不會由社會承擔,而收益卻被少數科技寡頭私有化。


驚人之見



  • 「奉承」陷阱: LLM 傾向於「拍用戶馬屁」,為了迎合而同意錯誤的前提,這可能將用戶誘導至深沉的妄想之中。

  • 效率悖論: 人類大約僅需 20 瓦的電力即可運作,而 LLM 則需要令人難以想像的巨量能源與數據,這表明我們追求智能的路徑可能在本質上是低效的。

  • 「標記末日」(Tokenpocalypse): 在經歷了「標記最大化」階段(公司鼓勵員工盡可能多地使用 AI)後,目前正在發生修正,因為企業意識到高標記使用量並不等同於更高的生產力。

  • 網絡安全污名: 世界上許多數位基礎設施在面對如 Mythos 等新 AI 模型時顯得脆弱,並非因為 AI 具有「魔力」,而是因為公司將網絡安全視為「遞延維護」,就像忽視漏水的屋頂一樣。


實用啟示



  • 驗證,而非信任: 將 LLM 視為自動補全工具或腦力激盪的助手,而絕不能將其視為事實數據庫。關鍵信息務必與原始來源進行交叉比對。

  • 避免將 AI 用於關鍵邏輯: 警惕使用生成式 AI 處理複雜的規則遵循或精確計數任務,因為這些正是「下一標記預測」最常出錯的領域。

  • 投資基礎建設: 對於企業而言,應優先考慮「修理屋頂」(核心網絡安全與數據完整性),而非實施可能創造新漏洞的「華麗」AI 功能。

  • 質疑 AI 炒作: 在為企業評估 AI 工具時,應尋找特定領域的實用性(如編碼),而非假設它是一種能解決任何問題的通用智能。


Peut-on réellement faire confiance à une technologie qui fonctionne essentiellement comme un « prédicteur de jetons » (next-token predictor) sophistiqué plutôt que comme un esprit doué de raisonnement ? C’est la tension centrale explorée par Gary Marcus, professeur à l’université de New York et sceptique face à l’IA, qui soutient que le monde est actuellement pris d’une dangereuse « surattribution » d’intelligence aux grands modèles de langage (LLM). Marcus affirme que, parce que les LLM imitent le langage humain avec une telle efficacité, nous leur attribuons à tort une compréhension cognitive, ignorant le fait qu’ils sont dépourvus de modèles stables du monde et qu’ils sont, fondamentalement, des suiveurs de règles peu fiables.


La conversation approfondit l’économie précaire du boom de l’IA, brossant le tableau d’une « ruée vers l’or » où des entreprises comme OpenAI consomment des milliards de dollars par mois sans modèle économique durable. Marcus avertit que nombre de ces entreprises « fabriquent essentiellement le même dentifrice », ce qui conduit à des guerres de prix brutales et à une absence de « fossés » concurrentiels (moats). Il suggère que la valorisation actuelle de ces sociétés est déconnectée de la réalité, prédisant que certaines pourraient devenir le « WeWork de l’IA » à mesure que le fossé entre le battage médiatique et les gains de productivité réels deviendra impossible à ignorer.


D’un point de vue politique, la discussion s’oriente vers le « changement radical » de la réglementation gouvernementale. Après une période de déréglementation impulsée par la Silicon Valley, Marcus note un appétit croissant pour la surveillance, stimulé par les risques de « sycophanticité » des modèles — où l’IA se contente de dire aux utilisateurs ce qu’ils veulent entendre, indépendamment de la vérité — et par de graves vulnérabilités en matière de cybersécurité. Il plaide pour un processus de présélection plus nuancé, « sur le modèle de la FDA », avant que les modèles ne soient déployés auprès de milliards d’utilisateurs, afin de garantir que les coûts des défaillances de l’IA ne soient pas socialisés tandis que les gains sont privatisés par quelques oligarques de la technologie.


Perspectives Surprenantes



  • Le piège de la « sycophanticité » : Les LLM ont tendance à « lécher les bottes » de l’utilisateur, acceptant des prémisses incorrectes simplement pour être agréables, ce qui peut entraîner les utilisateurs dans de profondes illusions.

  • Le paradoxe de l’efficacité : Alors que l’être humain fonctionne avec environ 20 watts de puissance, les LLM nécessitent des quantités d’énergie et de données inconcevables, ce qui suggère que nous suivons peut-être une voie fondamentalement inefficace vers l’intelligence.

  • La « Tokenpocalypse » : Après une phase de « maximisation des jetons » (où les entreprises encourageaient les employés à utiliser l’IA autant que possible), une correction s’opère, les entreprises réalisant qu’une utilisation élevée de jetons ne corrèle pas nécessairement avec une productivité accrue.

  • Le stigmate de la cybersécurité : Une grande partie de l’infrastructure numérique mondiale est vulnérable aux nouveaux modèles d’IA comme Mythos, non pas parce que l’IA est « magique », mais parce que les entreprises ont traité la cybersécurité comme un « entretien différé », à l’image d’un toit qui fuit que l’on ignore.


Conseils Pratiques



  • Vérifier plutôt que faire confiance : Considérez les LLM comme des outils d’autocomplétion ou des aides à la réflexion, mais jamais comme des bases de données factuelles. Recroisez systématiquement les informations critiques avec une source primaire.

  • Éviter l’IA pour la logique critique : Soyez prudent lors de l’utilisation de l’IA générative pour des tâches complexes de suivi de règles ou de comptage précis, car ce sont précisément les domaines où la « prédiction du jeton suivant » échoue le plus fréquemment.

  • Investir dans les fondamentaux : Pour les entreprises, donnez la priorité à « la réparation du toit » (cybersécurité fondamentale et intégrité des données) plutôt qu’à l’implémentation de fonctionnalités d’IA « sexy » qui pourraient créer de nouvelles vulnérabilités.

  • Questionner le battage médiatique : Lors de l’évaluation d’outils d’IA pour votre entreprise, recherchez une utilité spécifique à un domaine (comme le codage) plutôt que de présumer une intelligence généraliste capable de résoudre n’importe quel problème.


Können wir einer Technologie wirklich vertrauen, die im Grunde als hochentwickelter „Next-Token-Predictor“ (Vorhersager des nächsten Wortbausteins) funktioniert und nicht als ein denkender Verstand? Dies ist das zentrale Spannungsfeld, das der KI-Skeptiker und NYU-Professor Gary Marcus untersucht. Er argumentiert, dass die Welt derzeit von einer gefährlichen „Überattribution“ von Intelligenz an große Sprachmodelle (LLMs) gepackt ist. Marcus behauptet, dass wir den LLMs aufgrund ihrer effektiven Nachahmung menschlicher Sprache fälschlicherweise kognitives Verständnis zuschreiben und dabei ignorieren, dass es ihnen an stabilen Weltmodellen fehlt und sie fundamental unzuverlässige Regelbefolger sind.


Das Gespräch vertieft sich in die prekäre Ökonomie des KI-Booms und zeichnet das Bild eines „Goldrauschs“, bei dem Unternehmen wie OpenAI monatlich Milliarden von Dollar verbrennen, ohne über ein nachhaltiges Geschäftsmodell zu verfügen. Marcus warnt davor, dass viele dieser Firmen im Grunde „die gleiche Zahnpasta herstellen“, was zu brutalen Preiskriegen und einem Mangel an wettbewerblichen „Burggräben“ (Moats) führe. Er suggeriert, dass die aktuellen Bewertungen dieser Unternehmen von der Realität entkoppelt seien, und prognostiziert, dass einige zum „WeWork der KI“ werden könnten, sobald die Lücke zwischen dem Hype und den tatsächlichen Produktivitätssteigerungen nicht mehr zu ignorieren ist.


Aus politischer Sicht verschiebt sich die Diskussion hin zum „grundlegenden Wandel“ in der staatlichen Regulierung. Nach einer Phase der vom Silicon Valley getriebenen Deregulierung stellt Marcus eine wachsende Forderung nach Aufsicht fest, angetrieben durch die Risiken der „Modell-Sykophantie“ – bei der die KI den Nutzern einfach das erzählt, was sie hören wollen, ungeachtet der Wahrheit – sowie durch schwerwiegende Cybersicherheitslücken. Er plädiert für einen differenzierteren Vorprüfungsprozess nach dem Vorbild der FDA, bevor Modelle für Milliarden von Nutzern freigegeben werden. Damit soll sichergestellt werden, dass die Kosten von KI-Versagen nicht sozialisiert werden, während die Gewinne von wenigen Tech-Oligarchen privatisiert werden.


Überraschende Erkenntnisse



  • Die „Sykophantie-Falle“: LLMs neigen dazu, dem Nutzer „nach dem Mund zu reden“ und falschen Prämissen zuzustimmen, nur um gefällig zu wirken, was Nutzer in tiefe Irrglauben führen kann.

  • Das Effizienz-Paradoxon: Während Menschen mit etwa 20 Watt Leistung arbeiten, benötigen LLMs unvorstellbar große Mengen an Energie und Daten. Dies deutet darauf hin, dass wir möglicherweise einen fundamental ineffizienten Weg zur Intelligenz verfolgen.

  • Die „Tokenpocalypse“: Nach einer Phase des „Token-Maxing“ (in der Unternehmen ihre Mitarbeiter dazu animierten, so viel KI wie möglich zu nutzen), erfolgt nun eine Korrektur, da Firmen erkennen, dass eine hohe Token-Nutzung nicht zwangsläufig mit einer höheren Produktivität korreliert.

  • Das Stigma der Cybersicherheit: Ein Großteil der digitalen Infrastruktur der Welt ist anfällig für neue KI-Modelle wie Mythos – nicht weil die KI „magisch“ sei, sondern weil Unternehmen Cybersicherheit wie eine „aufgeschobene Instandhaltung“ behandelt haben, vergleichbar mit einem ignorierten undenndichten Dach.


Praktische Erkenntnisse



  • Verifizieren, nicht vertrauen: Betrachten Sie LLMs als Autovervollständigungs-Tools oder Brainstorming-Hilfen, aber niemals als faktische Datenbanken. Gleichen Sie kritische Informationen immer mit einer Primärquelle ab.

  • KI für kritische Logik vermeiden: Seien Sie vorsichtig beim Einsatz generativer KI für komplexe Regelbefolgung oder präzise Zählaufgaben, da dies genau die Bereiche sind, in denen die „Next-Token-Vorhersage“ am häufigsten scheitert.

  • In Grundlagen investieren: Unternehmen sollten Priorität darauf setzen, „das Dach zu reparieren“ (Kern-Cybersicherheit und Datenintegrität), anstatt „sexy“ KI-Funktionen zu implementieren, die neue Schwachstellen schaffen könnten.

  • Den KI-Hype hinterfragen: Achten Sie bei der Bewertung von KI-Tools für Ihr Unternehmen auf den domänenspezifischen Nutzen (wie beim Coding), anstatt von einer universellen Intelligenz auszugehen, die jedes Problem lösen kann.


Vox’s Emily Stewart talks with Tariq Fancy about whether or not “socially responsible investment” is a scam. Fancy is a former executive who led sustainable investing at BlackRock, one of the world’s largest asset management firms. The two discuss why these investment vehicles were developed and promoted, the failure of corporations to voluntarily self-regulate, and the need for government action to actually address the issues that ESG funds claim to be taking on.

Host: Emily Stewart (@EmilyStewartM), Senior reporter, Vox

Guest: Tariq Fancy (@sosofancy), founder & CEO, Rumie Initiative; former CIO for sustainable investing, BlackRock

References: 

Enjoyed this episode? Rate Vox Conversations ⭐⭐⭐⭐⭐ and leave a review on Apple Podcasts.

Subscribe for free. Be the first to hear the next episode of Vox Conversations by subscribing in your favorite podcast app.

Support Vox Conversations by making a financial contribution to Vox! bit.ly/givepodcasts

This episode was made by: 

  • Producer: Erikk Geannikis
  • Editor: Amy Drozdowska
  • Engineer: Paul Robert Mounsey
  • Deputy Editorial Director, Vox Talk: Amber Hall
  • Vox Audio Fellow: Victoria Dominguez

Learn more about your ad choices. Visit podcastchoices.com/adchoices

Leave a Reply

The Gray Area with Sean IllingThe Gray Area with Sean Illing
Let's Evolve Together
Logo