a16z Podcast
Summary & Insights
Is the current concentration of AI power an inevitable law of nature, or simply a limitation of today’s tools? Lucas Kaiser, a co-author of the seminal “Attention is All You Need” paper, suggests that the massive data centers and billion-dollar budgets currently required for AI are a property of the Transformer architecture, not a permanent feature of intelligence. While today’s industry trend is to simply “go bigger” to achieve better results, Kaiser argues that this brute-force approach is a temporary phase in the technology’s evolution.
The conversation highlights a growing divide between corporate product development and fundamental research. As giant labs like OpenAI shift their focus from pure research toward scaling products for consumers, a vacuum is created that open-source movements and academia can fill. The goal is to move away from generalist models that require the entire internet to function, moving instead toward “expert” models that can learn deeply from smaller, more diverse datasets.
Optimism for a distributed AI future is rooted in the biological reality of the human brain. Humans are not generalists who know everything; rather, society thrives on a network of distributed experts. Kaiser posits that the next great leap in AI will likely mirror this—moving toward ensembles of smaller, specialized models that are more efficient and accessible, eventually allowing individuals to own and train their own intelligence rather than relying on a corporate subscription.
Surprising Insights
- Hardware Leapfrogging: A single modern consumer GPU (like the RTX 5090) now possesses more computing power than the entire multi-GPU clusters used to design the original Transformer architecture.
- The “Brute Force” Trap: The current trend of scaling models “bigger, bigger, bigger” is a business strategy to bypass the need for fundamental research breakthroughs.
- Biological Proof: The existence of human expertise proves that it is mathematically and logically possible to be “smart” in a specific domain without being trained on all available global data.
- Transformer Limitations: While Transformers are powerful generalists, they are currently “stupid” when tasked with learning from small, specific datasets, signaling a need for a new algorithmic approach.
Practical Takeaways
- Experiment Locally: With the massive increase in consumer GPU power, independent researchers and hobbyists should experiment with smaller models and new architectures rather than feeling priced out by big tech.
- Focus on Domain Expertise: Since the future of AI likely involves “expert” models rather than just larger generalists, there is a significant opportunity to develop high-quality, specialized datasets for specific niches.
- Stay Open-Source: Support and engage with open-source AI movements, as these communities are more likely to pursue the fundamental research breakthroughs that challenge the current “scaling” status quo.
🛍️ Products & Resources Mentioned
- 🛠️ GearNVIDIA RTX 5090 GPU — Mentioned by Lucas Kaiser as a powerful tool for individual researchers to experiment and conduct machine learning research.View on Amazon →
Liệu sự tập trung quyền lực AI hiện nay là một quy luật tất yếu của tự nhiên, hay đơn giản chỉ là sự hạn chế của các công cụ hiện tại? Lucas Kaiser, đồng tác giả của bài báo mang tính bước ngoặt “Attention is All You Need”, cho rằng các trung tâm dữ liệu khổng lồ và ngân sách hàng tỷ đô la cần thiết cho AI hiện nay là đặc tính của kiến trúc Transformer, chứ không phải là đặc điểm vĩnh viễn của trí thông minh. Trong khi xu hướng của ngành công nghiệp hiện nay là đơn giản là “làm cho lớn hơn” để đạt được kết quả tốt hơn, Kaiser lập luận rằng cách tiếp cận dùng “sức mạnh thô” (brute-force) này chỉ là một giai đoạn tạm thời trong quá trình tiến hóa của công nghệ.
Cuộc thảo luận làm nổi bật sự chia rẽ ngày càng tăng giữa phát triển sản phẩm doanh nghiệp và nghiên cứu cơ bản. Khi các phòng thí nghiệm khổng lồ như OpenAI chuyển trọng tâm từ nghiên cứu thuần túy sang mở rộng quy mô sản phẩm cho người tiêu dùng, một khoảng trống đã được tạo ra mà các phong trào mã nguồn mở và giới học thuật có thể lấp đầy. Mục tiêu là chuyển từ các mô hình đa năng (generalist) yêu cầu toàn bộ internet để hoạt động, sang các mô hình “chuyên gia” có thể học sâu từ các tập dữ liệu nhỏ hơn và đa dạng hơn.
Sự lạc quan về một tương lai AI phân tán bắt nguồn từ thực tế sinh học của não bộ con người. Con người không phải là những kẻ đa năng biết mọi thứ; thay vào đó, xã hội phát triển dựa trên một mạng lưới các chuyên gia phân tán. Kaiser cho rằng bước nhảy vọt tiếp theo của AI có khả năng sẽ mô phỏng điều này—hướng tới các nhóm mô hình nhỏ hơn, chuyên biệt hơn, hiệu quả hơn và dễ tiếp cận hơn, cuối cùng cho phép các cá nhân sở hữu và huấn luyện trí thông minh của riêng họ thay vì phụ thuộc vào gói thuê bao của một tập đoàn.
Những góc nhìn bất ngờ
- Sự nhảy vọt về phần cứng: Một chiếc GPU tiêu dùng hiện đại duy nhất (như RTX 5090) hiện nay sở hữu sức mạnh tính toán nhiều hơn toàn bộ các cụm đa GPU từng được sử dụng để thiết kế kiến trúc Transformer nguyên bản.
- Bẫy “Sức mạnh thô”: Xu hướng mở rộng mô hình “lớn hơn, lớn hơn nữa” hiện nay là một chiến lược kinh doanh nhằm né tránh nhu cầu về những đột phá trong nghiên cứu cơ bản.
- Minh chứng sinh học: Sự tồn tại của các chuyên gia con người chứng minh rằng về mặt toán học và logic, ta hoàn toàn có thể “thông minh” trong một lĩnh vực cụ thể mà không cần được huấn luyện trên toàn bộ dữ liệu toàn cầu hiện có.
- Hạn chế của Transformer: Mặc dù Transformer là những mô hình đa năng mạnh mẽ, chúng hiện đang “kém hiệu quả” khi được giao nhiệm vụ học từ các tập dữ liệu nhỏ và đặc thù, điều này báo hiệu nhu cầu về một phương pháp thuật toán mới.
Bài học thực tiễn
- Thử nghiệm cục bộ: Với sự gia tăng mạnh mẽ của sức mạnh GPU tiêu dùng, các nhà nghiên cứu độc lập và những người đam mê nên thử nghiệm với các mô hình nhỏ hơn và các kiến trúc mới thay vì cảm thấy bị gạt ra ngoài bởi các ông lớn công nghệ.
- Tập trung vào chuyên môn lĩnh vực: Vì tương lai của AI có khả năng bao gồm các mô hình “chuyên gia” thay vì chỉ là các mô hình đa năng lớn hơn, nên có một cơ hội đáng kể để phát triển các tập dữ liệu chuyên biệt, chất lượng cao cho các ngách cụ thể.
- Gắn bó với mã nguồn mở: Hãy hỗ trợ và tham gia vào các phong trào AI mã nguồn mở, vì những cộng đồng này có nhiều khả năng theo đuổi các đột phá nghiên cứu cơ bản nhằm thách thức hiện trạng “mở rộng quy mô” hiện nay.
目前 AI 權力的集中是不可避免的自然法則,還是僅僅是當今工具的局限性?開創性論文《Attention is All You Need》的共同作者 Lucas Kaiser 指出,目前 AI 所需的大型數據中心和數十億美元的預算,是 Transformer 架構的特性,而非智能的永久特徵。雖然目前的產業趨勢是簡單地透過「規模擴大」來獲得更好的結果,但 Kaiser 認為這種暴力破解(brute-force)的方法僅是技術演進中的一個臨時階段。
這次對話凸顯了企業產品開發與基礎研究之間日益擴大的分歧。隨著 OpenAI 等巨頭實驗室將重心從純研究轉向為消費者擴展產品,這為開源運動和學術界創造了一個可以填補的真空地帶。其目標是擺脫那些需要整個互聯網才能運行的通用模型,轉而開發能夠從更小、更多樣化的數據集中深度學習的「專家」模型。
對分佈式 AI 未來的樂觀態度根植於人類大腦的生物現實。人類並非無所不知的通用主義者;相反地,社會在分佈式專家的網絡中繁榮發展。Kaiser 假設 AI 的下一次重大飛躍可能會模仿這一點——轉向由更高效、更易於獲取的小型專業模型組成的集成體,最終讓個人能夠擁有並訓練自己的智能,而非依賴於企業的訂閱服務。
驚人之見
- 硬體跨越: 單個現代消費級 GPU(如 RTX 5090)現在擁有的計算能力,已超過設計最初 Transformer 架構時所使用的整個多 GPU 集群。
- 「暴力破解」陷阱: 目前將模型規模擴大得「越來越大」的趨勢,實際上是一種規避基礎研究突破需求的商業策略。
- 生物學證明: 人類專業知識的存在證明了,在不需要接受全球所有可用數據訓練的情況下,在特定領域變得「聰明」在數學和邏輯上是可能的。
- Transformer 的局限: 雖然 Transformer 是強大的通用模型,但在處理從小型特定數據集學習的任務時卻顯得「愚笨」,這預示著需要一種新的算法方法。
實踐啟示
- 本地實驗: 隨著消費級 GPU 算力的巨大提升,獨立研究人員和愛好者應該嘗試小型模型和新架構,而不是感到被科技巨頭的高門檻所排擠。
- 專注於領域專業知識: 既然 AI 的未來可能涉及「專家」模型而非僅僅是更大的通用模型,那麼為特定細分市場開發高質量、專業化數據集將是一個重大機會。
- 支持開源: 支持並參與開源 AI 運動,因為這些社群更有可能追求挑戰目前「規模擴大」現狀的基礎研究突破。
La concentration actuelle de la puissance de l’IA est-elle une loi inévitable de la nature, ou simplement une limitation des outils d’aujourd’hui ? Lucas Kaiser, co-auteur de l’article fondateur « Attention is All You Need », suggère que les centres de données massifs et les budgets se comptant en milliards de dollars, actuellement requis pour l’IA, sont une propriété de l’architecture Transformer, et non une caractéristique permanente de l’intelligence. Alors que la tendance actuelle de l’industrie consiste simplement à « voir plus grand » pour obtenir de meilleurs résultats, Kaiser soutient que cette approche par force brute n’est qu’une phase temporaire de l’évolution technologique.
La discussion met en lumière un fossé croissant entre le développement de produits d’entreprise et la recherche fondamentale. Alors que des laboratoires géants comme OpenAI déplacent leur priorité de la recherche pure vers le déploiement de produits à grande échelle pour les consommateurs, un vide se crée, que les mouvements open-source et le monde universitaire peuvent combler. L’objectif est de s’éloigner des modèles généralistes qui nécessitent l’internet tout entier pour fonctionner, pour s’orienter vers des modèles « experts » capables d’apprendre profondément à partir de jeux de données plus petits et plus diversifiés.
L’optimisme quant à un avenir de l’IA distribuée s’enracine dans la réalité biologique du cerveau humain. Les humains ne sont pas des généralistes qui savent tout ; au contraire, la société prospère grâce à un réseau d’experts distribués. Kaiser avance que le prochain grand saut de l’IA reflétera probablement cela : une évolution vers des ensembles de modèles plus petits et spécialisés, plus efficaces et accessibles, permettant ainsi aux individus de posséder et d’entraîner leur propre intelligence plutôt que de dépendre d’un abonnement d’entreprise.
Perspectives Surprenantes
- Saut technologique matériel : Un seul GPU grand public moderne (comme le RTX 5090) possède désormais plus de puissance de calcul que l’ensemble des clusters multi-GPU utilisés pour concevoir l’architecture Transformer originale.
- Le piège de la « force brute » : La tendance actuelle à agrandir les modèles (« toujours plus grand ») est une stratégie commerciale visant à contourner la nécessité de percées en recherche fondamentale.
- Preuve biologique : L’existence de l’expertise humaine prouve qu’il est mathématiquement et logiquement possible d’être « intelligent » dans un domaine spécifique sans avoir été entraîné sur l’ensemble des données mondiales disponibles.
- Limitations du Transformer : Bien que les Transformers soient de puissants généralistes, ils sont actuellement « inefficaces » lorsqu’il s’agit d’apprendre à partir de jeux de données restreints et spécifiques, ce qui signale la nécessité d’une nouvelle approche algorithmique.
Points Clés Pratiques
- Expérimentez localement : Avec l’augmentation massive de la puissance des GPU grand public, les chercheurs indépendants et les passionnés devraient expérimenter des modèles plus petits et de nouvelles architectures, plutôt que de se sentir exclus par les coûts imposés par les Big Tech.
- Misez sur l’expertise sectorielle : Puisque l’avenir de l’IA passera probablement par des modèles « experts » plutôt que par de simples généralistes encore plus vastes, il existe une opportunité significative de développer des jeux de données spécialisés de haute qualité pour des niches spécifiques.
- Restez fidèle à l’open-source : Soutenez et engagez-vous dans les mouvements d’IA open-source, car ces communautés sont plus susceptibles de poursuivre les percées de recherche fondamentale qui remettent en question le statu quo actuel du « scaling » (mise à l’échelle).
Ist die derzeitige Konzentration der KI-Macht ein unvermeidliches Naturgesetz oder schlicht eine Einschränkung heutiger Werkzeuge? Lucas Kaiser, Mitautor des wegweisenden Papers „Attention is All You Need“, legt nahe, dass die massiven Rechenzentren und Milliarden-Budgets, die derzeit für KI erforderlich sind, eine Eigenschaft der Transformer-Architektur sind und kein dauerhaftes Merkmal von Intelligenz. Während der aktuelle Branchentrend darin besteht, einfach „größer zu werden“, um bessere Ergebnisse zu erzielen, argumentiert Kaiser, dass dieser Brute-Force-Ansatz nur eine vorübergehende Phase in der Evolution der Technologie ist.
Das Gespräch verdeutlicht eine wachsende Kluft zwischen der Produktentwicklung in Unternehmen und der Grundlagenforschung. Da riesige Labore wie OpenAI ihren Fokus von der reinen Forschung hin zur Skalierung von Verbraucherprodukten verschieben, entsteht ein Vakuum, das Open-Source-Bewegungen und die akademische Welt füllen können. Das Ziel ist es, sich von Generalisten-Modellen zu entfernen, die das gesamte Internet benötigen, um zu funktionieren, und stattdessen zu „Experten-Modellen“ überzugehen, die tiefgreifend aus kleineren, vielfältigeren Datensätzen lernen können.
Der Optimismus für eine dezentrale KI-Zukunft wurzelt in der biologischen Realität des menschlichen Gehirns. Menschen sind keine Generalisten, die alles wissen; vielmehr gedeiht die Gesellschaft durch ein Netzwerk verteilter Experten. Kaiser postuliert, dass der nächste große Sprung der KI dies wahrscheinlich widerspiegeln wird – ein Übergang zu Ensembles aus kleineren, spezialisierten Modellen, die effizienter und zugänglicher sind. Dies würde es Einzelpersonen letztendlich ermöglichen, ihre eigene Intelligenz zu besitzen und zu trainieren, anstatt auf ein Unternehmens-Abonnement angewiesen zu sein.
Überraschende Erkenntnisse
- Hardware-Sprung: Eine einzige moderne Consumer-GPU (wie die RTX 5090) besitzt heute mehr Rechenleistung als die gesamten Multi-GPU-Cluster, die für den Entwurf der ursprünglichen Transformer-Architektur verwendet wurden.
- Die „Brute-Force“-Falle: Der aktuelle Trend, Modelle immer „größer, größer und noch größer“ zu skalieren, ist eine Geschäftsstrategie, um die Notwendigkeit fundamentaler Forschungsdurchbrüche zu umgehen.
- Biologischer Beweis: Die Existenz menschlicher Expertise beweist, dass es mathematisch und logisch möglich ist, in einem spezifischen Bereich „intelligent“ zu sein, ohne mit allen verfügbaren globalen Daten trainiert worden zu sein.
- Einschränkungen der Transformer: Während Transformer leistungsstarke Generalisten sind, erweisen sie sich derzeit als „unfähig“, wenn es darum geht, aus kleinen, spezifischen Datensätzen zu lernen, was auf die Notwendigkeit eines neuen algorithmischen Ansatzes hindeutet.
Praktische Erkenntnisse
- Lokal experimentieren: Angesichts der massiven Steigerung der Consumer-GPU-Leistung sollten unabhängige Forscher und Hobbyisten mit kleineren Modellen und neuen Architekturen experimentieren, anstatt sich von den Big-Tech-Giganten preislich verdrängt zu fühlen.
- Fokus auf Domänenexpertise: Da die Zukunft der KI wahrscheinlich Experten-Modelle statt nur größerer Generalisten bereithält, gibt es eine bedeutende Chance, hochwertige, spezialisierte Datensätze für spezifische Nischen zu entwickeln.
- Open Source bleiben: Unterstützen Sie Open-Source-KI-Bewegungen und engagieren Sie sich darin, da diese Gemeinschaften eher dazu neigen, die grundlegenden Forschungsdurchbrüche voranzutreiben, die den aktuellen „Scaling“-Status quo infrage stellen.
MTS host Sophia Dew visits the Open Source AI Summit in San Francisco to ask researchers and founders across the AI stack a central question: can open source prevent AI power from concentrating in the hands of a few companies?
Lukasz Kaiser, co-author of Attention Is All You Need, argues that today’s concentration may be a feature of the current technological paradigm rather than a permanent feature of AI. Transformers reward enormous amounts of data and compute, but future breakthroughs could make smaller, more specialized models far more capable.
Across conversations with researchers and builders working on open models, infrastructure, and applications, Sophia explores why China has taken the lead in open-weight models, whether the U.S. needs more open-model startups, what it means for companies to own their own intelligence, and where openness alone falls short, particularly when access to compute remains concentrated.
Resources:
Follow Lukasz Kaiser on X:https://x.com/lukaszkaiser
Follow Sophia Dew on X:https://x.com/sophiadew
Follow MTS on X:https://x.com/mtslive
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
-
Tesla’s Road Ahead: The Bitter Lesson in Robotics
What does Rich Sutton’s “Bitter Lesson” reveal about the decisions Tesla is making in its pursuit of autonomy? In this episode, we dive into Tesla’s recent “We, Robot” event, where they unveiled bold plans for…
-
The story of Apple Pay with Jennifer Bailey
In 2024, marking 10 years since its launch, Apple Pay now boasts hundreds of millions of consumers in 78 markets, at checkout on millions of websites and apps, in tens of millions of stores worldwide,…
-
Lisa Su on the AI Ecosystem Behind AMD’s 50x Growth
Lisa Su has transformed AMD into a global leader in AI and high-performance computing. In this episode of the AI Revolution (AIR) series , Bob Swan, a16z Operating Partner and former CEO of Intel, sits…
-
The Deepfake Dilemma: The Technology, Policy, and Economy
Deepfakes—AI-generated fake videos and voices—have become a widespread concern across politics, social media, and more. As they become easier to create, the threat grows. But so do the tools to detect them. In this episode,…
-
A Big Week in Tech: NotebookLM, OpenAI’s Speech API, & Custom Audio
Last week was another big week in technology. Google’s NotebookLM introduced its Audio Overview feature, enabling users to create customizable podcasts in over 35 languages. OpenAI followed with their real-time speech-to-speech API, making voice integration…
-
From Swipe to Scale: How Tinder Became #1
In 1995, just 2% of couples met online. Today, that number has surged to over 50%, making online dating the top way couples connect. In this episode, a16z General Partner Andrew Chen chats with Tinder…
-
Human Data is Key to AI: Alex Wang from Scale AI
What if the key to unlocking AI’s full potential lies not just in algorithms or compute, but in data? In this episode, a16z General Partner David George sits down with Alex Wang, founder and CEO…
-
The Frontier of Spatial Intelligence with Fei-Fei Li
Fei-Fei Li and Justin Johnson are pioneers in AI. While the world has only recently witnessed a surge in consumer AI, our guests have long been laying the groundwork for innovations that are transforming industries…
-
Apple’s Big Reveals, OpenAI’s Multi-Step Models, and Firefly Does Video
This week in consumer tech: Apple’s big reveals, OpenAI’s multi-step reasoning, and Adobe Firefly’s video model. Olivia Moore and Justine Moore, Partners on the a16z Consumer team, break down the latest announcements and how these…
-
Grand Challenges in Healthcare AI
Vijay Pande, founding general partner, and Julie Yoo, general partner at a16z Bio + Health, come together to discuss the grand challenges facing healthcare AI today. The talk through the implications of AI integration in…
