a16z Podcasta16z Podcast
0
0
Summary & Insights

Could an AI possess “mathematical taste,” or is it simply a dogged executor of human-inspired ideas? This question sits at the center of a shift in mathematics where the bottleneck is moving away from the grueling process of proving a result and toward the high-level judgment of which problems are worth solving. OpenAI mathematicians Metav Swani and Mark Selke describe a new “renaissance of reachable results,” where AI is solving problems in sphere packing and group theory that have resisted human effort for decades, not through brute force, but through a reasoning process that mirrors the backtracking and intuition of a human expert.

The conversation highlights a fundamental difference between human and machine cognition: the ability to “reset” a mental context. While a human mathematician might become psychologically tethered to a failing approach, an AI can spawn multiple parallel sessions or pivot its strategy without the “pollution” of previous failed attempts. This doggedness, combined with a massive internal library of cross-disciplinary connections, allows the model to execute finicky, high-detail proofs—such as those involving linear programming bounds in high-dimensional space—that would be too risky or time-consuming for a human to gamble their career on.

As AI accelerates the production of new theorems, the role of the mathematician is evolving from the sole “prover” to a curator and communicator. The speakers suggest that the value in the field will shift toward “communal understanding”—the ability to internalize a result and fit it into the broader framework of human knowledge. While the “ceiling” for difficulty remains high—with the guests noting that some legendary mysteries like P versus NP may still remain out of reach—the accessibility of high-level math is expanding, allowing non-experts to absorb complex papers and apply theoretical results to practical problems faster than ever before.

Surprising Insights

  • Contextual Pollution: Humans struggle to abandon wrong paths because their intuition becomes “polluted” by the failed attempt; AI avoids this by effectively starting fresh clones of its reasoning process.
  • Emergent Reasoning: The model’s ability to backtrack and prune search trees emerged from general-purpose reasoning training, rather than being specifically taught via mathematical textbooks or formal languages like Lean.
  • The “Ugly” vs. “Elegant” Proof: While some famous human proofs (like those for 3D sphere packing) are notoriously “ugly” and hundreds of pages long, the AI-generated proofs for similar high-dimensional problems are often shockingly short and elegant.
  • The Efficiency of “Taste”: “Mathematical taste” can be viewed utilitarily as the ability to make better judgments that lead to faster solutions, a trait the AI is developing as it solves increasingly complex, multi-step problems.

Practical Takeaways

  • Use AI for Literature Synthesis: Instead of spending hours searching fragmented academic literature to see if a problem is still “open,” use advanced LLMs to find obscure references and determine if a result is within reach.
  • Accelerate Paper Absorption: Use AI to extract the “rough proof strategy” from complex PDFs to quickly understand the core logic before diving into the dense technical details.
  • Iterative Prompting for Depth: When using AI for complex reasoning, don’t settle for the first correct answer; ask the model to “push this further” or explore the limits of its own logic to uncover more sophisticated theoretical frameworks.
  • Shift Focus to Curation: For those in technical fields, focus less on the rote execution of proofs and more on the “architectural” side of research—deciding which directions are promising and how to communicate results to others.

Liệu một AI có thể sở hữu “gu toán học” (mathematical taste), hay nó đơn thuần chỉ là một cỗ máy thực thi bền bỉ những ý tưởng do con người gợi mở? Câu hỏi này nằm ở trung tâm của một sự chuyển dịch trong toán học, nơi “nút thắt cổ chai” không còn nằm ở quá trình chứng minh kết quả đầy gian nan, mà chuyển sang khả năng phán đoán cấp cao về việc vấn đề nào là đáng để giải quyết. Metav Swani và Mark Selke, những nhà toán học tại OpenAI, mô tả về một “cuộc phục hưng của những kết quả trong tầm với”, nơi AI đang giải quyết các bài toán về xếp chặt hình cầu (sphere packing) và lý thuyết nhóm vốn đã thách thức nỗ lực của con người trong nhiều thập kỷ. Điều này không đạt được thông qua sức mạnh tính toán thô (brute force), mà thông qua một quá trình lập luận mô phỏng lại khả năng quay lui (backtracking) và trực giác của một chuyên gia con người.


Cuộc thảo luận làm nổi bật một sự khác biệt cơ bản giữa nhận thức của con người và máy móc: khả năng “thiết lập lại” (reset) bối cảnh tư duy. Trong khi một nhà toán học có thể bị ràng buộc về mặt tâm lý với một phương pháp đang thất bại, AI có thể khởi tạo nhiều phiên làm việc song song hoặc xoay trục chiến lược mà không bị “nhiễm độc” bởi những nỗ lực sai lầm trước đó. Sự bền bỉ này, kết hợp với một thư viện nội bộ khổng lồ về các kết nối đa ngành, cho phép mô hình thực hiện các chứng minh chi tiết và khắt khe—chẳng hạn như các giới hạn lập trình tuyến tính trong không gian cao chiều—những việc mà một con người sẽ thấy quá rủi ro hoặc tốn thời gian để đánh cược cả sự nghiệp vào đó.


Khi AI đẩy nhanh việc tạo ra các định lý mới, vai trò của nhà toán học đang tiến hóa từ một “người chứng minh” duy nhất thành một người giám tuyển (curator) và người truyền tải. Các diễn giả cho rằng giá trị trong lĩnh vực này sẽ chuyển dịch sang “sự thấu hiểu chung” (communal understanding)—khả năng nội hóa một kết quả và đặt nó vào khung kiến thức rộng lớn hơn của nhân loại. Mặc dù “trần” về độ khó vẫn còn rất cao—các khách mời lưu ý rằng một số bí ẩn huyền thoại như P đối diện NP có thể vẫn nằm ngoài tầm với—nhưng khả năng tiếp cận toán học cao cấp đang mở rộng, cho phép những người không chuyên hấp thụ các bài báo phức tạp và áp dụng các kết quả lý thuyết vào các vấn đề thực tiễn nhanh hơn bao giờ hết.


Những hiểu biết bất ngờ



  • Sự nhiễu loạn bối cảnh (Contextual Pollution): Con người khó từ bỏ những con đường sai lầm vì trực giác của họ bị “nhiễu” bởi lần thử thất bại; AI tránh được điều này bằng cách khởi tạo các bản sao mới hoàn toàn cho quá trình lập luận của mình.

  • Lập luận mới nổi (Emergent Reasoning): Khả năng quay lui và cắt tỉa cây tìm kiếm của mô hình xuất hiện từ quá trình huấn luyện lập luận đa năng, chứ không phải được dạy cụ thể thông qua sách giáo khoa toán học hay các ngôn ngữ hình thức như Lean.

  • Chứng minh “Xấu xí” và “Thanh thoát”: Trong khi một số chứng minh nổi tiếng của con người (như bài toán xếp chặt hình cầu 3D) bị coi là “xấu xí” và dài hàng trăm trang, thì các chứng minh do AI tạo ra cho các bài toán cao chiều tương tự thường ngắn gọn và thanh thoát một cách đáng ngạc nhiên.

  • Hiệu quả của “Gu”: “Gu toán học” có thể được nhìn nhận một cách thực dụng là khả năng đưa ra những phán đoán chính xác hơn để dẫn đến giải pháp nhanh hơn, một đặc điểm mà AI đang phát triển khi giải quyết các bài toán đa bước ngày càng phức tạp.


Bài học thực tiễn



  • Sử dụng AI để tổng hợp tài liệu: Thay vì dành hàng giờ tìm kiếm các tài liệu học thuật phân tán để xem một bài toán còn “mở” hay không, hãy sử dụng các LLM tiên tiến để tìm các tham chiếu ít phổ biến và xác định xem kết quả đó có nằm trong tầm với hay không.

  • Đẩy nhanh việc hấp thụ bài báo: Sử dụng AI để trích xuất “chiến lược chứng minh sơ bộ” từ các tệp PDF phức tạp nhằm nhanh chóng nắm bắt logic cốt lõi trước khi đi sâu vào các chi tiết kỹ thuật dày đặc.

  • Đặt câu hỏi lặp lại để đào sâu: Khi sử dụng AI cho lập luận phức tạp, đừng hài lòng với câu trả lời đúng đầu tiên; hãy yêu cầu mô hình “đẩy điều này đi xa hơn” hoặc khám phá giới hạn logic của chính nó để khám phá ra những khung lý thuyết tinh vi hơn.

  • Chuyển trọng tâm sang giám tuyển: Đối với những người làm trong các lĩnh vực kỹ thuật, hãy bớt tập trung vào việc thực thi chứng minh một cách máy móc và chú trọng hơn vào khía cạnh “kiến trúc” của nghiên cứu—quyết định hướng đi nào là triển vọng và cách truyền đạt kết quả đến người khác.


AI 是否能擁有「數學品味」,抑或它僅僅是一個執著地執行人類靈感之想法的執行者?這個問題正處於數學界轉型的核心:目前的瓶頸正從枯燥艱辛的證明過程,轉向對「哪些問題值得解決」的高層次判斷。OpenAI 的數學家 Metav Swani 與 Mark Selke 描述了一場全新的「可達成結果之復興」——AI 正在解決球體填充(sphere packing)和群論(group theory)中讓人類苦思數十年的難題。這並非透過暴力運算,而是透過一種鏡像人類專家回溯(backtracking)與直覺的推理過程來實現。


這次對話強調了人類與機器認知之間的一個根本差異:重置思維脈絡的能力。人類數學家可能會在心理上被困於某個失敗的方法中,而 AI 則可以啟動多個平行會話,或在沒有先前失敗嘗試「污染」的情況下轉變策略。這種執著,結合其內部龐大的跨學科關聯庫,使模型能夠執行極其繁瑣且高精細度的證明(例如涉及高維空間線性規劃界限的證明),而這類證明對於人類而言,風險過高且耗時太久,不值得拿職業生涯去賭博。


隨著 AI 加速新定理的產出,數學家的角色正從唯一的「證明者」演變為策展人(curator)和傳播者。講者認為,該領域的價值將轉向「共同理解」——即將結果內化並將其納入人類知識更廣泛框架的能力。雖然困難度的「天花板」依然很高(嘉賓指出,像 P vs NP 這樣的傳奇謎題可能依然遙不可及),但高階數學的可近性正在擴展,讓非專家能比以往更快地吸收複雜論文,並將理論結果應用於實際問題。


驚人洞察



  • 脈絡污染(Contextual Pollution): 人類難以放棄錯誤的路徑,是因為其直覺被失敗的嘗試「污染」了;AI 則透過有效地啟動推理過程的新克隆(clones)來避免這一點。

  • 湧現推理(Emergent Reasoning): 模型回溯和剪枝搜索樹的能力源於通用推理訓練,而非透過數學教科書或 Lean 等形式語言的專門教學而來。

  • 「醜陋」與「優美」的證明: 雖然一些著名的人類證明(如 3D 球體填充)以「醜陋」且長達數百頁而著稱,但 AI 為類似高維問題生成的證明往往簡短得令人驚訝且極為優美。

  • 「品味」的效率: 「數學品味」可以從功利角度被視為做出更好判斷以 leading 至更快速解決方案的能力,而 AI 在解決日益複雜的多步驟問題時,正發展出這種特質。


實踐要點



  • 利用 AI 進行文獻綜合: 與其花數小時在碎片化的學術文獻中搜索以確認某個問題是否仍是「未解決」狀態,不如使用先進的大型語言模型(LLM)尋找晦澀的引用,並判斷該結果是否在可觸及範圍內。

  • 加速論文吸收: 利用 AI 從複雜的 PDF 中提取「粗略的證明策略」,在深入研究繁瑣的技術細節之前,快速理解其核心邏輯。

  • 透過迭代提示深化思考: 使用 AI 進行複雜推理時,不要滿足於第一個正確答案;要求模型「進一步推進」或探索其邏輯的極限,以揭示更精妙的理論框架。

  • 將重心轉向策展: 對於技術領域從業人員,應減少對證明機械式執行的關注,而更多地關注研究的「建築」面——決定哪些方向具有前景,以及如何將結果有效地傳達給他人。


Une IA pourrait-elle posséder un « goût mathématique », ou n’est-elle qu’un exécutant acharné d’idées inspirées par l’humain ? Cette question est au cœur d’une mutation des mathématiques, où le goulot d’étranglement se déplace : il ne s’agit plus tant du processus épuisant de prouver un résultat que du jugement de haut niveau permettant de déterminer quels problèmes valent la peine d’être résolus. Metav Swani et Mark Selke, mathématiciens chez OpenAI, décrivent une nouvelle « renaissance des résultats accessibles », où l’IA résout des problèmes d’empilement de sphères et de théorie des groupes qui ont résisté aux efforts humains pendant des décennies, non pas par la force brute, mais via un processus de raisonnement qui reflète le retour en arrière et l’intuition d’un expert humain.


La conversation souligne une différence fondamentale entre la cognition humaine et celle de la machine : la capacité à « réinitialiser » un contexte mental. Tandis qu’un mathématicien humain peut devenir psychologiquement lié à une approche défaillante, une IA peut générer plusieurs sessions parallèles ou pivoter sa stratégie sans la « pollution » des tentatives infructueuses précédentes. Cette opiniâtreté, combinée à une immense bibliothèque interne de connexions interdisciplinaires, permet au modèle d’exécuter des preuves pointilleuses et extrêmement détaillées — comme celles impliquant des bornes de programmation linéaire dans des espaces de grande dimension — qu’un humain jugerait trop risquées ou chronophages pour y engager sa carrière.


Alors que l’IA accélère la production de nouveaux théorèmes, le rôle du mathématicien évolue : de seul « prouveur », il devient curateur et communicateur. Les intervenants suggèrent que la valeur dans ce domaine se déplacera vers la « compréhension communautaire » — la capacité d’intérioriser un résultat et de l’intégrer dans le cadre plus large des connaissances humaines. Bien que le « plafond » de difficulté reste élevé — les invités notant que certains mystères légendaires comme P contre NP pourraient encore rester hors de portée — l’accessibilité des mathématiques de haut niveau s’élargit, permettant à des non-experts d’assimiler des articles complexes et d’appliquer des résultats théoriques à des problèmes pratiques plus rapidement que jamais.


Perspectives surprenantes



  • Pollution contextuelle : Les humains peinent à abandonner les mauvaises pistes car leur intuition est « polluée » par l’échec précédent ; l’IA évite cela en créant efficacement des clones vierges de son processus de raisonnement.

  • Raisonnement émergent : La capacité du modèle à revenir en arrière et à élaguer les arbres de recherche est issue d’un entraînement au raisonnement général, et non d’un apprentissage spécifique via des manuels de mathématiques ou des langages formels comme Lean.

  • La preuve « laide » vs la preuve « élégante » : Alors que certaines preuves humaines célèbres (comme celles pour l’empilement de sphères en 3D) sont notoirement « laides » et longues de centaines de pages, les preuves générées par l’IA pour des problèmes similaires en haute dimension sont souvent étonnamment courtes et élégantes.

  • L’efficacité du « goût » : Le « goût mathématique » peut être vu d’un point de vue utilitaire comme la capacité à porter des jugements plus justes menant à des solutions plus rapides, un trait que l’IA développe à mesure qu’elle résout des problèmes multi-étapes toujours plus complexes.


Conseils pratiques



  • Utiliser l’IA pour la synthèse bibliographique : Au lieu de passer des heures à fouiller une littérature académique fragmentée pour savoir si un problème est toujours « ouvert », utilisez des LLM avancés pour trouver des références obscures et déterminer si un résultat est à portée de main.

  • Accélérer l’assimilation d’articles : Utilisez l’IA pour extraire la « stratégie de preuve globale » de PDF complexes afin de comprendre rapidement la logique centrale avant de plonger dans les détails techniques denses.

  • Le prompting itératif pour plus de profondeur : Lorsque vous utilisez l’IA pour un raisonnement complexe, ne vous contentez pas de la première réponse correcte ; demandez au modèle de « pousser l’analyse plus loin » ou d’explorer les limites de sa propre logique pour découvrir des cadres théoriques plus sophistiqués.

  • Déplacer le focus vers la curation : Pour ceux qui évoluent dans des domaines techniques, concentrez-vous moins sur l’exécution mécanique des preuves et davantage sur l’aspect « architectural » de la recherche — décider quelles directions sont prometteuses et comment communiquer les résultats aux autres.


Könnte eine KI einen „mathematischen Geschmack“ besitzen, oder ist sie lediglich eine beharrliche Ausführerin von menscheninspirierten Ideen? Diese Frage steht im Zentrum eines Wandels in der Mathematik, bei dem sich der Engpass vom mühsamen Prozess des Beweisens eines Ergebnisses hin zur übergeordneten Beurteilung verschiebt, welche Probleme es überhaupt wert sind, gelöst zu werden. Die OpenAI-Mathematiker Metav Swani und Mark Selke beschreiben eine neue „Renaissance erreichbarer Ergebnisse“, bei der die KI Probleme in der Kugelpackung und Gruppentheorie löst, die sich menschlichen Bemühungen über Jahrzehnte widersetzten – nicht durch reine Rechenkraft (Brute Force), sondern durch einen Denkprozess, der das Backtracking und die Intuition eines menschlichen Experten widerspiegelt.


Das Gespräch hebt einen grundlegenden Unterschied zwischen menschlicher und maschineller Kognition hervor: die Fähigkeit, einen mentalen Kontext „zurückzusetzen“. Während ein menschlicher Mathematiker psychologisch an einen scheiternden Ansatz gebunden bleiben kann, kann eine KI mehrere parallele Sitzungen starten oder ihre Strategie ändern, ohne durch die „Kontamination“ früherer gescheiterter Versuche beeinflusst zu werden. Diese Beharrlichkeit, kombiniert mit einer massiven internen Bibliothek interdisziplinärer Verbindungen, ermöglicht es dem Modell, höchst detaillierte und knifflige Beweise auszuführen – etwa solche, die lineare Programmiergrenzen in hochdimensionalen Räumen betreffen –, für die es für einen Menschen zu riskant oder zeitaufwendig wäre, seine Karriere zu verspielen.


Da die KI die Produktion neuer Theoreme beschleunigt, entwickelt sich die Rolle des Mathematikers vom alleinigen „Beweisführer“ hin zu einem Kurator und Kommunikator. Die Sprecher legen nahe, dass sich der Wert im Fachbereich hin zum „gemeinschaftlichen Verständnis“ verschieben wird – der Fähigkeit, ein Ergebnis zu verinnerlichen und es in den breiteren Rahmen des menschlichen Wissens einzuordnen. Während die „Obergrenze“ der Schwierigkeit hoch bleibt – die Gäste merken an, dass einige legendäre Rätsel wie P gegen NP möglicherweise weiterhin unerreichbar bleiben –, erweitert sich die Zugänglichkeit von High-Level-Mathematik. Dies ermöglicht es Nicht-Experten, komplexe Fachartikel schneller als je zuvor zu erfassen und theoretische Ergebnisse auf praktische Probleme anzuwenden.


Überraschende Erkenntnisse



  • Kontextuelle Kontamination: Menschen haben Schwierigkeiten, falsche Pfade zu verlassen, da ihre Intuition durch den gescheiterten Versuch „kontaminiert“ wird; die KI vermeidet dies, indem sie effektiv frische Klone ihres Denkprozesses startet.

  • Emergentes Denken: Die Fähigkeit des Modells, zurückzugehen (Backtracking) und Suchbäume zu beschneiden, entstand aus einem allgemeinen Training im logischen Denken und wurde nicht spezifisch durch mathematische Lehrbücher oder formale Sprachen wie Lean gelehrt.

  • Der „hässliche“ vs. der „elegante“ Beweis: Während einige berühmte menschliche Beweise (wie jene für die 3D-Kugelpackung) als notorisch „hässlich“ und hunderte Seiten lang gelten, sind die KI-generierten Beweise für ähnliche hochdimensionale Probleme oft erschreckend kurz und elegant.

  • Die Effizienz des „Geschmacks“: „Mathematischer Geschmack“ kann utilitaristisch als die Fähigkeit betrachtet werden, bessere Urteile zu fällen, die zu schnelleren Lösungen führen – eine Eigenschaft, die die KI entwickelt, während sie zunehmend komplexere, mehrstufige Probleme löst.


Praktische Schlussfolgerungen



  • KI für die Literatursynthese nutzen: Anstatt Stunden mit der Suche in fragmentierter akademischer Literatur zu verbringen, um zu prüfen, ob ein Problem noch „offen“ ist, nutzen Sie fortschrittliche LLMs, um obskure Referenzen zu finden und festzustellen, ob ein Ergebnis in Reichweite liegt.

  • Die Aufnahme von Fachartikeln beschleunigen: Nutzen Sie KI, um die „grobe Beweisstrategie“ aus komplexen PDFs zu extrahieren, um die Kernlogik schnell zu verstehen, bevor Sie in die dichten technischen Details eintauchen.

  • Iteratives Prompting für mehr Tiefe: Geben Sie sich bei der Nutzung von KI für komplexe Denkprozesse nicht mit der ersten korrekten Antwort zufrieden; bitten Sie das Modell, „dies weiterzuführen“ oder die Grenzen seiner eigenen Logik zu untersuchen, um anspruchsvollere theoretische Rahmenbedingungen aufzudecken.

  • Fokus auf Kuration verschieben: Wer in technischen Bereichen tätig ist, sollte sich weniger auf die routinemäßige Ausführung von Beweisen und mehr auf die „architektonische“ Seite der Forschung konzentrieren – also darauf, welche Richtungen vielversprechend sind und wie Ergebnisse an andere kommuniziert werden.


a16z Infra Partner Lisha Li sits down with OpenAI mathematicians Mehtaab Sawhney and Mark Sellke to discuss how quickly AI’s mathematical capabilities are advancing, what recent results reveal about model reasoning, and what happens when AI begins making progress on problems mathematicians have struggled with for decades.

Mehtaab and Mark unpack several recent results from OpenAI’s models, including advances in sphere packing and the construction of a non-sofic group. They explain why the surprising part isn’t simply that models can search more possibilities or work longer than humans: in many cases, the reasoning traces look remarkably similar to the work of an expert mathematician, including choosing promising approaches, backtracking when they fail, and combining ideas from across the literature.

They also explore what this means for mathematics itself: how the role of human taste and judgment may change, whether AI could produce far more mathematics than humans can absorb, and why models that accelerate discovery may also make sophisticated results easier to understand.

Resources:

Follow Lisha Li on X: https://x.com/lishali88

Follow Mehtaab Sawhney on X: https://x.com/mehtaab_sawhney

Follow Mark Sellke on X: https://x.com/MarkSellke

Stay Updated:

Find a16z on YouTube: YouTube

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Show on Spotify

Listen to the a16z Show on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Leave a Reply

Let's Evolve Together
Logo