Summary & Insights
Can the very guardrails designed to protect AI from malicious actors actually handicap the “good guys” during a cyberattack? This paradox is at the heart of a shifting security landscape where traditional defenses are becoming obsolete. As AI models move from centralized clouds to endpoints and become integrated into thousands of enterprise applications, security teams are finding that the tools they’ve relied on for decades—built specifically to stop human hackers and malware—are fundamentally incapable of detecting agentic AI processes.
The transition toward “agentic” software—where AI can take autonomous actions to achieve a goal—is happening at a breakneck pace, with projections suggesting half of enterprise apps will be agentic by the end of this year. This shift renders signature-based detection and even modern behavioral analysis ineffective because AI doesn’t behave predictably. In a startling example of this friction, traditional “honeypots” (fake credentials left to lure hackers) are now triggering massive waves of false positives simply because helpful AI agents are finding those keys and attempting to use them to complete a legitimate task for an employee.
Despite these challenges, there is a silver lining: the same frontier models creating these vulnerabilities are providing defenders with unprecedented power. Security teams can now use thousands of automated agents to build threat taxonomies and defensive frameworks in weeks rather than years. While the attack surface has expanded to include the “total sum of human expression,” the ability to automate the defense side of the equation may eventually shift the balance of power back in favor of the defenders.
Surprising Insights
- The Guardrail Paradox: Safety filters intended to stop hackers often trigger “cyber refusals” for legitimate security defenders who ask the model to help them find or validate a vulnerability in their own system.
- Honeypot Noise: AI agents are so efficient at searching for resources that they frequently trip security “traps” (honeypots) by accident, leading to a surge in false positives for security operations centers.
- The Death of Signatures: Traditional signature-based security is essentially dead because AI-driven attacks don’t rely on static pieces of malware that can be identified by a known “fingerprint.”
- Biofilter Glitches: AI guardrails can be overly sensitive; for example, a security tool named “Vectra” once triggered bio-weapon filters because it shares a name with a veterinary drug.
Practical Takeaways
- Prioritize Model Flexibility: Avoid vendor lock-in with a single AI provider. Blue teams should maintain the ability to switch between models (e.g., moving from a highly guarded model to an open-weights model) to bypass refusals during active incident response.
- Shift to Endpoint Governance: Because AI agents operate autonomously across various apps, focus security controls at the endpoint level to monitor what software is installed and what it is actually doing, rather than relying on the publisher’s intended design.
- Re-evaluate “Normal” Behavior: Stop relying on static baselines of “normal” software behavior. Implement a more dynamic discovery process to understand how agentic software interacts with your specific environment.
- Automate the Defense: Use AI agents to accelerate the creation of threat research and security taxonomies to keep pace with the “machine speed” of modern attacks.
Liệu chính những “rào chắn” (guardrails) được thiết kế để bảo vệ AI khỏi những tác nhân độc hại lại có thể gây cản trở cho những “người tốt” trong một cuộc tấn công mạng? Nghịch lý này nằm ở trung tâm của một bối cảnh bảo mật đang thay đổi, nơi các biện pháp phòng thủ truyền thống đang dần trở nên lỗi thời. Khi các mô hình AI chuyển dịch từ các đám mây tập trung sang các thiết bị đầu cuối (endpoints) và được tích hợp vào hàng ngàn ứng dụng doanh nghiệp, các đội ngũ bảo mật nhận thấy rằng những công cụ mà họ đã tin dùng trong nhiều thập kỷ—vốn được xây dựng đặc biệt để ngăn chặn hacker là con người và phần mềm độc hại—về cơ bản là không đủ khả năng phát hiện các quy trình AI có tính tác nhân (agentic AI).
Quá trình chuyển đổi sang phần mềm “có tính tác nhân” (agentic software)—nơi AI có thể thực hiện các hành động tự chủ để đạt được mục tiêu—đang diễn ra với tốc độ chóng mặt, với các dự báo cho thấy một nửa số ứng dụng doanh nghiệp sẽ trở thành agentic vào cuối năm nay. Sự thay đổi này khiến việc phát hiện dựa trên chữ ký (signature-based detection) và thậm chí cả phân tích hành vi hiện đại trở nên vô hiệu vì AI không hoạt động theo một cách có thể dự đoán trước. Một ví dụ đáng kinh ngạc về sự xung đột này là các “honeypot” truyền thống (các thông tin xác thực giả được để lại để nhử hacker) hiện đang gây ra những làn sóng cảnh báo sai (false positives) ồ ạt, đơn giản vì các tác nhân AI hỗ trợ đang tìm thấy những khóa này và cố gắng sử dụng chúng để hoàn thành một tác vụ hợp pháp cho nhân viên.
Bất chấp những thách thức này, vẫn có một tia hy vọng: chính những mô hình tiên phong (frontier models) tạo ra những lỗ hổng này cũng đang cung cấp cho những người phòng thủ một sức mạnh chưa từng có. Các đội ngũ bảo mật giờ đây có thể sử dụng hàng ngàn tác nhân tự động để xây dựng phân loại mối đe dọa và khung phòng thủ trong vài tuần thay vì vài năm. Mặc dù bề mặt tấn công đã mở rộng ra bao gồm “toàn bộ tổng thể biểu đạt của con người”, nhưng khả năng tự động hóa phía phòng thủ cuối cùng có thể chuyển cán cân quyền lực trở lại có lợi cho những người bảo vệ.
Những góc nhìn bất ngờ
- Nghịch lý rào chắn: Các bộ lọc an toàn nhằm ngăn chặn hacker thường kích hoạt tình trạng “từ chối an ninh mạng” đối với những chuyên gia bảo mật chính thống khi họ yêu cầu mô hình giúp tìm hoặc xác minh một lỗ hổng trong chính hệ thống của họ.
- Nhiễu từ Honeypot: Các tác nhân AI tìm kiếm tài nguyên hiệu quả đến mức chúng thường vô tình vấp phải các “bẫy” bảo mật (honeypots), dẫn đến sự gia tăng đột biến các cảnh báo sai cho các trung tâm vận hành an ninh (SOC).
- Sự khai tử của chữ ký: Bảo mật dựa trên chữ ký truyền thống về cơ bản đã chết vì các cuộc tấn công do AI điều khiển không dựa trên các mẩu phần mềm độc hại tĩnh có thể được nhận diện bằng một “dấu vân tay” đã biết.
- Lỗi bộ lọc sinh học: Các rào chắn AI có thể quá nhạy cảm; ví dụ, một công cụ bảo mật tên là “Vectra” từng kích hoạt các bộ lọc vũ khí sinh học vì nó trùng tên với một loại thuốc thú y.
Bài học thực tiễn
- Ưu tiên tính linh hoạt của mô hình: Tránh việc bị phụ thuộc vào một nhà cung cấp AI duy nhất (vendor lock-in). Các đội phòng thủ (Blue teams) nên duy trì khả năng chuyển đổi giữa các mô hình (ví dụ: chuyển từ một mô hình bị kiểm soát chặt chẽ sang một mô hình trọng số mở – open-weights) để vượt qua các lời từ chối trong quá trình ứng phó sự cố thực tế.
- Chuyển sang quản trị đầu cuối (Endpoint Governance): Vì các tác nhân AI hoạt động tự chủ trên nhiều ứng dụng khác nhau, hãy tập trung các kiểm soát bảo mật ở cấp độ thiết bị đầu cuối để theo dõi phần mềm nào được cài đặt và nó thực sự đang làm gì, thay vì dựa vào thiết kế dự kiến của nhà phát hành.
- Đánh giá lại hành vi “bình thường”: Ngừng dựa vào các mốc cơ sở tĩnh về hành vi phần mềm “bình thường”. Hãy triển khai một quy trình khám phá năng động hơn để hiểu cách phần mềm có tính tác nhân tương tác với môi trường cụ thể của bạn.
- Tự động hóa phòng thủ: Sử dụng các tác nhân AI để tăng tốc việc nghiên cứu mối đe dọa và xây dựng phân loại bảo mật nhằm bắt kịp với “tốc độ máy” của các cuộc tấn công hiện đại.
旨在防止惡意行為者侵害 AI 的「護欄」(guardrails),是否反而會在網路攻擊發生時限制「正義之士」的行動?這一矛盾正是目前安全格局轉變的核心,傳統的防禦手段正逐漸失效。隨著 AI 模型從集中式雲端移至端點(endpoints),並被整合進數以千計的企業應用程式中,安全團隊發現他們數十年來依賴的工具——那些專為攔截人類駭客與惡意軟體而設計的工具——在根本上無法偵測「代理式 AI」(agentic AI)的運作流程。
向「代理式」軟體(即 AI 能採取自主行動以達成目標)的轉型正以驚人的速度進行,預測顯示到今年年底,半數的企業應用程式將具備代理能力。這種轉變使得基於特徵碼(signature-based)的偵測,甚至現代的行為分析(behavioral analysis)都變得無效,因為 AI 的行為並不具有可預測性。一個令人驚訝的衝突案例是:傳統的「蜜罐」(honeypots,即為了誘捕駭客而留下的假憑據)現在正觸發大量誤報,僅僅是因為熱心的 AI 代理在尋找資源以完成員工的合法任務時,發現並嘗試使用這些金鑰。
儘管面臨這些挑戰,但仍有一線希望:創造這些漏洞的尖端模型,同時也賦予了防禦者前所未有的力量。安全團隊現在可以使用數千個自動化代理,在短短數週而非數年內建立威脅分類法(threat taxonomies)和防禦框架。雖然攻擊面已擴展到涵蓋「人類表達的所有總和」,但自動化防禦的能力最終可能會將權力天平重新傾斜回防禦者這一方。
驚人洞察
- 護欄悖論: 旨在阻止駭客的安全過濾器,經常對合法的安全防禦者觸發「網路拒絕」(cyber refusals),即便後者只是請求模型協助尋找或驗證其自身系統中的漏洞。
- 蜜罐噪音: AI 代理在搜尋資源方面效率極高,導致它們經常意外觸發安全「陷阱」(蜜罐),進而造成安全運作中心(SOC)的誤報激增。
- 特徵碼之死: 傳統的特徵碼安全防禦基本上已失效,因為 AI 驅動的攻擊不再依賴於可透過已知「指紋」識別的靜態惡意軟體片段。
- 生物過濾器故障: AI 護欄可能過於敏感;例如,一個名為「Vectra」的安全工具曾觸發生物武器過濾器,僅僅是因為它與一種獸醫藥品同名。
實務建議
- 優先考慮模型的靈活性: 避免被單一 AI 供應商綁定(vendor lock-in)。藍軍(Blue teams)應保持在不同模型之間切換的能力(例如,在緊急事件響應期間,從高度受限的模型切換到開放權重模型),以繞過拒絕回應的問題。
- 轉向端點治理: 由於 AI 代理在各種應用程式中自主運行,安全控制應聚焦於端點層級,監控安裝了哪些軟體以及其真實行為,而非依賴發行商的預設設計。
- 重新評估「正常」行為: 停止依賴靜態的軟體行為「正常」基準線。實施更動態的探索流程,以了解代理式軟體如何與您的特定環境互動。
- 自動化防禦: 利用 AI 代理加速威脅研究和安全分類法的建立,以跟上現代攻擊的「機器速度」。
Les garde-fous mêmes conçus pour protéger l’IA contre les acteurs malveillants peuvent-ils, en réalité, handicaper les « gentils » lors d’une cyberattaque ? Ce paradoxe est au cœur d’un paysage sécuritaire en pleine mutation, où les défenses traditionnelles deviennent obsolètes. À mesure que les modèles d’IA quittent les clouds centralisés pour s’installer sur les terminaux (endpoints) et s’intègrent dans des milliers d’applications d’entreprise, les équipes de sécurité constatent que les outils sur lesquels elles se sont appuyées pendant des décennies — conçus spécifiquement pour arrêter les hackers humains et les logiciels malveillants — sont fondamentalement incapables de détecter les processus d’IA agentique.
La transition vers des logiciels « agentiques » — où l’IA peut prendre des mesures autonomes pour atteindre un objectif — s’effectue à un rythme effréné, les projections suggérant que la moitié des applications d’entreprise seront agentiques d’ici la fin de l’année. Ce changement rend la détection basée sur les signatures, et même l’analyse comportementale moderne, inefficaces car l’IA ne se comporte pas de manière prévisible. Dans un exemple frappant de cette friction, les « pots de miel » traditionnels (de faux identifiants laissés pour attirer les pirates) déclenchent désormais des vagues massives de faux positifs, simplement parce que des agents d’IA serviables trouvent ces clés et tentent de les utiliser pour accomplir une tâche légitime pour un employé.
Malgré ces défis, il y a une lueur d’espoir : les mêmes modèles de pointe qui créent ces vulnérabilités offrent aux défenseurs une puissance sans précédent. Les équipes de sécurité peuvent désormais utiliser des milliers d’agents automatisés pour élaborer des taxonomies de menaces et des cadres défensifs en quelques semaines plutôt qu’en quelques années. Bien que la surface d’attaque se soit étendue pour inclure « la somme totale de l’expression humaine », la capacité d’automatiser le côté défense de l’équation pourrait éventuellement faire basculer l’équilibre du pouvoir en faveur des défenseurs.
Insights Surprenants
- Le paradoxe des garde-fous : Les filtres de sécurité destinés à arrêter les hackers déclenchent souvent des « refus cyber » pour les défenseurs de la sécurité légitimes qui demandent au modèle de les aider à trouver ou à valider une vulnérabilité dans leur propre système.
- Le bruit des pots de miel : Les agents d’IA sont si efficaces pour rechercher des ressources qu’ils déclenchent fréquemment des « pièges » de sécurité (honeypots) par accident, entraînant une augmentation des faux positifs pour les centres d’opérations de sécurité (SOC).
- La mort des signatures : La sécurité traditionnelle basée sur les signatures est pratiquement morte, car les attaques pilotées par l’IA ne reposent pas sur des fragments statiques de logiciels malveillants pouvant être identifiés par une « empreinte » connue.
- Défaillances des bio-filtres : Les garde-fous de l’IA peuvent être hypersensibles ; par exemple, un outil de sécurité nommé « Vectra » a un jour déclenché des filtres d’armes biologiques parce qu’il partage son nom avec un médicament vétérinaire.
Conseils Pratiques
- Prioriser la flexibilité des modèles : Évitez la dépendance exclusive envers un seul fournisseur d’IA. Les « Blue Teams » doivent maintenir la capacité de basculer entre les modèles (par exemple, passer d’un modèle très restreint à un modèle à poids ouverts/open-weights) pour contourner les refus lors d’une réponse active aux incidents.
- Passer à la gouvernance des terminaux : Étant donné que les agents d’IA fonctionnent de manière autonome à travers diverses applications, concentrez les contrôles de sécurité au niveau du terminal pour surveiller quel logiciel est installé et ce qu’il fait réellement, plutôt que de se fier à la conception prévue par l’éditeur.
- Réévaluer le comportement « normal » : Cessez de vous appuyer sur des lignes de base statiques du comportement « normal » des logiciels. Mettez en œuvre un processus de découverte plus dynamique pour comprendre comment les logiciels agentiques interagissent avec votre environnement spécifique.
- Automatiser la défense : Utilisez des agents d’IA pour accélérer la recherche sur les menaces et la création de taxonomies de sécurité afin de suivre la « vitesse machine » des attaques modernes.
Können genau jene Leitplanken, die eigentlich KI vor böswilligen Akteuren schützen sollen, die „Guten“ während eines Cyberangriffs behindern? Dieses Paradoxon steht im Zentrum einer sich wandelnden Sicherheitslandschaft, in der traditionelle Verteidigungsmechanismen obsolet werden. Während KI-Modelle von zentralisierten Clouds zu Endpunkten wandern und in Tausende von Unternehmensanwendungen integriert werden, stellen Sicherheitsteams fest, dass die Werkzeuge, auf die sie sich jahrzehntelang verlassen haben – speziell entwickelt, um menschliche Hacker und Malware zu stoppen –, grundlegend unfähig sind, agentische KI-Prozesse zu erkennen.
Der Übergang zu „agentischer“ Software – bei der KI autonom handeln kann, um ein Ziel zu erreichen – vollzieht sich in einem rasanten Tempo. Prognosen gehen davon aus, dass bis Ende dieses Jahres die Hälfte aller Unternehmensanwendungen agentisch sein wird. Dieser Wandel macht signaturbasierte Erkennung und selbst moderne Verhaltensanalysen unwirksam, da KI sich nicht vorhersehbar verhält. In einem frappierenden Beispiel für diese Reibung lösen traditionelle „Honeypots“ (gefälschte Zugangsdaten, die Hacker anlocken sollen) nun massive Wellen von False Positives aus, schlicht weil hilfreiche KI-Agenten diese Schlüssel finden und versuchen, sie zu nutzen, um eine legitime Aufgabe für einen Mitarbeiter zu erledigen.
Trotz dieser Herausforderungen gibt es einen Lichtblick: Dieselben Frontier-Modelle, die diese Schwachstellen schaffen, verleihen den Verteidigern eine beispiellose Macht. Sicherheitsteams können nun Tausende von automatisierten Agenten einsetzen, um Bedrohungs-Taxonomien und Verteidigungsrahmen in Wochen statt in Jahren aufzubauen. Während sich die Angriffsfläche auf die „Gesamtsumme des menschlichen Ausdrucks“ ausgeweitet hat, könnte die Fähigkeit, die Verteidigungsseite der Gleichung zu automatisieren, das Machtgleichgewicht letztendlich wieder zugunsten der Verteidiger verschieben.
Überraschende Erkenntnisse
- Das Leitplanken-Paradoxon: Sicherheitsfilter, die Hacker stoppen sollen, lösen oft „Cyber-Verweigerungen“ (Cyber Refusals) bei legitimen Sicherheitsverantwortlichen aus, wenn diese das Modell bitten, ihnen dabei zu helfen, eine Schwachstelle im eigenen System zu finden oder zu validieren.
- Honeypot-Rauschen: KI-Agenten sind so effizient bei der Suche nach Ressourcen, dass sie häufig versehentlich Sicherheits„fallen“ (Honeypots) auslösen, was zu einem Anstieg von Fehlalarmen in Security Operations Centers führt.
- Das Ende der Signaturen: Die traditionelle signaturbasierte Sicherheit ist praktisch tot, da KI-gesteuerte Angriffe nicht auf statischen Malware-Fragmenten basieren, die durch einen bekannten „Fingerabdruck“ identifiziert werden können.
- Biofilter-Fehler: KI-Leitplanken können überempfindlich sein; so löste beispielsweise ein Sicherheitstool namens „Vectra“ einmal Bio-Waffen-Filter aus, weil es den gleichen Namen wie ein Veterinärpräparat trägt.
Praktische Empfehlungen
- Modell-Flexibilität priorisieren: Vermeiden Sie den Vendor-Lock-in bei einem einzigen KI-Anbieter. Blue Teams sollten die Fähigkeit behalten, zwischen Modellen zu wechseln (z. B. von einem stark reglementierten Modell zu einem Open-Weights-Modell), um Verweigerungen während einer aktiven Reaktion auf Vorfälle zu umgehen.
- Wechsel zur Endpoint-Governance: Da KI-Agenten autonom über verschiedene Apps hinweg agieren, sollten Sicherheitskontrollen auf der Endpunktebene konzentriert werden. Hierbei sollte überwacht werden, welche Software installiert ist und was sie tatsächlich tut, anstatt sich auf das beabsichtigte Design des Herstellers zu verlassen.
- „Normales“ Verhalten neu bewerten: Verlassen Sie sich nicht mehr auf statische Baselines für „normales“ Softwareverhalten. Implementieren Sie einen dynamischeren Discovery-Prozess, um zu verstehen, wie agentische Software mit Ihrer spezifischen Umgebung interagiert.
- Die Verteidigung automatisieren: Setzen Sie KI-Agenten ein, um die Erstellung von Bedrohungsanalysen und Sicherheits-Taxonomien zu beschleunigen, um mit der „Maschinengeschwindigkeit“ moderner Angriffe Schritt zu halten.
a16z’s Joel De La Garza is joined by Nick Warner of Neo and Max Pollard of Cotool to discuss what happens when cybersecurity tools built to defend against humans and malware suddenly have to contend with AI agents. As frontier models become more capable of finding and exploiting vulnerabilities, many of the assumptions underlying traditional security are beginning to break.
They explore why guardrails designed to stop AI-powered attackers can also prevent security teams from doing their jobs, why defenders increasingly need access to multiple models, and how agentic software creates an entirely new endpoint security problem. They also discuss why static signatures and even newer techniques like honeypots are struggling in a world where software can reason and act autonomously.
Recorded around Black Hat, the conversation looks at how security teams are adapting in real time and why the same AI capabilities creating new attack surfaces could ultimately give defenders their biggest advantage yet.
Resources:
Follow Nick on LinkedIn: https://www.linkedin.com/in/nicholaswarner/
Follow Max on LinkedIn: https://www.linkedin.com/in/mmpollard/
Follow Joel De La Garza on LinkedIn: https://www.linkedin.com/in/3448827723723234/
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Leave a Reply
You must be logged in to post a comment.