Security News

Cybersecurity news aggregator

INFO News SC Media

New AI security approach analyzes model internals

  • What: New AI security approach analyzes model internals
  • Impact: Researchers develop method to detect threats by examining AI activation patterns
Read Full Article →

AI/ML New AI security approach analyzes model internals July 29, 2026 Share By SC Staff As reported by Dark Reading, researchers are developing a novel approach to AI security that moves beyond analyzing only the inputs and outputs of large language models (LLMs). This new method aims to understand the internal workings of AI systems by examining their activation patterns, offering a more robust defense against malicious use. Offensive-security researchers are presenting a model-agnostic method for activation analysis at Black Hat USA 2026. This technique uses standardized rules for processing activation events, moving away from broad labels like "cybercrime" to a more granular scheme of cognitive elements (CEs). These CEs, such as "create content" or "personal information," can be combined into logical statements to detect specific threats like phishing attacks. The goal is to create an open system of identified CEs and rules, similar to Snort or YARA rulesets, to detect safety events. This approach, dubbed GAVEL (Governance via Activation-based Verification and Extensible Logic), aims to identify specific activation patterns associated with granular objects and predicates. Unlike token-based analysis, activation analysis is language-independent and can detect malicious intent even when prompts are altered to evade content filters. While still a research project, this method offers a potential new layer of defense for AI systems, complementing existing security measures. Source: Dark Reading An In-Depth Guide to AI Get essential knowledge and practical strategies to use AI to better your security program. Learn More SC Staff Related AI/ML Model Context Protocol releases major update to AI interaction technology SC Staff July 29, 2026 The latest release of MCP overhauls its request metadata processing, replacing a complex handshake workflow with a stateless protocol core. AI/ML Loss of control: The AI agent governance crisis Paul Wagenseil July 29, 2026 AI governance is an identity challenge, but legacy identity systems weren't made to handle non-deterministic software. AI/ML Seven controls enterprise teams need in place to avoid another OpenAI-Hugging Face incident Himanshu Shukla July 29, 2026 Seven controls help enterprises govern AI agents and prevent costly sandbox escapes. Get daily email updates SC Media's daily must-read of the most current and pressing daily news Business Email By clicking the Subscribe button below, you agree to SC Media Terms of Use and Privacy Policy . Subscribe You can skip this ad in 5 seconds

Share this article