<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
<channel>
  <title>Prasanna Malode — Articles</title>
  <link>https://prasannamalode.in/</link>
  <atom:link href="https://prasannamalode.in/feed.xml" rel="self" type="application/rss+xml"/>
  <description>DevSecOps transformation, cybersecurity leadership, AI governance and building high-performing engineering teams.</description>
  <language>en</language>
  <lastBuildDate>Mon, 14 Sep 2026 00:00:00 +0000</lastBuildDate>
  <item>
    <title>The 6 Outcomes of Information Security Governance: What Leaders Actually Need to Know</title>
    <link>https://prasannamalode.in/articles/information-security-governance-six-outcomes.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/information-security-governance-six-outcomes.html</guid>
    <pubDate>Mon, 14 Sep 2026 00:00:00 +0000</pubDate>
    <category>Cybersecurity</category>
    <description>Governance is the operating system for how an organisation makes security decisions. Six measurable outcomes it must deliver, how to test each one, and the loop that connects them.</description>
    <content:encoded><![CDATA[<h2>The Problem: Too Much Jargon, Not Enough Clarity</h2>
<p>Every leader I talk to tells me the same thing: "Security governance" sounds important, but they can't quite explain what it <em>actually</em> does. Is it compliance? Risk management? Strategy? All of the above?</p>
<p>The confusion is understandable. Security governance has become one of those overloaded terms that means everything to consultants and nothing to practitioners. But here's what I've learned after years managing security in complex organizations: <strong>governance isn't abstract—it's the operating system for how your organization makes security decisions and delivers business value.</strong></p>
<p>And it produces six clear, measurable outcomes. If your governance isn't delivering these, you have a governance problem—not a security problem.</p>
<h2>What Is Information Security Governance, Really?</h2>
<p>Let's start with a definition that sticks:</p>
<p><strong>Information security governance is the system by which your organization directs and controls security to:</strong></p>
<ul><li>Ensure security aligns with business objectives</li><li>Manage risk to acceptable levels</li><li>Use resources responsibly</li><li>Measure results and hold people accountable</li></ul>
<p>Notice what's missing from that definition? Firewalls. Encryption. Compliance checkboxes. Those are <em>security controls</em>. Governance is the framework that decides <em>which</em> controls matter, <em>why</em>, and <em>whether they're working</em>.</p>
<p>The difference matters. You can buy the best SIEM in the world and still fail at governance. You can hire brilliant engineers and still fail at governance. But fail at governance, and your security program drifts from business reality.</p>
<p>Governance is where business strategy meets security strategy. And that's exactly why boards and executives are paying attention to it now—not because regulators forced them to, but because companies with strong governance waste less money, lose fewer breaches, and move faster.</p>
<h2>The 6 Outcomes Every Leader Should Know</h2>
<p>When governance is working, you see these six outcomes in your organization. When you don't, you know where to look first.</p>
<h3>1. Strategic Alignment</h3>
<p><strong>What it means:</strong> Security decisions support business objectives, not exist apart from them.</p>
<p>This is the primary outcome. Everything starts here.</p>
<p>I've seen countless security programs that were technically excellent but strategically adrift. Mature architectures, strong controls, zero breaches—and yet the business viewed security as a cost center that slowed things down. That's a failure of alignment.</p>
<p>Strategic alignment means:</p>
<ul><li>Your security strategy <em>starts</em> with the business strategy, not the other way around</li><li>Every major security initiative connects to a business goal (revenue, customer trust, operational efficiency, regulatory compliance—whatever matters to your board)</li><li>The CISO can explain to an executive why you're investing in X and not Y in business terms they understand</li><li>Security is invited into strategic planning <em>before</em> decisions are made, not consulted after</li></ul>
<p>When alignment is strong, security becomes an enabler. When it's weak, security becomes a bottleneck.</p>
<p><strong>How to measure it:</strong> Can your leadership articulate 2–3 business objectives that security enables? If not, you lack strategic alignment.</p>
<h3>2. Risk Management</h3>
<p><strong>What it means:</strong> Risk is identified, understood, and managed to a level the organization finds acceptable.</p>
<p>Important note: "Acceptable" doesn't mean "zero." Most organizations I've worked with eventually realize they can't eliminate all risk—they can only manage it to levels that allow them to operate and compete.</p>
<p>Risk management as a governance outcome means:</p>
<ul><li>The organization has explicitly defined its risk appetite (how much risk are we willing to take?)</li><li>Risk decisions are made by business owners (who understand their assets) with security advisors, not by security alone</li><li>Risk assessments inform strategy, policy, and budget decisions</li><li>Residual risk (the risk that remains after controls are in place) is understood, documented, and accepted by management</li></ul>
<p>Without this outcome, you get:</p>
<ul><li>Security decisions driven by "best practice" instead of business need</li><li>Expensive controls that manage risks nobody cares about</li><li>Conflicts between security and business teams over what matters most</li><li>No framework for saying "no" to some controls and "yes" to others</li></ul>
<p><strong>How to measure it:</strong> Does your organization have a documented risk appetite? Can you name the top 5 risks to your business and how they're being managed? If not, your governance needs strengthening.</p>
<h3>3. Value Delivery</h3>
<p><strong>What it means:</strong> Security investments actually reduce business risk and enable business objectives. You get what you pay for.</p>
<p>This is where ROI enters the picture—and it's uncomfortable territory for many security teams. We like to talk about "you can't put a price on security." That's true philosophically. But you absolutely can and should put a price on <em>security investments</em>, and you should be able to show what they deliver.</p>
<p>Value delivery as a governance outcome means:</p>
<ul><li>Security spending is tied to strategy and risk reduction, not arbitrary budget percentages</li><li>Major security initiatives have business cases that justify the cost (initial purchase + ongoing operations)</li><li>You're actually measuring whether the investment delivered the promised benefit</li><li>Resources (people, tools, processes) are allocated to the areas of highest business impact</li></ul>
<p>When value delivery is weak, you see:</p>
<ul><li>Security budgets that are disconnected from risk</li><li>Tools and technologies deployed but underutilized</li><li>"We don't know if this control is working" situations</li><li>Senior management asking, "Why are we spending this much on security?"</li></ul>
<p><strong>How to measure it:</strong> For your three largest security investments this year, can you articulate what business outcome each delivered or is expected to deliver? If not, you have a value-delivery problem.</p>
<h3>4. Resource Management</h3>
<p><strong>What it means:</strong> People, processes, and technology are allocated efficiently to achieve security objectives.</p>
<p>This goes beyond "do we have enough budget?" It's about whether the right resources are in the right places doing the right things.</p>
<p>Resource management as a governance outcome means:</p>
<ul><li>You have the right people (with the right skills) in the right roles</li><li>You're developing talent and closing skill gaps, not hoping you'll find unicorn engineers</li><li>Process frameworks (how work gets done) reduce waste and enable consistency</li><li>Technology investments are matched to business needs, not acquired because they're shiny</li><li>Outsourcing/managed services decisions are made strategically, with clear accountability boundaries</li></ul>
<p>A governance failure in resource management looks like:</p>
<ul><li>Bottlenecks where one overworked person is critical to everything</li><li>Reactive hiring when crises hit instead of proactive planning</li><li>Tool sprawl and underutilized platforms</li><li>Outsourced functions with no clear oversight or accountability</li></ul>
<p><strong>How to measure it:</strong> If your top security person quit tomorrow, would the program continue running? If not, you have a resource-management problem.</p>
<h3>5. Performance Measurement</h3>
<p><strong>What it means:</strong> You have metrics that show whether governance and security objectives are being achieved.</p>
<p>This is measurement for decision-making, not measurement for its own sake. The goal is to answer: <em>Are we on track? Is the program delivering?</em></p>
<p>Most organizations measure the wrong things. They count blocked packets and security incidents as if these are the point. But the real governance questions are:</p>
<ul><li>Is risk trending in the right direction (down)?</li><li>Is the security program maturing in the areas that matter most?</li><li>Are we achieving our strategic objectives?</li><li>What's the business impact of our security investments?</li></ul>
<p>Performance measurement as a governance outcome means:</p>
<ul><li>Leadership-level metrics (KGI = key goal indicators) showing whether strategy is working</li><li>Operational metrics (KPI = key performance indicators) showing process performance</li><li>Risk indicators (KRI = key risk indicators) providing early warning when risk is climbing</li><li>Metrics tailored to the audience (board gets strategy + trend data, operational managers get control-level metrics)</li></ul>
<p>When measurement governance is weak, you get:</p>
<ul><li>Vanity metrics that look good but don't inform decisions</li><li>No connection between operational metrics and business outcomes</li><li>Board reporting that security doesn't understand or trust</li><li>Metrics that drive wrong behavior (e.g., "block everything" instead of "manage risk")</li></ul>
<p><strong>How to measure it:</strong> Can your board answer in one meeting: Are we safer this quarter than last quarter? If they can't, you lack governance-level measurement.</p>
<h3>6. Assurance Process Integration</h3>
<p><strong>What it means:</strong> Assurance functions (internal audit, risk management, compliance, physical security, etc.) work together instead of in silos, providing holistic confidence that governance and controls are effective.</p>
<p>Organizations often end up with overlapping or conflicting assurance activities. Audit checks one thing, compliance checks another, risk management owns a third. Nobody has the full picture.</p>
<p>Assurance process integration as a governance outcome means:</p>
<ul><li>Assurance functions coordinate so their work covers the full landscape without massive overlap</li><li>Findings and recommendations flow to governance bodies (audit committee, risk committee, steering committee)</li><li>There's a single source of truth about control effectiveness, not multiple conflicting reports</li><li>Business and security leadership can see where gaps exist and prioritize remediation</li><li>Assurance findings drive continuous improvement</li></ul>
<p>When assurance is fragmented, you see:</p>
<ul><li>The same control reviewed three different ways, three different conclusions</li><li>Business managers confused about what "assurance" they actually have</li><li>Audit finding reports that conflict with compliance reports</li><li>No mechanism for turning assurance feedback into action</li></ul>
<p><strong>How to measure it:</strong> When audit finds a gap, can you quickly determine whether compliance and risk management are aware and acting on it? If not, you lack assurance integration.</p>
<h2>How These Six Outcomes Connect</h2>
<p>Here's the elegant part: these outcomes aren't independent. They reinforce each other:</p>
<p><strong>Strategic alignment</strong> drives which risks matter most → <strong>risk management</strong> uses that to prioritize investment → <strong>value delivery</strong> measures whether those investments worked → <strong>resource management</strong> ensures you have what you need for the next cycle → <strong>performance measurement</strong> shows executives the results → <strong>assurance integration</strong> confirms controls are actually in place and effective → loop back to strategic alignment as business needs evolve.</p>
<p>Weak governance breaks that loop. For example:</p>
<ul><li>No strategic alignment → you can't decide which risks to focus on</li><li>Risk management without value delivery → you spend money without knowing if it matters</li><li>No performance measurement → you never know if anything is working</li><li>Fragmented assurance → business leaders don't trust the data</li></ul>
<h2>Reflection Questions</h2>
<p>Before you go, ask yourself:</p>
<ol><li><strong>Strategic Alignment:</strong> Can your business leaders explain what security enables for the organization?</li><li><strong>Risk Management:</strong> Does your organization have an explicit risk appetite, and do you know yours?</li><li><strong>Value Delivery:</strong> For your largest security investment, can you articulate what business outcome it delivered?</li><li><strong>Resource Management:</strong> If your top security leader left tomorrow, would the program continue?</li><li><strong>Performance Measurement:</strong> Can your board answer whether you're safer this quarter than last?</li><li><strong>Assurance Integration:</strong> When audit finds a gap, do your compliance and risk teams automatically coordinate?</li></ol>
<p>If you can't answer at least four of these, your governance framework has work to do.</p>]]></content:encoded>
  </item>
  <item>
    <title>The Demo You Didn&#x27;t Choose</title>
    <link>https://prasannamalode.in/articles/the-demo-you-didnt-choose.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/the-demo-you-didnt-choose.html</guid>
    <pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate>
    <category>AI Governance</category>
    <description>Every AI security vendor demo answers one question: can it handle a case chosen to be handleable. Here are the questions I ask instead.</description>
    <content:encoded><![CDATA[<p>Halfway through the third vendor demo of the month, I realized I had watched the same thirty minutes three times.</p>
<p>Different companies, different logos, genuinely different products underneath. But the structure was identical. An alert appears. The AI reads it. The AI explains it in fluent prose. The AI recommends a response. The response is correct. Somebody says the phrase "in seconds, not hours."</p>
<p>And every single time, the alert was one the vendor had selected.</p>
<p>That is not dishonest. Of course they pick a good example; I would too. But it means the demo answers exactly one question — can this system handle a case chosen to be handleable — and that is not a question I need answered. I need to know what happens on the cases nobody chose.</p>
<p>So I started asking different questions, and the quality of the conversation changed immediately.</p>
<h2>Ask what it does when it does not know</h2>
<p>This is the question I now ask first, and it is remarkably good at sorting vendors.</p>
<p>Every one of these systems will encounter inputs outside what it handles well. The only thing that varies is what happens next. Does it say so? Does it hand off with the raw data intact? Or does it produce a confident, fluent, wrong answer that looks exactly like its right ones?</p>
<p>A vendor who has thought hard about this will answer quickly and specifically, because they have had to build the behavior and probably argued about it internally. A vendor who has not will reach for the accuracy number — and the accuracy number is not the answer. A system that is right ninety-five percent of the time and indistinguishable in presentation the other five is a system your analysts must verify completely, which means it has saved them nothing.</p>
<p>Ask to see the uncertain case. Ask what the interface looks like when confidence is low. If there is no distinct behavior, you are buying something that will be trusted exactly until it burns someone, and then trusted by nobody.</p>
<h2>Ask what it can do, not what it can say</h2>
<p>Any product with tool access has an authority model, whether or not anyone has articulated it.</p>
<p>What credentials does it hold? What can it read? What can it change? Does it act with the permissions of the user who asked, or with a service account that aggregates everyone's? Is there a confirmation step before consequential actions, and who configures it?</p>
<p>I want these answers written down, not described. The gap between a vendor's mental model of their permission architecture and its actual implementation is where a lot of unpleasant surprises live, and writing it out tends to surface it.</p>
<p>The related question, for anything that ingests content from outside: what happens when that content contains instructions? If the answer is that the model is trained not to follow them, that is a mitigation, not a boundary. Ask what constrains the damage when the mitigation fails, because the useful answer is about capability limits, not about model behavior.</p>
<h2>Ask where the data goes and how long it stays</h2>
<p>Straightforward, frequently glossed over, and worth being tedious about.</p>
<p>Where is inference performed? Which subprocessors are involved? Is customer data used for training, by them or by anyone downstream? What is retained, where, for how long? What happens to it when the contract ends?</p>
<p>The subprocessor question is the one that catches people, because a vendor can answer "we do not train on your data" with complete sincerity while routing it to a model provider whose terms are a separate conversation you have not had. Ask about the whole chain.</p>
<p>And ask about incidents: if their provider has a breach, what is your notification path, and what is the timeline? You want that established before it is relevant.</p>
<h2>Ask how it fails over time</h2>
<p>Products get demonstrated at their best moment. Systems live in a world where things drift.</p>
<p>What happens when the underlying model is updated — do you get notice, can you pin a version, does behavior change under you? What happens when your environment changes, when you add a data source or restructure a naming convention? What is the retuning burden, who carries it, and how quickly do you find out that performance has degraded?</p>
<p>That last one deserves emphasis. Detection systems fail silently. A rule that stops matching does not raise an alarm; it just stops producing results, and absence of alerts is indistinguishable from absence of threats until something confirms otherwise. Ask how you would know. If the answer is that you would notice, the answer is that you would not.</p>
<h2>Ask for the boring version of the ROI claim</h2>
<p>"Reduces triage time by eighty percent" is a statement about a measurement someone performed. It is not necessarily false. It is just meaningless until you know the shape of it.</p>
<p>Measured against what baseline — a well-tuned process or a neglected one? Across which alert classes? At what organization size, with what data quality? Was the comparison run concurrently or before-and-after, and what else changed in between?</p>
<p>You are not trying to catch anyone out. You are trying to work out whether the environment where that number was produced resembles yours. Often it does not, and a vendor with a real deployment behind the number will happily tell you the conditions, because the conditions are where their credibility lives.</p>
<h2>The one that ends most conversations</h2>
<p>Run it on my data, on my cases, in shadow mode, for long enough to matter.</p>
<p>Not a proof of concept with a curated dataset. Live traffic, real alerts, real messiness, running alongside the existing process without authority, for a period long enough to include a bad week. Then compare what it said to what actually happened.</p>
<p>This is the only evaluation that answers the question you actually have. It is also the request that separates vendors with a working product from vendors with a working demo, faster than any question in the list above.</p>
<p>The good ones say yes and ask how soon you can start. They want the comparison, because they have run it before and they know how it comes out.</p>
<p>The rest explain why their architecture makes that difficult.</p>]]></content:encoded>
  </item>
  <item>
    <title>Software Supply Chain Security in 2026: Beyond Compliance to Strategic Risk</title>
    <link>https://prasannamalode.in/articles/supply-chain-security-strategic-risk.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/supply-chain-security-strategic-risk.html</guid>
    <pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate>
    <category>Cybersecurity</category>
    <description>From vendor questionnaires to transparency mandates. Three attack vectors, why SBOMs are required but still imperfect, and whether you could isolate a compromised vendor within four hours.</description>
    <content:encoded><![CDATA[<p>The software supply chain has become a primary attack surface. By 2026, the industry has shifted from "SolarWinds-style" awareness to understanding supply chain compromise as a persistent, sophisticated threat vector. Organizations no longer ask "Do we have vendor risk assessments?" but rather "Can we detect and contain a compromise in our supply chain within hours of discovery?"</p>
<h2>The Evolution: From Vendor Questionnaires to Transparency Demands</h2>
<h3>2020-2021: The Questionnaire Era</h3>
<ul><li>Security teams issued annual questionnaires to vendors</li><li>Vendors completed them (with varying honesty)</li><li>Assurances were documented, filed, and forgotten</li></ul>
<h3>2022-2023: The Compliance Shift</h3>
<ul><li>CISA began publishing supply chain guidance</li><li>NIST released SSDF (Secure Software Development Framework)</li><li>Procurement teams required vendors to demonstrate SBOM (Software Bill of Materials) compliance</li><li>Reality check: Most vendors couldn't produce accurate SBOMs</li></ul>
<h3>2024-2025: The Tooling Explosion</h3>
<ul><li>Automated dependency scanning (SCA tools) became standard</li><li>SBOM became a market category</li><li>Organizations began scanning their own software for vulnerable dependencies</li><li>First generation of supply chain management platforms emerged</li></ul>
<h3>2026: The Transparency Mandate</h3>
<ul><li>Customers demand real-time visibility into vendor security posture</li><li>SBOMs are table stakes; inaccuracy is contractual liability</li><li>Software provenance verification (code origin, build integrity) is emerging as critical</li></ul>
<h2>The Current Threat Landscape: Three Attack Vectors Dominating</h2>
<h3>Vector 1: Vulnerable Dependencies at Scale</h3>
<p>Modern software is built from thousands of open-source components. A single vulnerable library can affect millions of applications.</p>
<p><strong>Current Reality:</strong></p>
<ul><li>Average enterprise application contains 200-400 direct dependencies and 5,000-10,000 transitive dependencies</li><li>Zero-day vulnerabilities in popular libraries are discovered weekly</li><li>Time from disclosure to active exploitation: 2-14 days for libraries with significant downstream impact</li></ul>
<p><strong>2026 Response:</strong></p>
<ul><li>Organizations implementing continuous SBOM monitoring and automated vulnerability scanning</li><li>Shift from "patch cycle compliance" (patch within 30/60/90 days) to "zero-day response teams"</li><li>Emerging practice: organizations run simulations of "what if OpenSSL (or equivalent critical library) is compromised" scenarios</li></ul>
<h3>Vector 2: Compromised Maintainers &amp; Typosquatting</h3>
<p>Open-source ecosystems (npm, PyPI, RubyGems) have limited gatekeeping. Bad actors compromise maintainer accounts or create packages with names similar to legitimate libraries.</p>
<p><strong>Documented Incidents (2024-2026):</strong></p>
<ul><li>47 typosquatted packages with 2.8M downloads in PyPI ecosystem (Q2 2026)</li><li>12 documented cases of compromised npm maintainer accounts used to inject malware</li><li>Average time to detection: 14-21 days (relying on community reporting, not systematic monitoring)</li></ul>
<p><strong>2026 Response:</strong></p>
<ul><li>Organizations implementing package allowlisting (whitelist approved packages, deny everything else)</li><li>Dependency source verification (validate package signatures, audit package provenance)</li><li>Internal package mirrors to serve vetted dependencies</li></ul>
<h3>Vector 3: Build Pipeline Compromise</h3>
<p>If an attacker can compromise a software vendor's build pipeline, they can inject malware into legitimate releases.</p>
<p><strong>2026 Examples:</strong></p>
<ul><li>Q1 2026: Nation-state actor compromised CI/CD credentials at a mid-market SaaS provider, injected exfiltration code into production releases</li><li>Q3 2026: Ransomware group compromised a popular developer tool vendor's build environment</li><li>Multiple documented cases of GitHub Actions exploitation leading to artifact tampering</li></ul>
<p><strong>Defense Emerging:</strong></p>
<ul><li>Binary transparency and reproducible builds (verify that same source code produces identical binaries)</li><li>Signed artifacts and supply chain standards (in-toto, SLSA framework)</li><li>Real-time monitoring of build pipeline activity</li></ul>
<h2>SBOM: Now Required, Still Imperfect</h2>
<h3>Regulatory Mandates</h3>
<ul><li>U.S. federal agencies (EO 14028) now require SBOMs for software purchases</li><li>EU Cyber Resilience Act includes SBOM transparency requirements</li><li>Private sector contracts increasingly require SBOMs as procurement condition</li></ul>
<h3>The Quality Problem</h3>
<p>SBOM formats exist (SPDX, CycloneDX), but:</p>
<ul><li>Incomplete SBOMs (vendor didn't capture all dependencies) are common</li><li>Accuracy varies widely (some vendors automate SBOM generation; others manually list components and miss transitive dependencies)</li><li>SBOMs generated once at release become stale if vendors backport patches</li></ul>
<p><strong>Emerging Standard:</strong> Organizations requesting SBOMs in CycloneDX format with automated generation workflows and timestamp validation.</p>
<h3>SBOM Intelligence Workflows</h3>
<p>Organizations that have achieved SBOM maturity are implementing:</p>
<pre><code>Vendor SBOM → Parse &amp; Normalize → Cross-reference NVD/OSV → 
Risk Score (CVSS + transitive depth + exploitation likelihood) → 
Alert if risk threshold exceeded → Track patch timelines</code></pre>
<p>Time from vendor update to internal assessment: 2-4 hours for mature organizations.</p>
<h2>The Emerging Category: Software Supply Chain Intelligence Platforms</h2>
<p>New vendors (and some existing security platforms) are focusing on:</p>
<ul><li><strong>Continuous vendor dependency monitoring</strong> (track which open-source packages your vendors use)</li><li><strong>Industry-wide threat correlation</strong> (if Vendor A and Vendor B both use Library X, and X is compromised, both are at risk)</li><li><strong>Automated remediation workflows</strong> (update vulnerable dependencies, or isolate affected components)</li></ul>
<p>These platforms are gaining adoption because they solve the core problem: <strong>manual monitoring of software composition at the scale of modern enterprises is impossible.</strong></p>
<h2>The Insider Threat Dimension</h2>
<p>2026 has revealed a secondary supply chain risk: insiders at software vendors. Documented cases:</p>
<ul><li><strong>Q1 2026:</strong> Former employee at major cloud infrastructure vendor retained valid credentials, used them to modify infrastructure code that customers depend on</li><li><strong>Q2 2026:</strong> Developer at mid-market security vendor deliberately introduced backdoor in release candidate</li><li><strong>Ongoing risk:</strong> Disgruntled employees, foreign nation-state recruitment of software engineers</li></ul>
<p>Response:</p>
<ul><li>Vendors improving access controls and audit trails</li><li>Organizations demanding transparency into vendor security incident processes</li><li>Emerging practice: "vendor transparency agreements" that require disclosure of security incidents within 24-48 hours</li></ul>
<h2>Organizational Readiness Assessment</h2>
<h3>Questions to Answer:</h3>
<ol><li><strong>Do you have a current inventory of all commercial software your organization uses?</strong><ul><li>Yes: 31% of enterprises</li><li>Partial: 52%</li><li>No: 17%</li></ul></li><li><strong>Can you produce an SBOM for all software developed or integrated internally?</strong><ul><li>Yes: 18% of enterprises</li><li>Partial: 39%</li><li>No: 43%</li></ul></li><li><strong>Do you have an incident response plan specific to supply chain compromise?</strong><ul><li>Yes: 22% of enterprises</li><li>Partially: 35%</li><li>No: 43%</li></ul></li><li><strong>Can you identify and isolate a compromised vendor within 4 hours of discovery?</strong><ul><li>Yes: 8% of enterprises</li><li>Partially: 24%</li><li>No: 68%</li></ul></li></ol>
<p><strong>Assessment:</strong> Most organizations are not prepared for supply chain compromises at scale.</p>
<h2>The Roadmap: 2026-2028</h2>
<h3>Phase 1: Visibility (Months 1-6)</h3>
<ul><li>Inventory all commercial software</li><li>Request SBOMs from all vendors</li><li>Deploy SBOM scanning tools internally</li><li>Begin baseline assessment of dependency risk</li></ul>
<h3>Phase 2: Detection (Months 6-12)</h3>
<ul><li>Implement continuous monitoring of dependencies against vulnerability databases</li><li>Set up alerts for new vulnerabilities in used packages</li><li>Establish incident response workflows for critical dependencies</li><li>Audit build pipeline security</li></ul>
<h3>Phase 3: Response (Months 12-18)</h3>
<ul><li>Formalize vendor incident response agreements</li><li>Build runbooks for supply chain compromise scenarios</li><li>Establish rapid patching capabilities</li><li>Implement binary transparency verification for critical vendors</li></ul>
<h3>Phase 4: Resilience (Months 18-24)</h3>
<ul><li>Implement package allowlisting and source verification</li><li>Deploy internal dependency mirrors for critical packages</li><li>Achieve reproducible build verification</li><li>Establish metrics for supply chain security posture</li></ul>
<h2>Critical Success Factors</h2>
<ol><li><strong>DevSecOps integration</strong> (SBOM and dependency scanning must be part of build process, not a separate compliance task)</li><li><strong>Vendor relationships</strong> (transparency agreements, disclosure timelines, audit rights)</li><li><strong>Automation</strong> (manual SBOM review of 100+ vendors is unsustainable)</li><li><strong>Rapid response capability</strong> (if you detect a compromise, can you patch within 4 hours?)</li><li><strong>Forgiveness in early phases</strong> (expect to find concerning findings; prioritize by actual exposure)</li></ol>
<h2>Conclusion: Supply Chain as Strategic Risk</h2>
<p>Software supply chain security is no longer a technical compliance problem—it's a strategic risk that reflects your organization's resilience to advanced threats. Organizations treating it as such are implementing systematic approaches to visibility, detection, and response. Those waiting for regulatory mandates will be reactive when compromise occurs.</p>
<p>The 2026-2028 period will likely see significant supply chain compromises that expose organizations unprepared for rapid detection and isolation. The question is whether your organization will be a victim or demonstrating competence in crisis response.</p>]]></content:encoded>
  </item>
  <item>
    <title>When Review Breaks: Code at Machine Scale</title>
    <link>https://prasannamalode.in/articles/code-review-at-machine-scale.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/code-review-at-machine-scale.html</guid>
    <pubDate>Tue, 18 Aug 2026 00:00:00 +0000</pubDate>
    <category>AI Governance</category>
    <description>PR throughput doubled and review time per PR halved. Coding assistants uncoupled production from review capacity, and nobody sent an email.</description>
    <content:encoded><![CDATA[<p>The metric that should have worried me was the one everybody was pleased about.</p>
<p>Pull request throughput on one of our teams roughly doubled over a quarter. More code written, more code merged, velocity charts pointing in the direction velocity charts are supposed to point. The team was using coding assistants heavily and openly, and by every measure we tracked, it was working.</p>
<p>What we did not track was review time per pull request. When I went back and looked, it had fallen by about the same factor the volume had risen.</p>
<p>The reviewers had not become faster readers. They had been handed twice the work in the same number of hours, and they had adapted the way anyone adapts: by reading less carefully, approving more readily, and reserving real scrutiny for the changes that looked like they needed it.</p>
<p>Which is precisely the failure mode, because looking like it needs scrutiny is a property we are bad at judging, and AI-generated code is unusually good at not looking like it needs it.</p>
<h2>The assumptions underneath code review</h2>
<p>Code review is an old practice and it rests on premises we rarely say out loud.</p>
<p>It assumes the author understands the change. When a human writes a function, they have reasoned about it. The reviewer is checking that reasoning. When an assistant produces a function and a human accepts it, the reasoning may or may not have happened, and from the diff you cannot tell.</p>
<p>It assumes production rate and review capacity are roughly matched. Both used to be bounded by the same resource — human hours. That coupling is what made review sustainable. Assistants uncouple them. Production scales; review does not.</p>
<p>It assumes that surface quality correlates with underlying care. This is the heuristic that does the most quiet work in real reviews. Consistent naming, sensible structure, handled errors, comments in the right places — these tell you someone was paying attention, and reviewers calibrate their scrutiny accordingly.</p>
<p>Generated code has excellent surface quality. It is well formatted, idiomatically named, plausibly commented, and structurally clean, regardless of whether it is correct. The correlation the heuristic depends on is simply gone, and the heuristic does not announce that it has stopped working. It just keeps producing confident, unjustified reassurance.</p>
<h2>What actually shows up</h2>
<p>I want to be careful here, because the discourse tends toward the dramatic and my experience has been more mundane.</p>
<p>I have not seen assistants produce obviously dangerous code very often. Ask for a database query and you generally get a parameterized one. The blatant failures are increasingly rare, and the tooling gets better every release.</p>
<p>What I have seen is subtler and harder to catch.</p>
<p><strong>Context-blind correctness.</strong> Code that is right in general and wrong here — a retry that is sensible unless the operation is non-idempotent, a cache that is fine unless the data is per-tenant, an error swallowed in a way that is reasonable in a script and not in a payment path. The model does not know your system's invariants. It knows what code like this usually looks like.</p>
<p><strong>Plausible API usage that is subtly off.</strong> A library called with almost the right arguments, or with a default the author did not consider, or in a pattern deprecated two versions back. It compiles. It passes the happy path. It reads as though someone knew what they were doing.</p>
<p><strong>Volume dilution.</strong> A necessary two-line fix arriving inside a three-hundred-line change that also refactored, renamed, and reorganized, because that was easy to generate. The important lines are in there somewhere. Nobody will find them.</p>
<p><strong>Silent dependency growth.</strong> New imports appearing because the generated solution used them, without anyone deciding to take on the dependency. This one bothers me most, because it bypasses every process we have for evaluating what we pull into a build — not maliciously, just quietly, one convenient import at a time.</p>
<h2>Moving the checks to where scale is free</h2>
<p>The response cannot be to ask reviewers to try harder. That is asking humans to absorb a machine-scale increase through willpower, and it fails in a predictable direction.</p>
<p>What has to change is where the checking happens.</p>
<p><strong>Automate everything mechanically checkable, and make it blocking.</strong> Static analysis, dependency scanning, secret detection, license checking, coverage floors. These scale with volume at near-zero marginal cost. Every one of them that runs before a human opens the diff is attention returned to the human for the things only humans can do.</p>
<p><strong>Gate dependency additions explicitly.</strong> A pull request introducing a new package should require a different, deliberate decision than one that does not. This is easy to enforce and it closes the quietest gap.</p>
<p><strong>Enforce change size.</strong> Not as a style preference but as a review-integrity control. Large diffs do not get read; they get skimmed. If the tooling makes large diffs cheap to produce, the process has to make them expensive to merge.</p>
<p><strong>Direct human review at intent, not syntax.</strong> The questions that matter are the ones no linter can ask. Does this belong in our architecture? Is this the right place for this logic? What happens to this under concurrency, at scale, on failure? Does it honor the invariants this system actually has? Reviewers should spend their scarce attention there, which means everything below that line must be handled before they arrive.</p>
<p><strong>Keep authorship accountable.</strong> Whoever opens the pull request owns the code, regardless of what produced it. Not as a blame mechanism — as a clarity mechanism. "I did not write that part" has to not be an available answer, because the moment it is, the accountability model that review depends on has dissolved.</p>
<h2>The measurement I would add tomorrow</h2>
<p>If you take one thing from this, take this: track review time per pull request alongside throughput, and watch the ratio.</p>
<p>Throughput alone will tell you a happy story. It told us one for a quarter. The ratio is where you find out whether your review process is still doing its job or has quietly become a formality that produces approvals at the rate they are requested.</p>
<p>The tools are good and getting better, and I am not arguing against using them — we do, deliberately and at volume. But they changed the shape of one side of an equation that had been balanced for decades, and they did it without changing anything on the other side. Nobody sends an email when that happens. You find it in a ratio you were not looking at, in a quarter everyone thought went well.</p>]]></content:encoded>
  </item>
  <item>
    <title>The Same Binary Under Four Names</title>
    <link>https://prasannamalode.in/articles/same-binary-under-four-names.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/same-binary-under-four-names.html</guid>
    <pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate>
    <category>AI Governance</category>
    <description>When a CVE lands in rebranded platform software, &quot;who is affected&quot; becomes archaeology. Provenance has to be captured at build time.</description>
    <content:encoded><![CDATA[<p>A CVE lands on a Thursday afternoon in a library nobody thinks about — the kind that arrives in your dependency tree three levels down, through something you did choose, and has been sitting there for years doing its job quietly.</p>
<p>The first question is easy: do we ship it? Yes.</p>
<p>The second question is the one that costs you your evening: which of the things we have shipped contain it, and who has them?</p>
<p>If you build one product and ship it under one name, that question has a short answer. If you build platform software that goes out under partner branding, with per-customer feature sets, on release trains that diverged eighteen months ago, the answer is a research project — and every hour of that research project is an hour your customers are waiting for a straight answer to a simple question.</p>
<h2>What rebranding does to provenance</h2>
<p>The technical situation is ordinary. The same core, built with different configuration, packaged under different names, for different partners.</p>
<p>The record-keeping situation is where it goes wrong, because the identity of the artifact is no longer a single thing. There is what you call it internally. There is what it is called on the box. There is the version the customer sees, which may not match the version in your source control, because the partner wanted their own numbering. There is the build that went to one customer in March and the build that went to another in June, which share a base but not a patch level.</p>
<p>Ask "is customer X affected" and you are really asking a chain of questions. Which product name did they receive? Which internal build does that correspond to? Which source revision did that build come from? Which dependency versions were resolved at that build? And did anything get patched in between for that customer specifically?</p>
<p>Every one of those links is somewhere. That is the thing — the information almost always exists. It lives in a build system, a release ticket, a shipping record, a spreadsheet one person maintains. The links exist and they are not joined, so answering the question means a human walking the chain by hand, once per customer, under time pressure, while being asked for updates.</p>
<h2>The record has to be made at build time</h2>
<p>This is the lesson, and it is the same lesson as attribution in an LLM gateway, which is probably not a coincidence.</p>
<p>You cannot reconstruct provenance afterwards. You can only capture it at the moment the artifact is created, because that is the only moment when everything is simultaneously known — the source revision, the resolved dependency versions, the build configuration, the toolchain, the flags. An hour later, some of that is inference. A year later, some of it is archaeology.</p>
<p>So generate the bill of materials as a build output, not as a periodic exercise. Not a scan of the repository, which describes what the source declares, but a record of what the build actually produced and included. Those differ more often than people expect, particularly where transitive dependencies, vendored code, or build-time resolution are involved.</p>
<p>Then attach it to the artifact by hash. Not by version string, which is a label someone can reuse, and definitely not by filename, which is the thing that changes under rebranding. The hash is the only identifier that survives every renaming downstream of it, which makes it the only reliable join key you have.</p>
<h2>The join nobody owns</h2>
<p>Here is the gap I see most often, and it is organizational rather than technical.</p>
<p>Engineering knows which source revision produced which build. Sales, or operations, or whoever manages the partner relationship, knows which customer received which named release. These two facts live in different systems, maintained by different teams, with no shared key between them, and no single person whose job it is to be able to join them.</p>
<p>That join is the entire answer to "who is affected." Without it, your SBOM discipline gets you to "build 7.2.1-internal contains the vulnerable library," which is true and does not tell you who to call.</p>
<p>Fixing it does not require a platform. It requires deciding that shipment records reference build identifiers, and that build identifiers reference source and dependency state, and that somebody owns keeping that true. The technology is a foreign key. The hard part is agreeing that it matters before the Thursday when it does.</p>
<h2>The tempting shortcut</h2>
<p>There is a version of this where you decide that since it is all the same core, you can answer at the core level. Is the vulnerable library in the 7.2 line? Yes. Then everyone on 7.2 is affected. Notify everyone.</p>
<p>This feels rigorous. It is over-notification, and over-notification has a cost that is easy to underestimate.</p>
<p>Customers who receive advisories for things that do not affect them — because the vulnerable component was not compiled into their configuration, or the feature is disabled in their build — learn to filter your advisories. They allocate less attention to each one. Eventually one arrives that genuinely matters and it lands in the same mental bucket as the last six that did not.</p>
<p>Precision is not a nicety here. It is what preserves the value of the channel for the day you need it.</p>
<h2>What good looks like on a Thursday</h2>
<p>The version of this I would like to live in is unremarkable. The CVE lands. Someone queries the dependency across all build records and gets the affected build hashes in minutes. Those hashes join to shipment records and produce a customer list, with which branded name each of them knows the product by. The advisory goes out with the right product names on it, to the right people, the same day.</p>
<p>No archaeology. No spreadsheet. No engineer reconstructing a build from eighteen months ago to find out what was in it.</p>
<p>Every part of that is infrastructure you build on a quiet week, for an event you cannot schedule. Which is exactly why it tends not to get built — it has no deadline until the deadline is already past, and then it is the most obviously necessary thing you never did.</p>]]></content:encoded>
  </item>
  <item>
    <title>The Ban That Didn&#x27;t Work</title>
    <link>https://prasannamalode.in/articles/the-ban-that-didnt-work.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/the-ban-that-didnt-work.html</guid>
    <pubDate>Tue, 07 Jul 2026 00:00:00 +0000</pubDate>
    <category>AI Governance</category>
    <description>A one-sentence AI policy, a four-month DNS log, and why making the sanctioned path the fast path beat every prohibition we tried.</description>
    <content:encoded><![CDATA[<p>The policy was one sentence long and it was not unreasonable: no company code or configuration is to be entered into external AI tools.</p>
<p>It was communicated well. People acknowledged it. Nobody argued with it, because nobody disagrees with the principle in the abstract.</p>
<p>Four months later I was looking at DNS logs for an unrelated reason and found steady traffic to consumer AI domains from engineering subnets, all day, every working day. Not a spike. Not one person. A baseline.</p>
<p>My first reaction was the wrong one. I started thinking about enforcement — blocklists, egress filtering, a stern follow-up message. It took a conversation with one of the engineers to understand what I was actually looking at.</p>
<p>She was not being reckless. She had a debugging problem at four in the afternoon, a stack trace she did not recognize, and a tool that could help. The sanctioned alternative required a ticket to request access. The ticket queue was measured in days. Her problem was measured in minutes.</p>
<p>She had not decided to violate policy. She had decided to do her job, and the policy had not offered her a way to do both.</p>
<h2>Prohibition has a prerequisite</h2>
<p>Bans work when the prohibited thing is not very useful, or when the sanctioned substitute is at least as convenient. AI tools fail both conditions badly.</p>
<p>They are extremely useful for exactly the tasks engineers spend their days on — reading unfamiliar code, drafting boilerplate, explaining an error, writing the regex nobody wants to write, turning a vague requirement into a starting draft. And they are effortless to reach. No install, no procurement, no approval. A browser tab.</p>
<p>So a ban does not remove the behavior. It removes your visibility of the behavior. The usage moves from corporate laptops to personal phones, from the network you monitor to the one you don't, and from something you could have shaped into something you will now only learn about from the outside.</p>
<p>I have come to think of it as a rule: when a control makes a behavior invisible rather than absent, you have made your position worse than when you started. You have traded a governable risk for an ungovernable one and given yourself a compliance artifact that says otherwise.</p>
<h2>Make the sanctioned path the fast path</h2>
<p>The thing that actually reduced our external traffic was not a stronger prohibition. It was reducing time-to-access from days to zero.</p>
<p>Concretely: every engineer got access by default, on joining, through our internal gateway. No request, no approval, no ticket. The credential existed before they needed it.</p>
<p>That single change did more than every message I had sent. Not because the engineers suddenly cared more about data governance, but because the friction that had pushed them outward disappeared. Given two tools of similar quality, people use the one that is already there.</p>
<p>A few things made it work, and they are worth being specific about.</p>
<p><strong>The internal option has to be good.</strong> If the sanctioned tool is a lesser model behind a slow interface with a restrictive filter, people will use it for things that do not matter and go elsewhere for things that do — which inverts your intent perfectly. Budget for the good model. It is cheaper than the incident.</p>
<p><strong>Onboarding by default, not by request.</strong> Any approval step reintroduces the delay you are trying to remove. If a class of user needs gating, gate that class; do not gate everyone to manage the exception.</p>
<p><strong>Say what is logged, in plain language.</strong> People assume the worst about monitoring they have not had explained. A short, honest note about what is captured and retained produces far less avoidance than silence does. Vagueness reads as surveillance.</p>
<p><strong>Keep the guardrails about consequences, not vocabulary.</strong> Blocking on keyword lists frustrates legitimate work constantly and stops determined misuse approximately never. Scope what the credential can reach, log what was done, and put your effort into the tail risks that would genuinely hurt.</p>
<h2>What the gateway gives you that the ban never could</h2>
<p>Once the traffic is yours, the security work becomes ordinary work.</p>
<p>You have per-team attribution, so an exposed key is a scoped incident rather than a company-wide one. You have usage telemetry, which turns out to be a decent anomaly signal in its own right. You have a single place to change models, apply policy, and respond to a provider incident without touching forty applications.</p>
<p>And you have something less tangible that I did not anticipate: you find out what people are actually doing. Our usage patterns told me which teams had adopted these tools deeply and which had not touched them, which informed training, tooling decisions, and a couple of conversations about processes that were clearly painful enough that people were reaching for help.</p>
<p>None of that is available from a blocklist. A blocklist tells you a request was denied. It tells you nothing about the need that produced it, and the need does not go away when the request does.</p>
<h2>The part that is uncomfortable to say out loud</h2>
<p>Shadow AI is a symptom, and the diagnosis is usually about us.</p>
<p>People go around controls when the controls cost more than the rule is worth to them. That is not a character flaw; it is a rational response to a badly priced obstacle. When I find a widely violated policy, the most useful question is not who is violating it. It is what it costs to comply, and whether we ever measured that before we wrote the rule.</p>
<p>We had written a sentence that was easy to write and expensive to follow, and then treated the resulting gap as a discipline problem.</p>
<p>The fix was not a better sentence. It was making compliance the path of least resistance, so that doing the right thing stopped requiring anyone to be heroic about it at four in the afternoon with a stack trace they did not recognize.</p>]]></content:encoded>
  </item>
  <item>
    <title>The AI-Driven Threat Landscape: How Machine Learning is Weaponizing Attacks in 2026</title>
    <link>https://prasannamalode.in/articles/ai-driven-threat-landscape.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/ai-driven-threat-landscape.html</guid>
    <pubDate>Tue, 30 Jun 2026 00:00:00 +0000</pubDate>
    <category>Cybersecurity</category>
    <description>AI has moved from defensive tool to offensive weapon: polymorphic malware, adaptive phishing, zero-day discovery at scale. Why behavioural detection and containment now beat signatures.</description>
    <content:encoded><![CDATA[<p>The cybersecurity landscape of 2026 is defined by a fundamental shift: artificial intelligence has moved from a defensive tool to an offensive weapon. Nation-states, criminal organizations, and advanced threat actors are deploying machine learning to automate vulnerability discovery, personalize social engineering campaigns, and orchestrate attacks at scale. Organizations that continue to rely on static defenses face existential risk.</p>
<h2>The Turning Point: From Defense to Weaponization</h2>
<p>For years, security teams celebrated AI as a great equalizer—anomaly detection systems that could process millions of events, threat intelligence platforms that correlated indicators across networks, automated response playbooks that reduced MTTR. By 2026, this narrative has inverted.</p>
<p>Adversaries have inverted the equation. They're using AI to:</p>
<ul><li><strong>Generate polymorphic malware</strong> that evolves in real-time, defeating signature-based detection</li><li><strong>Automate social engineering</strong> through hyper-personalized phishing campaigns that adapt based on target responses</li><li><strong>Discover zero-days at scale</strong> by training models on publicly disclosed vulnerability patterns and extrapolating to unpatched software</li><li><strong>Predict defensive measures</strong> by modeling security team behavior and adjusting attack timing and vectors accordingly</li></ul>
<h3>Real-World Manifestations</h3>
<p><strong>Polymorphic Malware Campaigns:</strong> In Q2 2026, CISA reported a 340% increase in malware variants detected in the wild compared to Q2 2025. Manual signature creation became obsolete; malware now regenerates bytecode between infections to evade pattern matching.</p>
<p><strong>Personalized Spear Phishing:</strong> Threat actors using large language models trained on corporate data—scraped from LinkedIn, GitHub, company SEC filings, and data breaches—created phishing emails that mimicked internal communication patterns with 73% accuracy. Early detection became nearly impossible until behavioral signals shifted.</p>
<h2>The Defense Imperative: Behavioral Over Signature-Based</h2>
<p>Traditional pattern matching cannot scale against AI-generated threats. The industry's response has fragmented into two camps:</p>
<h3>Camp 1: Behavioral and Contextual Defense</h3>
<p>Organizations investing in behavioral analysis report 4.2x better detection rates for AI-generated attacks. These systems don't ask "Is this a known bad hash?" but rather "Does this sequence of actions match the expected operational baseline for this user/device?"</p>
<p><strong>Implementation Reality:</strong></p>
<ul><li>Requires 12-18 months of baseline collection before reliable alerting</li><li>Demands integration across endpoint, network, and identity layers</li><li>Necessitates ML expertise in-house or outsourced to managed security providers</li></ul>
<h3>Camp 2: Containment-First Architecture</h3>
<p>A secondary trend emphasizes architectural resilience over detection. Zero Trust, micro-segmentation, and assume-breach models shift the burden from "catch the attack" to "limit blast radius."</p>
<p>Organizations following this path report 68% faster incident containment but higher operational complexity.</p>
<h2>The Human Element: Still Critical, Now Vulnerable</h2>
<p>AI-powered social engineering has exposed a critical vulnerability in human-centric security programs. Awareness training remains static (annual modules, phishing simulations with obvious tells), while adversary campaigns adapt in real-time.</p>
<p><strong>New Approaches Gaining Traction:</strong></p>
<ul><li>Continuous, AI-powered micro-training tailored to individual user susceptibility</li><li>Real-time intervention (blocking suspicious communications before user decision)</li><li>Organizational culture shifts toward "verify always" in communication flows</li></ul>
<h2>The Supply Chain Cascade</h2>
<p>AI-driven vulnerability discovery has weaponized supply chain risk. Threat actors use ML to identify vulnerable library versions across open-source ecosystems, then craft SolarWinds-style compromises targeting downstream consumers.</p>
<p><strong>Defense Requires:</strong></p>
<ul><li>Shift from "approved vendor lists" to continuous monitoring of third-party security posture</li><li>Automated SBOM (Software Bill of Materials) analysis with ML-powered risk scoring</li><li>Incident response playbooks specifically for supply chain scenarios</li></ul>
<h2>Organizational Readiness: The Gap Widens</h2>
<p>CISO surveys from H1 2026 reveal a widening capability gap:</p>
<ul><li>68% of enterprises lack ML expertise in their security teams</li><li>54% cannot distinguish between benign AI-generated content and malicious AI-generated attacks</li><li>Only 12% have adapted incident response procedures for AI-driven threats</li></ul>
<p>Organizations that hired security ML engineers in 2024-2025 are significantly better positioned than those starting recruitment in 2026.</p>
<h2>Future Outlook: Arms Race Acceleration</h2>
<p>By 2027, expect:</p>
<ul><li><strong>Predictive attack modeling</strong> where adversaries simulate organizational responses before execution</li><li><strong>Autonomous swarm attacks</strong> where coordinated malware instances communicate and adapt without central command</li><li><strong>AI-vs-AI arms races</strong> where defensive systems and attack systems iterate faster than human oversight can manage</li></ul>
<h2>Critical Priorities for 2026-2027</h2>
<ol><li><strong>Baseline behavioral profiles</strong> for your critical assets now—do not wait for incidents to teach you normalcy</li><li><strong>Hire or contract ML security expertise</strong> immediately; the talent shortage will worsen</li><li><strong>Revisit incident response for AI-driven scenarios</strong>; your current playbooks assume human-speed attack progression</li><li><strong>Invest in architectural resilience</strong> if you cannot achieve behavioral detection maturity in 18 months</li><li><strong>Monitor your supply chain continuously</strong>; vendor questionnaires are obsolete</li></ol>
<h2>Conclusion</h2>
<p>The 2026 cybersecurity environment is not a problem to be "solved" but rather an arms race to be managed. Organizations that recognize AI as both a defensive and offensive weapon will build systems that degrade gracefully under AI-powered assault. Those that treat cybersecurity as a static compliance problem will become cautionary tales by 2027.</p>]]></content:encoded>
  </item>
  <item>
    <title>Prompt Injection Is Not a New Class of Vulnerability</title>
    <link>https://prasannamalode.in/articles/prompt-injection-confused-deputy.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/prompt-injection-confused-deputy.html</guid>
    <pubDate>Tue, 16 Jun 2026 00:00:00 +0000</pubDate>
    <category>AI Governance</category>
    <description>Prompt injection is a confused deputy problem we have known about since 1988. Threat model the tools, not the prompt.</description>
    <content:encoded><![CDATA[<p>Someone built a helpful thing. That is how these always start.</p>
<p>It was an internal assistant that answered questions about our documentation. Ask it how a process worked, and it would search the wiki, read the relevant pages, and answer. Useful, popular, built quickly by a good engineer in the spirit of making everyone's life easier.</p>
<p>Then, to make it more useful, it was given the ability to read from a ticketing system too. Also sensible. Most of the questions people asked involved work in flight.</p>
<p>The demonstration that ended the conversation took about a minute. A ticket was created containing, in its description field, a paragraph of ordinary-looking text that concluded with an instruction addressed to the assistant. When the next person asked a question that caused the assistant to search tickets, it read that paragraph, and it did what the paragraph said.</p>
<p>Nobody had exploited a model. Someone had typed text into a field that the system treated as instructions.</p>
<h2>Where the security industry went briefly strange</h2>
<p>For a while the discussion around prompt injection had an air of novelty about it, as though we had encountered a genuinely new phenomenon requiring genuinely new theory. Papers, taxonomies, a small industry of detection products.</p>
<p>I want to argue for a more boring framing, because I think the boring framing is the one that actually leads to fixes.</p>
<p>Prompt injection is a confused deputy problem. It has a name because we have known about it since 1988. A program with legitimate authority is induced by a less-privileged party to exercise that authority on their behalf. SQL injection is this. Cross-site request forgery is this. Server-side request forgery is this. The LLM variant is this.</p>
<p>What makes it feel new is that the boundary between instruction and data, which in SQL we could at least in principle draw with parameterized queries, cannot be drawn cleanly here. The model consumes one stream of tokens. There is no prepared statement. There is no escaping function that reliably works, and I would be cautious about anyone selling you one.</p>
<p>But "we cannot fix it at the parser" does not mean "we cannot fix it." It means we fix it where we have always fixed problems we could not fix at the parser: at the boundary of what the deputy is permitted to do.</p>
<h2>The question that actually matters</h2>
<p>Stop asking whether the model can be tricked. Assume it can. Every serious attempt to prevent that through instruction-hardening has been defeated, usually quickly, usually by someone doing it for fun.</p>
<p>Ask instead: when it is tricked, what can happen?</p>
<p>If the answer is "it writes a wrong answer to a user's screen," you have a quality problem. Annoying, not urgent.</p>
<p>If the answer is "it calls an API that modifies a record," you have a vulnerability, and the severity is defined entirely by what that API can reach.</p>
<p>If the answer is "it reads from a privileged source and then writes to somewhere the attacker can observe," you have data exfiltration, and the model is not the interesting part of the chain. The interesting part is that you connected a read capability and a write capability through a component that will follow instructions from anywhere.</p>
<p>That last one is the pattern to watch for, and it is why the retrieval-plus-tools architecture is where most of the real damage lives. Retrieval brings untrusted content into the context. Tools give the context consequences. Either alone is mostly fine. The combination is where you need to be deliberate.</p>
<h2>Controls that are already in your toolkit</h2>
<p>None of what follows will be unfamiliar. That is the point.</p>
<p><strong>Least privilege, applied to the tool layer.</strong> The credentials the assistant uses should be its own, scoped to exactly the actions it needs. If it only answers questions, it gets read access and nothing else. I have seen assistants running with a service account inherited from an existing integration, carrying permissions that made sense for a nightly batch job and no sense at all for a component that reads user input.</p>
<p><strong>User-context propagation.</strong> If the assistant acts for a person, it should act with that person's authority, not with a superset of everyone's. Otherwise it becomes an access-control bypass with a chat interface, and the first person to notice will not report it.</p>
<p><strong>Human confirmation on state change.</strong> Reads are recoverable. Writes, sends, deletes, and payments are not. A confirmation step on consequential actions is unglamorous and defeats the entire exfiltration-by-tool-call category, because the loop passes through a person who did not ask for that.</p>
<p><strong>Egress control on the output path.</strong> If the assistant can render images from arbitrary URLs, or follow links, or call a webhook with data it has assembled, it has a channel out. Treat that with the same suspicion you would treat any outbound connection from a component processing untrusted input.</p>
<p><strong>Provenance in the context.</strong> You cannot reliably teach a model to disregard instructions in retrieved content. You can mark which parts of the context came from trusted configuration and which came from a wiki page anyone can edit, and you can make policy decisions about what actions are permitted when untrusted content is in play. This is defense in depth, not a boundary. It reduces the severity distribution; it does not eliminate the tail.</p>
<h2>The part I would put on a slide</h2>
<p>Threat model the tools, not the prompt.</p>
<p>Write down every capability the system has. For each one, ask what an attacker would do with it if they controlled the model's output completely, because in the worst case they do. Whatever survives that exercise is your real security posture. Whatever you were planning to achieve with careful system prompt wording is not a control; it is a preference, and it will be honored until someone decides otherwise.</p>
<p>The engineer who built our assistant was not careless. He had built a document search tool that happened to be conversational, and the day it gained a second data source it silently became a different kind of system — one where content written by one user could steer behavior on behalf of another. That transition happened in a pull request that looked like a feature addition, because it was one.</p>
<p>The failure was not in the model. It was that we had no review step that noticed the architecture had changed. That is a process gap, and process gaps are something our field already knows how to close. We just have to recognize the problem as ours.</p>]]></content:encoded>
  </item>
  <item>
    <title>The Queue Is Still There</title>
    <link>https://prasannamalode.in/articles/the-queue-is-still-there.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/the-queue-is-still-there.html</guid>
    <pubDate>Tue, 26 May 2026 00:00:00 +0000</pubDate>
    <category>AI Governance</category>
    <description>We put a model in front of the alert queue and it wrote beautiful summaries. The queue did not get shorter. Here is what did move the number.</description>
    <content:encoded><![CDATA[<p>We put a language model in front of our alert queue, and for about two weeks everyone was pleased.</p>
<p>Every alert now arrived with a paragraph at the top explaining what it was. Plain English, decent quality, genuinely well written. The analysts said it was nicer to read. Leadership liked the demo. I liked that we had shipped something.</p>
<p>Then I sat with one of the analysts for a shift and watched what actually happened.</p>
<p>She read the summary. Then she opened the raw alert anyway. Then she pivoted into the SIEM to check the source host. Then she looked up whether that host had fired the same thing before. Then she closed it as a false positive, the same way she would have without the summary, having spent an extra six seconds reading a paragraph that told her nothing she could act on.</p>
<p>We had not reduced her workload. We had added a pleasant preamble to it.</p>
<h2>The number that does not move</h2>
<p>Here is the uncomfortable arithmetic of alert triage. If a hundred alerts arrive and ninety-four of them are noise, the work is not understanding the hundred. The work is getting to the six.</p>
<p>Summarization does not touch that. It makes each of the hundred marginally more pleasant to process, and a hundred slightly-nicer units of work is still a hundred units of work. The queue depth is identical. The time-to-the-six is roughly identical. You have improved the experience of doing the wrong amount of work.</p>
<p>This is the trap, and it is seductive precisely because summarization is the thing language models are unmistakably good at. It demos beautifully. It never looks broken. Nobody complains. It is the local maximum you can reach in a sprint, and reaching it feels like progress in a way that makes it hard to ask whether it was the right direction.</p>
<h2>What moved the number</h2>
<p>When we rebuilt the thing, we changed the question. Not "how do we explain this alert" but "what does the analyst do next, and can we have already done it?"</p>
<p>That reframing produced a different system.</p>
<p><strong>Enrichment before summary.</strong> Every one of those manual pivots — is this host in scope, has this signature fired before, who owns this asset, what changed on it recently, is the source address one of ours — is a lookup with a deterministic answer. Do them all, automatically, in parallel, before a human ever sees the alert. Most of the analyst's clock was not comprehension. It was fetching context that a script could fetch.</p>
<p><strong>Correlation before presentation.</strong> Forty alerts from one misconfigured host are one piece of information presented forty times. Grouping them is not glamorous and it is not AI, but it did more for queue depth than anything else we built. Do this first. If your queue is full of duplicates, no amount of model cleverness matters, because you are summarizing the same thing forty times.</p>
<p><strong>A recommendation with its reasoning attached.</strong> This is where the model earned its place. Not "here is what this alert says," which the analyst can read, but "this looks like the false positive class you closed eleven times last month, for these reasons, and here is the evidence." That is a claim she can accept or reject in two seconds. A summary gives her something to read. A recommendation gives her something to decide.</p>
<p><strong>A confidence signal that is allowed to say nothing.</strong> The single most valuable behavior we built was the system declining to guess. When enrichment came back thin or contradictory, it said so and handed over cleanly, with the raw data and no narrative. A tool that is confidently wrong twice will not be trusted the third time it is right, and trust is the whole asset. Once analysts stop reading the recommendation, you own an expensive queue decoration.</p>
<h2>Automation earned, not assumed</h2>
<p>The obvious next question is whether the model should just close the low-risk ones itself.</p>
<p>Eventually, for narrow classes, yes. But not on day one, and not on the basis of a vendor's accuracy claim.</p>
<p>What worked for us was shadow mode. The system made its recommendation. The analyst made the real decision. We compared them, by alert class, for long enough to have meaningful numbers. Where agreement was very high and the disagreements were consistently the model being appropriately more cautious, we promoted that specific class to auto-closure — with sampling, so a percentage still got human eyes, and with a monthly review of what the sampling found.</p>
<p>Everything else stayed advisory. Some of it still is.</p>
<p>This is slower than the deployment everyone wants. It is also the only version I have seen survive contact with a real incident review, because when someone eventually asks "why did nobody look at this," the answer needs to be a documented decision with evidence behind it, not a procurement decision with a confidence score behind it.</p>
<h2>What I would ask before building this</h2>
<p>If you are about to point a model at your alert queue, I would ask three questions first, in this order.</p>
<p>What fraction of your queue is duplicate or correlated? If it is large, fix that before anything else. It is cheaper, it is deterministic, and it will deliver more than the model will.</p>
<p>Which lookups do your analysts perform on nearly every alert? Automate those. They are the actual minutes. You do not need a model for most of them.</p>
<p>And then, only then: where is the genuine judgment call, and can a model make it well enough that a human can verify the answer faster than they could produce it?</p>
<p>That last clause is the whole test. A tool that makes a judgment I must independently redo has cost me time. A tool that makes a judgment I can check at a glance has saved me most of it.</p>
<h2>The honest summary</h2>
<p>The summaries were good. They were well written, accurate, and useless, and the reason they were useless had nothing to do with model quality. We had automated the part of the job that was not the job.</p>
<p>The queue is not hard because the alerts are hard to read. It is hard because there are too many of them and the important ones are indistinguishable from the rest until you do the work. Anything that does not reduce the count, or does not do the work, is decoration — however well it writes.</p>]]></content:encoded>
  </item>
  <item>
    <title>The Invoice as an Intrusion Detection System</title>
    <link>https://prasannamalode.in/articles/invoice-as-intrusion-detection.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/invoice-as-intrusion-detection.html</guid>
    <pubDate>Tue, 05 May 2026 00:00:00 +0000</pubDate>
    <category>AI Governance</category>
    <description>Token spend is unforgeable, baselined for free, and denominated in a unit finance already tracks. Read it as a detection signal.</description>
    <content:encoded><![CDATA[<p>The first sign was not an alert. It was a graph with a step in it.</p>
<p>Saturday, some time after two in the morning, token consumption on one of our internal keys went from a flat weekend baseline to roughly forty times that, and stayed there. No error spike. No failed authentications. No unusual source geography that anyone had configured us to care about. Just a service that normally spent almost nothing overnight, suddenly working very hard at something.</p>
<p>Nobody was awake to see it. We found it on Monday, in a cost dashboard, because someone was doing budget planning.</p>
<p>It turned out to be a runaway retry loop — a misbehaving job that failed, retried, failed, retried, with no backoff and no ceiling. Expensive, embarrassing, not malicious. But the thing that stayed with me was not the root cause. It was that the same graph, with the same shape, is what a stolen API key looks like.</p>
<h2>Spend is behavior, rendered in a unit finance already tracks</h2>
<p>We spend a lot of effort building telemetry for security. We ship logs, we normalize them, we write detections, we tune them. Meanwhile, sitting in a completely different part of the organization, there is a continuously updated, high-fidelity record of exactly how much work every credential has caused something to do.</p>
<p>Token consumption is a behavioral signal. It has properties that most of our detections would love to have.</p>
<p>It is unforgeable by the client. The consuming application does not get to report its own usage; the gateway counts it. Someone abusing a key cannot make the counter say something friendlier.</p>
<p>It is baselined almost for free. Most workloads are boringly regular. A summarization service handles roughly the number of documents that arrive. A coding assistant tracks the working hours of the people using it. Variance exists, but the shape of a week is stable enough that departures are visible without sophisticated modeling.</p>
<p>And it is denominated in a unit that non-security people already care about. This matters more than it sounds. You will never have to justify why you are monitoring cost.</p>
<h2>The shapes worth naming</h2>
<p>Not every anomaly is interesting. These are the ones I have learned to look at.</p>
<p><strong>The step function.</strong> Flat, then high, then flat at the new level. Automation, either yours misbehaving or someone else's. The overnight version is the one to care about, because legitimate step changes usually accompany a deploy, and deploys usually happen when people are awake. Correlate against your release calendar before you page anyone.</p>
<p><strong>The slow ramp.</strong> Consumption climbing a few percent a day for weeks. Almost never an attack. Almost always a product decision nobody costed — a feature quietly expanded to more users, or a prompt that grew a few hundred tokens of context each time someone improved it. Worth catching anyway, because this is how a budget dies without a single alarming day.</p>
<p><strong>The pattern break.</strong> Total volume unchanged, but the mix is different. A key that has only ever called a small fast model starts calling the largest one. A service that has always sent short requests starts sending very long ones. This one is the most interesting from a security standpoint, because it is the least likely to be accidental. Software does not spontaneously develop new preferences. Something changed the caller, or something else is using the credential.</p>
<p><strong>The weekend floor that lifts.</strong> Every environment has a natural quiet period. When the floor of that quiet period rises and stays risen, something now runs continuously that did not before.</p>
<h2>Why this catches things your other controls miss</h2>
<p>A leaked LLM gateway key is an awkward artifact. It is not a database credential, so your data-access monitoring does not see it. It usually does not touch your network perimeter in an interesting way, because the traffic is ordinary HTTPS to an endpoint you have deliberately allowed. It rarely triggers authentication alerting, because the authentication succeeds — that is the entire point of a stolen credential.</p>
<p>What it cannot hide is that it costs money to use. Whoever has it is using it for something, and using it is the only thing that makes it worth having.</p>
<p>That is an unusually good property for a detection to rest on. You are not watching for a signature that can be changed or an anomaly that can be smoothed out. You are watching the one behavior that is inseparable from the motive.</p>
<h2>Making it actually work</h2>
<p>Three practical notes, because the idea is easy and the implementation is where it goes quiet.</p>
<p>First, this requires the attribution I have written about elsewhere. A single shared key gives you one aggregate line, and aggregates hide everything. Forty times normal on one team's key is obvious. The same absolute increase spread across company-wide totals is a rounding error you will never notice.</p>
<p>Second, alert on rate of change, not on absolute thresholds. Absolute limits are a budget control, and they should exist, but they fire when you have already spent the money. A derivative-based alert — this key is consuming at N times its trailing baseline — fires while it is happening. Set the baseline per key, not globally, because a key serving a batch pipeline and a key serving an internal chat tool have nothing in common.</p>
<p>Third, route it somewhere a human reads at odd hours, or accept that you will find these on Monday. I am not going to pretend every organization should page for this. But decide consciously. Our Saturday event ran for roughly thirty hours because the answer to "who sees this" was nobody, and nobody had ever been asked.</p>
<h2>The wider point</h2>
<p>There is a habit in security of assuming that a good signal must come from a security tool. It leads to us instrumenting elaborately for detections that a billing system was already providing for free.</p>
<p>Your gateway is producing cost telemetry whether you look at it or not. Someone in finance is already looking at it, once a month, for entirely different reasons. The distance between their dashboard and a working detection is smaller than the distance between where you are and most of the detections on your roadmap.</p>
<p>Go and read that graph. You may find, as we did, that the interesting event happened weeks ago and has been sitting there patiently, correctly recorded, waiting for someone to ask what it was.</p>]]></content:encoded>
  </item>
  <item>
    <title>Shift Left Has a Ceiling</title>
    <link>https://prasannamalode.in/articles/shift-left-has-a-ceiling.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/shift-left-has-a-ceiling.html</guid>
    <pubDate>Tue, 28 Apr 2026 00:00:00 +0000</pubDate>
    <category>DevSecOps</category>
    <description>Some defects do not exist until production exists. The classes above the ceiling, the shift-right toolkit, and escape rate, the metric that connects both halves of the pipeline.</description>
    <content:encoded><![CDATA[<p>Three posts arguing for moving controls earlier, and now the correction: there is a floor of defects that no amount of pre-deployment work will ever reach, and pretending otherwise is how organisations end up with an immaculate pipeline and a breach.</p>
<p>Some bugs do not exist until production exists. Not "are hard to find before production" — <em>do not exist</em>. They are properties of a running system under real load with real data and real concurrent users, and a static analyser examining source code is looking at the wrong object entirely.</p>
<h2>The classes that live above the ceiling</h2>
<p>Four families, roughly.</p>
<p><strong>Load-dependent defects.</strong> The race condition that needs two requests to arrive within the same few milliseconds. The connection pool that only deadlocks at a concurrency level your staging environment never reaches. The rate limiter that is correct per-instance and useless across twelve instances behind a load balancer. Your test suite runs these paths serially. Production does not.</p>
<p><strong>Data-dependent defects.</strong> The authorisation check that is correct for every tenant in your seed data and wrong for the one enterprise customer with a nested org structure. The regex that is fine on ASCII and catastrophically backtracks on the input a real user pasted. You cannot fixture your way to this, because the defining property of the triggering data is that nobody imagined it.</p>
<p><strong>Configuration drift.</strong> Your Terraform says the bucket is private. The bucket is not private, because someone fixed an incident at 2am eight months ago through the console and never came back. Static analysis reads what you <em>declared</em>. Only the running environment knows what is <em>true</em>, and the gap between those two grows continuously.</p>
<p><strong>Emergent cross-service behaviour.</strong> Service A trusts requests from service B because they are both inside the perimeter. Service B was, in the meantime, given a new public endpoint. Neither repository contains the vulnerability. It exists only in the composition, and no analysis scoped to a single repo can see it.</p>
<p>That last one is worth sitting with. In an architecture with a few dozen services, a growing share of your real risk lives in the <em>relationships</em>, and relationships have no source file.</p>
<h2>Where risk actually gets created</h2>
<p>There is a second reason the ceiling exists, and it is about what changes.</p>
<p>A large fraction of production incidents are triggered by change — and a large fraction of those changes are configuration, not code. Feature flag flips, scaling parameters, IAM policy edits, DNS records, rate limit adjustments. These are changes that alter system behaviour in production and frequently never pass through the pipeline you spent three years hardening.</p>
<p>If your entire control surface is the pull request, you are not covering the surface where a large share of your risk is introduced.</p>
<h2>The shift-right toolkit</h2>
<p>The answer is not to abandon prevention. It is to stop treating deployment as the end of the security process.</p>
<p><strong>Progressive delivery.</strong> If a change reaches 1% of traffic before 100%, then a defect that only appears under real traffic appears with a 99% smaller blast radius, and it appears while someone is still watching. Canaries and feature flags are security controls, even though they are almost never budgeted as such.</p>
<p><strong>Runtime authorisation telemetry.</strong> Log every authorisation decision — subject, action, resource, verdict. Then alert on the shapes that should never occur: one principal accessing an anomalous number of distinct objects, a service account making a call it has never made before, a deny rate spiking on one endpoint. This is the only control that reliably catches IDOR-class flaws, because the signature of IDOR is not in the code, it is in the access pattern.</p>
<p><strong>Drift detection.</strong> Run the plan continuously, not just at apply time:</p>
<pre><code class="language-yaml"># .github/workflows/drift.yml
name: Infrastructure drift detection

on:
  schedule:
    - cron: &quot;0 */6 * * *&quot;
  workflow_dispatch:

jobs:
  detect:
    runs-on: ubuntu-latest
    permissions:
      id-token: write        # OIDC — no long-lived credentials
      contents: read
      issues: write
    steps:
      - uses: actions/checkout@v4
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::111122223333:role/drift-detector
          aws-region: eu-west-1
      - uses: hashicorp/setup-terraform@v3
      - run: terraform init -input=false
      - id: plan
        run: terraform plan -detailed-exitcode -lock=false -no-color
        continue-on-error: true
      - name: Raise drift issue
        if: steps.plan.outputs.exitcode == 2
        uses: actions/github-script@v7
        with:
          script: |
            github.rest.issues.create({
              owner: context.repo.owner,
              repo: context.repo.repo,
              title: `Infrastructure drift detected — ${new Date().toISOString().slice(0,10)}`,
              labels: [&#x27;drift&#x27;, &#x27;security&#x27;],
              body: &#x27;Live infrastructure no longer matches committed state. See run logs.&#x27;
            })</code></pre>
<p>Note the <code>id-token: write</code> and the absence of any stored AWS key. That is the pattern from part one — remove the secret rather than scan for it.</p>
<p><strong>Adversarial testing against the real thing.</strong> Bug bounty and pentest are not compliance line items; they are the only controls that test the composed system the way an attacker would. The finding that comes back from a good bounty programme is almost always one that no scanner could have produced, because it depends on chaining three things across two services.</p>
<h2>Escape rate: the metric that connects both halves</h2>
<p>Here is the number that makes this whole series cohere.</p>
<blockquote><p><strong>Vulnerability escape rate</strong> = defects first discovered in production ÷ total defects discovered</p></blockquote>
<p>Read it as a diagnostic, not a score:</p>
<ul><li><strong>Escape rate falling, total findings steady</strong> — prevention is working. This is the good state.</li><li><strong>Escape rate rising, total findings falling</strong> — the alarming state. It usually means your left-side tooling has been tuned into silence, and you are discovering things in production instead. Part one's fix rate metric will confirm it.</li><li><strong>Escape rate near zero</strong> — do not celebrate. It almost always means you are not looking in production, not that nothing is there. Zero escapes is a detection failure, not a prevention triumph.</li></ul>
<p>The operational discipline this creates: <strong>every production finding becomes a specification for a left-side control.</strong> For each escape, ask what would have caught it, and where it should have lived. Sometimes the answer is a new SAST rule. Sometimes it is a type-level guarantee from part two. Sometimes the honest answer is "nothing could have caught this before deploy", and that answer is valuable too, because it tells you to invest in detection rather than in another scanner.</p>
<p>Pair escape rate with <strong>security MTTR</strong> — time from discovery to deployed fix. Escape rate tells you how much gets through. MTTR tells you whether that matters. A team with a moderate escape rate and a four-hour MTTR is in far better shape than one with a low escape rate and a six-week deploy cycle, because the second team cannot respond to the zero-day that arrives on a Tuesday.</p>
<p>That framing lands well with engineering leadership, incidentally, because it is the same argument as DORA. Deployment frequency and lead time are security metrics. The ability to ship a fix in an hour is a security capability. This is the point where security stops asking for a slower pipeline and starts arguing for a faster one.</p>
<h2>What the whole series was arguing</h2>
<p>Four posts, one thesis:</p>
<ol><li><strong>Moving unfiltered machine output earlier is not shifting left.</strong> It is interrupting people sooner. Measure fix rate, give every stage a latency budget, and prefer secure defaults over blocking gates.</li><li><strong>Only class-level work compounds.</strong> Every instance you fix by hand returns. Climb the ladder: findable by machine, then loud on failure, then impossible to express.</li><li><strong>Acceptance is fine; silence is not.</strong> Give every accepted risk an owner, an expiry, and a CI job that enforces both. Renewal is a legitimate outcome. Forgetting is not.</li><li><strong>And there is a ceiling.</strong> Some defects exist only in the running system. Detect them deliberately, measure the escape rate, and feed every escape back into the left side as a new control.</li></ol>
<p>The version of this that fails is the one where "shift left" means buying tools and pushing the resulting work onto developers while calling it empowerment. The version that works treats prevention and detection as one loop with a measured leak rate, where the security team owns the noise it generates and the loop teaches itself.</p>
<p>Start with the fix rate. It is one number, you can compute it this week, and it will tell you immediately whether the rest of this applies to you.</p>]]></content:encoded>
  </item>
  <item>
    <title>Who Spent This?</title>
    <link>https://prasannamalode.in/articles/who-spent-this.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/who-spent-this.html</guid>
    <pubDate>Tue, 14 Apr 2026 00:00:00 +0000</pubDate>
    <category>AI Governance</category>
    <description>Attribution in an LLM gateway is not a reporting problem. It is a key-issuance decision, and it has to be made before the first request.</description>
    <content:encoded><![CDATA[<p>There is a particular kind of silence that follows a simple question in a meeting.</p>
<p>Ours came in month three of running an internal LLM gateway. Someone from finance had the usage report open and asked what should have been a trivial question: which teams were driving the spend? Engineering wanted to know before they asked for more budget. Finance wanted to know before they approved it.</p>
<p>The report had one number. A big one, technically accurate, and completely useless. Every request had gone through the same API key.</p>
<p>We had done the hard part well. We had stood up a proxy, centralized model access, kept raw provider keys out of application code, and given ourselves a single place to enforce rate limits. On any architecture diagram it looked correct. The one thing we had not done was decide, before the first request, who each request belonged to.</p>
<h2>Attribution is a design decision</h2>
<p>This is the part I got wrong, and I think a lot of teams get wrong in the same way: attribution feels like a reporting problem. It presents as a dashboard gap. It seems like the sort of thing you fix later by adding a column.</p>
<p>It is not a reporting problem. It is a key issuance problem, and key issuance happens at the beginning.</p>
<p>Whatever your gateway is — a commercial product, an open-source proxy, something you wrote yourself — it will let you mint virtual keys that sit in front of the real provider credentials. The question is what those keys map to. One key for the whole company. One per application. One per team. One per user.</p>
<p>Every answer downstream of that choice is fixed by it. Your usage data can only ever be as granular as your keys. If one key serves forty engineers, no amount of clever querying recovers which engineer did what. The information was never captured. You are not missing a report; you are missing a fact.</p>
<p>And the retrofit is genuinely painful, because it is not a schema change. It is a credential rotation across every integration you have built, coordinated with every team that depends on them, while the old key stays alive long enough that nobody's pipeline breaks. I have watched that project consume weeks that a fifteen-minute decision at the start would have saved entirely.</p>
<h2>Why security ends up caring more than finance</h2>
<p>The budget conversation is what forces the issue. It is not the reason the issue matters.</p>
<p>Consider what happens when a key leaks. Someone commits it, or pastes it into a support ticket, or it ends up baked into a container image that gets pushed somewhere public. You find out — maybe from a spend anomaly, maybe from a scanner, maybe from someone outside the company being kind enough to tell you.</p>
<p>Now you are in incident response. The first questions are always the same. What could this credential reach? How long was it exposed? What was done with it? Who else is affected if we kill it right now?</p>
<p>With one shared key, every one of those questions returns the worst possible answer. It could reach everything. It was exposed since whenever it was created. What was done with it is indistinguishable from normal traffic, because normal traffic looks exactly the same. And revoking it takes the entire company offline, which means you will feel pressure to wait, investigate more, be careful — while the credential stays live.</p>
<p>With per-team or per-user keys, the same incident is a different event. Blast radius is scoped by construction. The traffic on that one key is a narrow, readable slice. Revocation affects one team, who you can notify in a Slack message. You go from an organizational emergency to a Tuesday.</p>
<p>That difference was not created during the incident. It was created months earlier, by someone deciding how to issue keys.</p>
<h2>The uncomfortable part about per-user</h2>
<p>Per-user attribution is the most useful and the most contentious, so it deserves an honest paragraph rather than a recommendation.</p>
<p>Per-user keys give you the cleanest audit trail available. They also mean you are now holding a per-person record of interactions with a language model, and people use these tools to think out loud. Some of what they type is half-formed, or personal, or about problems they have not told anyone about yet. Depending on where your employees sit, that record may carry obligations you have not thought about.</p>
<p>I do not think there is a universal right answer. What I think is wrong is arriving at an answer by accident. Decide it deliberately, write down what you retain and for how long, tell people plainly what is logged, and be able to explain why. A team that knows the shape of the logging will work with it. A team that discovers it later will route around it, and then you have shadow AI on top of everything else.</p>
<p>If per-user is too heavy for your context, per-team is a defensible middle. It is coarse enough to avoid individual surveillance and fine enough that incident scoping and cost allocation both work. What is not defensible is one key for everyone, which is the option that feels like no decision at all.</p>
<h2>What I would tell myself at the start</h2>
<p>Before the first request goes through the gateway, answer three things.</p>
<p>What is the smallest unit you need to attribute to? Pick it based on the incident you would least like to handle, not the report you would most like to see.</p>
<p>Who issues keys, and how does someone get one? If the answer involves a ticket that takes four days, you have just designed your shadow AI problem. Make the sanctioned path the fast path.</p>
<p>What happens when a key has to die? Write the runbook while nothing is on fire. It should be short. If it is long, your key structure is wrong.</p>
<p>None of this is sophisticated. It is the same least-privilege reasoning we have applied to service accounts for twenty years, pointed at a new kind of credential. The only reason it gets skipped is that the gateway works fine without it, right up until the moment it doesn't.</p>
<p>The number in that report was correct. That was the problem. It was one correct number where we needed forty, and the only moment we could have had forty was before we started.</p>]]></content:encoded>
  </item>
  <item>
    <title>Ransomware in 2026: Evolution Beyond Encryption to Business Disruption</title>
    <link>https://prasannamalode.in/articles/ransomware-beyond-encryption.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/ransomware-beyond-encryption.html</guid>
    <pubDate>Tue, 24 Mar 2026 00:00:00 +0000</pubDate>
    <category>Cybersecurity</category>
    <description>Ransomware is now a professionalised extortion business with double extortion in 84% of cases. The four attack phases, the four defensive tiers, and the nine readiness questions.</description>
    <content:encoded><![CDATA[<p>Ransomware has evolved beyond encryption-based attacks into a multi-faceted extortion business. By 2026, the ransomware ecosystem includes: encryption-based payload delivery, data theft and public exposure ("double extortion"), operational technology disruption, and DDoS amplification. The business model has matured to the point where ransomware groups operate like software vendors—with service tiers, technical support, and transparent pricing. Organizations treating ransomware as a technical problem rather than a business resilience challenge are significantly underestimating their actual risk.</p>
<h2>The Ransomware Market: A Business Analysis</h2>
<h3>Market Size and Economics</h3>
<ul><li>Reported ransomware payments in 2025: $1.1 billion globally (actual total likely 2-3x higher due to unreported payments)</li><li>Average ransom payment: $450,000 (up from $200,000 in 2023)</li><li>Ransom payment ranges: $50,000 for small businesses to $60M+ for Fortune 500 companies</li><li>Recovery cost (downtime, remediation, legal): 3-5x the ransom payment</li></ul>
<h3>The Operational Model</h3>
<p>Modern ransomware groups operate as distributed criminal enterprises:</p>
<ul><li><strong>Development team:</strong> Builds and maintains encryption, exfiltration, and command-and-control infrastructure</li><li><strong>Affiliate program:</strong> Recruits experienced intruders to identify targets and deliver initial payload</li><li><strong>Negotiation team:</strong> Professional communicators who handle ransom negotiations</li><li><strong>Technical support:</strong> Assists victims with decryption if ransom is paid</li><li><strong>Public relations:</strong> Maintains leak site, manages reputation, provides transparency on victim statistics</li></ul>
<p>This operational maturity is new in 2026 and reflects the professionalization of ransomware as an industry.</p>
<h2>Attack Evolution: Four Phases of Modern Ransomware</h2>
<h3>Phase 1: Initial Access (Weeks 1-4)</h3>
<p><strong>Common vectors (in order of frequency):</strong></p>
<ol><li>Compromised credentials (phishing, credential stuffing, vendor compromise)</li><li>Exploited vulnerabilities (unpatched internet-facing applications)</li><li>Supply chain compromise (compromised software vendor or managed service provider)</li><li>Lateral movement from other compromised systems</li></ol>
<p><strong>Duration:</strong> Most initial access happens within 2 weeks. Organizations without rapid detection are already in Phase 2 by the time they discover compromise.</p>
<p><strong>Defense:</strong> Rapid detection of suspicious credentials or unknown access is the critical control. Organizations achieving &lt;24 hour detection for suspicious access report 65% lower probability of ransomware escalation.</p>
<h3>Phase 2: Persistence and Lateral Movement (Weeks 2-8)</h3>
<p>Once inside the network, attackers establish persistence and move toward high-value targets.</p>
<p><strong>Attacker objectives:</strong></p>
<ul><li>Install persistent backdoor (enabling re-entry if initial compromise is discovered)</li><li>Identify high-privilege accounts (domain admin, backup system access)</li><li>Locate critical data and backup systems</li><li>Map network topology to understand isolation points</li></ul>
<p><strong>Duration:</strong> This phase averages 4-6 weeks in 2026, down from 8-12 weeks in 2023. Attackers are becoming more efficient.</p>
<p><strong>Defense:</strong> Network segmentation, endpoint detection and response (EDR), and privilege access management (PAM) are critical. Organizations with mature endpoint monitoring report detection within 10-15 days of initial access.</p>
<h3>Phase 3: Pre-Encryption Reconnaissance and Data Exfiltration (Weeks 6-10)</h3>
<p>Attackers don't encrypt immediately. They first:</p>
<ul><li><strong>Identify high-value data</strong> (customer databases, intellectual property, financial records)</li><li><strong>Exfiltrate critical data</strong> (to enable "double extortion" ransom demand)</li><li><strong>Locate and compromise backup systems</strong> (to prevent recovery without payment)</li><li><strong>Identify operational technology systems</strong> (target for maximum disruption)</li></ul>
<p><strong>Data exfiltration volumes:</strong> Attackers now extract 100GB-10TB per target, up from 1-50GB in 2023.</p>
<p><strong>Timeline:</strong> This phase reveals the attackers' sophistication. Careful exfiltration takes 2-4 weeks; hasty exfiltration happens in days.</p>
<p><strong>Defense:</strong> Data loss prevention (DLP) systems, network egress monitoring, and behavioral analytics on file access are critical. Many organizations detect exfiltration only after encryption begins (too late).</p>
<h3>Phase 4: Encryption and Extortion (Days 1-30 Post-Encryption)</h3>
<p>Once exfiltration is complete, attackers encrypt all accessible systems and present ransom demand with three components:</p>
<ol><li><strong>Encryption ransom:</strong> "Pay X to receive decryption key"</li><li><strong>Data suppression ransom:</strong> "Pay Y to prevent publication of exfiltrated data" (double extortion)</li><li><strong>Timeline pressure:</strong> "Decide within 7 days or price increases" or "We will publish data on [date]"</li></ol>
<p><strong>Ransom amounts breakdown (2026 data):</strong></p>
<ul><li>Encryption-only ransom: 40% of total ask</li><li>Data suppression (double extortion) ransom: 40% of total ask</li><li>Negotiation buffer: 20% of total ask (attackers know victims will negotiate)</li></ul>
<h2>Double Extortion: The Game-Changer</h2>
<p>By 2026, approximately 84% of ransomware incidents include data theft and extortion alongside encryption. This fundamentally changes risk calculus.</p>
<h3>Traditional Encryption-Only Scenario</h3>
<ul><li>Victim: "We have backups. We'll just restore."</li><li>Attacker: Loses leverage</li><li>Outcome: Victim may not pay</li></ul>
<h3>Double Extortion Scenario</h3>
<ul><li>Victim: "We have backups. We'll just restore."</li><li>Attacker: "We'll publish 2TB of customer PII on our leak site and notify customers, regulators, and media."</li><li>Victim: Now faces regulatory fines, customer notification costs, reputational damage</li><li>Outcome: Victim pays even if they can restore from backup</li></ul>
<p><strong>Real-world cost impact:</strong> Organizations paying double extortion ransom spend an average of 60 days in "ransom negotiation" compared to 2-3 days for encryption-only attacks.</p>
<h2>Defensive Evolution: Organizations' Responses</h2>
<h3>Traditional Defense (Outdated by 2026)</h3>
<ul><li>Maintain backups offline</li><li>Incident response plan for encryption recovery</li><li>Employee awareness training for phishing</li></ul>
<p><strong>Reality check:</strong> These work for encryption-only attacks. They're insufficient for double extortion, lateral movement, and backup compromise attacks.</p>
<h3>Modern Defense (2026 Standard)</h3>
<p><strong>Tier 1: Prevent Initial Access</strong></p>
<ul><li>MFA enforcement (reduces credential compromise)</li><li>Vulnerability scanning and patching (addresses exploitable internet-facing assets)</li><li>Threat intelligence on supply chain vendors (warns of vendor compromise)</li><li>Secure email gateway (reduces phishing success rate)</li></ul>
<p><strong>Effectiveness:</strong> 20-30% reduction in initial compromise attempts</p>
<p><strong>Tier 2: Detect Compromise Early</strong></p>
<ul><li>Endpoint Detection and Response (EDR) with behavioral alerting</li><li>Network monitoring for suspicious lateral movement</li><li>Darkweb monitoring for stolen credentials (early warning of compromise)</li><li>Continuous privilege assessment (detect when privileged accounts are overused)</li></ul>
<p><strong>Effectiveness:</strong> Reduces average time to detection from 220 days (industry average) to 10-15 days</p>
<p><strong>Tier 3: Limit Damage If Compromised</strong></p>
<ul><li>Network micro-segmentation (contain lateral movement)</li><li>Data-centric security (classify and protect high-value data)</li><li>Backup immutability (backups that cannot be deleted by attackers)</li><li>Incident response runbooks specific to ransomware scenarios</li></ul>
<p><strong>Effectiveness:</strong> 70-85% reduction in systems encrypted if compromise is detected in Phase 2-3</p>
<p><strong>Tier 4: Recover Quickly</strong></p>
<ul><li>Disaster recovery infrastructure (pre-built recovery environment)</li><li>Tested recovery procedures (60-90 day recovery validated quarterly)</li><li>Communication infrastructure independent of primary network (ability to reach stakeholders if network is encrypted)</li></ul>
<p><strong>Effectiveness:</strong> Reduces recovery time from weeks/months to hours/days</p>
<h2>The Attack/Defense Arms Race: 2026 Specifics</h2>
<h3>Attacker Evolution</h3>
<ul><li><strong>Living off the land:</strong> Using legitimate system tools (PowerShell, Windows Admin Center) to avoid EDR detection</li><li><strong>Encryption-free attacks:</strong> Exfiltrating data without encrypting systems (forcing victim to pay without proof of capability)</li><li><strong>Operational technology targeting:</strong> Expanding from IT networks to manufacturing, utilities, healthcare OT systems</li><li><strong>Supply chain targeting:</strong> Compromising managed service providers and cloud integrators to target entire customer bases</li></ul>
<h3>Defender Response</h3>
<ul><li><strong>Behavioral analytics:</strong> AI-powered detection of suspicious patterns even when using legitimate tools</li><li><strong>Zero Trust for sensitive data:</strong> Data protection independent of network trust</li><li><strong>OT-specific monitoring:</strong> Behavioral monitoring of operational technology systems</li><li><strong>Supplier security:</strong> Continuous monitoring of vendors' security posture and incident response</li></ul>
<h2>Ransom Payment: The Uncomfortable Reality</h2>
<h3>To Pay or Not to Pay</h3>
<p><strong>Arguments for paying (what victims report):</strong></p>
<ul><li>Decryption key actually works (attackers maintain reputation for technical capability)</li><li>Data suppression is real (identified victims from leaked data lists)</li><li>Recovery timeline: paying may enable decryption faster than rebuild</li><li>Insurance covers most of ransom cost</li></ul>
<p><strong>Arguments against paying:</strong></p>
<ul><li>Funds criminal enterprise</li><li>No guarantee all data is deleted after payment</li><li>Encourages future ransomware development</li><li>May violate sanctions laws (if ransom group is connected to sanctioned nation-states)</li></ul>
<p><strong>2026 Reality:</strong> Approximately 67% of organizations that can afford ransom payment choose to pay. Organizations with mature backups and incident response are more likely to refuse payment.</p>
<h3>Insurance Dynamics</h3>
<p>Cyber insurance plays a complex role:</p>
<ul><li>Insurers require specific security controls to underwrite ransom risk</li><li>Insurers negotiate with attackers on behalf of insured organizations</li><li>Insurers increasingly require proof of incident response capability before paying claims</li><li>Premium increases (30-50% post-incident) impact long-term cost calculus</li></ul>
<h2>Organizational Readiness Assessment: The Critical Questions</h2>
<h3>Prevention Readiness</h3>
<ol><li>Can you detect and lock down a compromised credential within 4 hours? (Most: 2-3 days)</li><li>Do you have 100% MFA coverage on critical systems? (Most: 60-80%)</li><li>Is your vulnerability patching cycle &lt;30 days for critical assets? (Most: 30-90 days)</li></ol>
<h3>Detection Readiness</h3>
<ol><li>Do you have EDR deployed to 100% of endpoints? (Most: 70-85%)</li><li>Can you detect suspicious lateral movement in real-time? (Most: 48+ hours delay)</li><li>Do you monitor for exfiltration of sensitive data? (Most: No)</li></ol>
<h3>Response Readiness</h3>
<ol><li>Do you have backup systems that attackers cannot access? (Most: Partial)</li><li>Can you bring systems online from backup within 72 hours? (Most: 1-2 weeks)</li><li>Do you have pre-authorized incident response contacts? (Most: Ad-hoc)</li></ol>
<p><strong>Assessment:</strong> Organizations scoring "yes" to 7+ questions are in the top 10% of readiness. Most organizations score 3-4.</p>
<h2>The Realistic Roadmap: Ransomware Resilience by 2026-2027</h2>
<h3>Months 1-3: Quick Wins</h3>
<ul><li>Enforce MFA on all critical systems and cloud applications</li><li>Deploy EDR to 100% of endpoints (even if not all features are immediately used)</li><li>Conduct credentials review and revoke dormant accounts</li><li>Develop incident response runbook specifically for ransomware</li></ul>
<h3>Months 3-6: Detection Capabilities</h3>
<ul><li>Enable EDR behavioral alerting</li><li>Implement network egress monitoring for suspicious data transfer</li><li>Deploy DLP to monitor sensitive data access</li><li>Subscribe to darkweb monitoring for organization's stolen credentials</li></ul>
<h3>Months 6-12: Resilience Building</h3>
<ul><li>Implement network segmentation for critical systems</li><li>Test backup recovery procedures monthly</li><li>Establish backup immutability (snapshots that cannot be deleted)</li><li>Assign incident response roles and conduct tabletop exercises</li></ul>
<h3>Months 12-18: Maturity</h3>
<ul><li>Implement behavioral analytics for privileged account abuse</li><li>Conduct red team exercises simulating ransomware attack</li><li>Establish metrics for detection and response timing</li><li>Document lessons learned and improve procedures</li></ul>
<h2>Financial Reality: Cost-Benefit of Defense</h2>
<h3>Ransomware Defense Investment</h3>
<ul><li>EDR + SIEM + incident response capability: $500K-2M annually (depending on organization size)</li><li>Staff training and testing: $100K-500K</li><li>Backup infrastructure improvements: $200K-1M</li><li><strong>Total annual investment:</strong> $1M-3.5M for mid-sized enterprise</li></ul>
<h3>Cost of Attack (Without Defense)</h3>
<ul><li>Encryption ransom demand: $500K-10M (most pay 40-60% of ask)</li><li>Data suppression (double extortion): $200K-5M</li><li>Recovery costs (downtime, staff labor): $500K-5M</li><li>Regulatory fines and notification: $100K-50M+ (depending on industry and data exposure)</li><li>Reputational damage and lost revenue: Highly variable</li><li><strong>Total cost:</strong> $2M-65M+ per incident</li></ul>
<h3>ROI on Defense</h3>
<p>A single prevented or quickly-contained incident pays for multi-year defense investment.</p>
<h2>Conclusion: Ransomware is an Inevitability</h2>
<p>By 2026, every organization should assume it will experience at least one ransomware incident during the next 5 years. The question is not "if" but "when" and "how prepared are we?"</p>
<p>Organizations investing in layered defense (prevention, detection, response, recovery) are significantly reducing impact. Those treating ransomware as a compliance checkbox are setting themselves up for catastrophic incidents.</p>
<p>Ransomware has professionalized as a business. Organizations must professionalize their defense in response.</p>]]></content:encoded>
  </item>
  <item>
    <title>How I increased release frequency by 80% without sacrificing security</title>
    <link>https://prasannamalode.in/articles/cicd-80-percent.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/cicd-80-percent.html</guid>
    <pubDate>Tue, 10 Mar 2026 00:00:00 +0000</pubDate>
    <category>DevSecOps</category>
    <description>The CI/CD transformation framework, quality gates, and governance model that delivered measurable results for a global engineering organisation.</description>
    <content:encoded><![CDATA[<p>When I took over the release pipeline, releases were slow, unpredictable, and stressful. Engineers dreaded deployment days. Six years later we'd increased release frequency by 80% and cut deployment failures by 60%. Here's the honest story of how we got there.</p>
<h2>The starting point</h2>
<p>The first thing I did was measure. Not guess. Measure. We tracked lead time for change, deployment frequency, failure rate, and time to restore. The numbers were uncomfortable. That was the point.</p>
<blockquote><p>You can't improve what you don't measure. But more importantly, you can't get buy-in without showing the business what slow really costs.</p></blockquote>
<h2>Pre-QAT quality gates</h2>
<p>The biggest unlock was introducing Pre-QAT quality gates: automated checks before code even reached the test environment. This shifted quality left and eliminated entire categories of late-stage failures.</p>
<ul><li>Static code analysis via Coverity at commit time</li><li>FOSS compliance scanning on every build</li><li>Automated security checks (NMAP, OpenVAS) in staging</li><li>Memory leak detection integrated into the CI pipeline</li></ul>
<h2>The release governance layer</h2>
<p>Speed without control creates chaos. We built a release governance framework that gave every stakeholder visibility without creating bottlenecks. Executive dashboards showing MTTR, SLA compliance, and incident trends meant I never had to prep a slide deck for a C-suite review. The data spoke for itself.</p>
<h2>Tools that drove the change</h2>
<p>The stack that made it possible: Jenkins, Python, shell scripting, Coverity for static analysis, Artifactory and Nexus for artifact management, Docker and Kubernetes for containerised deployments, all running on Linux. None of these tools were exotic. The discipline around how we used them was what mattered.</p>
<h2>What actually moved the needle</h2>
<p>Honestly? Culture. The tools were 30% of the improvement. The other 70% was getting the team to own quality rather than throw it over the wall to QA. That took time, consistent reinforcement, and being willing to slow down briefly to accelerate long-term.</p>
<p>The principle I kept coming back to: make the right thing the easy thing. When quality gates caught issues automatically, engineers stopped seeing them as obstacles and started seeing them as safety nets.</p>]]></content:encoded>
  </item>
  <item>
    <title>The Vulnerability I Accepted, and the Expiry Date I Put on It</title>
    <link>https://prasannamalode.in/articles/risk-acceptance-with-an-expiry-date.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/risk-acceptance-with-an-expiry-date.html</guid>
    <pubDate>Tue, 24 Feb 2026 00:00:00 +0000</pubDate>
    <category>DevSecOps</category>
    <description>A 1,900-line suppression file is four years of defensible calls adding up to something indefensible. Risk acceptance as a record with an owner, an expiry, and a CI job that enforces both.</description>
    <content:encoded><![CDATA[<p>I once opened a suppression file that was 1,900 lines long.</p>
<p>It had been started four years earlier. It had eleven distinct comment styles in it, which is roughly one per engineer who had ever touched it. Some entries had a ticket number; most of those tickets were closed, and a few referenced a tracker the company no longer used. One entry said <code># temporary - remove after migration</code>. The migration it referred to had completed two years before.</p>
<p>Nobody had done anything wrong. Every single line had been added by a reasonable person under deadline pressure making a defensible call. The file was the sum of four years of defensible calls, and the sum was indefensible.</p>
<p>That is the real problem with risk acceptance. Not that individual decisions are bad. That they never end.</p>
<h2>Suppression is not a decision, it is the absence of one</h2>
<p>Watch what happens when a finding is suppressed in a typical setup.</p>
<p>Somebody adds a line. Maybe a comment. The scanner goes quiet. The PR goes green. And then — this is the important part — <em>nothing else ever happens</em>. No record of who decided, no record of why, no date, no review, no signal if the surrounding conditions change.</p>
<p>A decision that produces no artefact and triggers no future event is not a decision. It is deferral wearing a decision's clothes. And deferral accumulates silently, which means the cost lands on somebody who was not in the room, often years later, usually during an incident.</p>
<p>The fix is not to forbid acceptance. Acceptance is legitimate and necessary — plenty of findings genuinely are not worth fixing this quarter, and a security function that cannot say "fine, not now" is a security function that gets routed around.</p>
<p>The fix is to make acceptance <strong>structured, owned, and temporary</strong>.</p>
<h2>What an acceptance record has to contain</h2>
<p>Five fields. If any one is missing, it is not an acceptance, it is a suppression.</p>
<ol><li><strong>Owner</strong> — a named person, not a team. Teams do not get paged, do not leave the company, and do not feel accountable.</li><li><strong>Expiry</strong> — a date. Not "when we migrate". Not "next quarter". A date that a computer can compare against <code>now()</code>.</li><li><strong>Rationale</strong> — why this is tolerable <em>now</em>. Written for a stranger, because in eighteen months the reader will be one.</li><li><strong>Compensating control</strong> — what reduces the impact in the meantime. WAF rule, network boundary, feature flag, monitoring alert. If the answer is honestly "nothing", write "nothing" — that is valuable information.</li><li><strong>Trigger conditions</strong> — what would make this unacceptable before the expiry date. "If this service starts handling payment data." "If a public exploit lands." This is the field everyone skips and it is the one that catches the real disasters.</li></ol>
<p>Here is what that looks like as a file that lives in the repo:</p>
<pre><code class="language-yaml"># security/acceptances.yaml
- id: ACC-2026-007
  finding: CVE-2025-XXXXX
  component: github.com/example/imageproc
  severity: high

  owner: priya.raman           # a person
  accepted_on: 2026-01-12
  expires_on: 2026-04-12       # 90 days for high

  rationale: &gt;
    Vulnerable code path is the TIFF decoder. We only accept PNG and
    JPEG at the API boundary, enforced by content-type allowlist in
    gateway/filters.go:88. Upstream fix requires a major version bump
    that breaks our colour-profile handling.

  compensating_controls:
    - Content-type allowlist at the gateway (PNG, JPEG only)
    - Decoder runs in a seccomp-restricted sandbox
    - Alert on any non-allowlisted content-type reaching the decoder

  invalidating_conditions:
    - The gateway allowlist is widened to any additional format
    - A public exploit is observed in the wild
    - imageproc is called from any new service

  renewals: 0</code></pre>
<p>Two properties matter here. It is in version control, so it is reviewable and blameable like code. And it is machine-readable, so the expiry can be enforced rather than hoped for.</p>
<h2>Enforcement is what makes it real</h2>
<p>Every part of this is theatre without a job that fails the build when the date passes.</p>
<pre><code class="language-python">#!/usr/bin/env python3
&quot;&quot;&quot;Fail CI when any risk acceptance has expired or is incomplete.&quot;&quot;&quot;
import sys, datetime, yaml

REQUIRED = [&quot;owner&quot;, &quot;expires_on&quot;, &quot;rationale&quot;,
            &quot;compensating_controls&quot;, &quot;invalidating_conditions&quot;]
MAX_TERM = {&quot;critical&quot;: 30, &quot;high&quot;: 90, &quot;medium&quot;: 180, &quot;low&quot;: 365}
MAX_RENEWALS = 2

def main(path=&quot;security/acceptances.yaml&quot;):
    today = datetime.date.today()
    records = yaml.safe_load(open(path)) or []
    errors, warnings = [], []

    for r in records:
        rid = r.get(&quot;id&quot;, &quot;&lt;no id&gt;&quot;)

        missing = [f for f in REQUIRED if not r.get(f)]
        if missing:
            errors.append(f&quot;{rid}: missing required fields: {&#x27;, &#x27;.join(missing)}&quot;)
            continue

        expires = r[&quot;expires_on&quot;]
        accepted = r[&quot;accepted_on&quot;]
        term = (expires - accepted).days
        cap = MAX_TERM.get(r.get(&quot;severity&quot;, &quot;low&quot;), 365)
        if term &gt; cap:
            errors.append(f&quot;{rid}: term of {term}d exceeds {cap}d cap for {r[&#x27;severity&#x27;]}&quot;)

        if r.get(&quot;renewals&quot;, 0) &gt; MAX_RENEWALS:
            errors.append(f&quot;{rid}: renewed {r[&#x27;renewals&#x27;]} times — escalate, don&#x27;t renew&quot;)

        days_left = (expires - today).days
        if days_left &lt; 0:
            errors.append(f&quot;{rid}: EXPIRED {abs(days_left)}d ago — owner {r[&#x27;owner&#x27;]}&quot;)
        elif days_left &lt;= 14:
            warnings.append(f&quot;{rid}: expires in {days_left}d — owner {r[&#x27;owner&#x27;]}&quot;)

    for w in warnings:
        print(f&quot;::warning::{w}&quot;)
    for e in errors:
        print(f&quot;::error::{e}&quot;)

    return 1 if errors else 0

if __name__ == &quot;__main__&quot;:
    sys.exit(main())</code></pre>
<p>Wire it into CI so it runs on every push and on a schedule:</p>
<pre><code class="language-yaml"># .github/workflows/acceptance-audit.yml
name: Risk acceptance audit

on:
  pull_request:
  push:
    branches: [main]
  schedule:
    - cron: &quot;0 9 * * 1&quot;   # Monday morning, so expiries surface before they bite

jobs:
  audit:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: &quot;3.12&quot;
      - run: pip install pyyaml
      - run: python security/check_acceptances.py</code></pre>
<p>The scheduled run is the part that does the work. It means an expiry arrives as a Monday-morning notification with fourteen days of warning, rather than as a red build on an unrelated PR from someone who has no idea what the record is about.</p>
<h2>Renewal is allowed. Silence is not.</h2>
<p>People push back here, and the pushback is always the same: "so we just renew everything and nothing changes."</p>
<p>Two answers.</p>
<p>First, even pure renewal is an improvement, because renewal is <em>an event</em>. A human reads the rationale and decides again. Sometimes the rationale no longer holds — the compensating control was removed in a refactor, the service now handles data it did not before, an exploit was published. Without the expiry, none of those changes would ever be noticed. The whole value of the mechanism is that it forces a re-read.</p>
<p>Second, cap the renewals. Two is a reasonable limit. After that the item does not get renewed; it gets escalated to someone with budget authority, who either funds the fix or accepts it at a level where the accountability is real. That is not a punishment — it is the correct routing. An issue that has survived three review cycles without being fixed is not a triage problem, it is a resourcing problem, and it belongs with whoever controls resourcing.</p>
<p>Suggested terms, which you should tune to your own tolerance:</p>
<div class="table-wrap"><table><thead><tr><th>Severity</th><th>Max term</th><th>Renewals before escalation</th></tr></thead><tbody><tr><td>Critical</td><td>30 days</td><td>1</td></tr><tr><td>High</td><td>90 days</td><td>2</td></tr><tr><td>Medium</td><td>180 days</td><td>2</td></tr><tr><td>Low</td><td>365 days</td><td>2</td></tr></tbody></table></div>
<h2>Migrating a file you inherited</h2>
<p>If you already have the 1,900-line file, do not try to adjudicate all of it. You will not finish, and the attempt will burn whatever goodwill you were going to spend on the actual fix.</p>
<p>What works:</p>
<ol><li><strong>Freeze it.</strong> Rename it <code>legacy-suppressions</code>. Make CI reject any new entries. From today, every new acceptance uses the structured format. This alone stops the bleeding, and it is a one-day change.</li><li><strong>Sample it.</strong> Take 25 random entries and adjudicate them properly. You now have an estimate of what fraction of the file is stale, which is the number you take to leadership. In my experience the answer is well over half.</li><li><strong>Expire it in tranches.</strong> Assign expiry dates to the legacy entries spread over the next twelve months, weighted by severity. Do not dump them all on one date.</li><li><strong>Let deletion be the default.</strong> For legacy entries where no owner can be found, the default action at expiry is to remove the suppression and see what breaks. Half the time the finding is gone because the code is gone.</li></ol>
<p>Step four sounds reckless and is not. An unowned suppression is an unowned risk. Surfacing it is strictly better than continuing to pretend somebody is watching it.</p>
<h2>The part that is actually about culture</h2>
<p>The reason suppression files grow is that suppressing is frictionless and fixing is not. The mechanism above does not add friction to suppression — it takes about the same time to write the YAML as to add the comment. What it adds is <strong>a future obligation</strong>, and it makes that obligation visible at the moment of the decision.</p>
<p>That changes behaviour, in my experience, more than any approval workflow. When you write <code>expires_on: 2026-04-12</code> and your own name next to it, you are making a small promise to a specific future version of yourself. Most engineers, given that framing, will fix the easy ones immediately rather than schedule a conversation with themselves in ninety days.</p>
<p>The 1,900-line file was not a discipline failure. It was a design failure: a system where the cheapest path produced no record and no future event. Change the design, and most of the discipline problem evaporates.</p>]]></content:encoded>
  </item>
  <item>
    <title>Building a cybersecurity function from scratch: lessons from zero to ISO 27001</title>
    <link>https://prasannamalode.in/articles/iso-27001-from-scratch.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/iso-27001-from-scratch.html</guid>
    <pubDate>Tue, 10 Feb 2026 00:00:00 +0000</pubDate>
    <category>Cybersecurity</category>
    <description>From no InfoSec function to zero major audit findings. The people, process, and tools decisions that got us there.</description>
    <content:encoded><![CDATA[<p>When I took on cybersecurity, there was no InfoSec function. No DLP. No formal controls. Just a growing risk surface and a team that knew something needed to change.</p>
<h2>First 90 days: visibility before action</h2>
<p>I didn't start by buying tools. I started by understanding what we had: all assets, all access points, all data flows. You can't protect what you can't see. We ran Nessus scans, mapped our external attack surface with NMAP, and built an asset inventory from scratch.</p>
<blockquote><p>Every security investment should be traceable to a specific risk. If you can't name the risk, don't buy the tool.</p></blockquote>
<h2>DLP and enterprise controls</h2>
<p>Data Loss Prevention was the first major implementation. We mapped our sensitive data classifications, identified egress points, and implemented controls across email, endpoint, and cloud. The key was getting HR and Legal involved early so controls were enforceable, not just technical.</p>
<h2>The ISO 27001 journey</h2>
<p>We achieved ISO 27001 certification with zero major audit findings. The secret wasn't documentation. It was making the controls real. Auditors can tell immediately when a control exists on paper but not in practice.</p>
<ul><li>Incident response procedures were rehearsed, not just written</li><li>Access reviews happened quarterly, not annually</li><li>Security awareness was embedded in onboarding and weekly team meetings, not a once-a-year checkbox</li></ul>
<h2>The number that mattered</h2>
<p>Security incidents dropped by 50%. More importantly, when incidents did occur, we detected and contained them faster. The MTTI (Mean Time to Identify) improvement was as significant as the incident reduction itself. Detection speed is often undervalued in security programs. It shouldn't be.</p>]]></content:encoded>
  </item>
  <item>
    <title>Cloud Security in 2026: The Multi-Cloud Fragmentation Problem</title>
    <link>https://prasannamalode.in/articles/multi-cloud-security-fragmentation.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/multi-cloud-security-fragmentation.html</guid>
    <pubDate>Tue, 27 Jan 2026 00:00:00 +0000</pubDate>
    <category>Cybersecurity</category>
    <description>The same control means three different implementations across AWS, Azure and GCP. Why policy drift is the top audit finding, and what outcome-based governance looks like.</description>
    <content:encoded><![CDATA[<p>By 2026, the cloud security landscape has shifted from "securing cloud migrations" to managing security across hybrid and multi-cloud environments where no single cloud provider dominates. Organizations operating across AWS, Azure, GCP (and sometimes smaller platforms like DigitalOcean, Heroku) face a fragmentation problem: identical security requirements must be implemented using completely different APIs, interfaces, and operational procedures. This creates a critical vulnerability: security governance that works in one cloud often becomes inoperable in another.</p>
<h2>The Multi-Cloud Reality by Numbers</h2>
<h3>Adoption Patterns</h3>
<ul><li>89% of enterprises use at least two cloud providers</li><li>62% use three or more</li><li>Average enterprise manages security across 4.3 distinct cloud platforms</li></ul>
<h3>Why Multi-Cloud?</h3>
<ul><li><strong>Vendor lock-in avoidance</strong> (strategic decision to maintain optionality)</li><li><strong>Workload-specific optimization</strong> (different workloads fit different cloud economics)</li><li><strong>Mergers and acquisitions</strong> (different business units operated on different clouds pre-merger)</li><li><strong>Compliance requirements</strong> (specific workloads must run on specific cloud in specific regions)</li></ul>
<h3>The Cost</h3>
<ul><li>Security operations teams report 3.2x higher operational burden managing multi-cloud vs. single-cloud</li><li>Policy drift (same policy interpreted differently across clouds) is the #1 cause of audit findings</li><li>Incident response timelines increase 40-60% in multi-cloud scenarios due to context switching</li></ul>
<h2>The Core Problem: Cloud Provider Heterogeneity</h2>
<h3>Identity and Access Management</h3>
<ul><li><strong>AWS:</strong> IAM with policies, roles, and resource-based access controls</li><li><strong>Azure:</strong> RBAC with role assignments and managed identities</li><li><strong>GCP:</strong> IAM with roles, custom roles, and service accounts</li></ul>
<p>Same requirement ("Developer can deploy but not delete"): three completely different implementations.</p>
<p><strong>Real consequence:</strong> Organizations often implement overly permissive policies across all clouds rather than manage cloud-specific complexity. Security teams report this as single biggest driver of excessive permissions.</p>
<h3>Network Security</h3>
<ul><li><strong>AWS:</strong> Security Groups + NACLs + VPC Flow Logs</li><li><strong>Azure:</strong> NSGs + UDRs + Network Watcher</li><li><strong>GCP:</strong> Firewall Rules + VPC Service Controls</li></ul>
<p>Micro-segmentation policy that's straightforward in one cloud becomes a tangled web of rules in another.</p>
<h3>Data Protection</h3>
<ul><li><strong>AWS:</strong> KMS, S3 bucket policies, encryption at rest/in transit</li><li><strong>Azure:</strong> Key Vault, Storage encryption, Purview for classification</li><li><strong>GCP:</strong> Cloud KMS, Cloud Armor, Data Loss Prevention</li></ul>
<p>Data classification and encryption policies require cloud-specific implementation, creating opportunities for inconsistent enforcement.</p>
<h3>Threat Detection</h3>
<ul><li><strong>AWS:</strong> GuardDuty, Security Hub, CloudTrail</li><li><strong>Azure:</strong> Defender for Cloud, Sentinel, Activity Logs</li><li><strong>GCP:</strong> Chronicle (if purchased), VPC Flow Logs, Cloud Audit Logs</li></ul>
<p>Unified security monitoring across clouds requires either: (a) export everything to a central SIEM at massive cost, or (b) accept cloud-native silos.</p>
<h2>The Operational Consequences: Three Critical Failure Modes</h2>
<h3>Failure Mode 1: Policy Drift and Inconsistent Enforcement</h3>
<p><strong>Scenario:</strong> You implement a policy "All data must be encrypted at rest."</p>
<p>In AWS, your security team codifies this in AWS Config rules and enforces compliance. In Azure, someone documents this in a runbook. In GCP, no one knows if this is enforced because no one documented the requirement.</p>
<p>Result: 18 months later, audit discovers unencrypted data in GCP. Argument ensues about whether this was a "known exception" or oversight.</p>
<p><strong>Frequency:</strong> 61% of multi-cloud organizations report finding unexpected policy gaps during audits.</p>
<h3>Failure Mode 2: Incident Response Blind Spots</h3>
<p><strong>Scenario:</strong> Security team detects unusual data access in one cloud, initiates incident response.</p>
<p>AWS: CloudTrail shows the exact API calls, source IP, IAM principal, resource accessed. Response is structured. GCP: Cloud Audit Logs show similar information, but in different format, with different fields. Azure: Activity Logs can show activity, but native alerting requires tuning for each cloud's unique behaviors.</p>
<p>The 20-minute delay in getting consistent data across clouds can be the difference between containment and exfiltration.</p>
<p><strong>Impact:</strong> Organizations with mature single-cloud security incident response take 60-90 minutes to contain. Multi-cloud organizations average 140-180 minutes.</p>
<h3>Failure Mode 3: Compliance Radioactivity</h3>
<p><strong>Scenario:</strong> Your organization is audited against a compliance requirement (HIPAA, SOC 2, PCI-DSS).</p>
<p>Auditor asks: "Show me your access control policies."</p>
<p>You produce AWS documentation. Auditor validates. You produce Azure documentation. Auditor asks why it's different. You explain it's cloud-specific. Auditor notes this as "policy inconsistency" finding. You produce GCP documentation. GCP wasn't even in scope, so auditor asks why security posture isn't consistent.</p>
<p>Result: Findings that are really about "cloud provided different mechanisms" are documented as security control gaps.</p>
<h2>The Vendor Response: The CSPM Explosion</h2>
<p><strong>Cloud Security Posture Management (CSPM)</strong> platforms emerged in 2022-2024 to address this problem. By 2026, the category includes 40+ vendors, each claiming unified multi-cloud security.</p>
<h3>What CSPM Tools Do (Well)</h3>
<ul><li>Inventory resources across clouds (EC2 instances, Azure VMs, GCP instances in one interface)</li><li>Scan configuration against benchmarks (CIS Controls, NIST, PCI-DSS)</li><li>Detect misconfigurations (publicly exposed storage, excessive IAM permissions, unencrypted data)</li><li>Generate compliance reports</li></ul>
<h3>What CSPM Tools Don't Do (Yet)</h3>
<ul><li>Enforce policies in real-time across clouds (most require manual remediation or cloud-native automation)</li><li>Unify identity and access management (CSPM tools see IAM but don't enforce consistent policy)</li><li>Provide real-time threat detection (most are asset-inventory systems with scanning, not security operations platforms)</li><li>Replace cloud-native security tools (you still need AWS GuardDuty, Azure Defender, GCP Defender; CSPM sits on top)</li></ul>
<h3>The Operational Reality</h3>
<p>CSPM tools are useful for compliance reporting and finding obvious misconfigurations. But organizations report:</p>
<ul><li><strong>Alert fatigue:</strong> CSPM tools generate 1000s of findings per day. Triaging is manual and time-intensive.</li><li><strong>Remediation friction:</strong> Finding that AWS Config has a policy that Azure doesn't have equivalent of—now what?</li><li><strong>Cost unpredictability:</strong> Some CSPM tools charge per resource, making multi-cloud cost prohibitive at scale</li></ul>
<h2>Emerging Pattern: Cloud-Native Specialization</h2>
<p>Rather than trying to manage all clouds uniformly, some security-mature organizations are adopting:</p>
<p><strong>Pattern:</strong> Assign security subject matter experts to each cloud (1-2 deep AWS experts, 1-2 deep Azure experts, 1 deep GCP expert) and accept that security operations won't look identical across clouds. Instead, enforce consistent outcomes:</p>
<ul><li>Same encryption strength (even if mechanisms differ)</li><li>Same audit logging precision (even if cloud providers have different native logging)</li><li>Same incident response time (even if procedures differ by cloud)</li></ul>
<p><strong>Results:</strong> Better security outcomes than attempting forced uniformity, but requires deep cloud expertise.</p>
<h2>The Practical Roadmap: Multi-Cloud Security by 2026-2027</h2>
<h3>Phase 1: Visibility (Months 1-3)</h3>
<ul><li>Deploy CSPM tool(s) across all clouds</li><li>Inventory all resources and configurations</li><li>Establish baseline of current security posture</li><li>Accept that you will find many misconfigurations</li></ul>
<h3>Phase 2: Standards (Months 3-6)</h3>
<ul><li>Define security requirements at the outcome level, not the mechanism level</li><li>Example: "All customer data encrypted at rest" (not "all S3 buckets use KMS")</li><li>Create cloud-specific implementation guides for each requirement</li><li>Document why implementation differs across clouds</li></ul>
<h3>Phase 3: Automation (Months 6-12)</h3>
<ul><li>Implement Infrastructure as Code (IaC) templates for each cloud that enforce security standards</li><li>Automate CSPM remediation where possible (AWS Config rules, Azure Blueprints, GCP Forseti)</li><li>Establish drift detection (alert when manual changes deviate from approved IaC)</li><li>Build incident response playbooks with cloud-specific procedures</li></ul>
<h3>Phase 4: Integration (Months 12-18)</h3>
<ul><li>Export cloud security data (logs, alerts, findings) to central SIEM</li><li>Implement correlation rules that normalize cloud-specific alert formats</li><li>Establish unified incident response process with cloud-specific execution steps</li><li>Achieve consistent audit posture across clouds</li></ul>
<h2>Critical Success Factors</h2>
<ol><li><strong>Accept cloud heterogeneity</strong> (trying to force uniformity creates worse security than embracing cloud-specific approaches)</li><li><strong>Hire cloud specialists</strong> (not "cloud security generalists")</li><li><strong>Infrastructure as Code discipline</strong> (IaC is the only realistic way to enforce consistent security at scale in multi-cloud)</li><li><strong>Outcome-based governance</strong> (define what you want, not how each cloud achieves it)</li><li><strong>Automation wherever possible</strong> (manual multi-cloud security operations is unsustainable)</li></ol>
<h2>Cost Implications</h2>
<p>Organizations implementing disciplined multi-cloud security report:</p>
<ul><li>Security operations team size: 15-25% larger than single-cloud equivalent</li><li>Tooling costs: 20-40% higher (you need cloud-native tools + CSPM + SIEM)</li><li>Training investment: 2-3x higher (cloud expertise requires hands-on experience per cloud)</li></ul>
<p>Organizations attempting to minimize these costs by accepting policy drift or inconsistent monitoring typically experience security incidents that cost 10x more to remediate.</p>
<h2>Conclusion: Multi-Cloud Requires New Thinking</h2>
<p>The 2026 security landscape is irreversibly multi-cloud. Organizations that successfully navigate this are those that accept cloud provider differences, focus on outcome-based security standards, and invest in automation and cloud expertise. Those attempting to enforce uniform security policies across fundamentally different cloud platforms will struggle with alert fatigue, compliance gaps, and operational friction that degrades security effectiveness.</p>
<p>Multi-cloud is here to stay. The question is whether your organization will manage it strategically or reactively.</p>]]></content:encoded>
  </item>
  <item>
    <title>What managing a 16-person global team taught me about follow-the-sun ops</title>
    <link>https://prasannamalode.in/articles/follow-the-sun-ops.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/follow-the-sun-ops.html</guid>
    <pubDate>Tue, 13 Jan 2026 00:00:00 +0000</pubDate>
    <category>Leadership</category>
    <description>Handovers, shift scheduling, and the one ritual that keeps 24x7 teams sane and aligned.</description>
    <content:encoded><![CDATA[<p>Running a follow-the-sun operation sounds simple on paper: coverage across timezones, seamless handovers, 24x7 continuity. The reality is messier, and more human, than any playbook suggests.</p>
<h2>The handover problem</h2>
<p>The single biggest failure mode in follow-the-sun is the handover. Information gets lost. Context doesn't transfer. The incoming shift restarts investigations the outgoing shift already half-solved. We fixed this with a structured handover template. Not long, not bureaucratic, but specific about what's in flight, what's blocked, and what the next two hours need.</p>
<blockquote><p>A 5-minute handover done well saves 2 hours of re-investigation. Most teams skip it because they're busy. That busyness is exactly why they need it.</p></blockquote>
<h2>Performance monitoring across shifts</h2>
<p>You can't manage what you can't see across eight timezones. We built real-time dashboards giving visibility into SLA compliance, incident queue depth, and team velocity, regardless of shift. This wasn't about surveillance. It was about being able to unblock people quickly and spot patterns that no single shift could see alone.</p>
<h2>The one ritual that keeps teams aligned</h2>
<p>Weekly all-hands across shifts. Twenty minutes, same agenda every time: metrics review, one learning from the week, one shoutout. Simple. Consistent. It built the shared identity that makes distributed teams function like a team rather than isolated shifts doing disconnected work.</p>
<h2>Mentoring across distance</h2>
<p>Leading two managers and two DevOps architects remotely forced me to be more intentional about development than I ever was in person. Written feedback became more important. Regular 1:1s became non-negotiable. Career conversations happened on a cadence, not when someone was already considering leaving.</p>
<p>The single biggest thing I learned: presence is not about timezone overlap. It's about reliability. If your team knows you'll respond, unblock, and follow through, they feel supported regardless of where you are.</p>]]></content:encoded>
  </item>
  <item>
    <title>Bug Classes, Not Bug Instances</title>
    <link>https://prasannamalode.in/articles/bug-classes-not-bug-instances.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/bug-classes-not-bug-instances.html</guid>
    <pubDate>Tue, 23 Dec 2025 00:00:00 +0000</pubDate>
    <category>DevSecOps</category>
    <description>Every vulnerability you fix by hand is one you will fix again. The four-rung ladder for eliminating whole categories: findable, loud, then impossible to write.</description>
    <content:encoded><![CDATA[<p>Look at your security backlog and count how many tickets are the same ticket.</p>
<p>Not literally identical — different files, different services, different authors. But the same shape. Four path traversal findings. Nine missing authorisation checks on new endpoints. Six places where a secret went into a log line. Every one of them gets its own ticket, its own fix, its own review, its own regression test.</p>
<p>And then next quarter you get another four, another nine, another six. Because nothing you did changed the conditions that produce them.</p>
<p>This is the treadmill, and most security engineering runs on it.</p>
<h2>The distinction</h2>
<p>An <strong>instance</strong> is one occurrence of a defect in one place in one codebase. Fixing it removes exactly one defect.</p>
<p>A <strong>class</strong> is the set of conditions under which that defect can exist at all. Eliminating it removes every current instance and every future instance, including the ones that would have been written by an engineer who joins in two years and has never read your wiki.</p>
<p>Nearly all security tooling operates at the instance level. Scanners find instances. Tickets track instances. Metrics count instances. And because the tooling is instance-shaped, the work becomes instance-shaped, and you end up with a function that consumes headcount in proportion to how fast the rest of the company writes code.</p>
<p>That is not a security programme. That is a tax.</p>
<h2>The ladder</h2>
<p>For any bug class, there are four rungs you can be standing on.</p>
<p>Most teams live on rung one and have convinced themselves that buying a rung-two tool is transformation. It helps — rung two is genuinely better than rung one — but the leverage is above it, and the leverage is where the engineering is.</p>
<p>Let me make each rung concrete.</p>
<h3>Rung 1 → 2: make the machine find it</h3>
<p>The move here is from human vigilance to mechanical certainty. The signal that you are on rung one is a wiki page, an onboarding slide, or a code review checklist item. Any control that depends on somebody remembering is a rung-one control.</p>
<p><strong>Property-based testing</strong> is the most underused tool in this category. Instead of asserting specific outputs for specific inputs, you assert an invariant and let the framework attack it:</p>
<pre><code class="language-python">from hypothesis import given, strategies as st

@given(st.text())
def test_sanitiser_never_emits_angle_brackets(raw):
    out = sanitise_for_html(raw)
    assert &quot;&lt;&quot; not in out and &quot;&gt;&quot; not in out

@given(st.binary())
def test_parser_never_raises_unexpected(data):
    # The only acceptable failure is our own typed error
    try:
        parse_message(data)
    except ProtocolError:
        pass</code></pre>
<p>That second test is a fuzzer in four lines. It will find the unicode edge case, the zero-length input, the surrogate pair, and the thing you did not think of — which is the entire point, because the things you did not think of are where the vulnerabilities are.</p>
<p><strong>Continuous fuzzing</strong> is the same idea with a much bigger budget. If you parse anything untrusted — a file format, a wire protocol, a token, a config — and you are not fuzzing the parser, that parser is an unexplored attack surface. OSS-Fuzz is free for open source. For internal code, a nightly <code>cargo fuzz</code> or <code>libFuzzer</code> target on your parsing boundary costs a day to set up.</p>
<p><strong>Custom static analysis rules</strong> belong here too, but with the warning from part one: a rule earns its place by precision. A rule with 40% precision is rung one wearing a costume, because a human still has to adjudicate every hit.</p>
<h3>Rung 2 → 3: make it fail loudly</h3>
<p>Here the bug is still writable, but it stops being <em>silent</em>. Silence is what turns a bug into a vulnerability — the request that should have been denied but was served, the validation that was skipped without complaint.</p>
<pre><code class="language-python"># Rung 1: authorisation is a thing you must remember to call
@app.get(&quot;/orders/{order_id}&quot;)
def get_order(order_id):
    return db.fetch_order(order_id)   # forgot the check, nothing complains

# Rung 3: absence of a decision is an error
@app.get(&quot;/orders/{order_id}&quot;)
@requires_authz                       # framework raises at request time if no policy is declared
def get_order(order_id, actor):
    authorize(actor, &quot;order:read&quot;, order_id)
    return db.fetch_order(order_id)</code></pre>
<p>The mechanism is that the framework refuses to serve a route with no declared policy. Forgetting is now a 500 in staging rather than a data leak in production. You have not made the mistake impossible, but you have made it <em>immediately visible</em>, which collapses the time-to-detection from months to minutes.</p>
<p>Infrastructure has the same move. Default-deny egress on a namespace turns "this service can reach anything" into a loud, specific failure the first time something tries. It is annoying for a week and then it is a permanent boundary.</p>
<h3>Rung 3 → 4: make it impossible</h3>
<p>The top rung is where the defect cannot be expressed in the language you have given yourself.</p>
<p>The canonical case is SQL injection. You do not solve it by training people to be careful with string concatenation. You solve it by making the unsafe path unreachable:</p>
<pre><code class="language-go">// internal/db/db.go — the raw path is unexported
func rawQuery(q string) (*Rows, error) { ... }

// The only exported query function takes args separately.
func Query(tmpl Template, args ...any) (*Rows, error) { ... }</code></pre>
<p>Now injection is not a thing to be careful about. It is a compile error in every package outside <code>internal/db</code>, and a reviewer who has never heard of SQL injection cannot let it through.</p>
<p>The same pattern generalises through the type system:</p>
<pre><code class="language-rust">// Untrusted input cannot be confused with validated input,
// because they are not the same type.
struct RawInput(String);
struct SafePath(PathBuf);

fn validate(raw: RawInput, root: &amp;Path) -&gt; Result&lt;SafePath, ValidationError&gt; {
    let candidate = root.join(&amp;raw.0).canonicalize()?;
    if !candidate.starts_with(root) {
        return Err(ValidationError::Escape);
    }
    Ok(SafePath(candidate))
}

// Every file operation demands SafePath. There is no way to pass
// RawInput here. Path traversal is now a type error.
fn read_user_file(p: SafePath) -&gt; io::Result&lt;Vec&lt;u8&gt;&gt; { ... }</code></pre>
<p>Path traversal used to be a class that produced findings every quarter. Now it produces compiler diagnostics, for free, forever, in code written by people who have never thought about it.</p>
<p>And then there is the biggest single instance of this move available to the industry: <strong>memory safety</strong>. Somewhere around two thirds to three quarters of severe vulnerabilities in large C and C++ codebases are memory-safety issues. That is not a class you can train away or scan away. It is a class you can move to rung four by changing language at the boundary — which is exactly what Android, Windows, and the Linux kernel have been doing, with published results showing new memory-safe code producing a fraction of the vulnerability density of the code it replaced.</p>
<h2>The rewrite objection</h2>
<p>The immediate response to rung four is that it means rewrites, and rewrites do not get funded.</p>
<p>Mostly true, and mostly beside the point, because you do not need a rewrite. You need a <strong>boundary</strong>.</p>
<p>The rule that makes this tractable: <em>new code goes on the high rung; old code stays where it is and gets a clear edge drawn around it.</em></p>
<ul><li>New services use the safe query API. Old ones get a scanner and a backlog.</li><li>New parsers get written in a memory-safe language. The old parser gets fuzzed and sandboxed.</li><li>New endpoints require a declared policy. Old ones get an allowlist that only shrinks.</li></ul>
<p>This works because code has a half-life. A meaningful fraction of your codebase will be rewritten anyway in the normal course of shipping. If the default for new code is on rung four, your vulnerability density decays without anyone ever running a rewrite project.</p>
<p>The failure mode to watch is the allowlist that grows. Whatever form the boundary takes — an exception list, a legacy annotation, a grandfathered directory — it needs a ratchet. It can shrink, it cannot grow, and adding to it requires the same approval as any other risk acceptance. Which is the subject of the next post.</p>
<h2>Running this as practice</h2>
<p>Two habits are enough to make this real.</p>
<p><strong>First: after every incident, ask the class question.</strong> Not "how do we fix this bug" but "what rung is this class on, and what would it take to move it up one?" Put that in the postmortem template as a required field. It will be answered badly at first and then it will start being answered well.</p>
<p><strong>Second: cluster your backlog quarterly.</strong> Take every security ticket from the last three months and group by shape, not by service. Any cluster with five or more members is not five tickets — it is one class with five symptoms. Close all five and open one ticket to climb a rung.</p>
<p>The cluster exercise is worth doing once just for the shock value. The first time I did it, about 60% of a quarter's security backlog collapsed into four classes. We had been paying to fix the same four things over and over, and the ticket system had been carefully hiding that from us by giving each symptom its own ID.</p>
<h2>The summary</h2>
<p>Every vulnerability you fix by hand is a vulnerability you will fix again, because the conditions that produced it are still there and the code is still growing.</p>
<p>The only security work that compounds is the work that makes a category of defect harder to write than the correct alternative. Everything else is maintenance — necessary, but it does not accumulate, and it scales with your engineering headcount rather than against it.</p>
<p>Ask the class question. Draw the boundary. Let the ratchet do the rest.</p>]]></content:encoded>
  </item>
  <item>
    <title>How ITIL-aligned incident management cut our recurring incidents by 50%</title>
    <link>https://prasannamalode.in/articles/itil-incident-management.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/itil-incident-management.html</guid>
    <pubDate>Tue, 09 Dec 2025 00:00:00 +0000</pubDate>
    <category>ITIL / Ops</category>
    <description>The exact process changes, tooling decisions, and cultural shifts that drove the improvement.</description>
    <content:encoded><![CDATA[<p>ITIL gets a bad reputation: heavy process, slow bureaucracy, documentation nobody reads. Done right, it's the opposite. Here's what ITIL-aligned incident management actually looked like in practice for a global IT operations team.</p>
<h2>The recurring incident trap</h2>
<p>Before the transformation, we were great at closing incidents. We were terrible at preventing them from coming back. The same infrastructure issues, the same application failures, the same network anomalies, cycling through the queue month after month.</p>
<blockquote><p>Closing incidents fast is a metric. Preventing them from recurring is an outcome. Most teams optimise for the metric.</p></blockquote>
<h2>Separating incident from problem management</h2>
<p>The structural change was treating Incident Management (restore service fast) and Problem Management (find root cause) as separate disciplines with separate owners. Incident managers focused on MTTR. Problem managers focused on permanent fixes. The handoff between them became a defined process, not an afterthought.</p>
<h2>The ServiceNow implementation</h2>
<p>ServiceNow gave us the data we needed: incident trends by category, MTTR by team, SLA breach patterns. But the data was only useful because we reviewed it weekly and acted on it. Tooling without review cadence is just expensive storage.</p>
<h2>Results that mattered to the business</h2>
<p>Recurring incidents down 50%. SLA performance up 40%. But the result I'm most proud of: the team stopped dreading the queue. When you're not firefighting the same fires repeatedly, morale follows. That's the real return on an ITIL investment. Not the certification, not the process documentation, but a team that's energised rather than exhausted.</p>]]></content:encoded>
  </item>
  <item>
    <title>Your PR Has 400 Security Findings. Nobody Is Going to Fix Them.</title>
    <link>https://prasannamalode.in/articles/your-pr-has-400-security-findings.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/your-pr-has-400-security-findings.html</guid>
    <pubDate>Tue, 25 Nov 2025 00:00:00 +0000</pubDate>
    <category>DevSecOps</category>
    <description>Coverage is a vanity metric. Track fix rate per tool, give every pipeline stage a latency budget, and prefer secure defaults over blocking gates.</description>
    <content:encoded><![CDATA[<p>There is a specific moment I want to describe, because I think most engineers have lived it and nobody writes it down.</p>
<p>You open a pull request. It changes eleven lines in a config file. The checks run. Fourteen minutes later you come back to a wall of red: 400-odd findings from six different tools, most of them in files you have never opened, several of them duplicates of each other described in different vocabularies, one of them possibly real.</p>
<p>You scroll. You do not read. You look for the button that makes it go away.</p>
<p>That reflex — the scroll, the scan for the override — is the actual output of most shift-left programmes. Not fixes. A reflex.</p>
<h2>Coverage is a vanity metric</h2>
<p>Here is the story security programmes tell themselves. We added SAST. We added SCA. We added secret scanning, IaC policy, container scanning, license compliance. Coverage went from 20% of repositories to 95% of repositories. The quarterly slide is green.</p>
<p>Here is the story the codebase tells. Findings went up. Fixes did not. The delta went into a suppression file.</p>
<p>Coverage measures how many tools you bought and wired up. It measures procurement. It says nothing about whether a single vulnerability left the codebase.</p>
<p>The number that matters is the fix rate:</p>
<blockquote><p>Of the findings surfaced in a given period, what fraction were remediated in code — not suppressed, not marked won't-fix, not aged out?</p></blockquote>
<p>Track that monthly. Track it per tool. The first time you look at it per tool, you will find at least one scanner with a fix rate near zero, which means it has been generating pure noise and consuming review attention for however long you have had it.</p>
<p>That tool is not a partial win. It is negative. It is making your other tools worse.</p>
<h2>Why adding a scanner makes your other scanners worse</h2>
<p>Developer attention on a pull request is fixed. Call it a few minutes. That budget does not expand because you added a seventh tool; it just gets divided more thinly.</p>
<p>So each new scanner does two things at once. It adds some true positives, which is the reason you bought it. And it dilutes the average quality of the findings list, which lowers the probability that any individual finding — including the ones from your good tools — gets read.</p>
<p>Below some threshold of precision, the second effect dominates. You cross a line where the list stops being a list of problems and starts being an obstacle. Once developers learn that the security check is usually wrong, they stop evaluating findings individually. They evaluate the <em>category</em>: "that's the security check, it's always noisy." And at that point your genuinely critical finding is indistinguishable from a style nit about a test fixture.</p>
<p>You cannot fix this by adding a tenth tool that ranks the other nine. You fix it by turning tools off.</p>
<p><strong>The uncomfortable exercise:</strong> pick your noisiest scanner. Turn it off in the PR entirely. Move it to nightly with a named owner. Watch the fix rate on your remaining tools. In my experience it goes up, because the remaining findings are now read.</p>
<h2>The feedback latency budget</h2>
<p>The second structural mistake is putting checks in the wrong place. Not the wrong tool — the wrong <em>stage</em>.</p>
<p>The principle is that a check earns its stage by how fast it can answer, not by how important somebody thinks it is. Importance determines whether you run it at all. Runtime determines where.</p>
<div class="table-wrap"><table><thead><tr><th>Stage</th><th>Budget</th><th>What belongs here</th><th>What kills it</th></tr></thead><tbody><tr><td>Editor</td><td>&lt; 1s</td><td>Type checks, linters, secret patterns, known-bad API usage</td><td>Anything needing a build</td></tr><tr><td>Pre-commit</td><td>&lt; 10s</td><td>Fast SAST on the diff, entropy scan on staged files</td><td>Whole-repo analysis</td></tr><tr><td>Pull request</td><td>&lt; 10 min</td><td>SCA on changed dependencies, IaC policy, unit and contract tests</td><td>Full DAST, deep taint analysis</td></tr><tr><td>Nightly</td><td>hours</td><td>Full-repo SAST, DAST, fuzzing, image and base-layer scanning</td><td>Nothing — this is the dumping ground, and that is fine</td></tr><tr><td>Backlog</td><td>days+</td><td>—</td><td>Everything. This is where findings go to die.</td></tr></tbody></table></div>
<p>The common failure is to take a check that genuinely takes 40 minutes and put it in the PR anyway, because it is important. Importance does not make it faster. What it produces is a 40-minute PR cycle, which produces batching, which produces bigger PRs, which produces worse review — so your important check has now degraded the quality of human review across the whole repository.</p>
<p>If a check cannot meet its stage's budget, you have two honest options. Narrow its scope so it only analyses the diff. Or move it right and give it an owner who triages the output on a schedule.</p>
<p>The dishonest third option is to keep it in the PR and let people learn to ignore it.</p>
<h2>Guardrails beat gates</h2>
<p>The deepest version of this problem is that blocking checks are a weak instrument. A gate stops the build. It does not teach anything, it does not fix anything, and it creates pressure to find the bypass. Every organisation that leans hard on gates develops a culture of bypass — an <code>--no-verify</code> habit, a <code>security-exception</code> label that nobody audits, a shared admin account that can force-merge.</p>
<p>Secure defaults work differently. They are invisible. Nobody notices them, nobody argues with them, and they have a 100% adoption rate on new code.</p>
<p>Some concrete substitutions:</p>
<ul><li><strong>Instead of</strong> a scanner that flags plaintext HTTP in Terraform, <strong>ship</strong> a module where the listener block is not exposed and TLS is the only path.</li><li><strong>Instead of</strong> a check that fails on missing resource limits in Kubernetes manifests, <strong>set</strong> a LimitRange on the namespace so the default is correct.</li><li><strong>Instead of</strong> a secret scanner in the PR, <strong>remove</strong> long-lived secrets from CI entirely with OIDC federation, so there is nothing to leak.</li><li><strong>Instead of</strong> a SAST rule for SQL string concatenation, <strong>make</strong> the raw-query function private and export only the parameterised one.</li></ul>
<p>The pattern: move the control from <em>detection at review time</em> to <em>impossibility at authoring time</em>. That is the real meaning of shifting left, and it has nothing to do with scanners.</p>
<h2>What a healthy programme looks like</h2>
<p>A short checklist you can hold your own setup against:</p>
<ol><li><strong>Fix rate is reported per tool, monthly.</strong> Any tool below a threshold you set gets moved out of the PR or removed.</li><li><strong>False positive rate is measured</strong>, by sampling — take 30 random findings a month and have an engineer adjudicate them. You cannot manage what you have not sampled.</li><li><strong>Every stage has a latency budget</strong> and a check that exceeds it gets moved, not excused.</li><li><strong>Suppressions expire.</strong> (This is the subject of part three.)</li><li><strong>Every recurring finding class has a ticket to eliminate the class</strong>, not a ticket to fix the instances. (Part two.)</li><li><strong>The security team owns the noise.</strong> If a tool is noisy, that is a security-team defect, not a developer-discipline problem.</li></ol>
<p>Point six is the cultural one and the hardest. The default posture in a lot of organisations is that findings are the developer's problem and triage is the developer's job. That posture guarantees noise, because nobody who generates the noise pays for it. Put the cost where the decision is made: the team that turns on the scanner owns its precision.</p>
<h2>The honest summary</h2>
<p>Shift-left, as commonly practised, is the act of moving unfiltered machine output earlier in the process and calling the relocation an improvement. It is not. Moving a bad signal earlier just means you interrupt people sooner.</p>
<p>The thing worth moving left is not <em>detection</em>. It is <em>decision</em>: the design choice, the default, the library boundary, the credential architecture. Those are the things that, moved left, stop bugs from existing.</p>
<p>Detection is what you do about the ones that get through anyway. Keep it, tune it hard, measure it honestly, and stop pretending the size of the findings list is a measure of anything but the size of the findings list.</p>]]></content:encoded>
  </item>
  <item>
    <title>How we integrated Gen AI into CI/CD pipelines, and what surprised us</title>
    <link>https://prasannamalode.in/articles/genai-cicd-pipelines.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/genai-cicd-pipelines.html</guid>
    <pubDate>Tue, 11 Nov 2025 00:00:00 +0000</pubDate>
    <category>Gen AI</category>
    <description>Real-world implementation of AI-assisted code reviews, decision gates, and memory leak detection in production CI/CD.</description>
    <content:encoded><![CDATA[<p>Everyone's talking about AI in software development. Most of it is theoretical. Here's what we actually implemented, what worked, and what we'd do differently.</p>
<h2>Where we started</h2>
<p>We started with the lowest-risk, highest-value use case: AI-assisted developer guides. Instead of static documentation that went stale, we built context-aware guides that pulled from live pipeline state. Engineers got relevant, current information at the moment they needed it, not a wiki page last updated 18 months ago.</p>
<h2>AI in CI/CD decision gates</h2>
<p>The bigger bet was integrating Gen AI into CI/CD decision gates, using models to assess code quality risks before builds proceeded. The model was trained on our historical incident data, so it understood which patterns of change correlated with deployment failures in our specific codebase.</p>
<blockquote><p>Generic AI gives generic advice. The value came when the model understood our patterns, our history, our failure modes.</p></blockquote>
<h2>Automated code reviews</h2>
<p>AI-assisted code reviews flagged memory leaks, FOSS snippet matches (critical for our IP compliance requirements), and code quality issues in real time, before human reviewers saw the PR. Human reviewers then focused on logic, architecture, and context rather than mechanical checks. Review quality went up. Review time went down.</p>
<h2>What surprised us</h2>
<p>The biggest surprise was adoption speed. Engineers embraced it faster than any other tooling change we'd made, because it made their jobs easier rather than adding overhead. The second surprise: the AI caught a category of FOSS compliance issues we hadn't anticipated it would find. That became one of its most valuable ongoing functions.</p>
<p>What didn't work: using AI for release go/no-go decisions without human review. The model was right 90% of the time, which sounds good until it's wrong on a major release. Human judgment stays in the loop for consequential decisions.</p>]]></content:encoded>
  </item>
  <item>
    <title>Zero Trust in 2026: Implementation Reality vs. Industry Hype</title>
    <link>https://prasannamalode.in/articles/zero-trust-reality-check.html</link>
    <guid isPermaLink="true">https://prasannamalode.in/articles/zero-trust-reality-check.html</guid>
    <pubDate>Tue, 21 Oct 2025 00:00:00 +0000</pubDate>
    <category>Cybersecurity</category>
    <description>Zero Trust is a 3 to 5 year architectural commitment, not a product. Where enterprises actually are, the hidden costs vendors skip, and a realistic roadmap.</description>
    <content:encoded><![CDATA[<p>Zero Trust has transitioned from a theoretical framework to the de facto standard for enterprise security architecture. However, the gap between announced commitments and actual implementation reveals a more complex reality: organizations are discovering that Zero Trust is not a product or quick initiative, but a fundamental redesign of trust assumptions that requires 3-5 years of sustained effort.</p>
<h2>The Hype Inflection Point</h2>
<p>In 2020, Zero Trust was visionary. By 2024, every major vendor claimed Zero Trust capabilities. By 2026, customer sentiment has shifted toward skepticism: organizations are realizing they've purchased Zero Trust solutions without actually implementing Zero Trust architecture.</p>
<h3>Why the Disconnect?</h3>
<p><strong>Zero Trust (The Framework):</strong> Never trust, always verify. Assume breach. Implement least-privilege access. Continuously monitor and validate.</p>
<p><strong>Zero Trust (The Product Category):</strong> A collection of point solutions—PAM, EDR, NAC, SIEM, identity governance—sold with Zero Trust branding but lacking cohesion in implementation.</p>
<h2>The Implementation Maturity Model: Where Most Organizations Actually Are</h2>
<h3>Stage 1: Assessment &amp; Quick Wins (2021-2023)</h3>
<ul><li>Inventory what you have</li><li>Implement MFA across cloud applications</li><li>Deploy EDR to endpoints</li><li>Declare Zero Trust adoption (largely aspirational)</li></ul>
<p><strong>Percentage of enterprises here:</strong> 64%</p>
<h3>Stage 2: Identity &amp; Access Centerline (2023-2025)</h3>
<ul><li>Implement PAM for privileged accounts</li><li>Deploy identity-centric network access (beyond VPN)</li><li>Build access policies based on identity + device + context</li><li>Realize that legacy applications don't support modern auth</li></ul>
<p><strong>Percentage of enterprises here:</strong> 28%</p>
<h3>Stage 3: Micro-Segmentation &amp; Continuous Verification (2024-2026+)</h3>
<ul><li>Segment network based on business logic, not perimeter</li><li>Implement continuous trust evaluation</li><li>Monitor and validate every transaction</li><li>Rearchitect applications to support least-privilege service-to-service communication</li></ul>
<p><strong>Percentage of enterprises here:</strong> 6-8%</p>
<h3>Stage 4: True Zero Trust Maturity (2026+)</h3>
<ul><li>Every access decision evaluated in real-time</li><li>Behavioral analytics inform trust scores</li><li>Legacy systems decommissioned or wrapped in security overlays</li><li>Organizational culture shifted to "verify always"</li></ul>
<p><strong>Percentage of enterprises here:</strong> &lt;1%</p>
<h2>The Hidden Costs: What Organizations Aren't Discussing</h2>
<p>Vendors market Zero Trust as "security modernization." What they don't emphasize:</p>
<h3>Operational Friction (The Invisible Tax)</h3>
<ul><li><strong>Authentication latency:</strong> Continuous verification adds 200-500ms to critical workflows</li><li><strong>User experience degradation:</strong> Security policies that block "unusual" access patterns frustrate legitimate users</li><li><strong>Support burden:</strong> Help desk calls increase 40-60% during Zero Trust transition as users encounter unexpected blocks</li></ul>
<p>Organizations report that the first 12-18 months of serious Zero Trust implementation produce measurable productivity drag.</p>
<h3>Legacy Application Incompatibility</h3>
<p>Approximately 30-40% of enterprise applications cannot be updated to support modern Zero Trust requirements within a reasonable timeframe (security debt from acquisitions, vendors out of business, custom legacy code).</p>
<p>Options are all painful:</p>
<ul><li><strong>Retire the application</strong> (expensive, operational impact)</li><li><strong>Wrap it</strong> (add a proxy/gateway layer, adds latency and risk)</li><li><strong>Exclude it from Zero Trust</strong> (leaves a perimeter)</li></ul>
<p>Most organizations choose wrapping, creating hybrid architectures that lack coherence.</p>
<h3>The Identity Crisis</h3>
<p>Zero Trust is fundamentally identity-centric. This creates a hard dependency on:</p>
<ul><li><strong>Directory services quality</strong> (if your AD/Okta data is inconsistent, everything downstream is fragile)</li><li><strong>Identity governance</strong> (managing thousands of service principals, API keys, and credentials at scale)</li><li><strong>Onboarding/offboarding precision</strong> (one mismanaged account becomes a backdoor)</li></ul>
<p>Organizations that have not achieved identity governance maturity are not ready for true Zero Trust, but many proceed anyway.</p>
<h2>Emerging Implementation Patterns in 2026</h2>
<h3>Pattern 1: Cloud-First Zero Trust</h3>
<p>Organizations with primarily cloud workloads (SaaS, AWS, Azure) are achieving Stage 3 maturity in 24-30 months. Their advantage: greenfield design, no legacy constraints, vendors supporting cloud-native architectures natively.</p>
<p><strong>Success rate:</strong> 52% of organizations starting with cloud-first approach achieve sustained Stage 3 implementation</p>
<h3>Pattern 2: Perimeter Modernization (Not True Zero Trust)</h3>
<p>Many organizations implement modern perimeter security—SASE, cloud WAF, advanced firewall—and rebrand it as Zero Trust.</p>
<p><strong>Reality check:</strong> This is castle-and-moat with better locks, not Zero Trust. It addresses relevant threats (DDoS, mass scanning, data exfiltration) but not insider threats or compromised credentials.</p>
<p><strong>Risk exposure:</strong> High vulnerability to insider threats and lateral movement post-compromise.</p>
<h3>Pattern 3: Segmentation-Led Zero Trust</h3>
<p>Rather than starting with identity, some organizations begin with network micro-segmentation, then layer identity controls.</p>
<p><strong>Results:</strong> 35% faster time-to-value than identity-first approaches, but creates fragmented trust policies if not unified later.</p>
<h3>Pattern 4: Assume-Breach Architecture</h3>
<p>Some security-mature organizations sidestep "Zero Trust transformation" and instead adopt "assume-breach" operations:</p>
<ul><li>Assume credential compromise</li><li>Implement rapid detection and containment</li><li>Build recovery capabilities</li></ul>
<p>This differs philosophically (prevention vs. containment) but addresses the same threat scenarios as Zero Trust.</p>
<p><strong>Emerging trend:</strong> 23% of enterprises with very strong security practices are pursuing assume-breach instead of Zero Trust, citing more realistic threat modeling.</p>
<h2>The Organizational Dynamics</h2>
<h3>The 18-Month Plateau</h3>
<p>Organizations implementing Zero Trust report significant early wins (months 3-12), then hit a plateau around month 18 when:</p>
<ul><li>Legacy applications slow progress</li><li>User friction becomes political</li><li>Incident response for Zero Trust-related blocking becomes routine</li></ul>
<p>This is where many initiatives stall or get deprioritized. Organizations pushing through report the plateau resolves around month 24-30.</p>
<h3>The Skills Gap</h3>
<p>Zero Trust requires security architects who understand:</p>
<ul><li>Identity and access management at depth</li><li>Network architecture and segmentation</li><li>Application architecture and service communication patterns</li><li>Compliance implications of access policies</li></ul>
<p>This skill combination is rare. Salary premiums for "Zero Trust architects" are 30-40% above industry standard, and talent remains scarce.</p>
<h2>Realistic Roadmap for Enterprises Starting in 2026</h2>
<h3>Year 1: Foundation</h3>
<ul><li>Assess current state (identity, access, applications, network)</li><li>Achieve &gt;95% MFA coverage for cloud and critical systems</li><li>Deploy EDR to 100% of endpoints</li><li>Invest in identity governance tooling</li></ul>
<p><strong>Expected maturity:</strong> Early Stage 2</p>
<h3>Year 2: Identity &amp; Access</h3>
<ul><li>Implement PAM for privileged accounts</li><li>Deploy identity-centric network access for critical assets</li><li>Begin application inventory and auth requirements assessment</li><li>Build access policies based on role + device context</li></ul>
<p><strong>Expected maturity:</strong> Mid-Stage 2</p>
<h3>Year 3: Segmentation &amp; Monitoring</h3>
<ul><li>Implement network micro-segmentation for critical workloads</li><li>Deploy continuous compliance monitoring</li><li>Address legacy applications (retire, wrap, or modernize)</li><li>Mature incident response for Zero Trust scenarios</li></ul>
<p><strong>Expected maturity:</strong> Late Stage 2 to Early Stage 3</p>
<h3>Years 4-5: Maturity &amp; Culture</h3>
<ul><li>Extend segmentation to non-critical workloads</li><li>Implement behavioral analytics for trust scoring</li><li>Retire legacy wrappers through application modernization</li><li>Achieve organizational alignment on Zero Trust principles</li></ul>
<p><strong>Expected maturity:</strong> Stage 3 sustained</p>
<h2>Critical Success Factors</h2>
<ol><li><strong>Executive sponsorship</strong> that understands this is 3-5 year commitment, not a 12-month project</li><li><strong>Dedicated Zero Trust architecture team</strong> (not a side project for existing security staff)</li><li><strong>Application inventory and dependency mapping</strong> before policy design</li><li><strong>Regular user feedback loops</strong> to balance security with usability</li><li><strong>Forgiveness in early phases</strong>—over-blocking is better than under-blocking, but creates friction</li></ol>
<h2>Conclusion: The Honest Assessment</h2>
<p>Zero Trust is not a destination; it's a directional commitment that requires sustained investment and organizational change. Organizations that position it as a "security initiative" will struggle. Organizations that treat it as an architectural transformation with 3-5 year timelines and significant investment are seeing real results.</p>
<p>The vendors' silence on timelines and costs is the industry's biggest credibility gap. Expect organizations to demand more realistic roadmaps and pricing models in 2027.</p>]]></content:encoded>
  </item>
</channel>
</rss>
