Est.

Continuous Vendor Monitoring vs. Annual Assessment Programs

Combining assessments and continuous monitoring closes the gaps each tool leaves alone.

Features Editor · · 12 min read
Cover illustration for “Continuous Vendor Monitoring vs. Annual Assessment Programs”
Vendor Monitoring · September 21, 2026 · 12 min read · 2,718 words

When MOVEit's file-transfer software broke open, the vendors on the other end of that breach weren't unknown quantities. Most had assessments on file. Many had monitoring tools running somewhere in the stack. The breach cascaded across a wide range of downstream organizations anyway, because the two safeguards never spoke to each other. That gap, not the absence of either method, is the actual failure mode in third-party risk management today, precisely because the two safeguards never speak to each other.

Most vendor risk programs treat continuous monitoring and annual assessments as if they compete for the same budget line, as if choosing one means shorting the other. That framing is wrong on its face. Each method does a job the other one cannot do. An assessment establishes what a vendor's controls look like on paper, at a point in time, under scrutiny. Monitoring watches what happens to that vendor's exposure every day after the paperwork gets filed. Neither substitutes for the other, and a program running both without wiring them together is arguably worse than a program that admits it only has one. The rest of this piece lays out what each method actually measures, where each one fails in isolation, and what it takes to build the connective tissue that turns two separate tools into a single functioning loop.

What annual assessments measure, and what they cannot

An annual assessment exists to build a documented baseline: security posture, compliance certifications, financial standing, contractual commitments, all captured and signed off at a fixed moment. Without that baseline, a monitoring alert is just noise, a number with nothing to compare it against. Without it, a monitoring alert is just noise, a number with nothing to compare it against. A risk score drop means nothing if nobody recorded what "fine" looked like for that vendor.

Certifications feed into that baseline, but they get misread constantly. A SOC 2 Type II report covers a defined audit window, and it confirms that controls operated during that stretch. It says nothing about whether those controls are still running the day someone pulls the report to make a decision. That distinction, between a control that was operating and a control that is operating, is where a lot of vendor risk quietly slips through.

The AT&T vendor matter and its FCC settlement makes the point concretely. The contract required data deletion. Assessments and reviews conducted across 2016 through 2020 reported compliance. The data was never actually deleted. Nobody built a mechanism to check the declared control against what had actually happened on the vendor's systems, and the questionnaire almost certainly came back looking fine every single year. A declared control and an implemented one are not the same thing, and a questionnaire has no way to tell them apart.

Questionnaires are self-reported by design. A vendor answers based on stated policy and intended controls, not necessarily on what's true of their environment at the moment of submission, and there's no built-in mechanism that validates the answer against reality. None of this makes assessments useless. It makes them structurally blind to anything that happens after the review date closes, which is not a vendor shortcoming so much as a predictable property of the instrument itself. Assessments are indispensable for setting the baseline. They are the wrong tool for catching drift away from it.

The gap between assessments is not managed risk, it is unmanaged exposure

An annual cycle leaves 364 days between formal reviews. Vendor environments don't sit still for even one of them: staff turn over, infrastructure gets reconfigured, new integrations go live, ownership changes hands. Treating that stretch as "covered" because a review happened once is a category error, and it's one a lot of programs make without quite admitting it.

The timing math makes the exposure worse than it looks on paper. The average vendor takes 117 days to move from discovering an internal breach to disclosing it downstream, so even a vendor acting in good faith, one that catches its own incident quickly, may sit on that knowledge for close to four months before a client organization hears a word. Layer that onto an annual review cycle and the effective blind spot stretches well past a year in the worst case.

Third-party involvement in breaches isn't a marginal concern anymore, either. The Verizon 2026 Data Breach Investigations Report puts third-party involvement at 48% of all confirmed breaches in the first half of 2026, up 60% year over year. Separately, 45% of organizations report third-party related business interruptions over the past two years, a 68% year-over-year jump. These aren't edge cases. They describe the median experience of running a modern vendor portfolio.

Think of an annual review as a single frame pulled from a twelve-month film. The frame itself might be perfectly accurate. Everything else in the film, every frame before and after it, stays invisible until the next review comes around. What changes in that invisible stretch: external attack surface, newly exposed assets, certifications that quietly lapse, key personnel walking out the door, ownership changes including M&A and shifts in foreign control, new supply chain dependencies. All of it moves on a timescale of days and weeks. None of it waits for the calendar to turn over.

The cost of that blindness is measurable. The gap between assessments isn't administrative overhead nobody got around to fixing. It carries a price tag, and it's getting paid.

What continuous monitoring does, and where it breaks down alone

Continuous monitoring is the automated, ongoing collection and analysis of externally observable security data, surfacing changes in a vendor's posture as they happen rather than months later. It pulls on signals like exposed infrastructure and open ports, credential leaks, patching speed, botnet infections, DNS health, dark web exposure, financial indicators, and regulatory actions, all gathered from outside the vendor's own perimeter. Monitoring doesn't need the vendor's cooperation or an honest self-assessment to work. It catches a compromise even when the vendor itself hasn't noticed yet.

The performance gap is substantial. Organizations running real-time oversight detect vendor risk degradation roughly 4.5 months earlier on average than those that don't, a meaningful head start when breaches take an average of 277 days to identify and contain.

But monitoring left to run without a baseline breaks in its own particular way. A risk score drop needs context before it means anything. Was the vendor already sitting close to the edge? Is the drop a remediation underway, or a new hole opening up? Without a documented baseline to compare against, every alert forces a manual look, which erases the exact efficiency the automation was supposed to buy. Teams facing a wall of uncontextualized signals start letting the queue pile up, and the monitoring runs while the protection it was built to deliver quietly stops happening. It's not a hypothetical: in the broader security operations world, 27% of IT professionals already field more than a million security alerts a day, and TPRM teams applying monitoring without any tiering are headed for a structurally similar wall.

Watching every vendor with the same intensity is an alert factory, not a risk program. It's an alert factory. The fix is matching monitoring depth to vendor criticality rather than spreading it evenly, and the Ncontracts State of Third-Party Risk Management report found that 43% of organizations, the largest share, classify only 0 to 5% of their vendors as critical. That's a narrow slice, and it's exactly where depth of review and monitoring intensity ought to concentrate.

Each method's dependence on the other to function as designed

The relationship between the two methods runs in both directions, not down a hierarchy where one method reports to the other. Monitoring without a baseline produces signals nobody can interpret, and interpretation problems curdle into alert fatigue, and alert fatigue quietly kills the protection the program was funded to deliver. Assessment without monitoring in between produces a snapshot that starts going stale the moment it's approved: a vendor with a clean January review can be sitting on an active compromise by April, with nothing in the program built to notice.

"Connective tissue" means something specific here, not a vague call for better communication. Monitoring signals should trigger a reassessment, not stand in for one; a material posture change mid-cycle is grounds to pull that vendor's file back open early, not just fire off an alert into a queue. And assessment findings should set the thresholds that make monitoring alerts mean something: the same size score drop reads very differently for a vendor whose baseline was already marginal than for one whose baseline was strong.

Revisit MOVEit with that lens. Affected organizations had assessed the vendor. Some had monitoring tools active. The failure sat in the missing link: the monitoring signal, active exploitation of a known vulnerability, had no automated route back into the assessment record, the risk score, or the contract compliance status. The information existed. It just had nowhere to go.

Structurally, the connected version of this looks like a loop, not two tracks running in parallel: assessment sets the baseline, monitoring watches for drift, drift crossing a threshold triggers reassessment or a contract review, the reassessment updates the baseline, and the cycle starts again.

Risk tiering as the practical precondition for running both methods at scale

Scale makes the case on its own. Ponemon-Sullivan's 2026 figures put the average enterprise vendor count at 2,643, and at a typical six-week-plus assessment cycle per vendor, more than 950 of them go unassessed in any given cycle purely on the math. Treating every vendor the same isn't a philosophy at that point, it's an arithmetic failure. Separately, vendor portfolios across organizations of all sizes have been growing steadily, so even outfits running leaner portfolios are watching vendor count outpace analyst headcount.

Tiering fixes the mismatch by matching method intensity to actual exposure rather than applying one standard everywhere. Vendors with deep network or data access, the critical tier, get full annual assessment plus real-time, multi-signal monitoring, and on-site audits where the relationship warrants it. Vendors with limited access to confidential data get a structured periodic assessment plus daily or weekly attack-surface and breach-intelligence monitoring. Vendors with medium business impact get a lighter periodic review plus monthly automated scans. Vendors with no sensitive data access get a zero-touch automated assessment and quarterly public-record checks.

That bottom tier deserves defending on its own terms. Keeping low-risk vendors off an analyst's desk entirely isn't cutting corners, it's the correct allocation of a scarce resource: human judgment, spent where exposure is actually concentrated rather than spread thin across vendors that barely register as risk.

None of this works if the underlying vendor inventory is a mess. Roughly a third of organizations, 34% per one source, still run TPRM primarily through spreadsheets, and that approach introduces error rates of 15 to 20% into risk scoring. Tiering only functions if vendor classifications live in a system that can act on them automatically, not one where someone recalculates the list by hand every cycle.

Non-human identities and agentic AI as the monitoring gap that annual assessments were never built to see

Non-human identities, service accounts, API keys, OAuth tokens, AI agents acting on a system's behalf, now outnumber human identities by roughly 45 to 1 across the average enterprise. In cloud-native environments specifically, Entro Security research puts that ratio at 144 to 1, up from 92 to 1 in the first half of 2024, a 56% jump in a single year. That ratio marks a different landscape, not a rounding error in the identity landscape. It's a different landscape.

Annual assessments were never built to see any of it. NHIs skip past MFA, run continuously without anything resembling normal human behavior patterns, and stick around indefinitely without anyone managing their lifecycle. A questionnaire that asks "how do you manage privileged access?" has no way to capture a token issued two weeks after the vendor submitted that answer.

The Salesloft Drift breach of August 2025 is the case study that makes the abstraction concrete. The attacker compromised Salesloft Drift's backend and extracted OAuth access and refresh tokens tied to Salesforce integrations, hitting a large number of organizations in the process. The blast radius extended well beyond what direct attacks on Salesforce had previously produced. Nothing about the attack required sophisticated malware. It used valid credentials to do things those credentials were technically permitted to do, and that pattern, poor access governance rather than a clever exploit, has been widely characterized as the standard NHI kill chain.

The structural vulnerability sits in how coarsely these permissions get granted. Most AI agent deployments hand out access like "read/write to Salesforce," rather than something scoped like "read this one object for this one task, expiring in four hours." That gap, between what an agent actually needs and what it's actually been given, is what turns a routine credential compromise into a seven-hundred-organization event.

The gap isn't only technical. A global policy organization's analysis found that 51% of organizations have no clear owner for their AI identities at all, which means the failure starts before any monitoring tool ever gets a chance to catch it. There's a fourth-party dimension layered on top of that: a vendor's AI feature can route customer data through a sub-processor, a model provider, that the contracting organization never reviewed and may not even know exists, and that sub-processor can introduce obligations that were invisible to whatever assessment was filed at onboarding.

Catching any of this requires monitoring built for it specifically: tracking token lifecycles, auditing OAuth scopes, watching for new SaaS integrations or AI features a vendor adds mid-cycle. A periodic questionnaire, no matter how well designed, cannot detect any of this because it only captures what the vendor discloses at a fixed point in time, not what changes on their systems mid-cycle.

Why regulatory requirements now foreclose the annual-only option

Regulators have stopped treating annual attestation as sufficient, and the direction isn't subtle. DORA, and related frameworks have all moved toward continuous evidence of oversight rather than a once-a-year sign-off. DORA in particular has moved regulators toward expectations of continuous oversight of third-party risk, not just log them at the next scheduled review, which a point-in-time assessment simply cannot satisfy on its own.

NYDFS Part 500 requires written third-party risk policies with periodic reassessment, and it requires updating risk assessments whenever a material change in business or technology calls for it, language that treats monitoring as the operating mechanism rather than a nice-to-have layered on top. NIST's SP 800-161r1, the cybersecurity supply chain risk management guidance, treats ongoing monitoring as core to the requirement, not a supplemental control bolted on afterward. CIRCIA, with a final rule expected in September 2026, adds a 72-hour incident reporting clock, and an organization cannot report what it cannot see; that clock structurally demands real-time detection capability, not an annual check-in.

Auditors and examiners have followed the same logic. Scrutiny lands precisely in the gaps between assessment cycles, the exact periods a purely annual program has zero visibility into. Only 22% of organizations report having fully defined, operational metrics for measuring their TPRM program, Venminder data shows, which is a problem heading into examination season for a lot of programs. Most programs can't currently demonstrate the continuous oversight regulators are already asking for.

Building the connective tissue: how findings should flow between monitoring and assessment in practice

The design principle is simple to state and harder to build: a monitoring finding needs an automated path into four places, the vendor's risk score, the assessment record, the contract compliance status, and the business owner's escalation queue. Skip any one of those paths and monitoring and assessment stay two parallel tracks that happen to share a vendor name, exactly the failure mode MOVEit illustrated.

Not every signal deserves a full reassessment, and treating them all the same just recreates alert fatigue from a different angle. What the program needs are predefined thresholds that separate three outcomes cleanly: log it and keep watching, reassess the vendor now, or escalate straight to the business owner. Where that threshold logic doesn't exist yet, the monitoring tool and the assessment record will keep producing accurate information side by side, and the organization will keep discovering, after the fact, that neither one ever told the other what it knew.

Sources

  1. Continuous Vendor Monitoring Beats the Annual TPRM Audit
  2. Continuous Vendor Monitoring
  3. Article Not Found | Truvara
  4. underdefense.com
  5. labs.cloudsecurityalliance.org

More in Vendor Monitoring