Observation Window Evidence Gaps and How Auditors Handle Them
Auditors use sampling and re-performance to navigate evidence gaps in Type 2 reports.

A SOC 2 Type 2 report asks for something much harder to prove than a Type 1: that controls existed, and that they ran, consistently, for months on end. That gap between "we built this" and "we did this every time" produces evidence gaps as a routine feature of the audit, not a scandal. This piece walks through why those gaps form, how auditors test around them, and what separates a clean opinion with a footnote from a qualified one that makes procurement teams nervous.
Start with the basic distinction, because it explains everything downstream. Type 1 asks whether the control exists and is designed correctly, as of one date. Type 2 requires evidence stitched together across an observation window that can run anywhere from a few months to a full year to show consistent operation over time. PCAOB AS 1105 makes a point worth sitting with: observation as a technique only ever captures a single moment. An auditor watching someone run a backup on a Tuesday in March has learned something narrow and time-bound, nothing more, about what happened in June. The AICPA's Trust Services Criteria then raise the stakes by demanding proof of consistent operation on top of proof of setup. So the entire evidentiary burden shifts from "show me the policy" to "show me it fired, every time, for months," and that shift is exactly where gaps live.
How observation window length shapes the size and character of evidence gaps
Three months is the floor. The AICPA recommends at least six. Enterprise buyers typically want twelve. Organizations going through their first audit often pick the shorter window because it gets a report into a prospect's hands faster; mature organizations standardize on twelve months because it reads as more credible and it captures periodic controls, like quarterly access reviews, that a three-month window might simply miss entirely.
The AICPA does not hand down a specific number of months as gospel. What it requires is a period long enough to actually evaluate operating effectiveness. If a control only fires once a quarter, a ten-week window will not catch it firing twice.
Here's the part that catches people off guard. Many organizations run several months of their window before they've even engaged an auditor, arriving at fieldwork with a backlog of evidence already sitting there, waiting to be reviewed. That's efficient, in theory, but it also means the clock started ticking before anyone was checking whether the evidence was any good.
And the longer the window runs, the more chances a control has to quietly fall apart. A gap that gets closed in January can reopen in July with nobody noticing, because nobody was watching. Window length, in other words, is a decision with ongoing consequences, not a one-time scheduling choice. It determines how many times a control has to fire, which periodic reviews land inside the boundary, and how many sampling checkpoints the auditor will need to string together to call the period covered.
The specific ways evidence gaps form during a live observation period
Most gaps aren't dramatic. Nobody skipped the work; they just didn't keep the receipt.
That's the single most common failure mode: execution-without-retention. The access review happened. The training got completed. But there's no artifact proving it, because nobody thought to screenshot it, log it, or file the ticket. From an auditor's chair, unrecorded work and unperformed work look identical, which is an uncomfortable fact for anyone who assumed good intentions would carry the day.
Close behind that is the artifact-control mismatch: something exists, a log file, a ticket, a screenshot, but it doesn't actually demonstrate the control as written. A Slack message saying "reviewed access, looks fine" falls short of a documented review with named approvers and a decision trail, and auditors treat the former as insufficient. Timing mismatches follow a similar pattern: a control owner insists the activity happened, but there's no way to tie the record to the defined period. Undated screenshots leave auditors nothing to check the claim against, and they get flagged as exceptions, not because anyone's lying, but because there's no way to verify the timing.
The window-start gap deserves its own mention because it's so mechanically predictable. Observation period opens January 1. Quarterly access reviews don't get formalized until mid-February. That's a six-week hole sitting right at the front of the report, baked in before the audit even starts. Not coincidentally, a skipped or delayed access review is the single most frequent gap type auditors run into.
It helps to split gaps into two families. Design gaps are policy problems: "access is restricted to authorized users" sounds fine until you ask who approves access, who reviews it, and who revokes it, and the policy has no answer. Operational gaps are execution problems: most employees finished security training, but not all; the incident response plan exists but has never been tested; access reviews happen, just not on any predictable schedule.
Layer control drift on top of all this, the phenomenon where a control operating fine in month one degrades by month six because nobody's watching it anymore, and it becomes clear why drift is the leading driver of exceptions in Type 2 exams specifically. Industry estimates put the gap rate for first-time SOC 2 pursuers somewhere in the 40 to 60 percent range. That should reframe how anyone new to this process thinks about gaps: they're the expected condition of a first audit, not evidence that something went wrong.
The four testing techniques auditors use to assess what actually happened during the period
Auditors lean on four techniques: inquiry, observation, inspection, and re-performance. None of them stands alone; they're used in combination to corroborate the same activity from different angles.
Inquiry, meaning simply asking someone what happened, is the weakest of the four on its own. PCAOB AS 1105 says so explicitly: inquiry of company personnel, by itself, does not reduce audit risk to an appropriately low level. Every control owner believes their control works, and that reflects confidence rather than proof, so it doesn't hold up in a report on its own.
Re-performance sits at the other end of the spectrum. The auditor actually runs the control themselves, restoring a backup, for instance, to see if it works as described. It's the strongest form of evidence available for automated or technical controls because there's no interpretation gap between "someone says it works" and "it worked, right there, on the auditor's screen."
Sampling is the workhorse that spans the whole period. Nobody reviews every record across twelve months; that would turn a six-week engagement into a six-month one. Instead, auditors sample under one of two governing standards, AU-C Section 530 for non-issuers under AICPA rules, AS 2315 under PCAOB rules for public companies and broker-dealers. For tests of controls, the sampling question is narrower than it sounds: did this control run consistently, a question distinct from whether some account balance is accurate to the penny. That's a different objective from substantive testing, and it's worth keeping the two separate in your head, because they get invoked at different points in the process (more on that in the escalation section below). Sampling shows up most often for manual controls tested by inspection or re-performance, less often for automated, entity-level controls that either fire or don't.
Some auditors also run interim procedures, checking evidence during the window itself rather than waiting until the period closes. That's a meaningful safeguard, not a minor scheduling preference. It surfaces gaps while there's still time to fix them, instead of discovering a six-week hole in month eleven with nothing left to do but write it up.
The standard threading through all four techniques is "sufficient appropriate audit evidence." It's a deliberately flexible phrase, and that flexibility is the whole point: the bar has to be met even when the evidence trail has holes in it, which is most of the time.
Expanding the sample when a deviation surfaces
Finding one bad record in a sample doesn't end the conversation; it starts a bigger one.
Say an auditor pulls 25 employee records to check security awareness training completion, and one person didn't finish within the required window. Standard practice calls for a measured next step, rather than a shrug or an automatic failure. The auditor expands the sample, pulling another 15 records, to figure out whether that one miss was a fluke or a pattern. If the expanded sample comes back clean, the single deviation gets documented as an exception against the relevant criteria (CC2.2, in this example) and life goes on. The opinion isn't disqualified over one person who forgot to click through a training module.
That's the diagnostic function of expanded sampling: separate isolated lapses from systemic ones. One miss out of 40 looks like a person who was on vacation the week reminders went out. Five misses out of 40 looks like a training program with no enforcement mechanism behind it, which is a different problem entirely and raises the odds of a qualified opinion.
Worth noting: the auditor's job here centers on characterizing the size and shape of the problem, not on finding someone to blame. The deviation, the reasoning behind expanding the sample, and the result of that expansion all get written into the auditor's working papers, and where it's material enough, into the report itself.
What compensating controls can and cannot do for an evidence gap
A compensating control is a different mechanism that addresses the same underlying risk, not a note explaining why the real control didn't happen. That distinction matters more than almost anything else in this whole process.
Say quarterly access reviews slipped for six weeks at the start of the observation window. A memo saying "we were busy migrating systems" doesn't qualify as a compensating control. What might qualify is evidence that access requests were still individually approved and logged during that stretch, that no new access was granted without sign-off, and that the risk the quarterly review exists to catch, namely unauthorized or stale access, was still being managed through a different mechanism the whole time.
PCI DSS offers a useful precedent here, since it allows compensating controls explicitly but attaches real conditions: document why the standard control wasn't feasible, prove the substitute offers comparable protection, and show it actually works through regular testing. SOC 2 auditors apply a similar logic. A compensating control can support an unqualified opinion, but only if the auditor is actually satisfied it meets the intent of the relevant Trust Services Criteria, and that satisfaction has to be earned with evidence, not assumed because the organization says so.
What does that evidence look like? A documented description tying the compensating control back to the original requirement, monitoring records, system logs, or internal reports showing the substitute actually ran consistently, not just once as a demonstration, plus proof that someone reviews it regularly to confirm it's still functioning as intended.
Where this goes sideways: organizations treat compensating controls as an escape hatch, slapping together a weak substitute after the fact and asserting equivalence without backing it up. Auditors have gotten sharper about spotting this pattern, and they scrutinize these justifications harder than they used to. The burden of proof sits entirely with the organization, and it takes both technical rigor and an honest read of the actual risk, not a document written to sound convincing.
When neither sampling expansion nor compensating controls resolve the gap: escalation to substantive testing
Sometimes the compensating control doesn't hold up under scrutiny. When that happens, the auditor doesn't just note the deficiency and move on; they reassess the risk, look for alternative procedures, and flag the finding to stakeholders.
The fallback is substantive testing, a meaningfully different kind of exercise from testing controls. Instead of checking whether a control ran consistently, the auditor goes directly at the underlying transactions and data, reaching conclusions independent of whatever the control was supposed to guarantee. Even here, control design still gets examined; it just stops being the thing the opinion leans on.
Operationally, this is where things get expensive. Substantive testing takes more time from the auditor and more effort from the organization, which often has to produce volumes of underlying records it never organized with an audit in mind, and fieldwork expands accordingly. And this is the point where the conversation changes tone: it moves from how to document a gap toward what the gap means for the opinion itself, which is exactly where the next section picks up.
How auditors decide between a clean opinion with disclosed exceptions and a qualified opinion
An exception is a specific, testable instance where a control didn't operate as described, tied to a named control and a named criterion, not yet a verdict on the opinion. What happens after that fact gets recorded is where judgment enters the picture.
Exceptions don't automatically tank the opinion. The auditor weighs severity, how often it happened, and whether a compensating control adequately covers the residual risk. If the deviation was isolated and a compensating control genuinely addresses the exposure, the auditor can issue an unqualified opinion while still disclosing the exception transparently in the report. That's actually the most common outcome for organizations running a solid, if imperfect, control environment; a reasonably well-run compliance program isn't the same thing as a flawless one, and the report format has room for that reality.
A qualified opinion enters the picture when exceptions, alone or stacked together, are material enough that the organization can't be said to have met one or more of its service commitments under the Trust Services Criteria. PCAOB AS 3105 governs this and requires the report to spell out exactly what scope limitation or evidence shortfall triggered the qualification.
None of this is formulaic. Materiality is a judgment call, weighing whether the control was preventative or merely detective, how often it lapsed, and how sensitive the data or process behind it actually was. A skipped training module for one employee sits in a different category of problem than a customer database sitting exposed for six weeks. The opinion isn't a pass-fail switch; it's a spectrum, and where a given gap lands on it depends heavily on how well the organization documented its compensating measures and how honestly it communicated its risk posture throughout the engagement, not just at the end when the report gets drafted.
Adjusting the window boundaries to avoid embedding avoidable gaps
Here's a fix that costs almost nothing and gets skipped constantly: don't start the clock before the controls are actually running.
Auditors routinely advise setting the observation window start date only once controls are fully operational. Opening the window early bakes a known gap into the report from day one. Delaying the start date by a few weeks to sidestep a six-week gap at the front of the period is a far better trade than opening early and explaining the hole later. A short delay reads as prudent, while a documented gap covering the first month of the period reads as something else entirely.
The start date carries formal weight, too, rather than functioning as an informal internal choice. It gets documented formally in compliance records, and auditors reference it throughout fieldwork as a fixed anchor point. For organizations that run several months of their window before an auditor ever gets involved, this creates an obligation worth sitting with: evidence discipline needs to start on day one of the defined period, not on the day the auditor's calendar invite lands.
This is also where continuous evidence collection tools earn their keep, systems that capture proof throughout the window rather than reconstructing it at the end. The value lies in what the tool prevents: the situation where the work got done, but the proof it happened evaporated somewhere along the way. That's the operational argument for treating compliance as an ongoing workflow instead of a scramble assembled the week before fieldwork starts.
What organizations can do during the observation period to keep gaps from compounding
Treat the observation window as an active job, not a waiting room. That's the whole discipline, really. Control drift is the main driver of exceptions, and drift is fundamentally a monitoring failure, not proof the control itself was ever badly designed.
Periodic controls need a calendar and a named owner, full stop. Access reviews, training completions, vendor management checks: these need to be scheduled with someone accountable for both doing the work and capturing proof it happened. A control that fires but leaves nothing retrievable behind it is, as far as an auditor is concerned, indistinguishable from a control that never fired at all.
Letting the auditor run interim procedures during the window, rather than dumping everything on them at the end, surfaces problems while there's still runway to fix them. Nobody wants to discover a gap in month eleven of a twelve-month window.
Artifact quality beats artifact volume, every time. A handful of clean, clearly labeled records tied directly to the control as written will do more for an auditor than a mountain of loosely related logs that require guesswork to interpret. The artifact-control mismatch problem gets solved with specificity, not with more paper.
And when a gap does happen, and at some point one will, the move is to document it right away, figure out whether a compensating control is available, and tell the auditor before fieldwork forces the conversation. Auditors respond better to organizations that flag problems early and bring a remediation plan than to ones that get caught flat-footed at fieldwork with a gap and no explanation.
Evidence gaps are a predictable feature of asking an organization to prove, across months, that something happened every single time it was supposed to, rather than a sign the audit went wrong. The organizations that come through without report damage aren't the ones with flawless controls; they're the ones that treated the observation window as ongoing work, not as a countdown clock they'd deal with once it ran out.


