Back to Articles and Learning
Risk Management11 min read

Control Testing 101: How to Know if Your Controls Really Work

Elliot Poublan
Aug 19, 2026
Control Testing 101: How to Know if Your Controls Really Work

Most control libraries look complete. Owners are named. Frequencies are filled in. Residual risk is lower than inherent risk, so the story to the board is that controls are working. Then an incident happens, or an auditor samples ten payments, and it turns out the second authoriser was the same person using two logins.

A control on a register is a claim. Control testing is how you check the claim. Without it, residual risk is an opinion, RCSA is a calendar exercise, and the board pack is a comfort document.

In short

Control testing answers two separate questions: is this control designed to mitigate the risk, and did it operate as intended over the period? Both must pass before you can defend the residual score.

  • Design first - a faithfully operated bad control is still a bad control.
  • Evidence, not attestation - "we always do this" is not a test result.
  • Key controls get the effort - do not spread sampling thinly across hundreds of minor checks.
  • Fails must move residual risk - otherwise testing is theatre.

Why Untested Controls Are Not Assurance

In inherent vs residual vs appetite, residual risk is the position after controls. That sentence only holds if the controls are real. Listing them is not the same as knowing they ran.

This is also why most RCSA programmes fail. The first line rates controls as effective because the process exists on paper. Nobody samples. Nobody asks for the log. Residual stays green. The register and the operational reality drift apart until something breaks in public.

The Two Tests You Must Run Separately

If you take one thing from this article, take this. Design and operation are not a single RAG rating. They fail in different ways and they need different evidence. The fuller taxonomy of control types sits in our guide to control frameworks; this piece is about how you prove them.

Test Question Typical evidence
Design effectiveness If it ran as written, would it mitigate the risk? Process walkthrough, system configuration, control description vs risk, gap analysis
Operating effectiveness Did it actually run, as intended, across the period? Sample of transactions, system logs, reconciliations, exception follow-up

A payment-approval rule that is not enforced in the system fails design, even if people say they double-check. A rule that is enforced in the system but bypassed via shared credentials fails operation. Both look like "we have a dual-approval control" on the register. Only testing tells you which story is true.

How to Test: Methods That Actually Count

Not every method carries the same weight. Rank them by how much they prove:

  • Inquiry - asking the owner "does this control operate?" Weakest. Useful for context, never sufficient on its own for a key control.
  • Observation - watching the control being performed. Better for understanding design. One observation does not prove a year of operation.
  • Inspection - reviewing the evidence the control produces (approvals, logs, signed packs). The workhorse method for operating tests.
  • Reperformance - doing the control yourself on a sample and comparing to what the owner did. Strongest for reconciliations and calculations.
  • System interrogation - pulling configuration and audit logs from the application. Best for automated preventive controls, because the system either blocked the action or it did not.
Common Pitfall

Attestation dressed up as testing

A quarterly email that says "please confirm your controls operated" and a Yes/No reply is not a control test. It is an attestation. Attestations are cheap and they produce a reassuring spreadsheet. They do not tell you whether the second authoriser was independent, whether the reconciliation was completed, or whether exceptions were cleared. If your "testing programme" is a round of confirmations, you do not yet have a testing programme.

Sampling: Enough to Support a Conclusion

You cannot test every transaction. You also cannot test one item in June and call the year effective. The sample has to cover the period you are opining on, and it has to be large enough that a clean result is meaningful.

Practical rules for a mid-market firm:

  • Key automated controls - confirm configuration once, then sample exceptions or bypasses. If the system cannot be bypassed, a configuration test plus a small operating sample is often enough.
  • Key manual controls that run often - sample across the period (for example 25 items across four quarters), not a cluster in one month.
  • Infrequent controls - if the control runs monthly, test several months, not one. If it runs quarterly, test every instance in the year.
  • One exception is a finding - do not average it away. Investigate whether it is isolated or a symptom. Expand the sample if the first fail looks systemic.

SOX shops will have statistical sampling tables. Most mid-market GRC programmes do not need that precision. They do need a written rationale: why this size, why this period, why these items. If you cannot explain the sample, you cannot defend the conclusion.

What Evidence Counts

Good evidence is contemporaneous, attributable, and specific to the control objective. A screenshot of a policy is not evidence the policy was followed. A training completion list is not evidence the trained person performed the control. A committee pack that mentions "controls were reviewed" is not evidence unless the pack shows what was reviewed and what was concluded.

Store evidence against the control and the test, not in a shared drive labelled "audit 2026". When internal audit or a regulator asks, you should be able to open the control, see the last test, and open the sample pack without a treasure hunt. That is also what material controls under Provision 29 will demand: not a narrative that monitoring happened, but the artefacts that prove it.

Cadence: When to Test

Write the frequency on the control, then stick to it. A sensible default:

  • Material / key controls on high residual or high inherent risks - at least annually, often quarterly for high-volume processes.
  • After change - new system, new owner, new third party, or a related incident should trigger an out-of-cycle test.
  • After a fail - retest once the action is closed, not in twelve months' time.
  • Non-key controls - rotate. You do not need to test everything every year, but you do need a plan that cycles through the library.

Year-end scrambles are a symptom, not a methodology. If testing only happens when the external auditor arrives, the first line has already learned that controls are a December problem.

Worked Example: Dual Approval on Payments

Control: no payment over £10,000 is released without a second authoriser independent of the raiser. Type: preventive, intended to be automated. Risk: fraudulent or erroneous large payment.

Design test: walk through the payment system. Confirm the threshold is configured at £10,000, that the system blocks single-approver release, and that the second authoriser cannot be the same user ID. If anyone can disable the rule without a change ticket, design is already weak.

Operating test: extract all payments over £10,000 in the quarter. Sample 25. For each, confirm two distinct users in the approval log, neither of whom is a shared mailbox. Investigate any payment that posted without two IDs, any same-day override, and any "emergency" bypass.

If it fails: rate operating effectiveness as ineffective or partially effective. Open an action with an owner and a date (for example: remove shared credentials, lock the override path). Raise residual likelihood on the payment-fraud risk until the retest passes. Report the exception in the next risk committee pack, not in a footnote of the RCSA spreadsheet.

What to Do With a Fail

A fail that does not change anything is worse than not testing. The minimum loop:

  1. Rate design and operation separately - do not bury a fail inside an overall "amber".
  2. Log the exception - what failed, in which sample items, and why.
  3. Assign an action - named owner, due date, and what "done" looks like.
  4. Recalculate residual risk - if this was a key control, the residual position should move unless a compensating control is tested and holds.
  5. Retest - closure of the action is not the same as proving the control now operates.

This is also how you make risk ownership real. The owner is not the person whose name sits on the register. The owner is the person who has to explain the fail and fix it.

How Initia Risk Makes Testing Usable

Testing dies in spreadsheets because the test, the evidence, the action and the residual score live in four different files. Initia Risk keeps them on the same object: each control has an owner and a cadence, design and operating effectiveness are separate assessments, evidence sits on the test, fails open actions, and residual risk on the linked risk can reflect how the control actually performed.

That is the difference between a control library and an assurance process. For the assessment cycle that sits around this, see how to run an RCSA. For the scoring layer, see inherent vs residual vs appetite. For one-line definitions - control, key control, design effectiveness, operating effectiveness - see control and related terms in the glossary.

See Initia in action

Book a demo.

An exploratory call to discuss what works and what doesn't, what's still done on Excel, and what you're looking for in a tool.

By submitting, you agree to allow Initia Risk to store and process your personal data.

Initia Risk
As a professional risk manager with over 16 years' experience, I have seen many systems over the years. Initia Risk is without doubt one of the most user-friendly and intuitive platforms I have encountered.
Elaine Atkinson · Risk Manager