Back to Articles and Learning
Risk Management12 min read

Control Design vs Operating Effectiveness

Elliot Poublan
Oct 2, 2026
Control Design vs Operating Effectiveness

It is March. Your internal audit lead walks you through a finding on the payment release control. The control failed in the period. Four payments above the approval threshold went out without the second signature.

You pull the file. The control was tested in October and it passed. The tester sat with the payments manager, watched one payment go through the dual-approval screen, took a screenshot, wrote it up, and marked the control effective.

Both things are true. The control was well designed and it did not operate. Those are two different questions, tested two different ways, and confusing them is the most common reason a control testing programme gives a board false comfort. Control testing 101 covers how to run the programme. This piece is the distinction that programme rests on.

In short

Design effectiveness asks whether the control would address the risk if it operated exactly as described. Operating effectiveness asks whether it actually operated, consistently, across the whole period. You test design when a control is new or has changed. You test operation on a sample drawn from across the period.

  • Design is the control on paper and in concept. Is there a control, is it at the right point, can the person performing it actually do it, and would it catch the error it is meant to catch?
  • Operation is the control in the period. Did it run every time it was supposed to, by the right person, with a record left behind?
  • A walkthrough proves design. It is a single observation. It cannot prove operation across twelve months.
  • A control can pass design and fail operation. If it fails design, an operating pass only proves that an inadequate control ran consistently.
  • Provision 29 asks boards of UK premium-listed companies to declare on material controls. That declaration rests on operating evidence, not on process maps.

What Is a Design Effectiveness Test?

Design effectiveness is whether a control, performed as written, would prevent or detect the risk it is assigned to within a useful timeframe.

A design test is a thinking exercise supported by evidence. You are not counting anything. You are asking whether the control makes sense. In practice you check six things:

  1. The control exists and is written down. A control that lives only in someone's head is not testable and is not evidenced.
  2. It is aimed at a real risk. The risk statement and the control statement have to line up. A monthly reconciliation does not address an access risk. The types of control, and how they map to risk, are covered in control frameworks explained.
  3. It sits at the right point. A check performed after funds have left does not prevent the loss. It may still be a valid detective control, but it should be labelled as one.
  4. The performer has authority, information and independence. A second approver who cannot see the supporting documentation is performing a formality.
  5. The threshold and frequency are defensible. A control that triggers above £250,000 when 90% of payments fall below that leaves most of the risk untouched.
  6. It produces evidence. If performing the control leaves no trace, you can never test its operation. Design work is where you fix that, not a year later.

The usual evidence for a design test is a walkthrough. You take one transaction and follow it end to end with the person who performs the control, confirming that what happens matches the documentation. You also read the policy, the system configuration and the delegated authority.

You test design when a control is new, when a process or system changes, when the risk changes, and at least annually for material controls. You do not need to test design every month for a control that has not moved.

What Is an Operating Effectiveness Test?

Operating effectiveness is whether the control actually operated as designed, consistently, throughout a defined period.

This test counts things. You define the population, draw a sample across the period, and inspect evidence for each item.

  1. Define the period. Usually the financial year, or the period since the last test.
  2. Define the population. Every occurrence of the control. For a daily control that is around 250 items. For a monthly control it is 12.
  3. Establish the population is complete. This is the step most often skipped. A sample from an incomplete list tells you nothing. Pull the population from the source system, not from the spreadsheet the control owner maintains.
  4. Select a sample across the period. Not the first five, and not the five the owner sends you. Spread it, so a control that stopped working in August is visible.
  5. Inspect the evidence for each item. Approval recorded, by whom, on what date, against what supporting information.
  6. Record and classify exceptions. One missing approval out of twenty-five is not a rounding error. It is an exception that has to be explained, quantified, and either cleared or reported.

Sample sizes vary by frequency and by the assurance you need. A common starting point for a manual control is around 25 items for a daily control, 5 for a monthly one, and 1 for an annual one, with larger samples where the control is material or where exceptions have been found before. Automated controls are different: if you can prove the configuration did not change and that general IT controls are sound, a small sample can support a conclusion for the whole period. If you cannot, treat the automated control like a manual one.

Design vs Operating Effectiveness Compared

Aspect Design effectiveness Operating effectiveness
Question answered Would this control address the risk if performed as described? Did this control actually run as described across the period?
Unit of testing The control itself Individual occurrences of the control
Typical method Walkthrough, inspection of policy, system configuration review, enquiry Sampling, inspection of evidence, re-performance, system logs
Sample size One transaction, traced end to end Sized by frequency, materiality and prior exceptions
Timing On creation, on change, and at least annually for material controls Across the period, often at interim and year end
Evidence produced Process narrative, flowchart, walkthrough note, authority matrix Approval records, signed reconciliations, system logs, exception logs
Typical failure No control at the right point, threshold too high, no evidence produced Control skipped, performed late, performed by the wrong person, no record kept
What a failure means The control needs redesign before operating testing is worth doing The control needs remediation and the exposure needs quantifying
Supports Framework assertions, new process sign-off Board declarations, attestations, external audit reliance

The last row is the one that changes what you can say. A design conclusion supports a statement that you have a control framework. An operating conclusion supports a statement that your controls were effective. Those are not the same claim, and Provision 29 asks for the second one.

Worked Example: One Payment Approval Control, Tested Both Ways

The same kind of control, tested properly, looks like this. A UK-authorised payments firm with 40 staff runs client money through a single banking partner. One control carries the process: any outbound payment above £10,000 requires approval by a second authorised individual before release.

Design test. The tester reads the payments policy, pulls the delegated authority matrix, and checks the banking platform configuration. They walk one payment through with the operations manager. The control is at the right point, before release rather than after. The second approver can see the invoice and the client instruction from the approval screen. The system enforces the threshold, so it cannot be bypassed by choosing not to seek approval. One weakness surfaces: four people sit in the approver group, and two of them also initiate payments. The system does not prevent the same person doing both. The tester records design as effective with an exception, and recommends a system rule blocking self-approval.

Operating test. The tester asks the banking platform for a full extract of outbound payments for the year, not a report from the operations spreadsheet. The population is 1,412 payments, of which 318 are above £10,000. They select 25 across all twelve months. Twenty-three show a second approval by a different user, timestamped before release. Two, both in August, show the initiator and the approver as the same user ID. The firm's second approver was on leave and the platform rule had been relaxed to keep payments moving.

Design passed with a noted weakness. Operation failed for a defined window. The two findings are connected. The design test identified the possibility of self-approval. The operating test proved it happened, told you when, and let the firm quantify it: 41 payments released during that fortnight, all of which needed retrospective review. One test told the firm what could go wrong. The other told the board what did.

The Walkthrough That Gets Filed as Operating Evidence

Common Pitfall

One screenshot, marked effective

The most frequent weakness in a first-year testing programme is a file whose only evidence is a walkthrough note and one screenshot, with a conclusion that reads "control operating effectively". A walkthrough is one observation, usually of a transaction the control owner selected, usually performed while someone is watching. It proves the control can work. It says nothing about the other occurrences, and nothing about the two weeks the control was switched off.

Three related traps sit alongside it:

  • Sampling from a list the control owner produced. If the owner builds the population, the items that never hit the control are the ones most likely to be missing from it. Trace the population back to a source system.
  • Testing only the most recent month. The evidence is to hand. It is also the period when everyone knew testing was coming.
  • Treating enquiry as evidence. "The manager confirmed the reconciliation is performed monthly" is a statement about belief. Inspect the signed reconciliations.

The correction is straightforward. Label every control in the plan with which test it needs, when, and what counts as evidence, and do it before testing starts rather than when the file is reviewed.

Which Test Do You Need for Provision 29?

You need both, in sequence, and the conclusion you report is the operating one.

Provision 29 of the UK Corporate Governance Code asks boards of premium-listed companies to declare whether material controls operated effectively. It applies to accounting periods beginning on or after 1 January 2026, so firms in scope are building the evidence now for a declaration that lands in the annual report afterwards. Firms outside the Code often adopt the same discipline, because SM&CR accountability, Consumer Duty outcomes monitoring and DORA operational resilience testing all pull the same way: a named person, a defined control, and evidence that it ran. How to decide which controls are material is covered in how to identify material internal controls under Provision 29.

The sequence matters. Test design first. If a control is not designed to catch the risk, sampling 25 instances of it working perfectly proves only that an ineffective control was performed diligently. Fix the design, then test operation, then report the operating conclusion with design weaknesses noted alongside.

How Initia Risk Handles This

Initia Risk keeps design and operating tests as separate records against the same control, not one field called "tested". Each control carries its design assessment, its test frequency, its sample basis and its evidence requirement, so a tester opening the control knows which question they are answering, and a reviewer can see which question was answered.

Evidence attaches to the individual test item, not to the control as a whole. A sample of 25 leaves 25 evidence records, with dates, performers and any exceptions raised. The audit trail shows who tested, who reviewed, and when the conclusion changed. Exceptions become tracked items with owners. When a control fails, the risk it maps to, the regulatory obligations mapped alongside it, and the committee reporting all update from the same record.

Controls sit with first-line owners, with second-line oversight visible on top rather than copied into a parallel file. Attestation cycles run from the register, so a control owner confirms operation and attaches evidence in the same place the testing plan lives. Regulatory mapping links controls to the obligations they support, which is what turns a testing plan into a position you can defend when a board or a regulator asks which controls were material and how you know they worked. Board and committee reporting draws from the live register, so the pack reflects the testing position on the day it is produced.

Davies chose Initia Risk as its risk and compliance platform. Across the firms we work with, the pattern is the one in the payments example: the design work is usually reasonable, and the gap is operating evidence that survives scrutiny.

Key Takeaways

  • Design effectiveness asks whether the control would work as described. Operating effectiveness asks whether it did work across the period.
  • A walkthrough tests design. It is a single observation and cannot support an operating conclusion.
  • Test design when a control is new or changed, and at least annually for material controls. Test operation on a sample spread across the whole period.
  • Prove the population is complete before you sample from it, and pull it from a source system rather than from a list the control owner maintains.
  • A control can pass design and fail operation. If it fails design, fix the design before spending effort on operating tests.
  • Provision 29 declarations, SM&CR accountability and Consumer Duty outcomes monitoring all rest on operating evidence, not on process documentation.
  • Label every control in the plan with which test it needs, when, and what counts as evidence, before testing starts.

Frequently asked questions

Can a control pass design testing and still fail operating testing?
Yes, and it is the most common outcome in a first full testing cycle. Design testing confirms the control is capable of addressing the risk: the logic, placement, authority and evidence trail hold up. Operating testing then checks whether it ran every time it should have. That is where absences, workarounds, cover during leave and system-rule changes show up. The reverse is not useful. If a control fails design, an operating test only tells you that an inadequate control was performed consistently, which does not reduce the risk.
How large should an operating effectiveness sample be?
Sample size depends on how often the control operates, how material it is, and whether exceptions have been found before. A common starting point for manual controls is around 25 items for a daily control, about five for a monthly control, and a single item for an annual one, increased where the control is material or a prior test raised exceptions. Automated controls are different. If you can show the configuration did not change during the period and that general IT controls over change and access were sound, a small sample can support a conclusion for the full period. Without that evidence, sample an automated control as you would a manual one.
How often do you need to retest control design?
Retest design when the control, the process, the system or the underlying risk changes, and at least annually for controls you have identified as material. Between those points the design conclusion carries forward, because the control has not moved. The trigger that gets missed is a system upgrade or a reorganisation, both of which can quietly move a control or change who performs it. Tie design retesting to the change process so it happens because something changed, not because someone remembered.
Does a design failure need to be reported even if the control operated?
Yes. A design weakness means the risk was not fully covered even when the control performed exactly as written, so a clean operating sample does not close the exposure. Report it, quantify the residual risk, and track the remediation. Self-approval that is possible by design is still a finding in the months nobody used it, because it is what makes the later operating failure predictable.
What evidence is enough to conclude a control operated effectively?
Enough for an independent reviewer to reach the same conclusion without speaking to anyone. For each sampled item that usually means a record of what was checked, who performed the control, when they performed it, and what they did with the result, drawn from a system or a signed document rather than a verbal confirmation. If a control does not produce that trace, the fix belongs in design: change the control so that performing it creates a record, then test operation from the point the change took effect.

See Initia in action

Book a demo.

An exploratory call to discuss what works and what doesn't, what's still done on Excel, and what you're looking for in a tool.

By submitting, you agree to allow Initia Risk to store and process your personal data.

Initia Risk
“As a professional risk manager with over 16 years' experience, I have seen many systems over the years. Initia Risk is without doubt one of the most user-friendly and intuitive platforms I have encountered.”
Elaine Atkinson · Risk Manager