Operational Risk

Operational risk is the possibility of loss or disruption caused by failed people, processes, systems, third parties, or external events.

Operational risk is the possibility of financial loss, customer harm, or business disruption caused by inadequate or failed people, processes, systems, third parties, or external events. Examples include payment-processing errors, fraud, cyber incidents, system outages, vendor failures, legal-document defects, and breakdowns in transaction controls.

The Basel banking definition includes legal risk but excludes strategic and reputational risk. Other industries and organizations may use a broader scope, so an operational-risk report should state which taxonomy, legal entity, process, and loss types it covers.

Key Takeaways

  • Operational risk arises from how an activity is performed, not primarily from market prices or borrower credit quality.
  • A process can operate normally for years and still contain a severe control weakness or single point of failure.
  • Loss-event data is useful but backward-looking; process mapping, scenarios, control testing, and key risk indicators add forward-looking evidence.
  • Outsourcing an activity does not eliminate the operational risk or the organization’s responsibility for managing it.
  • Operational resilience is related but different: it focuses on continuing or restoring critical operations through disruption.
  • A useful assessment identifies the process owner, failure path, affected customers or cash flows, controls, residual exposure, escalation trigger, and recovery plan.
RiskMain sourceExample
Operational riskFailed people, processes, systems, third parties, or external eventsA software release duplicates customer payments
Business riskDemand, competition, pricing, costs, or strategyA product loses market share and margins contract
Market riskPrices, rates, spreads, currencies, or volatilityA bond portfolio loses value when yields rise
Credit riskBorrower or counterparty failureA borrower stops making required payments
Model riskIncorrect model design, inputs, implementation, or useA valuation model applies stale volatility data
Conduct riskProducts, incentives, behavior, or controls that create harmful outcomesA sales process rewards unsuitable recommendations
Reputational riskLoss of stakeholder confidenceCustomers leave after repeated service failures

These risks can occur together. A system outage is an operational event, but it can also create liquidity needs, customer remediation, legal costs, conduct concerns, and reputational damage.

Is “Operating Risk” the Same Term?

“Operating risk” is used inconsistently. Some writers use it as a synonym for operational risk. Others use it for variability in a company’s operating earnings, which is closer to business risk. This site uses operational risk for failures involving people, processes, systems, third parties, or external events and uses business risk for demand, pricing, competition, strategy, and cost-structure exposure.

Common Sources of Operational Risk

People

Human error, inadequate training, misconduct, staffing gaps, poor segregation of duties, and unclear accountability can create losses. Automation changes the failure mode but does not remove the need for ownership and review.

Processes

Manual handoffs, incomplete reconciliations, weak approvals, incorrect data, missing documentation, and poorly designed exception handling can cause errors to accumulate or remain undetected.

Systems and Data

Software defects, cyber incidents, outages, capacity constraints, access-control failures, data corruption, and unsuccessful change implementation can interrupt critical activity or produce incorrect records.

Third Parties

Cloud providers, payment processors, custodians, administrators, data vendors, and other service providers can create dependency and concentration risk. Contracts, service-level terms, audit rights, substitution options, and exit plans affect the residual exposure.

External Events

Natural disasters, utility failures, civil disruption, public-health events, and other external shocks can interrupt premises, staff, communications, suppliers, or infrastructure.

From Process Map to Risk Decision

Operational-risk lifecycle from mapping a critical activity through failure scenarios, control assessment, monitoring, response, recovery, and lessons learned.

A practical operational-risk review follows an evidence trail:

  1. Define the product, service, process, legal entity, customer group, and time horizon.
  2. Map critical steps, systems, data, people, premises, and third-party dependencies.
  3. Identify failure scenarios and the resulting financial, customer, legal, and operational consequences.
  4. Estimate inherent risk before controls.
  5. Test preventive, detective, corrective, and recovery controls.
  6. Evaluate residual risk against risk appetite, limits, and tolerance for disruption.
  7. Monitor indicators, incidents, control exceptions, and changes in the operating environment.
  8. Respond, recover, reconcile records, remediate harm, and incorporate lessons into the control framework.

The sequence is not a universal regulatory formula. The required evidence and approvals depend on the organization, industry, jurisdiction, and materiality of the activity.

How Operational Risk Is Assessed

Operational risk cannot be reduced to one reliable number. Organizations commonly combine several methods:

MethodWhat it contributesMain limitation
Loss-event dataFrequency, severity, causes, and recovery from actual incidentsRare severe events may not appear in internal history
Risk and control self-assessmentStructured review of processes, risks, controls, and residual exposureRatings can become subjective or optimistic
Key risk indicatorsEarly warning from volumes, errors, outages, backlogs, turnover, or exceptionsThresholds may not predict the next failure
Control testingEvidence that a control is designed and operating as intendedA passed sample does not prove the control cannot fail
Scenario analysisForward-looking analysis of severe but plausible eventsResults depend on assumptions and expert judgment
Business-impact analysisCritical activities, dependencies, maximum tolerable disruption, and recovery prioritiesCan become stale as processes and vendors change
External events and near missesFailures that could occur even if the organization has not experienced themComparability and data completeness may be weak

Frequency and severity are common dimensions, but they are not sufficient. Assessments may also consider customer harm, time to detect, duration, legal obligations, liquidity needs, concentration, substitutability, and the possibility that several controls fail together.

Worked Example: Payment-Processor Outage

Assume an online retailer depends on one payment processor. A six-hour outage prevents card authorization during a high-volume sales period.

The initial operational exposure is not only lost sales. The review should also consider:

  • orders that were accepted but not paid
  • duplicated or delayed transactions after service resumes
  • refunds, chargebacks, customer support, and remediation
  • cash-flow timing and reconciliation differences
  • contractual service credits and reporting obligations
  • concentration in one processor and one network connection
  • whether a manual or alternate payment route can handle peak volume

Preventive controls might include resilient architecture, release controls, capacity testing, and a secondary processor. Detective controls might include failed-authorization alerts and reconciliation breaks. Recovery controls include traffic switching, customer communication, backlog processing, and post-incident reconciliation.

Residual risk remains because the backup provider may share infrastructure, routing may fail, staff may not be trained, or transaction data may not reconcile cleanly. The decision is therefore whether the remaining exposure fits approved tolerances and whether recovery can meet customer, contractual, and liquidity requirements.

Operational Risk and Operational Resilience

Operational risk management asks what can fail, how loss can arise, which controls apply, and whether residual risk is acceptable.

Operational resilience asks whether critical operations can continue through disruption or be restored within an approved tolerance. It therefore emphasizes critical-service mapping, interdependencies, scenario testing, response, recovery, and learning.

The concepts reinforce each other, but they are not interchangeable. A control may reduce incident probability without ensuring fast recovery. A resilient process may continue operating while still producing losses, customer harm, or legal exposure that requires remediation.

Governance and Evidence

Management should assign each material risk and control to an accountable owner. Depending on the organization, review and challenge may involve business management, an independent risk function, compliance, information security, legal, finance, and internal audit.

Useful evidence includes:

  • current process and dependency maps
  • incident, near-miss, and loss-event records
  • control design and test results
  • system availability and change-management records
  • vendor performance, concentration, and exit-plan evidence
  • scenario assumptions and business-impact analysis
  • risk acceptance, exception, and remediation approvals
  • recovery tests and unresolved reconciliation items

The absence of a recorded loss does not prove that controls are effective. A near miss, repeated exception, delayed reconciliation, or dependency with no tested substitute can be material evidence.

Common Mistakes

  • Treating operational risk as only fraud, cyber risk, or back-office error.
  • Confusing operational risk with business-model or earnings variability.
  • Assuming insurance or outsourcing transfers all financial and control responsibility.
  • Recording incidents without analyzing root cause, control failure, and recurrence.
  • Using a low average loss to dismiss a severe tail scenario.
  • Reporting a color rating without the evidence, owner, methodology, and action threshold.
  • Testing components separately while ignoring shared infrastructure and correlated failure.
  • Measuring recovery time without testing data integrity and transaction reconciliation.

Questions to Ask

  • Which critical activity, customer outcome, asset, cash flow, or obligation can be affected?
  • What event starts the loss or disruption pathway?
  • Which systems, people, data, premises, and third parties are required?
  • What is the inherent exposure before controls?
  • Which controls prevent, detect, contain, correct, and recover from the event?
  • When were those controls last tested, and what exceptions remain open?
  • What residual exposure remains after credible control performance is considered?
  • Which indicator, limit, or disruption tolerance triggers escalation?
  • Who can accept the risk, fund remediation, or stop the activity?
  • How will records be reconciled and affected customers made whole after recovery?

Official Sources

The Basel and U.S. interagency materials apply to specified banking and supervisory contexts. Their definitions and practices are useful reference points, but they should not be treated as universal requirements for every company, investor, industry, or jurisdiction.

  • Business Risk: Exposure to demand, pricing, competition, strategy, concentration, and operating cost structure rather than process failure.
  • Model Risk: Adverse consequences from incorrect, misused, or poorly governed quantitative output.
  • Fraud Detection: The control activity that identifies suspicious behavior and connects signals to evidence and investigation.
  • Reputational Risk: Financial or operating harm caused when an incident changes stakeholder confidence and behavior.
  • Risk Appetite: The types and amounts of risk an organization is prepared to pursue or retain within its objectives and capacity.

FAQs

What is operational risk in simple terms?

Operational risk is the possibility that the way an activity is performed fails and causes loss, harm, or disruption. The source may be a person, process, system, third party, or external event.

Is operational risk the same as business risk?

No. Operational risk concerns failures in execution and supporting resources. Business risk concerns demand, competition, pricing, strategy, and cost structure. One event can create both.

Can operational risk be eliminated?

No. Controls, insurance, diversification, resilient design, and recovery planning can reduce or transfer parts of the exposure, but residual risk remains.

Educational Use

This article provides general financial education. It is not personalized banking, investment, cybersecurity, accounting, legal, regulatory, insurance, or risk-management advice.

Browse Risk Management