Written evidence submitted by Center for Evidence-Based Managment (OPR0003)

 

Overview

This paper is in response to a request for submissions of evidence to the Treasury Committee into the “IT failures in the financial services sector inquiry”, closing on 18 January 2019.

 

It consists of four parts.

 

Part A provides a general overview of the lifecycle of systems and the causes of system failure. Part B provides high level responses to the majority of questions posed by the committee. Part C provides a more detailed analysis of the issues with replacing core systems in financial institutions, addressing three of the questions raised in detail. The final section contains profiles of the authors including publications related to this important topic.

 

 

Martin Walker & Patrick McConnell

 

18th January 2019

 


Part A – Context

Every computer system consists of a set of fundamental elements and goes through a lifecycle that ultimately ends in its replacement or failure. In Part C, this is explored in greater detail in relation to the Core Systems of banks.

The Elements

To truly understand the nature of the system it is necessary to consider

 

The key point to understand is the evolution of any system is ultimately driven by people. Both people and technology need to be understood to have a real understanding of problems, causes and solutions.

 

Figure 1 - The Path to Entropy

Stability to Entropy

A stable system had the following attributes

 

         Ease of support (including quick and simple recovery from failure)

         Ease of change (including maintenance and testing)

         Low risk of system failure

 

Over the lifecycle of any system, age and growing complexity inevitably will make a system harder to support and to change. They can also increase the rate of failures and the general unpredictability of the system. This is the path to entropy.

 

A stable system is not simply an unchanging system because the elements around a system change frequently: users may create more data; the operation of related systems may change; the hardware may become obsolete and impossible to maintain. A system that does not change even though the world around it is changing will also move along the part to entropy.

Effects of Aging

The simple passage of time causes a system to age in a generally negative way as shown on Figure 1. Typically, a new system goes through an initial period after go-live where it becomes more stable. Bugs are fixed, more support procedures are put in place and both users and IT become more familiar with it. However, unlike good wine, complex computer systems do not improve with age, but turn sour fairly quickly, and with age the following negative effects take place:

 

 

In some cases, a system may be substantially re-written while retaining the same use and same name. In those cases, the “Aging” clock is wound back to (almost) the beginning.

Effects of Complexity

The specific meaning of “complexity” in systems is often misunderstood and confused with the concept of being “complicated”. The following helps explain the difference.

 

The more expensive and prestigious mechanical Swiss watches are characterised by a large number of parts, including numerous wheels and jewelled bearings. With over a hundred parts these watches are a truly incredible piece of precision engineering. The resulting device is often so complicated it requires a great deal of training to simply understand how to clean it, let alone repair it. The key concept to take from Swiss watches is the term Complicated. Complicated devices or system are generally hard to understand but capable of being understood and their operation is generally highly predictable. After all a Swiss watch that did not predictably tick every second would not be much use as a time piece.

 

Now imagine taking that watch to an international space station where astronauts from different countries have to communicate, but while expert scientists, none are expert watchmakers. Different people may call the same part by different names, sometimes different parts are called by the same name. A part may be contained within a larger part while at the same time the larger part is contained within a smaller part of a larger mechanism. In places, the mechanisms do not quite connect and require astronauts to observe one wheel rotating before manually rotating a nearby wheel to set the ‘time’. The watch itself is anything but predictable, sometimes it ticks at one second intervals, occasionally the interval changes and sometimes the tick is replaced by a different alarm sound. Different people looking at the same watch may also see different times (am/pm) because the usual visual clues are absent. This is a Complex system, hard to understand and very often hard to predict. And as a result, mechanical watches, though superb machines, are usually replaced by digital clocks in space.

 

A complex IT system is generally a bad thing. Over time increased complexity makes a system harder to understand, harder to change and harder to test. An increase in complexity can be driven by many factors but the key ones are a rate of change to the system that is too fast or simply a succession of changes that are badly implemented or confuse the original purpose of the system.

Implications

The implications of this are that age and complexity will ultimately make any system have errors and those errors will have varying degrees of impact and predictability[i]. It is easy to investigate every single notable failure, identify a specific cause and recommend a specific solution. This adds value and adding to procedures can reduce the risk of the errors occurring again.

 

However, it is no substitute for good management. Good management comes from understanding the lifecycle of systems, the factors that speed a system along the path towards entropy and making the decisions that slow down the path to failure. Ironically in much of banking we see a failure to make changes because of fear or complacency (described in more detail in Part C). Systems are not replaced or sufficiently upgraded because “it just works”, “it would be too risky to touch it” or “it would be too expensive.” Each of these excuses for inaction will increase the risk of failure. On the other hand, there is a love of novelty and “innovation” that causes some firms to take too much risk while neglecting the basics. Using the latest technologies (regardless of their value or maturity) is “fun”, adds value to CVs, create positive publicity, typically comes with the promise of easily solving a multitude of problems and wins prizes.

 

One final and critically important point about system failure is the “silent failure” of systems producing poor quality data for use in management decision-making. A failure in an ATM network gets immediate media and political attention because the impact is clear and negative. A system that produces the wrong data or makes it very hard to extract data required for decision making, will ultimately lead to the wrong decisions being made over a prolonged period of time. That build-up of bad decisions can ultimately destroy a bank and destabilise the financial system (and the wider economy) causing far more damage than a short -lived failure in a specific system.

 


Part B – High Level Responses to Committee’s Questions

The extent to which operational incidents are becoming more frequent, and how the prevalence of such incidents may change in future as consumers and firms come to rely more heavily on technology.

 

Comment             

Financial institutions have always suffered from IT failures, most of which do not come to the eyes of the public, press or politicians because they do not have an immediate impact on retail or corporate customers. Typically, banks have systems and processes in place to record all "incidents". Sometimes this is from the IT Management perspective, sometimes from the operational risk perspective, sometimes from both. Whether the extent of incidents is becoming more frequent cannot be determined by looking solely at headline incidents but by collecting data from banks and analysing it.

 

However, just as the quality of systems varies widely between banks, the systems and processes for recording incidents varies widely. Some are very high quality and comprehensive. If less well designed/poorly policed processes, incidents may not be consistently recorded and information about the causes of those incidents is unclear or wrong.

 

There are mechanisms at work within banking IT that could be leading to an increasing failure rate, as suggested by the survey of IT failures in banking, FCA Cyber and Technology Resilience: Themes from cross-sector survey 2017-18[19]. These include increased complexity, a fear of upgrading core systems and aggressive outsourcing/offshoring. As a great deal of interactions between banks and their clients (retail or corporates) becomes more automated manual or human based alternatives are removed e.g. bank branches are closed. This means the impact of system failure can become more wide-spread and hence more damaging.

 

Aging core systems, the additional complexity from dealing with clients via the Internet and lack of a manual fall back can prove to be toxic combination.

 

Personal Experience

MW – After taking charge of part of the system infrastructure at RBS Capital Markets, I found my systems were falling over every few days. Sometimes the systems stopped working, over times the systems malfunctioned without falling over. For example, on one occasion trades feeding into the system were duplicated. As I worked to resolve the underlying problems, I found another problem, the operational incidents were not being consistently recorded, neither senior management nor me as head of the department had a clear picture of just how unreliable the system had become.

 

Recommendations             

Developing a better evidence-based understanding of this area requires more analysis but also better data and more consistent collection of data on failures. An analysis of this area should include a review of how data regarding incidents is captured both technically and procedurally. The right data needs to be captured about incidents in a timely manner in a format it can easily be queried and analysed. The most fundamental issue is whether incidents really are recorded. Headline-making incidents do bring awareness of the problems to the attention of the problems to the public, press and politicians. Investigation of individual failures can provide valuable insights but there is a strong need for more evidence to be gathered to gain a true understanding of the problem.

 

The common causes of operational incidents in the financial services sector.

 

Comment             

One of the main causes of operational incidents is simply change. Every time a system is changed and either code released or hardware upgraded there is a risk of making a mistake.

 

Specific problems that cause problems in the change process are

  1. Creation of bugs that are not identified in testing. This may be due to deficiencies in the testing process (or testers). Another reason for buggy code being released is aggressive management pushing changes to be released before they are fully tested in order to meet deadlines.
  2. Errors in the process of releasing and deploying changes. The wrong code or configuration (i.e. not what was developed or tested) can be deployed. Changes can be deployed at the wrong time, in wrong ways or ignoring dependencies on other systems. Sometimes people may simply neglect to activate the new software.
  3. A popular trend in software development in banking and deployment is called “Dev Ops”. There are a variety of definitions of Dev Ops but a common model is based around making frequent changes to production systems and identifying errors in production. While this may improve customers’ experience, it also has great potential to increase the number of failures if applied to core systems that support retail or corporate customers.

Other common causes of errors include

 

Personal Experience

MW – The worst failure I experienced as an IT manager was when the trading system that I was responsible for stopped allowing traders to enter new trade records. This impacted over 100 front office staff worldwide plus support functions. There was no problem in the system but the underlying database (Oracle) that it used had been ‘locked. There was only one Oracle Database Administrator (DBA) in the bank. He was only holiday and mostly out of contact because he was sailing. After several hours the other DBAs working with Oracle managed to resolve the problem. We later learnt after the return of the Oracle DBA that he could have fixed the problem in 5 minutes.

 

Recommendations             

Such IT problems are not unique to banks. Some are almost impossible to avoid but there is a very wide variation in the quality of infrastructure, staff and management. Regulators should look for objective, quantifiable measures of performance to encourage the worst banks to compete with their peers. The management of complex infrastructure (including the human/political dimensions) would also benefit from more fundamental research.

 

The extent to which there exist “single points of failure” and/or other sources of concentration risk in the financial services sector.

 

Comment             

At bank and market level there are multiple "single points of failure". Most banks for instance have a centralised gateway for connecting to the SWIFT network or for access to payments networks. If these gateways fail, they can have traumatic impacts on the bank, its clients, counterparts and the broader economy. However, centralisation does not necessarily increase the chance of failure. A variety of systems spread across a firm, essentially doing the same things, but using different versions of software increases the chance of failure because they will each have their own change cycles and support procedures.

 

For business functions that are genuinely consistent, such as SWIFT messaging, centralisation can make a great deal of sense from the support perspective. A single system with more mature support processes, monitoring and experienced staff may still fail but it is likely to recover more quickly. There are also other techniques that can reduce risks such ensuring the system has “hot-hot” disaster recovery (there is a parallel system on-line that can immediately replace the failed system), ensuing systems have excess processing capacity to deal with spikes in demand, that they are simply well designed and that disaster recovery plans are frequently tested. Naturally any centralised system (at bank or market level) combined with low quality support is going to have higher impact if it fails.

 

One newer area of concern that has been raised is the concentration of power in the provision of “Cloud” services. There is nothing mysterious about a “Cloud” is it simply one or more data centres, where clever software is used to share processing across multiple servers to achieve economies of scale. Using Cloud providers is becoming increasingly popular in banks but the cloud market is becoming increasingly concentrated between, Amazon, Microsoft, IBM, Google and AliBaba. Though each of these are large reputable organisations, history shows that even the largest organisations may suffer technical failures or financial crisis.

 

Personal Experience

MW – Single points of failure can often have very mundane causes. A traumatic IT failure that impacted my systems was caused by the maintenance contractor for our data centres. Over the course of many years, the contractor failed to clean the filters on the cooling units. Eventually on one particularly hot day the cooling failed and system after system overheated and failed. The same problem applied in both the main data centre and the disaster recovery data centre.

 

As mentioned in the previous points, a common type of single point of failure, is not technical but human. For many legacy systems there may only be one person who truly understands how it works. In the worst cases nobody may truly understand the operation of a system, something that does not usually become apparent to management until the system fails.

 

Recommendations             

Systems at a bank or market level that have a high impact if they fail need to be properly maintained, supported and have realistic/well tested plans for disaster recover. These types of systems need regular and zealous review by regulators and internal management. This cannot simply be a box-ticking exercise carried out by internal or external auditors. They need to be carried out by people who understand systems of the type being reviewed.

 

The incidence of multiple old legacy systems and the nature of their connectivity, and the impact of retrofitting web based/mobile systems to legacy systems.

 

Comment

 

Note: For a more detailed analysis please refer to Part C of this submission - The Risks of Core Systems Replacements (CSR) in Large Financial Institutions.             

 

Almost every bank has multiple legacy systems (some 20+ years old). New "challenger" banks or Fintech firms, even where they have built new internal infrastructure, typically have a major dependency on the legacy systems of established banks that provide them with services such as access to central bank clearing systems. The term "Legacy System" itself can be highly misleading. Some legacy systems may be simple, well designed, well maintained and efficient. "Legacy System" can be used as a term of abuse, put to use by those who have internal political or external commercials reasons for wanting to replace an existing system. In general, the older the system, the greater the complexity, the less the understanding, the harder to support and the more expensive/riskier to modify. However, we should not lose sight of the fact that some fairly new systems can also end up complex, incomprehensible, hard to support and expensive/risky after a relatively short period of time. Particularly if they were poorly designed or implemented from the outset, not to mention built using inappropriate technologies.

 

Banks have now been building web-based interfaces (for both internal and external use) on top of legacy systems for over 20 years. In more recent years they have connected legacy systems to “Apps” based on smart phones and tablets and increasingly allowing external systems to interact with legacy systems via an Applications Programming Interface (API).

If designed correctly the software connecting to legacy systems to web-based applications, Apps and APIs can be highly secure and protect the legacy system from being overloaded. However poor design, poor implementation or inadequate testing can potentially create situations where interfacing to legacy systems can cause problems.

 

Personal Experience

Interfacing browser-based applications to legacy systems for either internal or external users has been routine for decades. The same basic principles of good software development apply as in any other area of banking IT.

 

Recommendations             

Open banking and greater use of APIs to connect bank legacy systems to Apps provided by third parties is an area needs to be monitored from both the general quality of implementation and perspective of adding complexity to the overall banking system.

 

The risks associated with integrating banks/systems, following takeovers and mergers, for example.

 

Comment             

Note: For a more detailed analysis please refer to Part C of this submission - The Risks of Core Systems Replacements (CSR) in Large Financial Institutions              

 

Personal Experience

MW During takeover/merger situations they is generally a great deal of pressure to act quickly to rationalise infrastructure and cut IT costs (often one of the key drivers for takeover/merger). However, this can have highly negative implications, decisions can be made without considering the full implications. Decisions made regarding infrastructure typically need to be made after decisions regarding the combined business models otherwise firms may find they have a “rationalised” infrastructure that does not support all the businesses in their combined business model.

Another risk is of prematurely declaring victory in an integration project. Many years after the series of series of takeovers that lead to the phenomenal growth of the RBS group there was duplication of systems and other issues because the “integration” had never fully been completed.

 

Recommendations

See Part C             

 

 

The quality of relevant technical documentation.

Comment             

The quality of technical documentation (from either support or maintenance perspective) is highly variable from the poor/non-existent). Even where documentation exists it can be very hard for external parties such as auditors or regulators to check the validity and completeness of documentation. Lack of good quality documentation creates major risks. It intensifies “Key Person Risk”, the risk of relying on a single person or a small group of experienced staff. It can also turn the outsourcing or offshoring of IT activities from a very hard process to an extremely hard and risky process.

 

Personal Experience

Major systems in banks can have excellent documentation that fully explains the operation of the system, the technical architecture, the independencies on other systems and the procedures to follow in the event of issues. In the worst cases a complex system may have no documentation whatsoever. Unfortunately, the undocumented systems are often the old or poorly designed systems that critically need documentation.

 

In addition to documentation about specific systems, a critically important form of documentation is the system architecture diagram, including the connectivity of systems and the degree of inter-dependency [2]. This again varies from the excellent to non-existent. I arrived at a major global bank and found there was no diagram of how the systems supporting a multi-billion dollar business connected together.

 

Recommendations             

At a minimum every bank should be able to demonstrate they know all of their systems, what they do, how they connect together and what the dependencies are [2]. Producing and updating system documentation is something that is unpopular with IT staff. What management, regulators and auditors should look for whether the creation and updating of important documentation is embedded into key IT change processes. Checking for the existence of documentation can add value but it is very hard for an outsider to check the quality or relevance of documentation.

 

The impact of outsourcing on operational resilience.

 

Comment             

Outsourcing (and even off-shoring in many cases) can create issues if the incentives of the bank and the outsourcer are not aligned. In the case of off-shoring there can be a similar misalignment of incentives and objectives between offshore management and on-shore management/end users.

 

Creating complicated contracts, including detailed Service Level Agreements (SLAs) is the common approach to do this but often does work. There are typically strong incentives for an outsourcer/offshore team to take on more work even if it cannot cope with existing work or maintain quality levels.

 

One of the most fundamental problems in outsourcing is knowledge loss. Whether this occurs depends on the outsourcing model applied. In many cases the existing team (including all their experience) is absorbed into the outsourcer and gradually assimilated into their organisations. However, in the high-risk scenarios, work is outsourced but the outsourcer increases profit margins by not taking on the existing team and either uses cheaper junior staff or moves work to low cost locations. Knowledge loss is a major risk for systems that are older/more fragile. A lack of understanding of the system makes it harder for a new team to know how to resolve issues. As mentioned above, older systems typically have less documentation so genuine understanding of the system disappears with the former team. However, there are scenarios, particularly for more utility software, where outsourcing can introduce more professional and robust practices in areas of support and maintenance.

 

Personal Experience

MW I saw a great deal of very rushed off-shoring. Rushed to the point of recklessness. Managers were given quotas in terms of number of heads off-shored, not in terms of cost-effectiveness. This created the scenario where rushed offshoring increased risk without actually cutting costs. In one bank I had discussions with, very large-scale offshoring and outsourcing over a prolonged period of time had resulted in a core middle management layer in London which lacked practical hands-on management of IT systems and project and only had experience of relationship management between teams.

 

Recommendations             

Fundamental and objective research needs to be carried out into the costs and risks of both outsourcing and offshoring. Management should have objectives of cost effectiveness and quality not just quotas for heads off-shored.

 

The ways in which consumers typically lose out as a result of operational incidents, including inconvenience and vulnerability to fraud.

 

Comment             

The impact of major incidents in areas such as payments is well documented.

 

Personal Experience

Operational incidents happen all the time. However, the amount of the disruption depends on the nature of the business process. An account opening system that fails for a few hours may have little to no impact. An FX auto-hedging system that fails can cause a bank financial loss within seconds. A failed interface between a core banking system and an ATM network will rapidly cause severe distress to those that need funds to purchase necessities.

 

Recommendations             

Firms need to make clear to IT teams, particularly those that have been off-shored and are physically and emotionally disconnected from their business users (and ultimately the customers of the bank) what the impact of issues with the system can cause. With the fragmented nature of many IT teams there is sometimes a more fundamental requirement to make them understand what the system they develop or support actually does.

 

 

Examples of best practice with respect to firms’ responses to and handling of operational incidents, including approaches to communicating with customers, identifying and addressing the causes of incidents, and handling customer complaints and compensation.

 

Comment

Comprehensively designed and thoroughly tested, business continuity planning (BCP) is the basis for dealing well with operational incidents. The Basel Committee on Banking Supervision set out comprehensive high-level principles for business continuity in 2006 report. Those principles provide a good framework to this day.

 

The fundamental requirement is for firms to understand the potential impact of operational incidents, including those caused by internal failures in the process of making and applying changes to systems. They also need to understand the size of a technical change or failure may be much smaller than the overall impact of the change on the financial institution, its customers and the wider economy. Only then can they design and test contingency plans that both realistic and effective.

 

Personal Experience

It has to be remembered that operational incidents can result from IT failure, people failure or combination. I have seen operational incidents blamed on IT systems when they were rather caused by the absence of appropriate IT systems, where humans had built complex semi-manual processes using data coming out of the core systems without any knowledge of the IT department. There are also examples where the absence of a system necessitated the creation of inherently risky manual processes, notably sending confidential client reports to the wrong clients.

 

Recommendations             

Being able to put in practice effect plans and also execute them does not just result from having a checklist. It is also highly dependent on having staff and management that understand what their systems do technically and in terms of business outcomes.

 

What should be learned from the operational incidents witnessed in recent years.

 

Comment             

The notable failures of bank IT systems in recent years have arisen from a whole range of causes. Failed data migrations, cyber-attacks, errors in change management and general human error.

 

The only truly common themes are that the more complex the infrastructure the easier to fail and the need for good management at all levels of the organisation. Good managers understand what they are managing, invest where investment needs to be made, do not put cost reduction ahead of risk management and are willing to take career risks (e.g. updating systems) that reduce the overall risk to an organisation.

 

Personal Experience

Large scale data migrations (between systems and/or legal entities) are one of the riskiest operations any bank can take. I managed many smaller data migrations in capital markets, each of which could have crippled individual businesses and in one case could have crippled the whole bank if they had gone wrong. In each the data migration was planned over the course of months involving all stakeholders and practiced multiple times including full dress rehearsals at the weekends. Even then, unforeseen events can happen so it is essential there is a full-time team available post migration to available to do post-migration checks and deal with any issues as soon as they appear.

 

Recommendations             

Ultimately there is no substitute for competent, accountable management in IT and IT investment that is driven by long term return and the responsibility of banks to provide systemically important services. Investment should not be driven by the search for short cuts, the glory of being at the cutting edge or to service short-term boosts to P&L.

 

 

The ability of the regulators to ensure firms are adequately guarding against service disruptions.

 

Comment

Post hoc analysis of specific incidents and surveys such as the FCA’s “Cyber and Technology Resilience: Themes from cross-sector survey 2017-18” are useful exercises but we are living in age where an understanding of the robustness and resilience of the financial systems IT systems could be more data driven and forward looking.

There is also a tendency for regulators (in multiple jurisdictions) and for the senior management of banks to overemphasise the importance of governance. Accountability, and clarity regarding roles and responsibility can add considerable value.

 

However, governance does not equal effective management. Large and complex governance structure add no value if those in governance functions have no idea what they are governing and if they have not objective sources of data to supplement self-reported data. Also, if governance is treated as something separate from operational management, by for example creating specialist governance functions (typically populated with consultants), it dilutes accountability and can make a bad situation worse.

 

Similarly, simply specifying a “three lines of defence” model for IT risks and looking for the presence of specific types of procedures is no guarantee that risks are understood and management effectively. Regulators themselves need to understand the nature of IT risks within organisation and know how to look for them. For instance, a firm that cannot provide a diagram demonstrating how all their systems connect to each other is probably at greater risk than one that does not have a formal three lines of defence model in IT.

 

Personal Experience

 

No comment

 

Recommendations             

 

Industry consistent logging of all incidents by sufficiently trained staff could provide the data to create predicative models and to do peer group comparisons. Even the daily, small-scale issues, can be indicative of wider problems in terms of system, staff and management quality. Other forms of data, such as trade reporting provided under EMIR (European Market Infrastructure Regulation) and the upcoming SFTR (Securities Finance Transaction Regulation) can provide valuable insights into the quality of both systems and processes. Poor quality data, as demonstrated by the number of errors, corrections and time delays in reporting, can be mined to identify overall system quality.

 

Otherwise regulators need people, or cross-disciplinary teams that can look at system infrastructure from the full set of dimensions listed in Part A.

 

In the FCA/PRA report it mentions the issues some firms have maintaining a view of the age of systems, third parties and end-of life assets. There is no magic involved in firms doing this. The better banks maintain comprehensive Configuration Management Databases (CMDB). These list systems and other key information such as vendors, age and dependencies on other systems within the bank. Better banks have been doing this for decades. Failure to maintain a CMDB is simply a symptom of bad management or over aggressiveness in M&A or rolling out new systems. The existence of a CMDB and links between updating the CMDB with the change management process is a fundamental practice that regulators should look for.

 

Whether the regulators have the relevant skills to hold appropriate parties to account in the event of significant operational incidents.

 

Comment             

While accountability is important, particularly at the board and senior management levels, holding people accountable in a complex organisation with complex infrastructure is probably less important than working with the organisation to fix problems and deal with underlying issues.

 

Personal Experience

No comment

 

Recommendations             

Regulators themselves are often responsible for complex infrastructure within their organisations. Perhaps it would be best that regulators spend some time looking inward and reflecting on the quality of the systems. If regulators cannot sufficiently understand their own issues there is the risk that regulators simply impose (or inspire) procedures and organisational structures within financial firms that add to complexity rather that help solve the problems.

 

 

Approaches to operational resilience in different jurisdictions.

Comment

This is ultimately dependent on like-for-like comparisons of the rate and type of failure between jurisdictions.             

 

Personal Experience

Good banking infrastructure (at a firm or market level) is not something specific to developed nations. Many developed nations have surprisingly primitive infrastructure compared to countries that may be considered to be less developed financially or technically. The internal banking infrastructure of a well-known British bank in its off-shore entities was typically considered much worse than the systems of local competitors.

 

Recommendations             

Regulators should look at other jurisdictions but they need to look far more broadly for best practices than the main financial centres.

 

 

The opportunities and risks presented by the application of new technology in the financial services sector with respect to operational resilience.

 

Note: For a more detailed analysis please refer to Part C of this submission - The Risks of Core Systems Replacements (CSR) in Large Financial Institutions              

 

Comment             

Information technology continually advances. Many mature technologies have grown more stable and there is an ever-growing range of tools for managing and monitoring system infrastructure. However, these are not typically the ones that gain the attention of senior business management or regulators.

 

However, IT is particularly prone to “fad” like thinking which offers almost magical solutions to a whole range of problems. Technology Fads are not necessarily forms of technology that do not work or completely lack value (though some are). They are technologies where those marketing or selling them make promises far beyond what they are capable of delivering. In many cases attempting to apply a technology in an area where it does not add value or where alternatives offer superior alternatives.

 

The various forms of technology grouped under the terms “Artificial Intelligence” (AI) have gone through several cycles of hype of a fifty-year period punctuated by prolonged periods of disillusion called “AI winters”. Some forms of AI have the potential to identify patterns in data relating to operational incidents and potentially predict problems but so do standard statistical packages. Blockchain is another over-hyped technology (but with little track record of practical implementation) that has had claims made in terms of resilience and security. Blockchain (depending on the variety) also raises many problems for financial services firms looking to use the technology. “Public Blockchains” have the fundamental problem for instance that nobody is accountable for the operation of the network, even “Private Blockchains” can “leak” data in ways that conventional systems would not. Some relatively fashionable technologies, such as the not-so-new “Robotic Process Automation” often create additional complexity and fragility for the overall system infrastructure.

 

Firms must always look for “Good” in technologies rather than “New”.

 

Recommendations             

New technologies should be experimented with but it is essential the senior management of banks focus on real tools that solve real problems rather than chasing fads, and demonstrating magical thinking.

 

What should be considered an appropriate level of tolerance for operational disruptions?

 

Comment             

It depends on the specific area of banking/finance and the impact of failure. Every IT system can potentially fail for a wide variety of reasons. The better ones have two characteristics, they fail less frequently and they are able to recover quickly. For key economic activities such as the operation of the payments network (particularly in an economy that is rapidly moving away from the day-to-day use of cash) it can be argued the standard expected should be on a par with those expected in aerospace and other sectors where IT failures can cause fatalities.

 

Personal Experience

MW - In many areas of banking, small scale regular failures or filling functional gaps with people is endemic. This works against creating a general culture of quality. When this is combined with a lack of documentation and a lack of understanding of the interdependencies between systems, it means small scale problems in one area can cause problems for more core systems, which otherwise have higher quality thresholds.

 

Recommendations             

Given the systemic importance of banks it is important they seek to raise their overall quality level in every area from development to support. The fundamental way to drive this is an open and honest attitude by IT and bank management to understanding their short-comings and willingness to work with independent and external researchers in this area.

 


 


Part C - The Risks of Core Systems Replacements (CSR) in Large Financial Institutions

This part addresses three questions raised by the Committee

 

 

In answering the questions as a single topic, the section addresses the considerable risks of integrating or replacing the critical Core Systems of financial institutions, particularly banks, that result from a takeover or merger of institutions and/or a strategy of technology-led innovation. A number of case-studies, such as the Co-Operative Bank, TSB and Danske Bank, are cited as examples of the considerable systemic risks that can arise when a firm attempts to pursue a technology strategy in support of an acquisition or growth business strategy.

Technology Evolution not Revolution

Financial institutions, in particular retail banks and insurance companies, employ business models that are driven by economies of scale. The more deposits a bank takes, the more loans it can make; the more policies an insurance company underwrites, the more able it is to spread its risks; and the more ‘assets under management’ an investment or pension fund manager manages, the more stable the returns will be.  In any business that is subject to economies of scale, application of information technology (IT), will usually lower the marginal cost of acquiring new customers, and selling products and services to them.  There is then a virtuous circle; better technology allows more customers to be acquired and serviced at lower costs; leading to greater asset levels; and all other things being equal, higher profits; which then allows newer and better technology to be purchased.  Newer technology allows firms to service ever more customers at lower marginal costs, and so the virtuous circle spirals upwards, always assuming, of course, that the greater risks that are being taken are being managed prudently.

 

And, as a result of providing competitive advantage through its ‘smart’ application, technology also promotes consolidation of information-intensive industries, such as banking, through mergers and acquisition of firms that do not keep up with cost-reducing technology innovations. And there is no indication that, what Nobel Prize winner Robert Merton calls, the “financial innovation spiral” [5] will run out of steam any time soon.  In Merton’s spiral, innovations in financial products give rise to innovations in processes to support those products [6], which, in turn, are made more efficient by innovations in technology, and new technology helps firms to create new products; and so on.

 

Financial institutions, especially large banks, are usually “early adopters” of technology innovations rather than “innovators” themselves [1]. Pioneers, such as Walter B. Wriston, ex-Chairman and CEO of Citibank, applied innovative technology to propel his bank to industry leadership for decades through the introduction of automated teller machines (ATMs) in the 1980s. In the 1990s, investment banks, such as Barclays, built enormous trading floors in financial centres around the world with the latest technologies at traders’ fingertips to create global trading giants, operating 24 hours per day.  In the 2000s, the same investment banks employed new, faster computers and high-speed networks to create high frequency trading (HFT) capabilities that traded semi-autonomously, an early example of machine learning (ML).

 

While, financial institutions have a long history of absorbing technological innovation and prospering. it is not, however, their core competence, which is the management of the very many risks that arise from their core business of lending and underwriting risks.  However, boards must also be aware of the technology risks to which they are exposed and failure to manage a firm’s technology risks has led to the, sometimes spectacular, failure of boards’ business strategies (as illustrated in some of the case studies below).  Technology is changing rapidly, which makes strategic decision-making very difficult and risky for the boards of any financial institution [2]. Because of the rapidly changing technology landscape, a strategic decision taken today may look very different when a business strategy is finally implemented five or more years in the future. Without a crystal ball, a board must understand the technology trends that are driving technology innovation and the opportunities and risks posed by those trends.

 

Whereas technology professionals are always (overly) enthusiastic about the prospects of reducing IT costs due to technology innovation, directors and executives must be more circumspect. There is a trade-off between enthusiastically pursuing the opportunities presented by each new technology that appears and managing the day-to-day, often boring job of ‘business as-usual’ (BAU) ensuring that the ‘core functions’ of the firm are delivered to customers, efficiently, whenever and wherever needed.

 

The Information Technology (IT) community are notoriously addicted to technology innovations, constantly adopting the next new technology as the “silver bullet” that will solve all existing problems. For example, ‘blockchain’ is just the latest of such cure-alls and Artificial Intelligence (AI) is, as it has been for the last forty years, the “next big thing”.   This does not mean that these technologies do not have their uses - they most certainly do, but they will a much smaller niche that proponents claim when they first arise. There is even a name for this incessant chasing after technology unicorns – the Gartner Group’s “Hype Cycle” [3].  In this cycle the capabilities of a particular technology innovation are exaggerated until they reach a ‘peak of inflated expectations’ before dropping precipitately into a ‘trough of disillusionment’.  Most technology innovations do not survive this fall but some do and the most useful gradually crawl up the ‘slope of enlightenment’ to eventually become productive.

 

In reality, successful IT Innovations require many stars to be aligned to emerge from a clever idea (Invention) to be a truly useful game changer across the financial industry as in the equation below [4] and illustrated in Figure 2:

 

IT Innovation = (I)5 + (II)

 

That is - (Invention* Integration * Industry Standards * Infrastructure * Investment) + (Incremental Improvement)

 

Figure 2 Innovation – Evolution not Revolution

 

No matter how brilliant an invention, an IT innovation will never instantly replace existing technology but: must first be integrated with existing technologies (such as the Internet); must conform to constantly evolving industry standards (such as the UK Faster Payments Service); and must be based on solid high-performing infrastructure (such as 4G wireless communications). Most of all, IT innovations require money (investment) and, importantly, time to emerge and to fit into the overall IT environment.  And when eventually implemented, IT innovations must be constantly changed, or incrementally improved, in order to integrate with the next wave of innovations, industry standards and infrastructures, such as 5G wireless technologies.

 

But there is nothing smooth or even inevitable about the adoption of innovations.  Many new technologies fall by the wayside [1], because they turn out not to be as useful as first thought, or are sometimes superseded by other innovations, or “technology disruptors” [7]. An innovation by itself will not succeed unless it first fills a real need in an industry, and also industry processes are adapted to utilise the innovation.  Innovation then is not merely invention, but invention plus a long process of ‘adoption’.  And over the long period of adoption, the boring business of providing core services and operating core systems must be continued.

 

Although technology is constantly changing and evolving, business must also have a high degree of continuity, otherwise useful and cost-effective products will never get to be delivered to customers.  In business, there is always a trade-off between flexibility and stability: flexibility to develop new products and to enter new markets, yet stability to generate sustainable profits and a long-term return on investment for stakeholders. 

 

Today, both flexibility and stability are delivered by technology. For example, today new products are developed and sold over the Internet while smaller, faster, cheaper computers and high-speed networks provide more stable data centre platforms. Unfortunately, manging the trade-off between flexibility and stability is not easy as the two factors are inter-dependent.  A search for flexibly will inevitably reduce stability (as innovations have to be integrated into existing systems) and vice versa.

Core Systems

Technology innovation, then, takes considerable time and, in the meantime, firms must continue to do business day-in, day-out with the technology that they already have – their ‘core systems’. The trick is to gain competitive advantage through innovation, while maintaining the stability of their core systems and processes.  This, however, is not an easy trick to pull off.

 

Core systems are those IT systems that maintain the essential ‘books and records’ of a financial institution and maintain details of assets (e.g. loans, or investments) and liabilities (e.g. deposits or insurance premia) and, importantly, customers’ information. The systems are ‘core’ because they support the firm’s main business processes, such as lending, underwriting insurance policies or liability management.

 

A firm may have several ‘clusters’ of business applications that could be classed as ‘core’.  In banking, for example, retail banking systems will often service millions of customers, but also key sets of applications, such as treasury or financial trading, will service tens of thousands of customers and others, such as accounting, will interact with all other systems in a firm, and hence will be difficult to change, without potentially impacting millions of customers. 

 

Figure 3 shows a somewhat idealised ‘core banking system’, with a central computer, or group of computers, that maintain ‘customer accounts’, containing details of the transactions that the bank has transacted with its customers, usually grouped by the products or services sold by the institution. These accounts will typically also hold the ‘current balances’ of customers’ accounts updated by the transactions undertaken. Within, or closely associated with, the customer accounts will be some sort of ‘customer information file’ (CIF), which contains fixed details about customers, such as names, addresses and account numbers.

 

Figure 3 Typical simplified Core Systems (Banking)

 

Each transaction undertaken with customers and interactions with other institutions, such as payments made to other banks, must be reflected in the ‘financial accounts’ of the firm, usually based around some form of ‘general ledger’ (G/L). From these two main sources of information, multiple reports will be produced for the board, management and regulators and also statements will be prepared for customers.

 

The complexity in this relatively primitive IT model arises from having to support the diverse business processes needed to handle each different type of product sold by the firm.  For example, executing a payment on behalf of a customer is very different to extending a loan to a customer, or selling an insurance policy.  In the traditional model, the ‘customer accounts’ becomes a central focal point updated by different product processing applications and used by multiple reporting systems.  So, any major change to information in the customer accounts database(s), such as adding new regulatory classifications, requires testing of the entire system to ensure that the changes are correctly processed.

 

In short, flexibility deteriorates over time as business grows. To counteract this loss of flexibility, and also to maintain stability by reducing the number of changes needed, data is often held outside of the ‘core system’. Figure 4 illustrates such an evolution.

 

 

 

 

Figure 4 Typical Evolution of Core Systems (Banking)

 

In Figure 4, data stores are created outside of the core and used for reporting purposes by merging ‘core data’ with additional separately maintained data, such as marketing information.  Surrounding the original core system, a plethora of essential, ‘tightly coupled’ IT applications, such as marketing and risk management systems, are developed over time.  Pragmatically, such a scheme works, increasing flexibility for specific purposes but, on the other hand, makes the overall ‘system’ much more complex, ultimately decreasing overall flexibility, because few people completely understand the entire system in order to make changes, such as those required by regulators.

 

In addition, since the innovation of automated teller machines (ATMs) in the 1970s, banks’ core systems have moved from operating in an overnight ‘batch’ model to real-time operations.  As a result, often based on expensive, high-performing computer hardware, supplied by companies such as IBM, such systems have become incredibly complex and costly to run.   The emergence of the Internet, in the early 2000s, and the need for uninterrupted 24*7 operations has made banks’ core systems even more complex. And increasing complexity has reduced flexibility, at the same time as requiring improved stability, to support additional customers, around the clock.

 

Over time, a firm’s core systems become, if not obsolete, then expensive and difficult to maintain, especially if a bank’s growth in customer numbers and new products has overtaken its ability to grow its core systems.  IT systems have in-built limitations usually resulting from the restrictions of the technology on which they were first built.  And, even though a core system may have been designed to use the most efficient hardware and databases available at the time, it will not be long before that hardware and software is overtaken by later innovations.  Eventually, after a number of years, all core systems begin to atrophy, with management reluctant to make significant changes fearing the stability of the system may deteriorate and constantly working around the lack of flexibility, by building even more peripheral systems outside of the ‘core’. 

 

The dangers of relying upon such complex ‘systems of systems’ were graphically illustrated in the failure of key customers services for several weeks at the Royal Bank of Scotland (RBS) in 2012, when a relatively simple initial mistake was amplified many times as key databases were being updated independently across interacting and semi-independent systems resulting in lost transactions and incorrect customers’ balances.

 

Another major problem with system longevity is that core systems are often written in computer languages that are no longer taught widely, such as COBOL, and it becomes increasingly difficult to find knowledgeable staff to work on such systems.  Sometimes the computer hardware on which the core systems operate is coming, even has come, to the end of its life and is no longer being actively supported by the manufacturer.  At that point, such core systems are often classified as ‘legacy’ systems, to be supported until they can be replaced but increasingly risky to rely upon.

Legacy Systems

While there is no agreed definition of ‘legacy system’ and some lower risk legacy systems may not be part, or only a minor part, of a firm’s core systems, a legacy system can colloquially be considered as a core system that is ‘out of its warranty period’.

 

When a household appliance, such as a television set or dishwasher, is ‘out of warranty’, it does not mean that the machine will not continue to work quite satisfactorily for a long time but that any repairs or replacement parts will be expensive. As long as nothing goes wrong, an appliance, and for that matter a core system, will tick along doing a perfectly good job.  In such a situation, one would be loath to replace an ageing appliance, but there is always a niggling doubt about when to replace and one would, if prudent, pay for an annual service. If something goes wrong, however, repairs can be expensive and will get more expensive year after year. And as years go by, the manufacturer releases new models that have better functionality and annoyingly are cheaper than the original. Soon parts for the original model become hard to source and if a problem occurs there is often a considerable wait for parts from overseas. And as new models proliferate, engineers who understand the old product move on and repairs will become even more expensive as new staff, untrained in the old product, are involved.

 

With legacy systems, the informal out of warranty period can stretch into decades rather than years, and firms will be loathed to make major changes, even if required by new industry initiatives (such as Faster Payments Service) or regulatory requirements (such as GDPR). In practice, the changes made will be minimal, just enough in order to tick compliance boxes. Sometime a replacement may become imperative, such as, for example, when computer hardware or software is discontinued, as suppliers must also replace their core systems and cannot continue to support unprofitable legacy systems. Replacements can also become pressing when it is no longer possible to hire staff who are familiar and competent with old hardware, software or computing languages.  And as legacy systems age, it becomes more difficult to keep up with competitors who are able to bring new products to market faster, using newer, better and less expensive hardware and software.

 

There are many reasons why an ageing legacy system really should be replaced by a new core system. However, we are dealing with human nature and as long as legacy systems are kept alive, albeit on life support, managers will usually find a supposedly better use for the money needed to replace a core system.  There is unfortunately little kudos in spending a lot of money now to provide benefits that won’t be apparent until 5 or 10 years in the future, when some money spent this year will guarantee this year’s dividends for investors or next year’s bonuses for management.  When, as reported by PWC in 2017, UK CEOs stay in their jobs for just under 5 years on average, and core systems replacement programmes take 5 year or more, a CEO would have to be very brave, or foolhardy, to kick off such a risky change programme on their watch, unless absolutely necessary. 

 

However, it is the board that is responsible for the long-term sustainability and profitability of their firm and as such must consider all of the important risks to their business and technology strategies. And one of the key risks is that of replacing their core systems.

Core Systems Replacement – a Strategic decision

Deciding to replace a firm’s core systems is possibly one of the riskiest strategic technology decisions that a board can make. The problem is one of maintaining a firm’s current IT model operating ‘business-as-usual’ effectively while, first a new IT model is designed and implemented, and then business data and functionality is ‘migrated’ to the desired target model.  It is often referred to as the apocryphal ‘changing the engine while the plane is in flight’, the risk being that business-as-usual may be severely disrupted during the several years of strategy execution, and, as a result, the overall strategic objectives may not get realised.

 

Documenting a ‘Decade of IT failures’ in 2015, the Institute of Electrical and Electronics Engineers (IEEE) observed [8] 

As hard as it is to build IT systems in the first place, it’s arguably even more difficult to maintain them properly over time. In many government agencies, decades of neglect have resulted in a tangled mess of poorly understood and poorly implemented systems that limit operational effectiveness and efficiency. In the past decade, we’ve seen numerous attempts to combine the functionality of such legacy systems into a single modern replacement system [emphasis added].

 

Even though core systems are normally sourced from third-party suppliers, rather than being developed in-house, any replacement project will inevitably be a multi-year effort that requires detailed planning and, often forgotten, detailed risk analysis. Depending on the size of the firm, such an exercise can take 5-7 years to complete from when a decision is taken to move to a new core system, and finally decommissioning the old systems[ii].

 

An important strategic question facing the boards of large institutions is ‘when to change core systems’?  Note it is a question of ‘When’ not ‘If’ a major replacement is needed.

 

Financial institutions spend millions, in some cases billions, of dollars each year in operating their core IT systems.  At the same time, technology is becoming simultaneously cheaper, more powerful and with improved capabilities. As a result, firms have an enormous investment in the status-quo and there are always many reasons, often valid, to prioritise other investments over core system replacements.

 

However, even though older technology also becomes less expensive over time, inexorably, newer technology will be less expensive to operate in the long term and more importantly will be more flexible.  At some point, older technology will become a drag on, rather than an enabler of, performance and, at that stage, a board will be forced to make the investments necessary to replace their core systems.  A far-sighted board will already have made some plans to consider such a situation before replacement becomes inevitable.

 

In 2015, a survey of large financial institutions by the technology research group, Finextra [10], found that “61% of respondents agree or completely agree that core banking replacement will need to take place in the next three to five years” and outside of Europe, the percentage was higher (89%).  However, the survey found that

43% of respondents say they’ve found it hard to get boardroom sponsorship and agreement across business, IT and strategy groups to pursue a modernisation agenda [and …] 36% of all respondents agree or completely agree the risk, cost and failure rate of core replacement programs as preventing them from undertaking such a project. Around half of all those who agreed there was a need also said they are prevented from meeting this need.

 

In other words, there is a problem but it is ‘too difficult’ to tackle just now!

 

In 2016, similar results were found in a survey undertaken by researchers at the University of Surrey [11], which found that

some 40-50% of all IT assets are in urgent need of modernisation. In the case of the UK alone this represents a technology debt estimated to be in the tens of billions of pounds, somewhat equivalent to our national pensions deficit.

 

While changing a firm’s core systems is inevitable, there is rarely a good time to time to begin a full-blooded CSR programme. A focus on short-term profitability as opposed to longer-term sustainability will inevitably mean that it will tempting for senior management to ‘kick the CSR can down the road”. Initiating a CSR strategy takes courage from a board and also needs resolve from them to continue with the strategic plan, as major roadblocks are inevitably encountered during very CSR implementation programme.

Case Studies of Core Systems Replacement Programmes

There many instances of firms failing to implement changes to their core systems but as the IEEE noted that it is difficult to study IT failures outside of publicly accountable institutions, as “private companies tend to bury their IT failure, so except for the rare lawsuit, their operational failures rarely make it into the news”. 

 

However, there are a number of case studies in the financial industry that give some insights into how not to address CSR programmes (and one that worked despite the odds).

 

Co-Operative Bank

 

The unfortunate case of the very public failure of the Co-op Bank’s (COB) strategy of ‘growth by acquisition’ was well documented in the final report of the independent inquiry undertaken by Sir Christopher Kelly [9]. The COB provides some valuable lessons on how NOT to execute a CSR strategy.  It should be noted that the problems encountered by the bank were not merely the result of a bad IT strategy but a failed business strategy overall – and an example of how not to approach corporate strategy in general.

 

In 2007, the Co-operative Bank was not a large institution, with less than 100 branches throughout the UK, but its business model was focused and innovative, providing a full range of services, including Internet banking, to its customers and members. However, it posted relatively modest profits, due to the bank’s higher than average cost/income ratio.  But pure profit was not the bank’s number one priority.  In its 2007 annual financial statement, the bank noted that its ‘vision’ was to be “to be the UK’s most admired financial services business” and its strategies were driven by

A focus on customers who share our belief in the Co-operative difference, developing strong, loyal relationships

 

Technically, the Co-op Bank was not, prior to 2007, a true co-operative as it was almost wholly owned by the Co-operative Group, but that, in turn, was owned and controlled by its members, as opposed to shareholders. Like all other UK banks, the COB was impacted by the GFC but not so much that it was in placed in danger of failing. In 2007, the board’s major strategic initiative was to invest £250 million in expanding its corporate banking business including acquiring vital new technology. 

 

2008 was a monumental year in banking history and the start of what has become known as the global financial crisis (GFC).  But while other UK banks, such as Royal Bank of Scotland, Lloyds TSB and HBOS, had to be rescued by the UK government, the Co-op Bank managed to avoid the pile-up. In fact, 2008 was a fairly good year for the Co-op Bank, and aside from improved profits before adjustments of some £85 million, there was surprisingly, despite the failing UK mortgage market, a small decrease in bad debt provisions.   In the economic circumstances, the firm could be very proud and this was reflected in a 61% increase in customers switching current accounts to the Co-op from the ‘big four’ banks. 

 

At the end of 2008, with minimal fanfare, the COB board announced what was, in effect, a change of business strategy – a merger with the Britannia Building Society (BBS), which was at the time the second largest building society in the UK with over 250 retail branches.  The merger, which followed a change in the law permitting co-operatives and mutual societies to merge, was to take place in early 2009, if approved by Britannia’s members (i.e. shareholders).

 

It should be noted that although nominally a merger, this was in effect an acquisition, retaining the Co-op brand and underwritten by the Co-op Group. The original plan for the merger was to re-brand all Britannia branches as Co-op by 2013 and, after rationalisation of the combined group, to operate some 600 banking branches across the UK.  After a positive vote by members, the CEO of the Britannia, Mr. Neville Richardson, became CEO of the merged group. In what was in effect a “growth by acquisition” strategy [12], both parties saw benefits in the merger. The COB saw the opportunity the merger gave to significantly increase ‘scale’ at what appeared to be little cost and the Britannia board saw the opportunity to acquire current account and internet banking capabilities that it had decided would be impractical to develop themselves. 

 

Before considering the events that followed the merger, it is worth pointing out that ‘growth by acquisition’ is one of the most difficult of all strategies to make work as “most acquisitions fail” to achieve their stated objectives [13]. The reason typically is that it is extremely difficult to make the benefits, arising from the hoped-for synergies between the merged entities, worth the costs of integration.  This is a major execution risk for all acquisition strategies.  And one of the major reasons for the inability to make acquisitions/mergers work is the uncertainty involved – the acquirer (here Co-op) usually does not have the full set of information needed to properly assess the proposed acquisition.  And often, this is a result of pressures to ‘seal the deal’ quickly, before competitors are alerted.

 

At this point it should be noted that, even without significant capital constraints, the merger of two large financial institutions is not a trivial matter requiring large-scale integration of operating procedures and computer systems and rationalisation of two quite different corporate cultures.  Though often overlooked, the integration of two very different IT systems poses a major strategic technology risk to the board’s strategy.

 

Before the merger became a possibility, the Co-op Bank had intended to grow organically and Kelly reported [9] that the bank’s

Main focus was on developing deeper relationships with existing customers and reducing its cost base. In 2008 it had just begun a large-scale replatforming programme to address its high cost (and increasingly high risk) IT systems [emphasis added].

 

It should be noted that the term ‘replatforming’ is not widely used in IT management circles, who tend to use the term Core Systems Replacement (CSR) instead.   For COB, the term ‘replatforming’ implied moving all of their core systems, over time, to a new generation of hardware and operating system software, in other words it was a full CSR strategy. 

 

An example of the problems in replacing core systems can be seen by the fact that the systems used by COB were developed in-house in the 1970s, and hence were older than many of the managers hired to replace them! Kelly noted the problems that typically arise as a consequence of

The age and complexity of the [existing] system, and the many interfaces between its components, meant that the Bank’s technology platform was unstable, expensive to maintain, complex to adapt and ill-equipped to support its business requirements. There were particularly severe problems with the functionality of the online business banking platform [emphasis added].

 

The COB board had in fact considered replacing its core systems in 1996 and 2001, but Kelly reported that

On both occasions the board had decided that the costs and difficulties were too great. It revisited the issue again in 2006-2007 when it was developing a new retail banking strategy. This time it concluded that the weaknesses in the legacy systems, and the need to support a more customer-centric approach, meant that it could no longer defer taking action [emphasis added].

 

A decision to replace core systems is very risky but does have upsides, if the project is successful. As Kelly noted, the COB board (as it turns out over-optimistically) liked the idea of

Leapfrogging the competition and gaining an advantage, or better protecting its position, through improved customer relationship management and quicker delivery of new products.

 

However, this proved easier to say than to do as the complicated history of COB’s IT initiatives shows

 

In 2006, the COB board had agreed a plan, called ‘Project Olympus’, to outsource development of its core systems and some business processes to IBM. But as a socially aware company and fearful of the reputation risk involved in the resulting redundancies and later anticipating a possible merger with Britannia, the COB board cancelled this first CSR project in late 2008.

 

After the acquisition of Britannia in 2009, the new board decided to change their IT strategy and effectively bring the CSR programme back in-house with help from an external ‘systems integrator’ – Infosys.  The bank did, however, retain IBM to deliver an Internet component of the new core system, while its second CSR project was being implemented.  In other words, the bank had decided to change its technology ‘horses in mid-stream’, switching from outsourcing to insourcing and deciding to install a new core system. 

 

In 2009, under a banner of a new ‘transformation programme’, the Finacle banking software package, developed by Infosys, was selected for the core systems replacement software.  While the new CEO was a very experienced executive, he was not an expert in banking, a problem compounded when he brought with him a number of ex-Britannia executives, including a new ‘Director of Integration and Change’ to lead the CSR programme.  While these executives were undoubtedly experienced in operating IT systems for a building society, they were less knowledgeable about the complexity of a regulated banking environment. In choosing to put all of their eggs in the Infosys/ Finacle basket, the COB board and management (possibly from inexperience) completely underestimated the risks involved, a problem compounded when the experienced CIO resigned in 2008.  Kelly concluded that

The Bank's management seems to have underestimated the amount of work reconfiguration would entail both for Infosys and for its own staff. The project would also require significant effort to migrate data from the legacy systems, to build interfaces with other bank applications, to train staff to use the new systems, and so on. Internal presentations and communication, including papers prepared for the Bank board, described Finacle as a 'bank in a box'. This description may have given a misleading impression [emphasis added].

 

The acceptance, at face value, by the board of such overly confident and simplistic assertions by management on such an important topic speaks volumes as to the quality of corporate governance in the new bank.  There should have been more challenge, especially when such a dramatic about-turn in technology strategy was being considered.

 

Unfortunately for the bank, the transformation programme was a disaster.  The technical deficiencies in executing the CSR programme were manifold and a case study in bad project risk management (PRM).   As might be expected from such chaos, there was an enormous cost blow-out in the programme from a projected £ 184 million in 2008 to an estimated £ 460 million just three years later.  Kelly concluded that the COB board was at fault for letting the situation get out of hand, and warned other boards in a similar situation of the

high risks and the size and unpredictability of its likely costs in relation to the Bank’s profitability and capital might all have been expected to cause the Bank’s Executive and board to pay it particular attention [emphasis added].

 

After a number of management changes to try to put the programme back on track, in 2012, the transformation programme was first paused and then subsequently cancelled, at a write–down cost of some £ 300 million, which blew a sizeable hole in the bank’s capital.

 

At this stage, the COB board decided on yet another change in strategy, a plan to acquire some 600 retail branches being spun off by Lloyds Banking Group (LBG), as a result of its forced take-over of the stricken Halifax/Bank of Scotland (HBOS), in an initiative known as Project Verde.  However, the plan was never going to fly, as the then UK banking regulator (FSA) placed quite severe restrictions on the capital that could be injected into the deal by the Co-op Group and also required detailed plans on how the IT systems could be integrated without significant operational risk.  It should be noted that the FSA had already placed the Co-operative Bank on its ‘watch list’ over concerns about its deteriorating capital position and had asked for more detailed plans on how the technology integration of Project Verde would be accomplished.

 

By June 2013, with increasing provisions for loan impairments, the bank’s capital fell below some of its regulatory minimums, prompting Moody’s to significantly downgrade its ratings and its auditor to cast doubt on the bank’s ability to continue as a ‘going concern’.  Technically, the bank had failed, although it was permitted to keep operating, with reduced capital requirements, until a resolution could be affected.  Since 2017, the bank has been put up for sale, but buyers are not rushing to acquire it.

 

It should be noted that the failure to manage its technology strategy was not the main reason for the near-demise of the bank, but was a major factor in the board’s inability to recover from the crisis that beset the bank.

 

In summary, without a clear strategy, plans and objectives, it is fairly meaningless to consider strategic technology execution risks to the bank’s strategy, since there were no measures against which the execution programme could reasonably be gauged.  In effect, the board had no coherent strategic plan, but a series of changes in direction. In such a situation, whatever is attempted will be replete with risks. However, one key decision made by the COB board was particularly risky.  In a merger/take-over it is customary, for very good reasons, to select one of the two core systems as the ‘new core’ and then to move all of the customer records onto that platform.  In the COB case, however, the board made the decision to select a new (i.e. a third) core system and then move both sets of customers and accounts independently onto the new system at the same time.  This made the new core system extremely complex to implement as it had to handle two different sets of integration requirements.  It is little wonder that the two IT programmes, and COB’s overall IT strategy, failed and much of the responsibility for that failure rests with the COB board.

 

TSB Bank

 

The botched implementation of the Core Systems Replacement project at TSB Bank, in April 2018, confirms a truism that no two CSR programmes are quite the same – firms begin at different starting points, have different goals and very different management capabilities.

 

The technology strategy of the TSB board made very good sense. The core systems that TSB inherited from Lloyds Bank Group (LBG) when it was spun off in 2013, as a condition of government support following the banking crisis of 2009, were a typically mishmash of systems accumulated over the years by the mergers of some of TSB’s precursors, including Halifax Bank of Scotland (HBOS).  The new firm was burdened by a set of systems that were already ‘legacy’ when the firm was floated, a situation made worse by the fact that LBG was charging TSB a lot for maintenance of these systems.  In effect, TSB had lost, or never had, control of its technology strategy.

 

In 2015, TSB was acquired by the Spanish conglomerate, Banco Sabadell, with a plan to migrate TSB’s disparate technology platforms to one based on Sabadell's in-house core system, called Proteo, by the end of 2017. This replacement strategy, while ambitious, made sense, as it would better achieve the strategic objectives of the take-over.

 

However, the implementation programme was not well planned, nor crucially were the considerable risks appropriately identified upfront and managed carefully throughout strategic execution. The implementation had two major change programmes, both of which were extensive and in retrospect underestimated in scale and complexity. The first major change programme involved an extensive series of essential changes to Proteo, now re-named Proteo4UK, to handle products specific to TSB branches and regulations particular to the UK.  The second major change programme involved migrating (i.e. moving) all customer account and historical data from the various legacy systems to the new Proteo4UK core system.

 

Migration of customer data from one system to another is always complex and error prone. In the case of TSB, some 1.3 billion customer records had to extracted from various legacy systems, converted to the target format(s) used by the new system and then loaded up to the new system’s database(s).  Such a migration is problematic because any errors or mistranslation in such a move may not be immediately detected but emerge whenever the data is first used by customers. And given the complexity of both the existing and new systems, mistranslated data, even of a minor nature, may cause the new system to fail.

 

There are two main strategies for migrating data: (1) ‘big-bang’: where all of the data to be migrated is converted to the new systems in one effort, usually, because of volumes, over several hours or even days when the target system is unavailable to customers, i.e. there is a distinctive ‘before’ and ‘after’; and (2) ‘staggered’ or ‘staged’ : where data is progressively moved from legacy systems to the new  core system. It should be noted that, for either approach, migration of data is almost always automated, using computer programs specifically written to perform the complex conversions. Like all computer programs there is much room for misunderstanding of the data structures and content, and, of course, programming errors or ‘bugs’.

 

It should be noted that both big-bang and staggered approaches to migration are valid, but both have significant and very different risks and costs associated with them.  It is obvious that for a big-bang strategy, everything has to go (almost) perfectly the first (and only) time it is attempted, as there is no going back. In other words, there are enormous technology risks in any big-bang approach. On the other hand, a staggered approach will allow data errors and programming bugs to be detected and quickly fixed on a smaller subset of customers before the bulk of customer data is migrated in a series of ‘mini big-bangs’.  Of course, while less risky, staggered migration takes longer to complete and hence will inevitably cost more.

 

For the TSB migration, the board and management decided on a ‘big bang’ approach to take place over the weekend of 21st April 2018, with all data being migrated in one coordinated effort.  It should be noted that, as previously reported in a letter from the TSB Union to the Treasury Committee [15], a previous big-bang migration attempt had been cancelled at the last minute in November 2017, because

... work on large parts of the new platform had not been completed on time, testing was still to be done and the testing that had been completed showed the system was not stable enough for customer role [sic] out.

 

The decision to proceed with the same strategy just five months later raises many questions about the strategic technology risk management processes of the TSB board.

A report by IBM [16] into the failures at TSB, which was focused on fixing the immediate and serious problems caused by the botched migration, concluded, inter alia, that

A combination of new applications, advanced use of microservices, combined with use of active-active data centres, have resulted in compounded risk in production … The complexity results in a broad range of technical and functional problems that are hard to diagnose [emphasis added].

 

In other words, the risks inherent in the implementation were not managed effectively. And, hinting that the implementation risks were not fully identified, the IBM report observed that

To address this risk profile, IBM would expect world class design rigour, test discipline, comprehensive operational proving, cut-over trial runs and operational support set-up: [However …]

• IBM has not seen evidence of technical information available to TSB, e.g. architectures, configuration and design documents, monitoring information, test

outcomes, etc.

Performance testing did not provide the required evidence of capacity and the lack of active-active test environments have materialised risk due to issues

with global load balancing (GLB) across data centres.

• IBM has not seen evidence of the application of a rigorous set of go-live criteria to prove production readiness. [emphasis added]

 

While many of the decisions that led to the botched implementation were technical in nature, not least capacity planning and IT programme management, especially as regards testing, it was ultimately the responsibility of the board and senior management to identify the considerable risks in the big-bang strategy that led to the problems foisted on customers.

 

In short, the TSB fiasco can be traced directly back to a failure of Strategic Technology Governance at the level of the TSB board and senior management[iii], in particular a failure of Strategic Technology Risk Management.

 

Danske Bank

In September 2018, Danske Bank, the largest bank in Denmark and one of the largest in the Nordic region, published a report [14] that the bank’s board had commissioned into lapses in Anti-Money Laundering (AML) policies at the bank, in particular within its Estonian subsidiary. The report was devastating in its criticism of AML processes in the Estonian branch, stating that, over a period of several years, “all lines of defence failed” to manage money laundering risks. Soon after the publication of this report, the CEO of Danske resigned, causing the details of the underlying scandal to become public knowledge (although some the issues involved had been aired publicly on a number of occasions previously). It was also revealed that the bank had become the subject of criminal investigations by US authorities.

 

While the events that are covered in the Danske report relate to failures to manage AML risks, the situation is more complex than merely deficient AML controls in a remote branch.  There was a failure to manage a number of different types of risks at both the local and group (i.e. headquarters) level, including: strategic risks; technology risks; legal risks; and other operational risks. As befits a sophisticated modern financial institution, Danske operates a group-wide enterprise risk management (ERM) framework covering multiple types of risk (credit, market, operational etc.).  The fact that the failure to manage the AML risks took several years to come to light, casts doubts on the efficacy of the ERM framework and/or its implementation and of risk monitoring by the Danske board.

 

The root causes of Danske’s problems go back to 2006 when Danske acquired the Finnish- based Sampo Bank, including the bank’s subsidiary in Estonia named AS Sampo Pank, which in 2008, was turned into a branch of Danske Bank.  In announcing the acquisition, Danske noted that Sampo Bank would be integrated “into Danske Bank’s IT platform and organisation”.  Such a strategy, of course, makes sound economic sense as such an integration would be able to maximize synergies and cost savings between the different organizations. 

 

However, in its 2008 annual report, the Danske Board announced a low-key change of strategy: 

on the basis of a cost analysis, the Group decided to discontinue the migration of Banking Activities Baltics to its shared IT platform. But this will not stop future investments in the Baltic banks and their IT departments and local IT systems. The Group will also continue the business integration of the banks and will market Danske Bank products where relevant [emphasis added].

 

In particular, the Estonian branch had its own IT platform, which meant that the branch was not covered by the same customer, transaction and risk monitoring systems as the Danske Bank Group headquartered in Copenhagen.  This also meant that Group risk management and compliance functions did not have the same insights into the operations and risks in the Estonian branch as with other parts of Group. In addition, many documents at the Estonian branch, including information about customers, were written in Estonian or Russian.

 

This lack of transparency proved to be calamitous because a significant number of accounts managed by the Estonian branch, importantly pre- and post-acquisition, were so-called ‘non-resident’ accounts, many of which were based in Russian and other ex-Soviet countries. These approximately 10,000, non-resident accounts were profitable, with payments to and from these accounts accounting for some 2% of profits for Danske in 2013, while accounting for only about 0.5% of the group’s total assets.

 

In 2017, Danske Bank was outed in the press, in the so-called ‘Russian Laundromat’ scandal, as being a regular conduit for rich individuals to launder money from Russia and other ex-Soviet republics (especially Azerbaijan and Moldova) to the UK and other western countries, often via tax-havens. Many of the accused money launderers used the non-resident accounts in the Estonian branch and, to comply with AML legislation, should have been should identified as ‘suspicious’, but were not. As noted in the independent report, “all lines of defence failed” to detect, highlight, report and stop suspicious money laundering activity.

 

This failure, over almost a decade, to detect and act upon potential breaches of AML legislation cannot be placed solely against the bank’s failure to replace the core systems used in the Estonian branches.  There were also many red flags raised, for example, by an internal whistle-blower and importantly by a large US correspondent bank, but those warnings were ignored by the bank’s senior management.

 

However, the failure of the central risk management and compliance functions at Danske Group headquarters to fully monitor activities in the Estonian branch must in part be due to the lack of coherent information flowing from the branch’s local IT systems.  The failure to follow through with the stated strategic intent of fully integrating the technology used by the branches acquired in the Sampo takeover was a serious error of judgement, which must be laid at the door of the board and senior management.

 

Commonwealth Bank

 

Not all CSR programmes are a failure and some despite the odds are successful.  On a narrow technical definition of success, the CSR programme initiated in 2007 by the Commonwealth Bank of Australia (CBA) to replace its ageing core retail banking systems, was far from successful.  The CSR programme ran massively over budget and was late in delivery by a number of years. However, from a strategic perspective, the CSR programme was a success, leaving CBA in what appears to be a strong and sustainable strategic position in the local Australian banking market.  And, although the programme was late and significantly over-budget, the strategic execution was commendable without major customer disruption.

 

In order to understand CBA’s CSR programme, it is important to understand why the CBA board took the strategic decision to replace its core retail banking systems in the first place. It is a story of two very different strategic technology decisions over 15 years: first outsourcing IT; then bringing the core systems back in-house to replace them.

 

In 1997, the CBA board took the decision to outsource the bulk of its retail banking systems, including systems development and support, to EDS, a large US technology conglomerate.  The core systems outsourced by CBA were around 30 years old at the time and in dire need of replacement, but replacement was not included in the outsourcing contract. The original contract included reassignment of many of CBA’s IT staff to EDS, in fact leaving a gap in CBA’s internal expertise as regards systems functionality and capability.  According to the bank’s CIO at the time, the outsourcing contract was to last some ten years and could be renewed after that time.

 

In 2007, the CBA board was faced with a stark strategic choice, either to renew the EDS contract for another ten years, or to attempt some other strategy. By this stage, the board had already decided to break its outsourcing contract into multiple IT components, so-called ‘multi-sourcing’, changing the role of the prime contractor, EDS.  Details of the board’s deliberations have not been published but one option considered was to replace the core banking system, then being maintained by EDS, with a new core system. 

 

The risks were significant. While the EDS contract was operating reasonably well, although as might be expected, there were some problems, the contract did not (in practice) cover major upgrades in functionality.  And in the meantime, the bank’s core banking system was growing ever older and, as the use of the Internet expanded, increasingly deficient in customer functionality.  It was past time for a change, so it possible to conclude that the board had few options other than replacing its core system, with all of its attendant strategic and technology risks.

 

In 2008, the CBA board announced a new strategic ‘core modernisation’ programme in which its core retail banking system would be replaced by functionality to be provided by SAP, a large German software supplier, operating on IBM mainframes with ‘integration support’ to be provided by Accenture, a large consultancy.  The projected cost of the programme, called core business modernisation (CBM), was estimated at AU$ 580 million and delivery was scheduled to take place over 4 years.

 

But, as the CIO admitted at the time, this choice was not without risks as the SAP software package was not widely used in banking at that time and the programme would be a major undertaking for the bank as it had already released many key IT staff as part of the outsourcing contract.  At the outset, no specific strategic IT objectives were articulated, but in announcing the programme the CEO outlined the rationale as

The modernisation of our core systems is a fundamental strategic plank to the successful future of the Group. It will demonstrate to our customers our determination to be different by providing them with excellent customer service.

 

In short, the programme recognised the need for improved customer service, in particular the need for real-time services.

 

In a presentation to shareholders in 2011, which lauded the performance and flexibility of the new system being implemented, the CIO provided details of the state of the CSR programme and its various staged/staggered phases:

1)     2009: migration of 53 million (presumably historical) “customer records” to the new SAP system, developing new deposit products and integration of Internet banking systems into the new core platform;

2)     2010: migration of some 11 million retail accounts to the new system; and

3)     2011: business accounts to be migrated, and new banking products to be built directly onto new system.

 

The 2012 CBA annual report, reported that the cost of the CSR programme had reached a total of some AU $1.1 billion, the multi-year project was “nearly complete”, and the bank aspired to” become a global leader in the application of technology to financial services”.  In the 2013 annual report, the board noted that the CSR project had been completed, with an additional cost of some AU$ 200 million for that year.  In short, the CSR programme had been completed at a cost of over 100% of the original budget and at least 18 months (almost 40%) over schedule.  Nonetheless, it had been completed, was working well and had given the bank a lead over its local competitors in technology. Despite the cost blow-out, the CSR programme was considered to be a great success by CBA management.

 

But what was the secret of the eventually successful CSR strategy? Undoubtedly part of the success was that the CBA board remained focused on the execution of the IT strategy, and was prepared to spend more than budgeted to complete the programme. Crucially, the board did not deviate from its strategic plan, although the bank had in the interim acquired BankWest, the Australian subsidiary of Halifax/Bank of Scotland (HBOS) which had failed in the aftermath of the GFC. Obviously, the stability and flexibility of the SAP banking package was crucial to the programme’s success but, probably more so, the strategic plan for migrating data in stages, by product type, was also very important.

 

National Australia Bank

At roughly the same time as Commonwealth Bank embarked on its CSR programme, one of the other ‘four pillars’ of Australian banking, National Australia Bank (NAB) began a similar odyssey but took a very different path.  In mid-2008, the then CIO of NAB announced a small programme, estimated at AU$ 30 million, to build the first phase of its envisaged core systems replacement based on the i-flex product from Oracle.  This was not an unreasonable ‘toe in the water’ approach to evaluating new technology.  At the same time, the CIO noted that NAB had earmarked AU$ 1 billion for IT expense over the next five years, but gave no further details at that time.

 

NAB faced the common core systems problem in that there were over 100 legacy systems, some more than 40 years old, and over 20 data centres of various vintages, mainly as a result of aggressive expansion plans in 1980s and 1990s.  In 2009, however, the CIO was replaced and the CSR programme, now dubbed NextGen, was started.  Interestingly, the programme was not discussed in NAB’s annual reports, other than noting that the projected costs had been considered. 

 

In mid-2013, NAB published a rare market update on the NextGen programme, which announced that the bank was “well advanced on 10-year transformation program”.  While some Internet functionality had been delivered much of the activity in the years since 2009, had been focused on rationalising the firm’s IT infrastructure and by 2014, it was hoped the plethora of small data centres would be replaced by a new “state of the art” data centre in Melbourne. The presentation noted that “productivity gains substantial – more to come … building a sustainable operating model for competitive advantage”.

 

As NAB’s management provided few updates on the progress of the NextGen programme to stakeholders, much of the public information about the programme comes from interviews with management and press releases following the many staff re-organisations.

 

In late 2013, another restructuring of the IT department took place with the new head of the ‘enterprise services and transformation division’ noting that NextGen “is the largest and most comprehensive transformation agenda of all the major banks, certainly in this country” and promising that the bank’s 4 million customers would be migrated within the next 3 years. However, in early 2014, the projected timeframe had already slipped and migration was now planned to be completed in 2017.  By mid-2014, another restructuring of the programme had occurred, coinciding with the arrival of a new CEO, with the new head of ‘enterprise services and transformation’ admitting that migration of customers to the new core system would be slowed down, but giving no reasons other than that the programme was very complex.

 

In late 2014, the new CEO, Andrew Thorburn, changed the trajectory of the NextGen programme, admitting that the programme had “taken too long and cost too much” and that going forward, the focus would be on delivering additional functionality as “We have learned that we can’t run multi-year commitments on technology projects …we have to focus on a few things that are significant and do them well.”

 

In writing down some AU$ 106 million in software costs in 2014, the CEO’s change of direction was an admission, if not of defeat but then of a significant retreat from the original NextGen concept. Two years later, in 2016, the new ‘incremental’ approach allowed the bank to roll out the final piece of its retail banking system, i.e. mortgage origination.  However, in 2017, a number of key capabilities, in particular a ‘single customer view’, remain to be delivered.  No details of the cost-overrun have ever been given by the bank and hence it is difficult to gauge the success or otherwise of the programme.  It may however, despite the many problems, be considered at least a partial success, and one which NAB’s management fervently hopes will yield benefits in the future.

 

Although senior management, in particular the CEO, became more involved in the CSR programme in 2015, the history shows that the NAB board was not deeply absorbed in the programme, maybe because of the problems that were being experienced in its overseas subsidiaries, in particular the Clydesdale and Yorkshire banks in the UK. In hindsight, this was a major strategic governance risk, as not enough focus had been placed on the CSR programme.  And while it is not possible to state definitively that the NextGen programme would have been more successful if there had been more board-level focus, the difference between the success of CBA and the messy end of NAB’s programme suggests that leadership from the board is a key factor in the success of any CSR programme.

Key lessons from the cases

Some key lessons can be learned from the failures and limited successes of the various CSR strategies described in the cases above. First, it is important to recognise that technology is critical to the long-term sustainability of every bank and in particular core systems must be considered strategically, not tactically.  The Co-operative Bank (COB) board did not fully recognise the importance of its IT systems to its ambitious acquisition strategy.  Likewise, at the outset. the NAB board did not sufficiently engage in the NextGen programme and it ran away for some time. And the TSB board did not give the due consideration to the technology risks it was running in its, very sensible, CSR programme.

 

Second, the decision when (not if) to replace core systems must be reviewed regularly and, unfortunately as the COB case shows, there is no right time to replace core systems as the current situation (as regards existing technology) and targets (as regards future technology) were constantly changing.  Not only did the COB and NAB boards leave their CSR strategies until too late, the focus of their IT strategies kept changing as new business strategies emerged. 

 

Boards should note that, any major CSR programme will take considerable time to implement, certainly in the same timeframe as any corporate strategy, or possibly longer.  This means that any discussions on business strategy and IT support for that strategy must be held at the same time.  Furthermore, both sets of risks must be considered in tandem, as they impact the success of each other. The COB, NAB and TSB boards seemed to consider IT strategy to be a technical rather than strategic business issue and did not focus on, nor rein in, IT plans sufficiently.  On the other hand, although the CSR strategic plan ran over-budget and over time, the CBA board maintained their focus on completing the programme.

 

Importantly, any strategic changes to core systems must be closely coordinated with business-as-usual operations, as the longer a CSR programme runs, the more difficult it is to delay necessary changes, such as regulatory requirements, to existing core systems. This appears to have been part of the problems encountered by NAB as the programme was delayed by a change in focus to infrastructure rather than applications, meaning that existing applications had continue to be changed over a prolonged period.

 

Lastly, any large-scale IT programme and especially for a CSR strategy, engaging specialist help is essential. In CSR programmes, such assistance is often termed a ‘systems integrator’, usually in the form of a large consulting firm, such as Accenture in the case of CBA’s CSR strategy.  On the other hand, the COB board’s decision to employ Infosys as both the main software supplier and also the systems integrator created conflicts of interest that may have jeopardised the risk management of its CSR programme. However, it is important that overall responsibility for the programme is not outsourced and remains with the board and designated senior management of the firm. Furthermore, it is imperative that the systems integrator risk function becomes a part of the overall strategic risk management organisation.

Recommendations on Core Systems Replacement

Recommendations on Corporate Governance

A firm’s core systems will typically follow a lifecycle characterised by: an initial period of peak productivity and flexibility following its inception; which is followed by many years of solid, cost-effective, if increasingly inflexible, service, day in, day out; followed by a period when systems are placed in ‘maintenance mode’, where only necessary changes are made; finally ending with a period of instability and eventually  (often forced) replacement. The full lifecycle may take decades to complete and intervening periods may be extended by fortuitous external changes, such as a supplier providing more powerful computers that are compatible with existing hardware.

 

Because the lifecycle of any core system is potentially so long, only the board of a firm can have the big picture view of the total life of a firm’s technology. Since the board has overall responsibility for a firm’s business strategy, the directors are also responsible for ensuring that the technology strategies needed to support their business strategy are in place and being properly managed. And the board is also responsible for ensuring that any risks being taken are within the firm’s risk appetite, including the board’s appetite for technology risk.

 

And in order to ensure that risks are being managed by the firm’s management must set up and monitor the firm’s risk management framework.  Part of that framework will be the identification and management of a firm’s technology risks, especially the risks related to the firm’s core systems.  And in order to understand those risks, directors and senior management must have a comprehensive understanding of:

1)     where their core systems are, in terms of the core systems lifecycle;

2)     what are the technology and business risks in their core systems and how will those risks change over the next few years;

3)     what is the timeframe and business pressures to undertake a CSR programme;

4)     what impact will other technology and business initiatives, such as an acquisition, have on the firm’s existing core systems;

5)     And if a CSR programme is being considered, what are the risks in undertaking such a programme;

6)     And if a CSR programme is already underway, what are the implementation risks, and are they being managed properly.

 

Obviously, such an understanding can only be gained by ensuring that the board has sufficient directors suitably qualified and experienced in the governance, management and development of complex technologies.

 

It is recommended, therefore, that boards of directors of banks should:

1)     Ensure that the board has a sufficient number of directors with technology expertise and experience, including directors that are experienced in operating and/or developing core systems;

2)     Develop a formal (if outline) strategy for CSR, even though the programme may be years in the future to allow all new IT initiatives to be benchmarked against the long-term CSR strategy to ensure that conflicts do not arise and risks are assessed;

3)     Evaluation of Core Systems (CS) risks is made a key item on board agendas at every meeting and, at least once a year, a formal full board off-site is convened to discuss core systems and technology strategy and risks;

4)     Create a high-level function with specific responsibility for developing and maintaining a Core Systems Inventory (CSI) and, in conjunction with Corporate Risk Management functions, assessing the risks in the current environment and any proposed initiatives.

 

Recommendations on Regulation

Since the failure of a CSR programme in a large financial institution may have consequences for the overall system, increasing systemic risk, regulators must ensure themselves that boards undertaking such a programme have a comprehensive understanding of the risks involved and have set up the necessary risk management frameworks and organizations.

 

It is recommended, therefore, that any firm undertaking a CSR programme should automatically be identified as having a higher risk rating than comparable institutions and should be allocated specialist supervisors to monitor the effectiveness of the firm’s risk management framework, with respect to its CSR programme.

 

But while such active supervision will cover firms that are embarking on a CSR programme, it does not identify firms that should be undertaking a long overdue CSR programme, i.e. the laggards.  To pick up potential systemic risks from firms that are not addressing the need for a CSR program, it is recommended that regulators should mandate that firms produce a report that details the risks in their core systems environment. Ideally, such a report should be completed by an independent qualified analyst on a regular basis (e.g. bi-annually) and be audited by a firm’s internal IT audit department.  It is also recommended that supervisors should require that firms produce concrete plans to mitigate any serious core systems risks, if necessary, initiating a CSR programme.

 

Overall, it is recommended that the risks of undertaking, or not undertaking, the necessary Core Systems Replacement programmes should become a major focus of both boards and their financial regulators.

 

Part D- Authors

Martin Walker

Director for Banking and Finance at the Center for Evidence-Based Management

 

Author of multiple papers and articles on Capital Markets technology, Blockchain and cryptocurrencies. Previous experience includes, Global Head of Securities Finance and Treasury IT at Dresdner Kleinwort, Global Head of Prime Brokerage Technology at RBS Markets, and Senior Business Analyst in Equity Finance Technology at Merrill Lynch. Provided written and oral evidence to the Treasury Committee inquiry into Digital Currencies.

 

Qualifications


Books

Evidence-Based Management: How to Use Evidence to Make Better Organizational Decisions (2018) - Kogan Page by Eric Barends and Denise M. Rousseau

 

Patrick McConnell

Honorary Fellow at Macquarie University Applied Finance Centre

 

Almost 30 years as a manager, consultant, and academic with an extensive knowledge of the application of Information Technology to global financial markets. Worked with board and senior level management of large international financial institutions in the UK, USA, Europe and Australia to develop strategies for applying IT to their expanding businesses and with all levels of operational management to deliver business change. Published author, speaker and chairman conference on Risk Management and IT issues.

 

In his career, Dr McConnell has, in addition to innovative IT projects, worked on several Core Systems Replacement programmes, such as for example UK Big Bang in 1986, and most of these large programmes were considered to be successful, but some less so.

 

Qualifications


Books

 


References

[1] Rogers E. M., 2003, Diffusion of Innovations 5th Revised edition, Simon & Schuster International, New York

[2] McConnell, P. J., 2017, Strategic Technology Risk, Risk Books, London

[3] Gartner Group, 2018, Hype Cycle, https://www.gartner.com/en/research/methodologies/gartner-hype-cycle

[4] McConnell P.J., 2018, “The Governance of Strategy and Strategic Technology Risks” in Fintech: Growth and Deregulation, Risk Books, London

[5] Merton R. C., 1990, “The Financial System and Economic Performance”, Journal of Financial Services Research

[6] Tufano P., 2002, “Financial Innovation”, in Handbook of the Economics of Finance volume 1, Elsevier BV, Amsterdam

[7] Christensen C. M., 2000, The Innovator’s Dilemma, Harvard Business School Publishing, Boston, MA

[8] IEEE, 2015,” Lessons from a Decade of IT Failures”, Institute of Electrical and Electronic Engineers, http://spectrum.ieee.org/static/lessons-from-a-decade-of-it-failures

[9] Kelly C., 2014, ‘Failings in Management and Governance: report of the independent review into the events leading to the Co-operative Bank’s capital shortfall, published at the request of the Co-op Group and the Co-op Bank’, www.thekellyreview.co.uk

[10] Finextra, 2015, ‘Invigorating Banking’, Finextra Research, London, www.finextra.com

[11] University of Surrey, 2018, Escaping Legacy Removing a major roadblock to a digital future”,
https://www.surrey.ac.uk/centre-digital-economy/publications/reports

[12] McConnell, P. J., 2016, Strategic Risk Management, Risk Books, London

[13] Rankine D. and Howson P., 2014, Acquisition Essentials: A step-by-step guide to smarter deals, Pearson United Kingdom.

[14] Danske Bank, 2018, “Report on the Non-Resident Portfolio at Danske Bank’s Estonian branch”,
https://danskebank.com/news-and-insights/news-archive/press-releases/2018/pr19092018

[15] TBU, 2018, “TSB Integration Shambles Letter to Rt Hon Nicky Morgan MP from Mark V Brown General Secretary TBU, 6th June 2018

[16] IBM, 2018, “IBM CIO Leadership Office -Update for TSB Board”, https://www.parliament.uk/documents/commons-committees/treasury/Written_Evidence/tsb0003.pdf

[17] https://www.handbook.fca.org.uk/handbook/REC/3/16.html

[18] https://www.fca.org.uk/news/speeches/cyber-and-technology-resilience-uk-financial-services

[19] https://www.fca.org.uk/publications/research/cyber-technology-resilience-themes-cross-sector-survey-2017-18

[20] https://www.fca.org.uk/publications/discussion-papers/dp-18-4-building-uk-financial-sector-operational-resilience

 

Submitted January 2019


[i] One of the authors of this paper worked on a system for analysing TV viewing. Every year the system fell over when the data was loaded covering the period when the clocks changed between GMT and BST.

 

[ii] Often the costs of decommissioning are underestimated, or even forgotten, and may be significant if IT programmes are delayed.

[iii] At the time of writing in early 2019, a report into the failures by law firm, Slaughter and May, which was initiated by the TSB board, has not yet been published although an interim report was promised by the end of 2018. That report may, or may not, investigate any governance failures of the board.