Written evidence submitted by Professor Nigel Harvey (ALG0054)
1. Witness Background
1.1 Nigel Harvey is Professor of Judgment and Decision Research in the Department of Experimental Psychology at University College London and Visiting Fellow in the Department of Statistics at the London School of Economics and Political Science. He is a past-president of the European Association for Decision Making (EADM) and is on the editorial boards of the International Journal of Forecasting and the Journal of Behavioral Decision Making. Much of his work has dealt with the use of judgment in forecasting, monitoring and controlling the behavior of systems. His current research focusses on how people combine their judgment with forecasts produced algorithmically, on the effectiveness of the resulting forecasts, and on ways of improving that effectiveness. It is his knowledge of this area that prompted this submission of written evidence to the inquiry.
2. Executive summary
3. The role of accountability in determining how algorithms are used in modern society
3.1 Recently, algorithms have been applied to a much greater range of tasks than before. Their application to advertising and use of recommender systems in e-commerce has triggered concern that they lack transparency and may be biased. Though research into algorithms in e-commerce is still limited, systematic study of their use in other contexts dates back to the 1950s. These contexts include medicine, control in nuclear and chemical process industries, financial trading, and forecasting.
3.2 It is likely that the rapid application of algorithms to e-commerce has arisen because a) they allow advertisements and recommendations that are tailored to users to be generated quickly and cheaply and b) those implementing them are not accountable for the effects of their use.
3.3 When decision makers are accountable for the effects of using algorithms, their uptake has often been slow, even when evidence indicates that they typically improve decision making. In clinical decision making, Meehl [1] showed that algorithmic methods for combining diagnostic information to produce prognosis and treatment decisions outperformed clinical judgment. Recent meta-analyses support his findings [2, 3]. Despite Meehl’s conclusions, clinicians still prefer to use judgment over 50 years later [4]. Research on forecasting has produced similar conclusions. A 2007 survey indicated that 25% of business forecasters used purely algorithmic approaches [5]; a similar survey carried out in 2014 showed an increase but only to 29% [6].
4. Algorithm aversion: Usually but not always undesirable
4.1 Why are decision makers who are accountable to those who are affected by their decisions reluctant to use algorithms? Twenty-five years ago, Bainbridge [7] discussed the ‘ironies of automation’ - ways in which algorithms may increase rather than eliminate problems for decision makers. When algorithms are introduced, human operators no longer automatically maintain the skills they need to operate the system themselves. To deal with emergency situations beyond the capability of the algorithms, precautions must be taken to ensure that operators retain the skills that are needed when algorithms fail. Complete dependence on algorithms can be avoided either by occasionally operating in the real situation without them or by practice using simulators: the Airbus A320 fly-by-wire system was not sufficient to enable US Airways flight 1549 to ditch in the Hudson River [8]but, fortunately, Captain Sullenberger’s well-maintained skills were.
4.2 There are other sound reasons for algorithm aversion. Just as people’s proficiency at performing tasks varies, so does that of algorithms. Although evidence broadly supports Meehl’s conclusion that algorithms outperform human judgment, there remain some tasks in which experts outperform the most sophisticated algorithms currently available. For example, whereas algorithms help most radiographers to detect cancers in mammograms, their input hinders the most expert radiographers [9]. Performance of those experts would benefit from algorithm aversion.
4.3 Despite these cases where algorithm aversion appears to be sensible, there remains a broad consensus consistent with Meehl’s views: people’s aversion to algorithms is generally not beneficial and it prevents performance in a wide variety of tasks from being improved. Recent research indicates that people treat errors produced by algorithms more harshly than those produced by people: they more quickly lose confidence in an algorithm after it makes a mistake than they do after a person makes the same mistake [10]. Similarly, they take less account of (exactly the same) advice when told it has been produced by an algorithm than when told it comes from a person [11]. These findings may be related to the fact that decision makers describe themselves as having much more in common with human advisors than with algorithmic ones [12].
5. Reducing and avoiding algorithm aversion: Opportunities and risks
5.1 One way of reducing algorithm aversion is to give users some degree of control over the final decision by allowing them to make adjustments to the algorithmic output. Giving decision makers such control makes them more satisfied with the decision-making process and more likely to opt for an algorithmic approach in future [13]. However, giving them this control over the final decision may lead to a worse outcome. For example, sales forecasters making predictions for the future demand for their firm’s products typically make adjustments to algorithmic forecasts. Small adjustments, perhaps intended to signal ‘ownership’ of the forecast, typically reduce accuracy. Large adjustments tend to increase accuracy only when they are in a downward direction [14]. Hence, while allowing adjustment provides an opportunity for increased use of algorithms (and, hence, increased accuracy), its effect on the final decision also carries a risk that the adjustment will reduce accuracy.
5.2 Other factors that increase acceptability of algorithms include their endorsement by a trusted expert source, users’ understanding of how they work (transparency), users’ involvement in their design, and an active choice by users to adopt them [15].
5.3 This implies that making use of algorithms mandatory may not be a good way of avoiding algorithm aversion. Evidence supports this inference. When forced to use algorithms, decision makers have a tendency to manipulate the inputs to them to ensure that they obtain the output that they prefer. For example, auditors manipulated an algorithmic aid designed to help them to determine the appropriate sample size for an audit. They decided what sample size they wanted to use and worked backwards to find the inputs to the algorithm that would produce it. As a result, their sample sizes tended to be too small for the risk level and there was low consensus between auditors [16, 17].
5.4 Mandatory use of non-transparent algorithms also carries a risk that operators will experiment to find out how they work. For example, the Chernobyl disaster appears to have been caused by the plant staff violating safety regulations in order to carry out an experiment to discover the operating limits of the algorithm controlling the plant [18].
6. Introducing accountability to reduce proliferation of algorithms that produce undesirable effects
6.1 Those implementing algorithms to generate web-based advertising and to support recommender systems in e-commerce are not currently accountable to those affected by the undesirable effects of their activities. Given what we know about the effects of accountability [19], we should expect that introducing it would increase efforts to ensure that effects of algorithms are socially acceptable.
6.2 There are a number of ways in which web-based marketing could be self-regulated. I will suggest three possibilities here but it is for others to decide how practicable they are.
6.3 Large organizations often contain a compliance team, sometimes attached to their legal department, to ensure that members of the organization operate within an agreed code of practice (e.g., http://www.corporatecomplianceinsights.com/effective-corporate-compliance-programs-ron-kral-candela/). Such teams provide training, deal with complaints, and investigate possible infringements that may leave the organization open to legal action. This approach might be appropriate for introducing a degree of accountability into large organizations (Google, Facebook).
6.4 Often problems arise because the algorithms (e.g., search engines) developed by large organizations are manipulated for commercial, political or other reasons by small organizations or individuals. They use their knowledge of how the algorithm works, often in conjunction with coding skills that implement machine learning, to produce non-representative outcomes. For example, the first results produced by a search term may be for extremist sites that are not representative of the population at large. Clearly, the approach to accountability suggested in the previous paragraph would not be appropriate for dealing with third party activities such as these.
6.5 Large scientific and medical publishing companies now include research integrity teams (e.g., https://www.biomedcentral.com/about/who-we-are/research-integrity-group). The need for them arose because they discovered that authors and editors were occasionally using unethical and corrupt practices to increase the chances that their papers would be accepted for publication. The research integrity teams investigate the suspicious activities of these third parties and, where necessary, take action. Large software companies (Microsoft, Google) could be asked to set up algorithm integrity schemes that would monitor misuse of the systems that they have developed and then act in a manner agreed with a regulator.
6.6 Professional societies issue codes of ethics and conduct. Members who do not abide by the code are excluded from the society. The Association for Computing Machinery (ACM) has issued a Software Engineering Code of Ethics and Professional Practice (http://www.acm.org/about/se-code). The ACM is the main international society for software engineers and their code is clearly relevant to algorithm misuse. However, many of those responsible for producing the algorithms that have undesirable effects will not be members of the ACM and will not want to be.
April 2017
References
1. Meehl, P. E. (1954/2013) Clinical vs statistical prediction. Brattleboro, VT: Echo Point Books and Media.
2. Grove, W. M., Zald, D. H., Hallberg, A. M., Lebow, B., Snitz, E. and Nelson, C. (2000). Clinical versus mechanical prediction: A meta-analysis. Psychological Assessment, 12, 19–30.
3. White, M. J. (2006). The Meta-Analysis of Clinical Judgment Project: Fifty-Six Years of Accumulated Research on Clinical Versus Statistical Prediction. The Counseling Psychologist. 34, 341–382.
4. Grove, W. M. (2005). Clinical versus statistical prediction: The contribution of Paul E. Meehl. Journal of Clinical Psychology, 61, 1233-1243.
5. Fildes, R. and Goodwin, P. (2007). Good and bad judgment in forecasting: Lessons from four companies. Foresight, Fall, 5-10.
6. Fildes, R. and Petropoulos, F. (2015). Improving forecast quality in practice, Foresight, Winter, 5-12.
7. Bainbridge, L. (1983). Ironies of automation. Automatica, 19, 775-779.
8. Greenspun, P. (2010). Review of Fly-by-wire: The geese, the glide and the miracle on the Hudson by William Langewiesche, http://philip.greenspun.com/book-reviews/fly-by-wire
9. Povyakalo, A.A., Alberdi, E., Strigini, L. and Ayton, P. (2013). How to discriminate between computer-aided and computer-hindered decisions: A case study in mammography. Medical Decision Making, 33, 98-107.
10. Dietvorst, B. J., Simmons, J. P. and Massey, C. (2014). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144, 114-126.
11. Önkal, D., Goodwin, P., Thomson, M., Gönül, S. and Pollack, A. (2009). The relative influence of advice from human experts and statistical methods on forecast adjustments. Journal of Behavioral Decision Making, 22, 390-409.
12. Prahl, A. and van Swol, L. (2017). Understanding algorithm aversion: When is advice from automation discounted? Journal of Forecasting, In Press.
13. Dietvorst, B. J., Simmons, J. P. and Massey, C. (2016). Overcoming algorithm aversion: People will use imperfect algorithms if they can (even slightly) modify them. Management Science, Published online articles in Articles in Advance Nov 2016.
14. Fildes, R., Goodwin, P., Lawrence, M. and Nikolopoulos, K. (2009). Effective forecasting and judgmental adjustments: An empirical evaluation and strategies for improvement in supply-chain planning. International Journal of Forecasting, 25, 3-23.
15. Kaplan, S.E., Reneau, J.H. and Whitecotton, S. (2001). The effects of predictive ability information, locus of control, and decision maker involvement on decision aid reliance. Journal of Behavioral Decision Making, 14, 35-50.
16. Kachelmeier, S.J. and Messier, W.F. (1990). An investigation of the influence of a non-statistical decision aid on auditor sample size decisions. The Accounting Review, 65, 209-226
17. Messier, W.F., Kachelmeier, S.J. and Jensen, K.L. (2001). An experimental assessment of recent professional developments in nonstatistical audit sampling guidance. Auditing, 20, 81-96.
18. United Nations Scientific Committee on the Effects of Atomic Radiation (2002-2017). The Chernobyl accident: UNSCEAR’s assessments of the radiation effects. (http://www.unscear.org/unscear/en/chernobyl.html).
19. Lerner, J. S. and Tetlock, P. E. (1999). Accounting for the effects of accountability. Psychological Bulletin, 125, 255-275.