Dr Lauren Harrington - Written evidence (POL0002)
House of Lords Public Service Committee
The use of transcripts in the criminal courts
I am a Postdoctoral Research Associate in the Department of Language and Linguistic Science at the University of York. I have spent the last 5 years researching transcripts used in the criminal justice system, with a focus on the transcription of forensic audio materials by experts and the use of AI-based Automatic Speech Recognition (ASR) systems to produce transcripts of police-suspect interviews. I am about to start a Postdoctoral Fellowship continuing and consolidating the research carried out during my PhD, and for the last 14 months I have been working on a project also exploring the use of AI-based speech technologies in forensic speaker comparison.
This submission draws upon my own research, carried out in collaboration with fellow academics, namely at the University of York[1], and with leading experts in forensic speech and audio analysis. The parts of this submission focused on forensic audio transcripts have received input from Dr Richard Rhodes, an expert practitioner in forensic speech and audio analysis, and a Director of the Forensic Voice Centre, one of the UK’s main private forensic speech and audio laboratories.
Within this submission, I will comment on the production, evaluation and presentation of forensic audio transcripts only, along with the use of AI for both forensic audio and police-suspect interview recordings.
● The transcription of forensic audio materials, in which the speech tends to suffer from issues related to intelligibility, has many challenges that must be considered in order to produce reliable transcripts. Experts are regularly instructed to produce transcripts of this type of audio, and use a wide range of highly-specialised skills and generally employ carefully-managed procedures that include techniques to mitigate the effects of cognitive bias.
● There is currently no regulation for experts in this particular field of forensic science and, to my understanding, even less regulation or guidance for non-experts who likely produce the majority of transcripts of this type of audio material.
● The answer to all of the issues related to transcribing forensic audio materials or police-suspect interviews will not be solved by Artificial Intelligence. There may be a place for AI within the process of transcribing better quality audio recordings within law enforcement, such as police-suspect interviews, but this should be part of a carefully-managed and scientifically-informed procedure that takes into account the associated risks.
In this response, I am addressing the preparation of forensic audio transcripts.
Something that is key to establish prior to answering the questions within this Call for Evidence is what is meant by the term ‘forensic audio’. ‘Forensic audio’ could be understood to mean any audio that is used for evidential purposes. In the field of forensic speech science, we often consider ‘forensic audio’ to be poor quality, indistinct recordings which contain (potentially) evidential material and are very difficult (or nearly impossible) to understand without the help of a transcript. In such cases, experts in forensic speech and audio analysis can use their experience, knowledge of language and speech perception, and skills in phonetic analysis to produce an independent transcript of the audio recording using a carefully-managed procedure. Experts may be involved during either the investigative process or the evidential process (or both), and will produce a report alongside the transcript detailing the limitations of the audio and stating that transcripts are interpretations and could be amended or reviewed with further relevant information.
To my knowledge, there are three main scenarios in which an expert may be instructed to produce an evidential transcript:
● The audio quality of the recording is very poor and some part(s) of the speech content is challenging to understand
● There is a dispute over the speech content of an audio recording and an expert is instructed to produce an independent transcript and comment on the existing interpretations put forth (this is called questioned content analysis)
● The court requires an independent, impartial transcript of an evidential audio recording
Examples of different types of recordings[2] that may be sent to experts in forensic speech and audio analysis include:
● Telephone calls to the emergency services (often where an event or the speech of interest is taking place in the background)
● Other types of telephone calls (e.g. fraudulent insurance calls)
● Covert or undercover officer recordings (made using hidden recording devices)
● Recordings made (non-)deliberately by a participant (deliberate or inadvertent recording of evidentially-salient speech)
● Recordings from surveillance devices (e.g. CCTV or Ring doorbell footage)
● Smartphone recordings (e.g. video or audio materials related to an offence that have been downloaded from a phone handset or posted to social media platform, or voice messages from Whatsapp or other similar applications)
Typically, the expert does not deal exclusively with very poor, indistinct audio. There will likely be cases where the audio quality is limited but the speech is relatively clear throughout much of the recording. Within this submission, I am therefore interpreting ‘forensic audio’ to mean evidential audio recordings where there are issues surrounding intelligibility, which may be related to technical factors such as the audio quality or may be related to the speech or the nature of the interaction itself.
The issue with a lack of specific definition for ‘forensic audio’ means that the decision to engage an expert is less clear-cut than other types of expert forensic speech and audio analysis, such as forensic speaker comparison casework. As previously mentioned, audio is often sent to experts for transcription (1) when it is hard to understand what is being said within the recording, and/or (2) to get an independent, impartial transcript. However, in context (1), it may be that a police officer has listened to the audio and written down what they think they have heard, and their interpretation is not contested so the transcript is submitted into evidence. Whether the audio gets sent to an expert can depend on factors such as (a) whether the police officer or Crown Prosecution Service worker on the case is aware of experts in forensic speech and audio analysis and the type of work they carry out, (b) whether the audio is considered to be ‘poor quality enough’ to send to an expert, or (c) whether the instruction of an expert is financially viable. It is very likely that experts are not involved in the vast majority of cases where there is an audio recording containing evidentially-relevant speech.
With regard to non-expert transcripts of forensic audio materials, I do not personally have a complete understanding of how these are produced or by exactly whom. However, I am part of a team that has recently sent out a Freedom Of Information request to all 43 territorial police forces in England and Wales on the topic of Record of Taped Interviews (ROTIs, i.e. police interview transcripts). Multiple forces indicated that police transcribers (ROTI clerks) are now producing transcripts of other types of recordings such as body-worn camera footage, mobile phone recordings, and RING doorbell footage. It may be, therefore, that an audio recording that would be sent to an expert by one police force is transcribed by a ROTI clerk within another police force. It is also likely that many transcripts of forensic audio materials are initially produced by police officers working on cases.
In this response, I am addressing the qualifications of experts in forensic speech and audio analysis who produce forensic audio transcripts.
With regard to experts in forensic speech and audio analysis, there are no formal qualifications or prescribed training routes. Forensic laboratories will often develop their own in-house training, but there are no national training courses for forensic speech and audio analysis experts. The closest training available is the postgraduate Forensic Speech Science course at the University of York, though this is not designed to be a vocational qualification and those who complete the course are not ‘signed off’ to do casework. Many experts in forensic speech and audio analysis hold PhD degrees in linguistics or similar areas of study, and have undergone significant on-the-job training/CPD before they can be considered ‘expert witnesses’.
At present, there are no compliance requirements under the UK Forensic Science Regulator (FSR) for forensic practitioners carrying out expert transcription (which is defined as ‘Questioned content analysis’ in the Code of Practice[3]). The FSR defines which Forensic Science Activities have a statutory requirement to adhere to the Code of Practice, and forensic speech and audio analysis is currently ‘out of scope’. This means that there is no requirement for experts in this area to prove that they are competent or that their methods work before they produce evidential transcripts.
In this response, I am addressing the production process of forensic audio transcripts produced by experts in forensic speech and audio analysis.
With regard to experts in forensic speech and audio analysis, there is no national guidance or standardisation for the production process of transcripts. Each laboratory or private practitioner has developed their own methods based on relevant research in the fields of forensic speech science, psychology and linguistics. However, a survey of international expert transcription practices[4] revealed a general level of consensus on many parts of the methods employed to produce transcripts of forensic audio materials. All experts employ good quality equipment, produce multiple drafts and employ mitigation techniques against priming and/or cognitive bias. There are some differences in approaches regarding how to produce drafts (i.e. whether they should be produced by one person and checked by another, or whether multiple independent versions should be created and then compared) but in almost all cases where possible, two or more transcribers will be involved in the process.
In this response, I am addressing the content and accuracy of forensic audio transcripts.
A forensic audio transcript may contain damning evidence of what people have done or what they plan to do; it may actually contain the offence (e.g. a threat to kill or a fraudulent financial transaction), or a clear firsthand account of what is happening in the lead-up or aftermath of an offence. In cases like this, where the speech itself is being used for evidential purposes, it is extremely important that the transcript is an accurate representation of what was said and by who. With regard to the extent of accuracy of forensic audio transcripts, this is a very difficult question to answer. Particularly in cases where there is only a poor quality audio recording, it is practically impossible to know the ground truth, i.e. what was really said. Furthermore, every recording, every speaker, every transcriber and every transcript is different, and therefore generalisations across all forensic audio transcripts are impractical.
The benefit of a transcript produced by an expert in forensic speech and audio analysis is a reliably-produced linguistically-informed account of the part(s) of the content that is possible to interpret, and an acknowledgment that the speech in some part(s) of the recording may be unintelligible. The expert will provide a report alongside the transcript, explaining the limitations of the audio recording and of the transcript itself. Forensic audio materials can be extremely challenging to work with and present a number of dangers if not managed in a careful way. When we are understanding speech, we use ‘bottom-up’ information, which is the actual sounds produced, and ‘top-down’ information, which is our contextual knowledge. When the audio quality of a recording is extremely poor and it is hard to make out the speech sounds, we rely more heavily on the ‘top-down’ contextual information. So it is very easy to convince yourself that you have heard a certain word or phrase because it fits with your expectations and knowledge of the case. Experts methods typically involve a carefully designed information management system[5] where the expert starts with no information about the case, and parts of the information that are deemed relevant will be revealed later in the process[6], typically after multiple drafts have been produced.
With regard to non-experts producing transcripts of forensic audio materials, we have very limited understanding of what non-expert forensic audio transcripts are going into courts and how they are being used. Experts in forensic speech and audio analysis, such as Dr Richard Rhodes, often come into contact with transcripts of forensic audio materials that have been produced by police employees; experts may be instructed to review an existing transcript after producing their own independent version, or they may receive transcripts of audio recordings that have been sent to the experts for other casework purposes, e.g. audio enhancement or forensic speaker comparison.
In Dr Rhodes’ experience, the quality of police transcripts varies greatly. Some transcripts are good and some police employees are good at transcribing. However, other transcripts are not as good and have clearly been influenced by the expectations of the police concerning the crime under investigation. For example, Dr Rhodes was instructed to produce an independent transcript of an audio recording in which the police thought that the speakers were discussing drug weights; the production of an expert transcript showed that the conversation was actually about how much weight the speakers in the recording could lift in the gym. It is not uncommon for experts to encounter errors like this within police-produced transcripts. Other issues include a lot of information being missed and, more concerningly, a transcription being provided when the audio quality is not sufficient for transcription and/or when that section of the audio does not actually contain speech. Problems can arise when these interpretations of the speech within an evidential audio recording are left unscrutinised and are allowed to be submitted as evidence.
In this response, I am addressing the standardisation of forensic audio transcripts and national guidance for the production process of forensic audio transcripts.
There is no guidance for experts in forensic speech and audio analysis regarding the level of detail or how to produce transcripts of forensic audio materials. There may therefore be some differences in the level of detail when it comes to things like disfluent speech (e.g. self-interruptions and repetitions) or how a particular non-speech noise (e.g. coughing) is represented within expert transcripts of forensic audio. However, there is generally consistency within the format and content of transcripts - time points will be included to help the listener identity the relevant part of the audio, speech will be attributed (if possible) to individual speakers within the recording, and the document will be structured in a way that is clear and easy to follow.
Experts who have encountered police-produced transcripts of forensic audio materials report great variability in the structure and formatting of the transcripts[7]. Many aspects of the transcript are rarely done in a clear or systematic way, including:
● How often time points are included
● Whether utterances are attributed to certain speakers
● How overlapping speech is represented
● What type of document is produced (Microsoft Word, Microsoft Excel)
● Whether the speech is summarised or transcribed verbatim
This is likely a product of a lack of support for police employees carrying out transcription. To my knowledge, there is no national guidance to support the production process for police employees who transcribe similar types of forensic audio recordings. Guidance and standardised methods for police employees conducting transcription of forensic materials would likely solve many of the issues outlined above, and training for employees could raise awareness of the challenges of transcription and the potential dangers of psychological phenomena like priming.
In this response, I am addressing the timeliness of forensic audio transcripts (with regard to both instruction and production).
Experts may be instructed to produce a transcript during the investigation of a case, or in the lead up to a trial. In Dr Rhodes’ experience, it is often the case that experts are instructed late in the legal process, as a result of a party requesting an independent transcript. For example, it may be that an investigation has proceeded on the basis of a police transcript but once they are preparing for court, the Crown Prosecution Service decides that an independent transcript is also required.
In most cases, an expert will take weeks to months to complete the full process of producing a transcript of forensic audio materials (preparing materials, drafting, checking/finalising, writing an accompanying report). For expert transcription of forensic audio, the process itself is very time-intensive and requires multiple experts to conduct multiple stages of drafting and checking. In terms of the time taken to produce a forensic transcript, indistinct audio can take between 60 to 100 times the duration of the audio, according to Dr Richard Rhodes of the Forensic Voice Centre. That means that one minute of poor quality audio (e.g. a recording made with a covert recording device installed under the seat of a car) could take an hour or longer to transcribe. In addition, there will be time required for preparatory work, documentation, report-writing, checking, etc.
In this response, I am addressing the production of forensic audio transcripts.
There are a number of improvements that could be made to the process of producing forensic audio transcripts. Firstly, there is currently no clear definition of ‘forensic audio’ or specifically of ‘forensic audio that needs to be sent to an expert’, which can cause confusion for those working within law enforcement. A clear definition of when transcription requires expert input, along with a standardised triaging process for evidential audio recordings, would likely be very helpful. This could also bring the expert into the process much earlier, which could help the investigation (although it is recognised that the costs will not be practical in all cases) and avoid delays to trials where a late request for an expert is made.
Secondly, if transcription of forensic audio materials is carried out by police employees, then a standardised and informed process should be employed. Transcribers should (a) be trained in the transcription of forensic audio materials, (b) have access to good quality equipment and working conditions (see question 4b), (c) be independent from the case, and (d) (initially or completely) have no knowledge of the context or case conditions. This process should draw on the practices of experts in forensic speech and audio analysis, who use appropriate equipment and carefully-managed methods which, amongst other things, mitigate the effects of priming and biases.
Thirdly, more regulation for experts in forensic speech and audio analysis would ensure and develop the quality of evidence produced. Forensic speech and audio analysis is currently a recognised but unregulated Forensic Science Activity under the UK Forensic Science Regulator. This means that experts in this area (unlike those in many other forensic disciplines) are not required to obtain and maintain accreditation to international forensic standards[8], nor are they required to follow all aspects of the Forensic Science Regulator’s Code of Practice[9], which includes requirements such as:
● Checking of casework and general quality controls/monitoring
● Formal competence and training procedures
● Validation of methods (i.e. proving that the method works and is fit-for-purpose)
● Mitigation procedures against cognitive bias
● Information- and cyber-security
● Appropriate equipment and environmental conditions
Even though it is not currently mandatory for experts in forensic speech and audio analysis, there are many benefits to following the Code of Practice: it provides a framework for the production of reliable forensic evidence, and compliance assures good quality and best practice. Experts should strive to do so in order to demonstrate a willingness to comply with regulation and therefore assure the court of the impartiality and quality of their evidence and the validity of the methods used to produce it.
In this response, I am addressing the use of AI for both forensic audio transcripts and police-suspect interview transcripts.
Within this section I will refer to the automatic systems used to produce transcripts as Automatic Speech Recognition (ASR). ASR is an AI-based technology that uses machine learning algorithms to convert spoken language into written text. In this particular context, “AI” and “ASR” could be used interchangeably, although ASR is the more specific term for the transcription technology being discussed.
Firstly, I will address poor quality forensic audio materials. Multiple studies have concluded that ASR performance is not adequate for this type of recording[10]. Background (or foreground) noise, speakers talking at the same time, and speakers being far from the microphone are all factors which commonly feature in these types of audio recordings, and are factors that have a significant negative impact on ASR performance. In a study[11] using a Facebook Live recording of a band rehearsal, the sound of musical instruments was incorrectly transcribed as speech and most of the actual speech content was omitted or mistranscribed by the ASR systems. In a study[12] using a mobile phone recording of a conversation in a noisy pub, transcripts produced by commercially-available ASR systems contained many nonsense transcriptions - for example, “onion bhaji” was transcribed by one system as “Nicki Minaj”, and “dragon curry” was transcribed by another system as “drug coverage”. These systems are not designed to deal with these types of forensic audio materials. As such, there is currently no place for AI in the transcription of poor quality forensic audio materials.
There is potentially a place for AI within the production of transcripts of better quality audio recordings within law enforcement, specifically police-suspect interviews. However, there will be considerable variability across police-suspect interview recordings, related to factors like:
● Audio quality and background/foreground noises (e.g. typing on computers, rustling bags)
● The speakers themselves (e.g. genders, ages, accents)
● The type of speech (e.g. shouting, whispering, mumbling)
● The nature of the interaction (e.g. friendly, emotional)
● The topic of conversation (e.g. crime type)
● The vocabulary used (e.g. proper nouns such as people’s names or place names, slang terminology)
The performance of an ASR system will naturally vary across different recordings, which needs to be taken into account when assessing the viability of ASR for a particular recording. Further to this, there are some important factors to be aware of when considering the use of ASR within the context of police-suspect interviews, including:
(1) Spontaneous and unpredictable speech can be challenging for ASR systems. When you tell Siri or Alexa to “set a timer for 5 minutes”, you will be speaking clearly and the combination of words in your request will be predictable for the ASR system. But ‘real speech’ is often not so predictable: people produce “ums” and “ahs”; they repeat words or phrases; they change what they are saying mid-sentence or even mid-word. At a basic level, ASR systems calculate the most probable sequence of words to fit a set of speech sounds. If someone stops mid-word and then continues speaking, the ASR system may misinterpret the incomplete word, and instead offer a plausible interpretation of the phonetic content of the utterance. An example from experimental data I am currently working with is shown below:
What was actually said: “He does a cooki- he does cooking now.”
The ASR transcription: “He does a cookie. He is cooking now.”
Although these transcriptions are not highly dissimilar, they do represent a change in meaning; the ASR transcript introduces the concept of a “cookie” which was not mentioned at all in the actual speech content. These kinds of errors - those that make sense grammatically and phonetically - are the ones which will likely be hardest to spot in an ASR transcript, yet could introduce significant changes to the meaning.
(2) Overlapping speech can cause problems for ASR systems. In multi-speaker interactions such as police interviews, there will often be times when multiple speakers are talking at once. Overlapping speech is challenging to deal with even for humans, though the human brain is able to focus on single conversations in noisy environments (often referred to as the ‘cocktail party effect’). Overlapping speech can be very challenging for ASR systems[13] and can lead to omissions of parts of the speech from the transcript or mistranscriptions due to confusion over the combination of sounds produced. In my experience with ASR performance, speech within overlapping sections (and the neighbouring utterances) has been missed completely from the ASR transcript, despite being easy for a human ear to hear.
(3) ASR performs better for some people than others. There is lots of research into bias within ASR systems, i.e. where ASR systems systematically perform better for one group than another. This research has generally explored demographic factors like age, ethnicity and accent; generally, ASR systems are found to perform worse for children, ethnic minorities, and accented speech[14]. My own research[15] has explored the effect of regional accent on ASR performance, finding that ’non-standard’ phonetic features in West Yorkshire English often lead to mistranscriptions - some errors include “muddy” transcribed as “moody”, and “it’s a bit cut and chop with staff” transcribed as “it’s bitter cutting chocolate stuff”.
My research has also demonstrated that individual’s voices can present challenges for ASR systems[16]. In a recent study exploring the performance of OpenAI’s Whisper ASR system, it was found that ASR performance can be significantly impacted by the rate at which someone is speaking as well as the speaker’s long-term vocal tract settings. For example, one of the speakers in the study had a very ‘mumbly’ voice quality which proved to be challenging for the ASR system, which produced an error rate of 30% for this speaker (i.e. almost one third of the transcript was wrong). Therefore, it is important to consider that ASR performance can vary significantly across different speakers.
(4) ASR should never be used unsupervised. This is particularly true in high stakes contexts like the production of evidential transcripts. ASR providers will likely report very low error rates for their systems, but these error rates do not take into account the severity of the error and the potential impact it could have on a case. A system could make an error as simple as missing out one word, but that turns the statement “I did not do it” into “I did do it”, which is then interpreted as a confession to the crime. Without someone checking the transcript’s content, errors like this could be contained within an evidential transcript that is presented to a jury, and lead to miscarriages of justice.
Furthermore, it is quite often the case that the output of an ASR system contains errors that cause parts of the transcript to be nonsensical - a few examples from experimental data I am currently working with are included in the table below:
What was actually said | The ASR transcription |
Touch wood we’ll be alright. | Touch wood with we alright. |
While they fix houses, yeah. | Why the fix house, [UH]. |
Just post the keys through, but yeah. | Just post for keys for per yeah. |
In order for a transcript to be useful to a jury, it needs to be readable, a quality which will likely be undermined by nonsensical transcriptions. A transcript produced by AI should therefore only ever be used as a first draft that is incorporated within a managed transcription process.
(5) Checking an ASR transcript for errors is not as simple as it sounds. If ASR is to be used for the production of transcripts, it should be part of a carefully-managed process whereby a human checks the ASR output for errors. However, our group at the University of York is currently researching how reliably people can correct ASR transcripts. Speech perception (i.e. listening to and understanding speech) is a complex psychological process and there is one phenomenon in particular that can influence our perception of speech. This is called ‘priming’ and, in this particular context, it is where exposure to a transcript can affect what people think they have heard, because they are expecting to hear the words on the page.
We have recently carried out an experiment where our participants were given an ASR transcript and had to identify and correct any errors that they found. Results of the experiment demonstrate that many errors are left uncorrected by participants, including some errors where a significant change in meaning has taken place. For example, the ASR system transcribed “Yeah I know” when the speaker actually said “But yeah, no”. There is a significant change in meaning within this example - the ASR transcript contains an acknowledgement of something by the speaker (“I know”), whereas the actual speech content is a denial of something (“no”). Yet only 40% of the participants corrected the ASR transcript. It is therefore clear that the process of checking a transcript for errors is not as simple and easy as it may seem, and requires further research into how to effectively identify errors within ASR transcripts.
Returning to the question, AI could potentially feature within the process of producing transcripts for better quality audio recording within law enforcement, such as police-suspect interviews, but there are many factors to consider, which likely will be (or have been, in cases where ASR is already in use by police forces) overlooked by those unfamiliar with speech recognition technology and linguistic science. ASR will not work for every case or every speaker or every recording. There needs to be explicit recognition of (a) how and when ASR could be used, and (b) the risks and ethical concerns involved. There also needs to be a prescribed method for the incorporation of ASR into the transcription process that is informed by scientific research and proven to be reliable and valid. Without thoroughly tested and robust methods for incorporating ASR into the transcription process, police forces risk the production and subsequent evidential use of unreliable transcripts.
Regarding the use of AI for other types of evidential recordings, responses from a Freedom of Information request recently sent to all 43 police forces in England and Wales demonstrate that multiple forces are exploring the use of ASR systems for the transcription of non-interview recordings, such as body-worn video footage and 999 calls. Whether ASR could work effectively for these types of recordings depends on the audio quality as well as all of the factors listed above. We would recommend that extreme caution is taken in such contexts, particularly in the case of body-worn video footage, where evidentially-relevant speech is likely to be produced far away from the microphone.
We are also aware that automatic transcription is offered within some evidence storage/transfer systems and that it is possible to produce an ASR transcript at the click of a button - we would strongly caution against the use of this feature to generate evidential transcripts, given that an ASR transcript should form part of a managed transcription process and not be used in isolation. Offering this functionality without proper caution or training for users of the software could lead to incidents of unreliable ASR transcripts priming the police officers working on a case and potentially making their way into evidence.
In this response, I am addressing the skills needed to produce a forensic audio transcript.
In order to produce good-quality and reliable transcripts of forensic audio (and audio in general), the transcriber should have the following skills:
● Experience with and knowledge of typical audio qualities and typical content found in the materials they are transcribing
● Familiarity with the language, dialect and accent being spoken
● Familiarity with the type of speaker (e.g. gender, age)
● Competency in transcription, monitored through regular proficiency testing
The transcriber will also require:
● Good quality equipment (such as headphones and software)
● Suitable working conditions (ideally not a noisy shared office)
● Sufficient time to do the job properly
● A colleague to check over their work
● Direct training in doing transcription
● An established procedure in place
● Appropriate breaks
For expert transcription of forensic audio materials, particularly in the case of poor quality audio recordings, the transcriber will need an understanding of:
● Speech perception and psychological phenomena such as priming and cognitive bias
● Speech production
● Auditory-phonetic and acoustic analysis skills
● An awareness of the limitations of their own/human abilities
● Processes to preserve the independence/integrity of the transcript (e.g. understanding of and procedure for information management and bias mitigation)
● Skills/training/experience required to provide reports for legal settings and attend court and give evidence on transcript/audio evidence
In this response, I am addressing Questions 5, 5a and 6 on the evaluation of forensic audio transcripts.
The evaluation of a transcript’s quality should not be a task that is carried out at the end of the process, but implemented at various stages throughout its production. The production process should incorporate some level of checking, and these checks should be carried out by independent transcribers who do not have access to information about the case or suggested interpretations. It is important that the person checking the transcript is not primed by this information. In the case of expert transcription of forensic audio materials, checking processes are often incorporated into the methods through practices such as the production of multiple drafts and the involvement of multiple transcribers[17].
To the best of my understanding, in practice, it is often the trier-of-fact (i.e. the jury) who is seen as responsible for evaluating the quality of a forensic audio transcript. However, the purpose of a forensic audio transcript is often to assist the jury in understanding the speech content of the (often poor quality) audio recording. It cannot simultaneously be the case that the transcript is there to help the jury understand the speech content and that the jury should evaluate the quality of the transcript. Particularly in the case of poor quality forensic audio materials, the jury can struggle to make out any of what was said in the audio recording without the aid of a transcript, but once they are exposed to that transcript, they can be ‘primed’ to hear its contents and struggle to consider other interpretations[18]. The jury cannot be expected to carry the burden of the evaluative task - the jury is therefore not in a position to provide an effective evaluation of the transcript.
Furthermore, in the case of multiple transcripts, I would advise that it is not best practice for the jury to be presented with competing interpretations. Due to the nature of priming, there is unlikely to be an effective way for a jury to resolve this type of issue, and there is a real danger of a bias towards the first interpretation presented, which will often be put forth by the prosecution. The most effective solution to these issues would likely be that the evaluative process takes place before a transcript is presented to a jury, and any disputes should be resolved prior to this stage. It should be noted that judges will also be susceptible to the effects of priming, and therefore are not best placed to resolve issues of disputed content.
I am unable to comment further on the solution to this issue, but the point I would like to highlight is that it is necessary that those involved within this process are aware of the risks associated with transcripts and their capacity to prime listeners.
Question outside of expertise.
Question outside of expertise.
Question outside of expertise.
In this response, I am addressing the usefulness of forensic audio transcripts.
This is a challenging question to answer, given that the ‘usefulness’ of a forensic audio transcript depends on many factors. The simple answer is that a transcript can be very helpful, but it can also be incredibly misleading and potentially contribute to miscarriages of justice.
The usefulness of a specific forensic audio transcript will depend on the case and the audio and the transcriber (and so on), so it is difficult to generalise across transcripts. However, a good quality transcript that contains an accurate representation of the speech content can be incredibly useful to a jury, particularly in cases where there are issues related to intelligibility and the speech content is very challenging to comprehend without an aid. Dangers are introduced if a transcript contains errors and causes the listener to hear content that is not there in the audio recording. Given the intelligibility issues and therefore greater reliance on contextual knowledge, listeners are at a higher risk of being primed when dealing with forensic audio materials, which means that they can be very easily influenced to hear a suggested interpretation, even if it is not accurate.
Another factor that contributes to a transcript’s usefulness is the readability of the document. In order for a transcript to serve its purpose of assisting the jury, a jury member should be able to intuitively follow the discourse within the transcript without difficulty. This is particularly true given that a substantial number of adults that are eligible for jury duty in the UK have low proficiency in literacy (1 in 6; 18%)[19]. A transcript that has a low level of readability (e.g. inconsistent and/or messy formatting and a lack of important detail such as time points or who is speaking) will therefore not be very useful to a jury, and could even distract from the content of the transcript itself.
In this response, I am addressing the issues related to the transcription of forensic audio materials.
We cannot know for sure how often miscarriages of justice may take place as a result of a forensic audio transcript, but there are very real risks of it happening. There are documented cases where a mistranscription has led or almost led to a wrongful conviction[20], and given the police transcripts that experts like Dr Richard Rhodes have seen (see example in question 2), it is not an unlikely phenomenon, especially if an independent transcript is not requested at any stage.
Perhaps the most concerning issue, with regard to forensic audio transcripts, is the case of untested hypotheses put forward by people who are not experts but who are invested in the case. Hypotheses regarding what was said, particularly if the speech content is playing a substantial evidential role within a case, should not be accepted without an appropriate level of scrutiny. But that scrutiny should be carried out by independent parties (such as experts in forensic speech and audio analysis) and not the jury, who can be particularly susceptible to priming through knowledge of the case, exposure to transcript(s), and poor audio quality paired with poor audio playback procedures.
The example given by Dr Rhodes in question 2, where the police believed a recording contained a conversation about drug weights, could feasibly have made its way in front of a jury if someone had not instructed an expert to produce an independent transcript. It should be noted that the solution is not for experts to transcribe everything, but some level of independent checking needs to take place in cases where the content of an audio recording plays a significant evidential role.
In this response, I am addressing the presentation of forensic audio transcripts in court.
Ensuring good quality playback equipment in the case of forensic audio materials is paramount. The audio recording is the primary evidence - the task of the jury is to consider the audio, and potentially what was said, and therefore they need to have appropriate access to the recording. Dr Rhodes has experience of poor quality audio recordings being played to the jury on loudspeakers within the courtroom[21] - the reverberation caused by the room and the quality of the speakers meant that it was extremely difficult to understand any of the speech in the audio recording. If the jury is not able to hear the content of the audio recording, they may rely more heavily upon the transcript, which undermines its role as an ‘aid’ and not the primary evidence.
In order to allow the jury to do their job properly, they should listen to the audio recording using good quality headphones and in a quiet environment, and be able to listen to the audio multiple times and in the jury room as well. Some kind of interactive transcript and audio player software would be ideal in this case, such that it is easy to locate and play the relevant parts of the audio when required or desired.
Further to this, interpretations of the speech need to be carefully managed in a way that avoids the damaging effects of priming. This is particularly true in the context of questioned utterance cases, where there is dispute over exactly what was said. There may be multiple versions of the transcript, e.g. a prosecution version that contains ‘incriminating’ evidence and a defence version where the ‘incriminating’ section is completely innocuous. I do not have a direct recommendation for what would constitute best practice in this situation, but the jury (or judge) is not best placed to resolve the issues of disputed transcript content[22]. The risks of priming are too great, and the burden of evaluating transcripts should not fall upon decision makers.
DiChristofano, A., Shuster, H., Chandra, S., & Patwari, N. (2022). Performance Disparities Between Accents in Automatic Speech Recognition. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2208.01157
Feng, S., Kudina, O., Halpern, B. M., & Scharenborg, O. (2021). Quantifying Bias in Automatic Speech Recognition. In arXiv [eess.AS]. arXiv.
Fraser, H., & Kinoshita, Y. (2021). Injustice arising from the unnoticed power of priming: How lawyers and even judges can be misled by unreliable transcripts of indistinct forensic audio. Criminal Law Journal, 45(3), 142–152.
Fraser, H., Stevenson, B., & Marks, T. (2011). Interpretation of a Crisis Call: Persistence of a primed perception of a disputed utterance. International Journal of Speech Language and the Law, 18(2). https://doi.org/10.1558/ijsll.v18i2.261
French, P., & Fraser, H. (2018). Why “Ad Hoc Experts” should not Provide Transcripts of Indistinct Audio, and a Better Approach. Criminal Law Journal, 298–302.
Harrington, L. (2023). Incorporating automatic speech recognition methods into the transcription of police-suspect interviews: factors affecting automatic performance. Frontiers in Communication, 8, 1165233.
Harrington, L. (2024). Towards improving transcripts of audio recordings in the criminal justice system (Doctoral dissertation, University of York).
Harrington, L. & Hughes, V. (2023). Automatic speech recognition: system variability within a sociolinguistically homogenous group of speakers. Proceedings of the 20th International Congress of Phonetic Sciences, Prague, Czech Republic 2023.
Harrington, L., Love, R., & Wright, D. (2022, July). Analysing the performance of automated transcription tools for covert audio recordings. Conference of the International Association for Forensic Phonetics and Acoustics, Prague, Czech Republic.
Harrington, L., & Rhodes, R. (2024). Survey on forensic transcription practices. The International Journal of Speech, Language and the Law, 31(2), 236-266.
Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., & Goel, S. (2020). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences of the United States of America, 117(14), 7684–7689.
Loakes, D. (2022). Does automatic speech recognition (ASR) have a role in the transcription of indistinct covert recordings for forensic purposes? Frontiers in Communication, 7. https://doi.org/10.3389/fcomm.2022.803452
Loakes, D. (2024). Automatic speech recognition and the transcription of indistinct forensic audio: how do the new generation of systems fare? Frontiers in Communication, 9. https://doi.org/10.3389/fcomm.2024.1281407
Markl, N. (2022). Language variation and algorithmic bias: understanding algorithmic bias in British English automatic speech recognition. 2022 ACM Conference on Fairness, Accountability, and Transparency, 521–534.
Martin, J. L., & Wright, K. E. (2023). Bias in automatic speech recognition: The case of African American language. Applied Linguistics, 44(4), 613-630.
Rhodes, R., Earnshaw, K., & Nuttall, B. (2024). Managing bias risks: information management strategies for forensic speech and audio analysis. The International Journal of Speech, Language and the Law, 31(2), 328-349.
Wassink, A. B., Gansen, C., & Bartholomew, I. (2022). Uneven success: automatic speech recognition and ethnicity-related dialects. Speech Communication, 140, 50–70.
Wu, J., Chen, Z., Chen, S., Wu, Y., Yoshioka, T., Kanda, N., Liu, S. & Li, J. (2021). Investigation of practical aspects of single channel speech separation for ASR. arXiv preprint arXiv:2107.01922.
[1] I would like to acknowledge the contributions of Dr James Tompkinson, Dr Jessica Wormald, Dr Vincent Hughes, and Dr Philip Harrison.
[2] For more detail on this topic, see Harrington & Rhodes (2024).
[3] Speech and audio analysis, including ‘Questioned content analysis’, is covered by Forensic Science Activity ‘DIG 401’ within the FSR Code of Practice.
[4] Harrington & Rhodes (2024). The survey received responses from 8 UK practitioners. It should be noted that there are only around 10-15 practicing experts in forensic speech and audio analysis in the UK.
[5] See Rhodes, Earnshaw & Nuttall (2024).
[6] See also Harrington & Rhodes (2024).
[7] See Chapter 5 “Focus interview: non-expert transcripts in England and Wales” of Harrington (2024).
[8] Such as ISO/IEC 17025.
[9] The newest version of the FSR Code of Practice can be found here.
[10] See Loakes (2022; 2024) and Harrington, Love & Wright (2022).
[11] Loakes (2022)
[12] Harrington, Love & Wright (2022)
[13] E.g. Wu et al. (2021).
[14] Feng et al. (2024); Koenecke et al. (2019); Wassink et al. (2022); Martin & Wright (2023); Markl (2022); DiChristofano et al. (2022); Harrington (2023).
[15] Harrington (2023).
[16] See Harrington & Hughes (2023).
[17] See Harrington & Rhodes (2024).
[18] See the work of Helen Fraser, e.g. Fraser & Kinoshita (2021).
[19] See OECD (2025), Survey of Adult Skills 2023 Technical Report, OECD Skills Studies, OECD Publishing, Paris, https://doi.org/10.1787/80d9f692-en.
[20] French & Fraser (2018).
[21] See also Chapter 5 of Harrington (2024).
[22] See Fraser, Stevenson & Marks (2014) and Fraser & Kinoshita (2021), among other publications by Helen Fraser, detailing experimental findings that demonstrate how the ordering of transcripts can significantly impact speech perception.