final logo red (RGB)

 

Public Services Committee

Corrected oral evidence: Interpreting and translation in the Courts Service

Wednesday 13 November 2024

11 am

 

Watch the meeting

Members present: Baroness Morris of Yardley (The Chair); Lord Bach; Lord Carter of Coles; Lord Laming; Lord Mott; Lord Prentis of Leeds; Lord Shipley; Baroness Stedman-Scott; Lord Willis of Knaresborough.

Evidence Session No. 4              Heard in Public              Questions 58 - 70

 

Witnesses

I: Larysa Booth, Head of Sales, LingvoHouse Translation Services; Reggie Mitchell, Strategic Partnerships at Speechmatics; Frankie Williams, Chief Legal Officer, DeepL.


16

 

Examination of witnesses

Larysa Booth, Reggie Mitchell and Frankie Williams.

Q58            The Chair: Welcome to the fourth public session of our inquiry into interpreting and translation services. I want to start by thanking our witnesses for joining us today and asking them briefly to introduce themselves and say where they are from.

Larysa Booth: I work with LingvoHouse Translation Services. We provide translation localisation and interpreting to regulated businesses, the NHS, local councils and legal companies as well as international businesses. We cover over 300 languages.

Frankie Williams: I am chief legal officer at DeepL, a Cologne-based AI company specialising in translation and writing solutions. We have been operating since 2017 when AI was not even a buzzword. Up to now, we have provided services to businesses globally in text translation, but today, at our customer conference in Berlin, we are launching our speech solution called DeepL Voice, which is speech to text translation.

Reggie Mitchell: I work for Speechmatics, a speech to text company based in Cambridge and with offices in London. Right now, we do 53 languages of speech to text. That is our primary product. We have translation in a further 35. Our specialism has always been on the speech to text side, working in the courts and with various Governments, but we focus on the real-time aspect of translation, making sure that you get that very accurate first pass of transcription so that it can be translated by another model to prevent cascading errors of fact.

Q59            Lord Willis of Knaresborough: Can I say how much I have looked forward to the three of you coming today? I hope we will now get some positives in the way that AI can counter a major problem for translation in the courts in this country, indeed even in police stations.

What we have heard during these sessions really confuses us. The chair of the Bar Council, Sam Townend KC, said: “The growth of AI tools in the legal sector is inevitable”, and that it is a good way to go forward. But the Minister of Justice said:It is not appropriate to introduce AI in the courts”. That is the opposite. That is where we are, and it is why we hope that this morning’s session will give us some enthusiasm for what we are going to say to government.

It worries me enormously that we have a lot of speech to word using AI, but, if you go into the courts, you need speech to speech so that you can interpret what somebody is saying in a foreign language, say Urdu, into English in return. So when you answer questions today, or give us advice, perhaps you could keep that in your heads and say whether we are anywhere near that.

To each of you, what is your vision for AI interpreting and translation services, and what are the immediate opportunities available, because we have to say to government, “This is what we think you should do to improve and guarantee these services”?

Starting with you, Frankie, you have a new tool coming out today that has had massive publicity from our committee.

Frankie Williams: Exactly, thank you. As you said, the current use cases for AI in businesses, particularly the legal industry, are pretty well established for contract analysis, document creation and chatbots. Because of the specificity of legal language, training AI to understand that languagein particular, the volume of textis a really good use case for the technology because of the volume of training data.

The opportunity for speech to speech is developing. We have two tools: one is speech to text for meetings, which we have launched today; and the other tool is a conversational technology on our app. That is speech to speech for a limited number of languages, but it is not completely concurrent; you have to pause and then the translation happens. I guess that is the same with human interpreters.

One of the issues that has been raised in previous committee hearings is the breadth of languages required for interpretation in the courts. Obviously, it is harder to find interpreters for the less-used languages, and there is also less commercial imperative for businesses to train technology to understand those languages. There will also be more limited training materials to train the models to be accurate and to understand the nuance of those languages, particularly dialects. In speech, there is the added complexity of accent and regional variations.

At the moment, the use of technology for language translation in text is there. We are able to translate documents with a very high degree of accuracy. The more text you put in, the better the translation will be because you get more context. So translating word by word is less accurate than translating paragraphs, and likewise larger documents.

There are lots of challenges in interpretation in the courts, most of which are probably the same ones you have talked about with human interpretation, which are the languages and quality assurance—

Lord Willis of Knaresborough: In terms of DeepL, what do you do to address some of those core issues? Are you looking to universities and their research teams to do that, or do you have your own research?

Frankie Williams: We own all our technology, so our models are proprietary; they have been built. We have an enormous research team who train those models. We also have our own data centres, so our tech staff are owned from end to end. We have focused on quality rather than quantity when it comes to language. We have 33 languages at the moment and we release them when we are comfortable that the quality is sufficiently good that they serve the purpose that our customers need them to serve. The challenge will always be that, from a business perspective, the languages you need will probably be different from the languages where there is a shortage in the courts.

We use a combination of humans and technology, in that we have a huge bank of human translators to train the models we use, so it is a combination of machines and people.

Lord Willis of Knaresborough: Are they the same people who perhaps work in the courts via the—

Frankie Williams: I would have thought not. Certainly, until now they have been concerned primarily with text translation rather than interpretation, which is obviously a very different skill.

Reggie Mitchell: As I see it, there are two different types of technology for speech to speech translation. It depends on what is doing the translation in the brain. Wherever you go, you will have speech in, so you need a model that will take the speech in and turn that into text. Wherever you go, you will have a model that needs to take the text out and back into speech. We focus on the first part, but focusing on the middle—in other words, what is actually doing the translationis probably where we will see advances over the next five years.

Right now, Speechmatics and DeepL work on text to text translation models. What will change is that you will see more LLMs—large language models—coming in, things like ChatGPT, being able to interpret vast amounts of text in lots of different languages. Basically, that means that it will become more accurate and much easier to get real-time interpretation.

As with DeepL, we have a product called Flow. You can put that in so that, essentially, you have a multilingual speech model. That is taking speech into text. Then you have a multilingual LLM, such as ChatGPT, as the translation or brain. That brain is just going to get better and better. New ways might be found to do that, but, essentially, all the innovation will be based on those kinds of AI structures.

That means that it becomes more viable and cheaper to run. Although there are problems with the availability of the data, on our side we have developed systems that mean that you need less data to get it up to a better level, and that encompasses lots of different accents and dialects. Effectively, it is a way to break that data reliance so you do not need to get so many more hours of a particular dialect to identify it properly. Having data from hundreds and hundreds of other languages helps the model rather than hinders it, which is what it did in the past.

Right now, we have this in English to Spanish. Models out there can encompass huge amounts of languages—up to 30 languages in one model. They are enormous and very expensive to run, and they are not necessarily very accurate at the moment, but they will continue to get more accurate.

If we want a solution right now for interpretation in the courts, I would not use the LLM systems. We have a bot now that I think is called Mixhalo. That replaces translation systems and interpretation. Anywhere where there is an interpreter, it can translate it into five different languages at a time. It probably uses a combination of different translation models, including DeepL. That is the state of the art right now, but in five years it will be much better than that; it will be able to understand those languages much better in a way that no human can. No human can handle 10 languages at the same time, but this can.

Lord Willis of Knaresborough: Frankie’s organisation is in the mid-30s in terms of languages. According to the internet, you are using about 53. You talk about Spanish, French, German or European languages, the more common languages, to which it will be relatively easy to find a solution, but they are the languages where you have plenty of interpreters. It is in the area where you do not have those that we are trying to solve something for the courts. I am still not clear what you think we should say to government to expand that arena and to make sure that the courts that rely on spoken language when taking evidence and responding can actually do that.

Reggie Mitchell: One approach would be a kind of unified response across government. I think it would be very difficult to do, but there is precedent for it in India. The Indian Government released information and data. I cannot remember how many dialects there are across India, but there are at least 10, possibly 20. This means that people across the world now have access to this technology and are able to train across those languages. Even in England we can use that and make language models for you.

In the British Government right now there is a requirement for speech recognition and translation across many different sectors. The fact that each of these different tenders requires its own process and specific user interfaces is quite unhelpful. Effectively, we need a more unified data strategy to make that more viable. For instance, GCHQ has data on these very difficult languages. It cannot share that at all, but if it was possible to share less sensitive data from less sensitive organisations, it would be a way of making those languages more viable for AI to be used.

Lord Willis of Knaresborough: Larysa, you have given us one particular solution that is worth noting.

Larysa Booth: Our company was established in 2008. When we started, some technology was already available—so-called translation memory. When we started, whenever the client ordered the translation, after we produced the translator’s work in the translation and had done the proofreading, translation memory was able to remember everything that had been translated and store it as a translation memory for the customer. That allowed us to provide a discount to the client for any future translation for a number of repetitions. Whenever a new document came in from a specific client, the project was loaded into the system and the software was able to pick up all the repetitions based on previous translations. That was what we were using before.

Then machine translation became available and we started using that. Whatever we do as a company and provider of translation services is driven by the needs of our customers. We need to be one step ahead not just in knowing how it works but in being able to advise the customer. Technology has developed so far and advanced rapidly, specifically in the past four to five years. Now, a lot of our customers already use machine translation in-house, so they know that it saves them time and cost, but there is still an element of human verification and proofreading of the actual translation, especially when it comes to sensitive subject matter such as legal or medical.

It is very important that there is a human in the loop when we are talking about any technology. Talking about opportunities, the software is able to enhance efficiency, meaning that we can speed up translation. There is data suggesting that, whenever machine translation is involved, the process is speeded up by about 40%. That means that the client will get their translation 40% faster, and it reduces the client’s costs.

Another opportunity is scalability. For instance, if a client sends us two documents and each has about 50 pages, that is a lot of text to process, but because of the available technology we are able to do the translation so much faster but still have a professional who oversees the whole process, verifies the translation and does quality assurance checks and spell checks. Then we as a company certify the translation, providing an affidavit of accuracy confirming that that translation is correct.

Q60            Lord Shipley: That is very helpful. Thank you for that. I want to go back to something you said and then ask others to comment. You said that in translation you can achieve a very high degree of accuracy. I wonder whether you can put a number on that or define a bit better what “very good” means. I heard on a Radio 4 programme a couple of weeks ago that the publishing industry has some problems with getting a high degree of accuracy in translation, because, although AI will translate, it does not do nuance. As for the courts, most of what we deal with is interpreting, but there is translation, nevertheless, and quite an amount of it. How accurate is it?

Frankie Williams: We do a lot of blind testing of the accuracy of our translations using language specialists. We tend to test it against the other technology solutions. I cannot give you a specific number for the percentage of accuracy. I know Reggie has some specific numbers, but we ensure accuracy by having language specialists stress test the quality of our translations, and then we use those human language specialists to help us to train the models where inaccuracies exist.

The interesting question is: to what extent can you test accuracy of human translation in a different way from how you test accuracy of computer translation? We do a huge amount of testing and continued human-guided training of the technology to correct any inaccuracies we find, and continue to develop that nuance. Obviously, language is not a science—that is the other element of it—and we need the human training element that we input into the technology to build up that muscle for the nuance of language.

Lord Shipley: Reggie or Larysa, I do not know whether you want to add to that.

Reggie Mitchell: I can comment a little bit on accuracy. Similar to DeepL, Speechmatics tries to get very good accuracy, particularly in transcription. Transcription is in effect an easier process to determine whether something is accurate. Human transcriptionists can get up to around 99% accuracy. They will get an um or an ah wrong; we know that. AI is very similar. Particularly in English, we can get up to 95% to 98% accuracy on the transcription side.

On the translation side, from what I have read, effectively you can have two translations next to each other that are both accurate to an extent, but they will not match each other at all. There is a metric called BLEU scores. From what I have read, two humans, both very skilled, will match each other only around 60% of the time because of the nuances.

In terms of nuance and AI dealing with it, some of us also miss irony sometimes, but we are developing a tool or research that we call paralinguistic modelling. This is a way in which AI will be able to detect sentiment and sarcasm. It is not a completely solvable problem, but, as with transcription, we can be quite accurate, even if we guess it, 75% to 80% of the time. We are getting to the point of human levels of understanding. There is a way forward on nuance, but if you are looking for a verbatim system, I agree that you need a human at least to check it, and certainly on our way forward it will require a lot of human intervention to make those models more accurate.

Larysa Booth: From a practical point of view as a translation provider, we have enough resources to use different software that helps us to translate and provide quality assurance of those translations. Some manage better and, based on practical experience, we have information that, if the text is not too nuanced and it is more or less a standard business document, the level of accuracy can reach 85% to 90%. With more nuanced text, there is a question here. There is still a lack of data, which means that more professional linguists should oversee the entire process. Some cope less well.

Whenever we have customers sending machine translation to us, I normally ask what software they have used. That helps us to understand how much proofreading we need to do with that specific text. If anything like DeepL is used, we know that the quality will be quite good, so linguists would just do the proofreading and everything else. Some provide not very good results. Quite often, the translation needs to be redone from scratch by the translator, because the quality is just not there. The translator would still do the proofreading work. It would take longer to change everything and edit the machine translation than doing it from scratch. As a result, it would cost the customer more.

Q61            Lord Prentis of Leeds: I have two related questions. First, I want to say to all of you how interesting your submissions have been. They go to the heart of what we are looking at. To some extent, development will depend on markets, not necessarily the needs of the UK courts, which itself is an issue.

My question is to Reggie. It is not on translation but on interpreting. You have set out very clearly a way forward—speech to text and back to speech, and the other developments that are taking place—but where would you put yourselves in relation to full implementation? Are you at the design stage? Are we talking about five years, 10 years or 15 years when we see what could happen? The related issue is that other countries and organisations will be looking at AI. The most obvious one is the UN. Do you know whether the UN is looking at AI for interpreting?

Reggie Mitchell: There are two parts to your question: the UN’s development effort and selling it to other Governments. Starting with the development effort, I could show you a demonstration of our AI interpreter in Spanish at this point. We are looking to roll that out within the next year. Feasibly, this can take up to 53 languages, and we are adding a couple more. Those include more common languages such as Mandarin and Japanese, and more obscure ones. Due to the efforts of various independent bodies, for instance, a Uighur model is provided by Uighurs themselves so that they do not lose their language, effectively. We can bring those out. Those niche ones might not be as good, but certainly the top 40 will be usable, particularly in a police setting but maybe less so in a verbatim setting in a court. They will be useful, and that is really exciting for me.

Talking through a Spanish interpreter at this point is very exciting. I know that our partners around the world have incorporated even better translation models than the ones that we are using, so that becomes very feasible very quickly—one to two years.

Lord Prentis of Leeds: One to two years?

Reggie Mitchell: Yes.

Lord Prentis of Leeds: Okay. Do you work with other organisations that may need AI instead of interpreting or for interpreting?

Reggie Mitchell: Yes. Speechmatics effectively only does the language models, the speech to text models, and then we move down into other systems. We provide a back-end that people will be able to modify. The great thing about AI is that you can input huge amounts of legal data. Some of these models have very large context windows. You can put in millions of words and they will be able to remember the context of the text that you have put in. You can train them for legal, you can train them to be an assistant for a dentist, for instance. That is one of the projects that we are working on at the moment. That is very exciting for us because of the adaptability of these models.

At a very large level, Speechmatics is trying to manage the entire speech interaction. Other speech systems might have vocabularies of about 200,000 words. From a base level, we have a vocabulary of 2 million, and that means you can take in huge numbers of company names, huge amounts of what some people call street language, and drugs and all that kind of thing.

Lord Prentis of Leeds: Do the Government work with you at all?

Reggie Mitchell: Yes. I could provide some examples afterwards.

Lord Prentis of Leeds: That would be useful.

Reggie Mitchell: A major one I can share is that the European Parliament uses us for transcription across a lot of its committees. We are under NDA under certain UK government contracts, so I cannot say.

Your second question was about the UN and where else we could work around the world. I do not think we are working with the UN currently, but, as I said, we are working with the European Parliament trying to encompass all those languages. We recently brought out Gaelic and Maltese for it, which you could not say are incredibly commercially viable, but for some reason we were able to build a business case for those. It is not so expensive for us to build them if the customer can provide the data. The more data we can find, the easier it is for us to build these models than perhaps some other machine-learning companies.

Q62            Baroness Stedman-Scott: How will the use of AI tools affect the role of interpreters and translators? In particular, how far can interpreters and translators make use of such tools to support their work currently? What training or education is currently available to interpreters and translators, and how accessible is such training? Frankie, you mentioned limits. To save time, I have given you the whole range of the question.

Frankie Williams: Our tools are used a lot by translators.

Baroness Stedman-Scott: Themselves?

Frankie Williams: Yes, exactly. Our API can be used in computer-assisted translation tools to do a first pass of the translation, which, as Larysa said, massively speeds up the process of translation. Until now, we have been doing text translation, so we have not yet had a use case for speech interpretation for professional interpreters, but certainly for translators the use of technology is becoming much more widespread to speed up that process. As I said, we also use human translators to train our models. So it is a two-way street in the use of technology by translators.

Baroness Stedman-Scott: What training and education is currently available to people?

Frankie Williams: Our models are pretty intuitive. You drop the text in and it spits the text out. Formats are maintained, so what you get out looks exactly the same as what you put in. It is just translated.

Baroness Stedman-Scott: It is straightforward then.

Frankie Williams: It is really straightforward: I can use it.

Baroness Stedman-Scott: Thank you.

Larysa Booth: Yes. As a translation company, our job is to ensure that our customer gets exactly the translation they require. I am talking about human translation or machine translation plus human post editing. By default, we provide human translation and human proofreading. Whenever a customer comes to us, they do not mind if translation is done by a machine and proofread by a linguist, or if they want machine translation to be used specifically because of the volume of the document or a cost implication. We will deliver exactly what they need.

We have purchased a licence to use a specific software. There are several available: memoQ, Smartcat, Stratus and many others. We work with this specific software. We need to make sure that if a customer orders human translation the linguist does not do machine translation. For us to be able to control and oversee the whole process, the linguist gets access to our platform and logs in. We limit the use of machine translation, and we know that this translation will be done by the linguist rather than by machine. Then we use a second linguist, a professional linguist with specific experience, who will do the proofreading.

In the past, there were times when the linguist would use their own translation tools, and there was no way for us to control whether it was done by machine or the linguist was doing the translation themselves. This is why they are able to log into our software and we are able to oversee the translation process.

We know what stage the translation is at. We know when translation will be delivered. We know if translation has finished and a translator is doing the proofreading. Once that step is completed, translation goes to our project manager, who oversees the whole process, and they apply an additional level of quality assurance checks. There are ways to do that using the software where it can flag any inaccuracies in the translation. It does not mean that the project manager speaks all the languages, but we use project managers who are native speakers in specific languages.

Whenever we get translation into Spanish, for example, we make sure that the native-speaking Spanish project manager oversees the project. That allows us to apply an additional level of quality assurance checks, knowing that translation comes from a professional linguist, proofreading comes from a professional linguist, and a third level of quality assurance is done by a native-speaking project manager. In that way we can ensure that translation fully reflects the original.

Q63            Baroness Stedman-Scott: Reggie, how will the use of AI tools affect the role of interpreters and translators, and how far can our translators and interpreters make use of these?

Reggie Mitchell: I am afraid that Speechmatics does not have much of a view on this, because we do very much further up the stack. We would prefer to do the speech to text. I know that interpreters will use the first step of transcription to aid them with their translation down the line, so you are able to have that human in the loop and the AI model to assist an interpreter.

Baroness Stedman-Scott: Do you ever see a time when your technology will make the use of translators and interpreters redundant?

Reggie Mitchell: No, not entirely redundant. The answer to that is that it requires humans to make it worth while at all.

Baroness Stedman-Scott: That is clear.

The Chair: That is a good answer.

Frankie Williams: I would say the same thing. We use humans to train our models, which is really important to ensure that the nuance is right.

Baroness Stedman-Scott: Thank you.

Q64            Lord Carter of Coles: Building on the accuracy question, because people have constantly drawn our attention to it, as these technologies develop, do you think there will be a regulatory framework to give quality assurance? How is the industry thinking about that?

Frankie Williams: It is a really interesting question. I do not know how the translation industry is think about it. Obviously, there is a much broader regulatory framework developing for AI, and that is front and centre of what we are thinking about. The EU AI Act is already in place, and Governments around the world are looking at how to regulate the technology. On the regulation of machine translation, I am not aware of any developments.

Lord Carter of Coles: Where would the overall regulation of AI sit in government in Britain, in which government department?

Frankie Williams: I do not know off the top of my head, sorry.

Reggie Mitchell: There are a number of different ways of looking at that. You can look at the internal and the external regulation. In terms of internal, Speechmatics attempts to give its customers assurances mostly on data protection. For a number of different clients, we offer completely offline systems so they are air gapped. This step makes sure that they cannot be hacked or broken down. It also means that there is absolutely no way that we could see their data. That is very useful for us from the GDPR perspective, because it completely removes the onus on us to have to manage that data.

I know that an ISO standard of AI training is coming out. I do not have much more information on it, but people can sign up to a framework voluntarily. In our contracts, there are also surface-level agreements, which would normally be for what you might call online deployments. We operate servers on the internet and then people upload, which means that the service needs to be running when it says it should. That is one framework. The accuracy is very unlikely to degrade unless there is something seriously wrong with the model. It is never a case of why the accuracy has suddenly dropped down. It is more a case of why the system is not working at all.

I do not have so much information on regulation for AI.

Lord Carter of Coles: Thank you. Larysa, do you have anything to add?

Larysa Booth: We probably need to remember that any AI requires new skills from translators as well as interpreters. AI is developing fast and anything new coming on to the market meets a little bit of resentment, not necessarily from translators at the moment, because there is more data available, but from interpreters. When interpreters do not have sufficient skills to manage and deal with AI interpreting, there should be more training for professional interpreters so that they can feel confident about the language and the software. They need to be able to understand how to oversee the whole process if AI is involved and how to make sure that they manage to provide that quality assurance level of checks to meet a client’s requirements. It is a question of training for translators and interpreters.

Q65            Lord Bach: Good morning. My question follows on nicely, because quality assurance has been asked about already today, quite rightly, but let us see it in the court context now, if we could. We have heard that the MoJ takes a certain attitude towards it—for reasons that I can understand, but I do not know whether you can—which relate, for example, to the number of languages and mistakes that are made, not necessarily all the time but particularly with dialects. There is a famous story that one academic talks about when AI tools translated the exact opposite of Arabic dialects. The court setting with the various individuals and their personalities, and the relationships between them all, perhaps putting the MoJ case, make it too difficult and perhaps unfair to have a strictly AI system. I would like your comments on that proposition first.

Also, who is responsible in the commercial world for mistakes made by AI? Secondly, if it came into court in a significant way, who might take responsibility in the courts?

Frankie Williams: On quality assurance, the same issues would exist in the use of AI for interpretation as they would for humans. Often, the only two people in the court who speak the language are the two people who are speaking it to each other, so there is not necessarily an ability to do checking. If the interpretation was being done through a machine, it would be possible to do more quality assurance checking after the event because it would have been run through a machine, if that makes sense.

One question raised in previous sessions was about the limitations of what is being asked of human interpreters. The oath they take is that they are interpreting to the best of their skill and understanding. That has its own limitations. For quality assurance, the same issues arise whether you are using machines or people.

On the second question about accountability, from a commercial perspective the contract between us and our customer allocates risk for inaccuracy, as you would expect. In a public services setting, that is different. I do not know who would take responsibility for liability if an interpretation was inaccurate through a human interpreter.

Lord Bach: It would be something to be worked out.

Frankie Williams: Yes, exactly.

Lord Bach: Larysa, what do you say on this?

Larysa Booth: Courts quite often require real languages or dialects, and this is driven by the diversity of the communities living in specific areas. AI is able to provide support for more common languages. This is the barrier. The solution here, as we have already discussed, would be to use a hybrid model where AI is used to aid and support the linguist, but it is still important to keep the linguist there to oversee the whole process and to make sure that any mistakes are corrected and the court gets the result based on the confirmation that comes from the professional interpreter rather than from the software.

Lord Bach: That is very helpful, thank you. Reggie.

Reggie Mitchell: I know that transcription is used very heavily in legal systems in America, and to make that happen we have specifically added in features such as utterances into our system. Previously, speech recognition systems would rub out every time someone “ummed” or “ahhed”, and that is not acceptable in a court setting. We have got to the point there where we are reasonably accurate in transcription.

There is not an enormous leap to be made to make that happen in translation as well. You need to do testing and trialling, and that is where you can do more human intervention to make sure that the system works. Once you have tested it 10 times, the likelihood is that for the other 90 times it will work very well. I do not think you will need a human there all the time. It slightly negates the point of having the AI in the first place; otherwise, you could just have the human interpret them at a different pace. Humans should be testing it, and testing it regularly. You could get to a point where the quality assurance was there.

From a legal perspective, we talked outside the committee about how you can sometimes have insurance that covers you in case of a mistake being made in the translation. If an insurer can come in, make that calculation and be happy with it, and it works from a commercial perspective that they are happy to cover the AI company, it would be possible for the AI company to take that contract on board, but we would need to work with the insurance company and the Government to trial those systems to make sure that they consistently perform as they should.

Frankie Williams: The difference in a court setting is that a mistake does not give rise to a financial damage; it gives rise to—

Lord Bach: Something much worse sometimes.

Frankie Williams:a potential miscarriage of justice. Exactly.

Lord Bach: Thank you for that.

Q66            Lord Laming: Each of you has placed an emphasis, understandably, on the accuracy of the end product, and you have indicated that there has to be human involvement in that to make sure it is right. We have had examples of how every word matters in this. If an advertisement that said “alcohol-free drinks here” was interpreted as “free alcohol drinks here”, that might lead to a commercial loss, but in the courts it could lead to the conviction of somebody and the deprivation of their liberty. If a human being has to check the accuracy of every statement that you produce in your business, you can only go at the speed of the human being, and you depend on the human being to read every word. I do not understand what the savings are then.

Reggie Mitchell: I would agree. We need to get to a point where translation consistently produces an output that we can rely on; otherwise, it is not worth introducing to the courts. We are reaching that point, and with testing we can make that happen.

Frankie Williams: This technology is moving incredibly fast. One of the issues across the board is trust in the systems. People are sceptical, understandably, about the use of technologies. That trust will come as people get more familiar with them and as the industry becomes regulated. That is an important part of the regulation. Perhaps that will serve to create more public trust.

On your point about specific languages, with our models we are seeking to make them as accurate in the context as possible. With our system, individual customers can set up a glossary. They can have a house style and use a specific selection of customised words that will then be used consistently in translation.

For the legal setting, that is particularly important, because we have a huge number of everyday words that are used really specifically in a legal context. I was trying to think of examples of this: minor” and “majority” in the context of age; “frustration” in the context of contract; and even the use of “briefs” when it comes to barristers’ instructions. The use of those words is really specific in a legal context. The ability to train the models to understand that, in a particular context, a word will mean a particular thing, as the models become more sophisticated, will eradicate those risks of material mistakes more and more.

Q67            Lord Willis of Knaresborough: We have obviously come to a point, have we not, where we have said that interpretation is the key thing? We feel that doing “written” would be relatively easy at a given time. I talked to BLANC earlier this week just to look at the issue of them translating it. If you look at human translation in a court, the chances of things going wrong there are significant, particularly with what I call the minority languages, where you depend on somebody who speaks part of it interpreting it correctly but you have no way of checking that. You do have a way of checking if you have used AI, because there is a physical recording of it to go over. I do not quite understand why we are not a little more adventurous in this translation argument than simply saying, “It can’t be done”.

Frankie Williams: I agree with that view. If something is passed through a machine, you are much more able to do a second pass and check accuracy than if you are in a court setting where the conversation is not necessarily recorded or heard by anyone else, and certainly not heard by anyone else who speaks the language. There is a real opportunity to be able to do more checking if you have a record.

Reggie Mitchell: I also agree. I am quite bullish on the language translation being able to get to a point very quickly, but I am keen to emphasise the slight complexities of the project and the need for testing as well as the potential for the technology.

Lord Willis of Knaresborough: Of course, yes.

Q68            Lord Carter of Coles: This has been tremendous. Bringing it back to the courts, we want to understand the barriers, but they are probably general. I think we are right to accept that it is coming, and it is a question of time. What advice would you give to us to give to the Government as to how to lay out that road map and make it an easier journey for your industry to deliver this quickly?

Frankie Williams: That is a difficult question. In terms of the barriers, one of the things that has been consistent in the conversations before this committee has been an inconsistency of access to videoconferencing facilities at all in the courts. There has to be a base level of access to technology to make these things work. Step 1 is to make those technologies available, and it does not need to be terribly complicated. I can do speech to speech using my mobile phone at the moment using our technology. That basic access to technology is important.

Again, a lot of it comes down to trust. It amazes me that people will drop all sorts of things into ChatGPT without thinking about the consequences, but there is still an amount of mistrust in technology and AI. I hope that the regulation that is coming will help that. As people become more familiar with it, that will be reduced. The basic access to workable technology is probably the first barrier.

Reggie Mitchell: Data is another barrier. If we can get more data from the languages that we want to design translation systems for, that will lower the barrier to entry for a huge variety of AI companies. That makes it more difficult for us but also more competitive. That is a good thing. My experience tells me that, with these government tender processes, the process itself will probably take one to two years even if we get to the point where we are ready to put out a tender. By that point, the technology will have moved on significantly anyway. I would say: do lots of research and lots of testing, particularly on the target languages you are looking to work with.

Q69            Lord Carter of Coles: Do you think the Government, particularly the Ministry of Justice, have the capability to write a tender to do this? Are they well enough informed? How will they get themselves informed? We are right up on the frontier now.

Reggie Mitchell: I do not think I would have the capability to write such a tender. In general, sometimes salespeople say, “Let our technicians have a look at your tender and let them shape the realities and shape your expectations”, because sometimes government expectations are not feasible for the kinds of costs involved. Then it becomes less competitive if one person has preferential treatment in the tender process. It will be very difficult to make it work. I am sure there is an opportunity for a consultancy to come in and think about it.

Lord Carter of Coles: Oh dear.

The Chair: You have got yourself another job here.

Lord Carter of Coles: Larysa, do you have anything to add to that?

Larysa Booth: I absolutely agree with everything that has been said. We have customers who have been with us for over 10 years. To help translators make their job a bit easier and to make sure that we deliver a polished result to the client, we create glossaries for our clients of terms based on the documents that we translate for them. The clients can provide the style guide as to tone of voice of the translation, style and everything else. If we train the AI properly for sufficient periods based on the data that the client provides, AI will provide a better result every time. That means that the clients will get the result faster. The translator would need to do just final QA checks. That is based on the fact that the client is happy for us to do that, because there are still those who want the full human service, and it is normally driven by court requirements, where there has to be a translator or interpreter with a specific qualification that the court would accept.

Q70            The Chair: Can I ask a question that has been raised with us previously and which I am not sure we have covered specifically? We have all talked about the sensitive information that would come through in courts and the consequences. I was not quite sure whether you saw that as a big barrier. Are you already accustomed to dealing with sensitive data at that level? Are you confident that the sensitivity of the data in a court would not be a major barrier to taking this work forward?

Reggie Mitchell: Yes, we are accustomed to dealing with very sensitive data. We had one client who was dealing particularly with child abuse scandals and taking evidence from children. It is quite easy to segregate that and to run it completely offline. As a base standard, all our data is deleted anyway.

The Chair: So it is not stored.

Reggie Mitchell: It is not stored. For that kind of data, we can create an air gap so that they can store it on a local system that is never online. No one from our company has access and, basically, no one has access even internally. It is quite easy to safeguard that.

Lord Prentis of Leeds: In your example, were you providing interpreting services or translation?

Reggie Mitchell: That was transcription. In fact, there was no translation service inside. It is not so much that the translation is more difficult. The only thing you have to make sure of, if you are going to run an offline system, is that you update it regularly with the software provider. Effectively, you download a new model from the internet and you can run that offline. Not every translation model can do that, because some of them are too big, so some have to be run on huge computers run by Google or Amazon somewhere. There will be providers able to produce these offline services, especially for government and defence purposes.

The Chair: I would assume that that answer was okay for everybody.

Frankie Williams: We have the same issue with our customers. They want complete confidentially. We do not store our translations. They go through the system and they are not stored by us. Lots of our customers are law firms and Governments, and they have exactly the same needs on confidentiality.

Larysa Booth: Clients come back to us a few years later asking us to resend the translation that we did three or four years ago. It is stated in the terms and conditions that we do not store the data. Every time a translation is delivered, the client is asked to make sure that they keep it, because the data will be deleted.

The Chair: Okay, that is really helpful. Thank you for a very informative and enjoyable session. You have increased our understanding a hundredfold, I can assure you of that. Thank you very much for your time.