Written evidence submitted by GreenNet (IPB0063)
Introduction
- Thank you for the opportunity to contribute comments to the Committee's inquiry into the Draft Investigatory Powers Bill published by the Home Office in November 2015. The inquiry covers technical aspects including “feasibility and costs of meeting the obligations [likely to be] imposed by the Bill”; “the impact on communications service providers and related businesses”; and “the likely consequences for citizen/consumer use of ICT services”.
- In fact, it is very difficult for the public or MPs to assess any these impacts at this stage owing to the lack of detail in the draft bill (or the overarching documents), and the absence of draft secondary legislation such as codes and the very probable lack of detail in those too when they are published in February/March. However, some technical background may help in trying to clarify relevant questions that have not yet been addressed. I hope the length of this submission might be justified by it proving a useful reference.
- GreenNet is a small not-for-profit organisation founded before commercial public internet service providers; Prof Peter Willetts of City University states that GreenNet was arguably the first ISP in Britain.[1] GreenNet has considerable technical expertise in providing a wide range of internet services to the non-profit sector, as well as actively countering network abuse, and works to support free software and open standards.
Current status
- Generally, we deliberately do not refer to specific clauses within the paper, owing to the inchoate nature of many of the concepts introduced in its more contentious parts. Variously described on gov.uk as a draft bill and a policy paper, it might be best thought of as a green paper to which the best response is to independently gather evidence to decide what measures are possible and their costs and benefits. In our opinion, the Joint Committee set up to provide recommendations will have a difficult job responding to the paper in a meaningful way, although it can of course draw on findings of the previous committee on the Communications Data Bill and the Investigatory Powers Review from the Independent Reviewer of Terrorism Legislation, David Anderson QC.
- We do however agree with evidence already given to the Committee by James Blessing of ISPA, Prof Ross Anderson, Dr Joss Wright and other experts about the complexity and scale of interference with communications implied by the bill and its context, lack of effectiveness against serious crime, disproportionate burdens that might placed on the UK technology industry and a general absence of meaningful definitions in the draft bill. In particular, we are also concerned about the proposals for legalisation of cracking / hacking / CNE / “equipment interference” on a wide scale, and the gathering of intelligence on legitimate political activity.
- Further, we concur with the written evidence of Adrian Kennard, managing director of connectivity provider Andrews & Arnold who has also raised some of the conceptual misunderstandings of technology that may be encountered either in discussion with the Communications Capabilities Directorate[2] of the Home Office or trying to interpret documents that they have released. As he explains, there are many important areas that the document leaves open, including lack of applicability of the proposals to modern network protocols and applications, “gagging orders” incumbent on service providers (CSPs), the extremely broad definitions such that almost all internet users could qualify as CSPs, data protection issues, powers to compel generation of “internet connection records” (ICRs, presumably equivalent to “web logs” as used in discussion in the media and previous Joint Committee) and lack of any answer about how they might be generated. We were represented at the same meeting that Mr Kennard refers to[3], and can confirm that all these areas are yet to be discussed, at least in public.
- Recommendation 18 of the Anderson review is that “There should be no question of progressing proposals for the compulsory retention of third party data before such time as a compelling operational case may have been made, there has been full consultation with CSPs and the various legal and technical issues have been fully bottomed out. None of those conditions is currently satisfied.”[4] Unfortunately the timing is somewhat constrained, and the discussion has not yet properly begun, let alone “bottomed out”. In one moment of clarity, c46(5)(c) shows the ICRs are definitely intended to include third-party data.
- The stated constraint of the DRIPA sunset clause is really self-imposed. If such a deadline passes, it may mean CSPs will have to purge excess data that is no longer required by law, but it should be remembered that DRIPA is subject to a legal challenge currently referred to the CJEU, which has previously ruled that the EU Data Retention Directive of which DRIPA is an extension is incompatible with the European Charter of Fundamental Rights; also that in many EU countries it has been struck down as unconstitutional, without obvious harmful effects; and that these powers were new in 2005. We would like to see statistical evidence of if and when these data retention powers have been strictly necessary (Anderson Annex 10 has some qualitative examples of use of subscriber data, but no clarity about the effect of the data retention regime).
- It seems that the technical advice the Home Office has received has been almost entirely from those companies (eg Detica, Huawei) that are providing supporting solutions to gather, store and access communications data, and quotes from the largest carriers (eg BT, Virgin). There has been relatively little discussion with the internet industry as a whole; the Joint Committee on the draft Communications Data Bill said “Before re-drafted legislation is introduced there should be a new round of consultation with technical experts, industry, law enforcement bodies, public authorities and civil liberties groups. This consultation should be on the basis of the narrower, more clearly defined set of proposals... CSPs should be given a clear understanding of the exact nature of the gap which the draft Bill aims to address”. We do not believe this has happened.
- It might be useful focussing on a simple example such as a missing person case and identifying a mobile device used to tweet or post a message – something still a little more complicated than a RIPA request to a telco for a reverse number lookup. Nowhere have we seen it clarified what extra data would be logged under the new legislation, and how it would help in such a case; nor much evidence that it has been considered.
- We can speculate that logging source and destination TCP ports of individual packets on the mobile carrier side of CGNAT was not successfully enabled by the amendment made by the Counterterrorism and Security Act 2015[5]. (This is implied towards the end of the ICR factsheet and “operational case”, although again without detail. To “resolve” an IP address to a mobile user you basically need a correlation from service logs including port from eg a web host like Twitter, plus CGNAT, DHCP and RADIUS logs from the mobile provider; or more likely destination IP addresses extracted from the NATted network.) There is a lot of educated guesswork needed to work out the practical intention of the legislation – the “what” and “why” as well as the “how”.
- The bill certainly implies deep-packet inspection (DPI, related here to “web logs”, “ICRs”, “sessionization”, s71(9)) of a type that is definitely not used by ISPs for eg traffic sampling and management (see below for how “DPI” covers a multitude of sins). Instead, the parsing of “communications data” from applications requires vastly more computing power, since it involves reconstruction of what the endpoint would see and interpretation of that. “Web logs” to most technical people would mean logs stored by a web server and so this appears to be misuse of an existing term when used in discussion of the previous draft bill. According to the Anderson report (9.53) it may include DNS (Domain Name System) server lookups and the body of HTTP requests obtained on a third party (not normally logged). This may be the reason the phrase “internet connection record” was invented, but people already seem to be wrongly assuming a domestic broadband supplier can track a “connection” to a remote website without storing TCP/IP packet headers. It would be a start to clarify if there is any difference between “web logs” and “internet connection record” as used in the respective draft bills.
- There is a strong idea underlying the drafting of the document that legislation can be “future-proofed” by describing things at a conceptual level. However, it should at least be possible to at deduce roughly how such legislation applies to the current technical environment, and at the moment it is not. Not being able to do so is equivalent to not being able to discuss that law enforcement ever uses techniques like CCTV, photofit or DNA databases yet having to accept that their use of such unknown techniques might be legal.
- Not only does the legislation not proceed from defined current needs and then try to generalise for future (unknown) technological development, this technical “flexibility” adds to uncertainty over purpose and legal principle. Like the Draft Communications Data Bill, this draft enables the executive to compel virtually anyone to spy on anyone else and make it a criminal offence for them to discuss the practice or process in perpetuity. The main safeguard is that compulsion must be decided to be for reasons of “national security” or prevention or detection of a “serious crime”, both of which are also defined not by law, but by the same executive, or “preventing disorder” which may have political implications.[6]
Not merely lack of clarity
- The draft has been prepared with limited time to meet a schedule allowing normal parliamentary scrutiny. Little in the way of technical explanation has yet made it into the public domain, although small clues have been revealed via supporting documents and the Anderson report. The particular novelty in the draft is “internet connection records”, which seem to involve storage of certain types of metadata by (Layer 1-3) carriers. What these enable is effectively the same as clause 1 of the Draft Communications Data Bill, a power to compel generation and retention of data without an effective limit (the logic is that law enforcement will look back at old records of a newly identified suspect, thus the intention must be to log data on a group of people the vast majority of whom are innocent).
- Anderson states “What is meant by web log in this context has caused some uncertainty, and independent experts to whom I have spoken criticise the term, and those who use it, on the basis of imprecision”[7]. Indeed, “Internet Connection Record” is similarly confusing to those in the industry. The definition the Home Office submitted to the review was “include websites visited up to the first ‘/’ of its [url]”. In fact this would capture the scheme or protocol only (eg “http:”, “https:”, “ftp:”). Note also that what appears to the left of the first single slash may also include sensitive data including usernames and passwords (//username@password:hostname/). The footnoted Home Office understanding makes this even more confusing, including “http 'GET' messages” which in fact consist of “GET” followed by the part of the URL to the right of the first single slash (the hostname being in a Host: HTTP header); and DNS server logs – whether this is referring to recursive or authoritative DNS servers isn't clear, either way, such logs do not normally include hostname lookups. However, “web logs” or ICRs are “not limited to” those mentioned, making the term even more meaningless.
- In general, it appears that the Home Office documents have been drafted with some understanding of the value of IP addresses in investigation, but without any clear idea of technical limitations. For example, strings are given in Operational Case for the Retention of Internet Connection Records (Home Office, 2015) that are presumably intended to be IPv4 addresses but are not, when they could use IP addresses reserved for documentation such as 192.0.2.0 or 203.0.113.0. The use of port numbers as “identifiers” does not appear coherent to me; the same string “62.25.961.0” appears to be used as both the public IP address and the internal NATted address.
- There is also still a lack of congruence between the draft and what we know agencies are already doing thanks to Snowden and other whistleblowers. For example, how does the “request filter” map to XKEYSCORE? Another example is the type of “communications data” that seems to be referred in the Investigatory Powers Review which may well include a Host: header parsed from a (possibly encrypted?) network stream. However, it does not mention the use of Cookie: headers to identify a device or individual used as “Target Description Identifiers” in the Snowden archive.[8] While this is treated as “Communications Data” (CD), it actually could allow full access to a web session and can only be mapped to an individual by looking up a separate database. We are not sure whether this is contemplated or not, as it requires advanced capabilities and may not be intended for mass surveillance.
“Communications data” v “metadata” v “content”: background & internet glossary
- The terms “metadata” and “communications data” are sometimes used interchangeably. Metadata has a simple standard definition of “data about data” and is frequently used as clearer and more convenient.[9] Some content is metadata, and all metadata is content in one sense. The distinction between metadata and “Communications Data” (CD, “who, where, when and how” as I understand it) means that CD is sometimes a superset of metadata (adding a time, which may be present in server logs, but not in intercepted packets or reconstructed sessions, and possibly adding information derived from correlating with an external data source), and sometimes a subset (metadata can include a title, abstract or subject line, as well as a full URL path or query, which has usually been described as excluded from CD.) CD (not requiring ministerial or judicial authorisation) is also a form of content. Ross Anderson has pointed out that a .ics (calendar) file would be ordinarily thought to be content, but actually includes when, where and who of a meeting; and such details can also be extracted from natural language.
- CD has also been referred to as “what's written on the outside of the envelope” in a postal service. However, it should be realised that the way the internet is structured is as series of several, possibly many, envelopes, one inside the other, more like “pass the parcel” or a Russian Doll, each envelope having a distinct function. The following uses a little artistic licence (we can produce diagrams if needed), but is meant as a summary of how the internet is typically used.
- As an example of these multiple layers of metadata, take the process of sending an email using email client software like Microsoft Outlook or Mozilla Thunderbird. Each layer's function may be accomplished in a number of different ways, or protocols. First, various preparatory things happen within the innermost application layer (Layer 7 according to the OSI model, Open Systems Interconnection): text or attachments are MIME-encoded, a “RFC 5322” header adds the Subject line, Date, To and so on. The network part of the email software then requests the OS negotiate a TCP (layer 4) “connection” to the Simple Mail Transfer Protocol (SMTP) service on the appropriate email server (a server also being one meaning of “host”); then an exchange of SMTP commands begins with something like “HELO” (Layer 7 protocol). The email header and body data is broken into TCP segments (Layer 4) that specify the service (using port numbers that tell the operating systems which software process is involved). Individual web (HTTP) requests to fetch a page or file or submit a blog entry work in a similar way to SMTP. These TCP segments are then wrapped inside a Layer 3 network header including the destination Internet Protocol (whether IPv4 or IPv6) address of the server, as well as the source IP address to which responses should be returned. These IP packets are routed using the this 20-byte IP header, initially by your device's operating system (OS, eg Windows, Mac OS X, Android), then via a series of routers (run by various organisations), first maybe your home broadband router, then maybe four to twenty “hops”. In each hop between your computer, each router and the server, the destination IP address is looked up in a routing table and sent via a local network (LAN) or some other connection to the next hop. To get to the next hop, the unchanged Layer 4 datagram, already wrapped into a Layer 3 packet is placed inside a Layer 2 protocol to become a frame (transmission unit) of typically 100 to 1500 bytes. At each router hop, the previous layer 2 envelope is discarded, the layer 3 header updated, and the data re-enclosed in a fresh layer 2 envelope before sending on (or possibly returning an apologetic ICMP layer 3 packet if it cannot be routed). Layer 2 encloses the IP header and body by prefixing them with a further header, typically including MAC (addresses), to enable network switches to switch the frame to (the network segment for) the correct (usually trusted) device on the local network, such as a router, server, smartphone or printer. The frames are transmitted between switches and other devices by holding them very briefly in a buffer, again topping and tailing them with a preamble and suffix, and sending via a Layer 1 physical protocol such as “Wi-Fi”, ADSL broadband and an Ethernet standard like 10GBASE-LR – your SMTP data may very well be sent to your email provider using all three of these outer layer 1 mechanisms in series, but never changing anything inside layer 3. Then once your encoded email is unwrapped through all these layers 1-7 by the email host and reassembled, it will temporarily be queued, probably undergo an authorisation and spam check and be logged, and the SMTP wrapper (the Envelope To, usually equivalent to the To header you see) used to determine the next host to send on to, through several Layer 7 mail relays (MTAs), until it is held for your email contact to collect using another Layer 7 protocol such as IMAP, finally to be displayed on the recipients' screens.
- “Layer 1” (the outer skin as it were, the first we encounter when unwrapping) is the Physical Layer that transports bits (binary digits, 0 or 1) at a fixed data rate across a medium between two points. Such protocols are typically defined by IEEE and include wireless-g, wireless-n, V.90 (56K dial-up modems), Ethernet standards like 1000BASE-T, and ADSL. A variety of other standards such as ISDN, USB 2.0 and RS-232 and GSM mobile standards can also be included here. (Outside or below layer 1 might be considered to be a “layer 0”, the actual analogue hardware such as cables and phone lines.)
- “Layer 2” is the (Data) Link Layer. Protocols vary according to purpose: some as mentioned include a switching function like the Ethernet frame, or a channel multiplexing function like (generally for UK ADSL) the ATM cell. PPP (point-to-point protocol, typically for dial-up and mobiles) as suggested by the name does not include a switching function and is used as a thin wrapper in layer 2. The MAC in Ethernet-based Local Area Networks (LAN) is a 6-byte address to deliver the frame to, usually specific to the hardware (such as a particular Network Interface Card); it can indicate a manufacturer, although it can also be a virtual device or forged. Like other layers, layer 2 does some error detection.
- Inside that, “Layer 3” is the Network Layer (also known as internet layer) and nearly synonymous nowadays with Internet Protocol (IPv4 and IPv6, collectively just referred to as IP). Layer 3 is responsible for routing over an internet or network of networks, that is between networks. The units of data here (that have been further encapsulated by layer 2) are usually called packets (but strictly speaking datagrams). “The Internet” is sometimes seen as the largest and by far the most popular IP-based Wide Area Network (WAN). Protocols at this level and above are defined by IETF (Internet Engineering Taskforce) in the form of RFCs (Requests for Comment). (There are other Layer 3 internetworking protocols such as Novell IPX, but for simplicity we can ignore those.)
- “Layer 4” is the Transport Layer. The two most frequently encountered protocols here are UDP (User Datagram Protocol, used for many things but notably for DNS domain name lookups) and TCP (Transmission Control Protocol). Both of these two Layer 4 protocols (but not others) include the concept of port that usually specifies a socket or process on the source and target machine; or can be used for NAT (Network Address Translation) to funnel a lot of devices on a LAN or mobile network onto a single public IPv4 address. TCP includes additional features to UDP such as reliable error correction, establishing a TCP connection, and turning a stream of any size into TCP segments.
- “Layer 5/6”, the session and presentation layers, don't strictly speaking apply to TCP/IP, but is where it makes sense to imagine the most frequently used form of public-key encryption when present. The relevant protocols are Secure Sockets Layer (SSL), now being replaced by the more secure TLS (Transport Layer Security). The software to do this may exist inside an application or involve a common library provided by a device's currently active operating system. Other forms of encryption may be inside Layer 7 (STARTTLS), or inside the inner body, such as OpenPGP (Pretty Good Privacy). Thus SSL/TLS may be applied outside any Layer 7 protocol, for example to HTTP to create https, SMTP to create SMTPs and so on. After it is applied, it is sent using layer 4 just as if it were not encrypted.
- “Layer 7”, the Application Layer, is the “innermost” layer of the OSI model, and includes many different protocols for sending potentially large files (such as FTP or HTTP), email (SMTP, POP, IMAP), instant messages (XMPP), voice, and/or establishing interactive sessions or streams (RTSP). This will include commands and response codes (such as 404 not found). However, layer 7 is not usually the end of the metadata.
- Data transmitted inside layer 7 protocols typically consist of layer 7 headers followed by a corresponding body. The header may include To and From addresses (for email), last modified dates and state cookies (for web/http(s)), information about how the body is to be decoded or treated and its length, information about the software involved and debugging information for technical staff to read, and so is really yet another wrapper of metadata inside layer 7.
- Then finally there is the message body, what you want to see – except it may be encoded using MIME, and be composed of multiple MIME parts, like the pass-the-parcel might turn out to be multiple gifts; and the file might be in one of several standard formats, ultimately containing the information you want. Presumably file extensions and MIME types are not “communications data” (despite answering a “what type of communication?” question). In the case of a web page, this will be marked-up using some version of HTML (such as HTML5) which itself includes a hierarchical model of objects including a “head” (again more metadata such as a title) and a “body” (the text sections and formatting) and refers to other objects accessible via URLs, such as images, fonts or client-side scripts that run in your browser to perform practically any function.
- The layer 7 application may actually not be end user data but be there to facilitate the process of communication, such as if the protocol is BGP or DNS (Domain Name System). Routing information is the only ultimate content in these cases, and is actually part of the process of transmission (it has been claimed [by law enforcement wanting to interfere in a domain registry] but not tested that this does not make it subject to the e-Commerce “mere conduit” exemption from liability since it is also content, but then so is all metadata; if this claim is true however, it would follow that it is not “communications data” despite the indications given to Anderson).
- The above wrappers apply to the simple case of server-client communication, but also to more complex systems involving encapsulation or tunnelling. For example, after unwrapping layers 1-4, decrypting and unwrapping layer 7, inside it may be another layer 2, inside which is another layer 3, inside which is another layer 4, inside which is the layer 7 data actually intended for the application. This is the situation with a VPN (Virtual Private Network) using a protocol like L2TP, and enables a secure connection like an office LAN when you're actually working from the other side of the world. This is even repeated multiple times in the anonymising network Tor (The onion router); layers 3-7 are encrypted and repeated at least three extra times. So just as when you unwrap a layer in pass-the-parcel, you are not sure if you will find another layer of metadata or spoil the surprise.
- It should be apparent from all the preceding that an internet (or the Internet) is not a “space”[10] in any sense (except perhaps in the specialised mathematical sense that some identifiers have a namespace, similar to variable or module names in computer code). The internet is better conceived of as a single medium (or protocol) with characteristics of fidelity (like air) that encapsulates and enables an indefinite number of other media. Each of these media has different characteristics and is based on one or more Layer 7 protocols; so a moderated e-mail discussion list, Twitter and an instant messenger application and an online game have some commonalities and different cultures and weaknesses, but are all provided over [within] TCP/IP.)
- If you realise from the above that the internet is just a great big onion, it becomes clearer why DPI (Deep Packet Inspection) is so called. A Layer 3 connectivity provider transports IP datagrams, not TCP packets (segments), not email or cat pictures. A Layer 2 connectivity provider (such as an office LAN, internet exchange or fibre provider) transports frames. A Layer 7 (for some reason called “over the top” in these discussions) provider might provide email or a microblogging platform. Without encryption, it's possible for a Layer 2 or Layer 3 provider to monitor or interfere with the “deeper” layer 7. Sometimes this is done for anti-abuse reasons, but experts see it as at best “kludgy” and it is often said that if DPI is the solution, you're asking the wrong question (for prioritising traffic on a congested line there are better solutions like DSCP). Hence surveillance will usually be easiest at the appropriate layer. However there are very distinct forms of DPI; those that were used in the 2010s for traffic management often consist of a layer 3 reading port numbers and packet sizes from layer 4 – they don't look for or store “communications data”.
- “Cloud” is a very flexible term, largely used in marketing to mean any service provided using the public internet, including Software as a Service (SaaS) applications such as Google Docs and Platform as a Service (such as Amazon's Elastic Cloud). It can imply that data may be in dynamic or multiple countries.
- “Dark web” or “dark net” can have numerous meanings, including services that are not public for security reasons, and those that are provided anonymously, such as web sites available using .onion addresses. Web 2.0 is likewise not a technical term, but generally refers to web sites with either more user-produced content or a greater amount of code running on a web client (browser).
- IP resolution is a common means of tracing abuse, but cannot always be successful. In our experience, results can vary from red-faced students being brought before their headmaster, to the trail ending at a server that was compromised years ago by malware and command-and-control servers set up using Tor and bulletproof hosting.
- In the past, dial-up internet access often used a “pool” of IP addresses from a block assigned to the service provider, and to trace a home user you would need the time of the incident under investigation plus complete RADIUS logs showing which account was assigned the IP address at the time. Now in the UK, very often ADSL/broadband users have a static IP that is unlikely to change for years, but mobile users are behind NAT (Network Address Translation), where a single IP address may be used by multiple unrelated users at a single time, the port number is necessary to log as well as the IP address, and the amount of data produced includes far more personal information as collateral than RADIUS logs showing when someone was connected to the internet and a simple proxy for device identity. In some cases, identifying a device is not sufficient to identify an individual. As IPv6 (with 2128 possible device addresses) succeeds IPv4 (a mere 3000 million or so), NAT will become less common, but other complexities (not just in amount of data) will make resolution as a tool trivial to evade.
- Even ignoring everyday encryption and deliberate evasion, the notion of communications data can be hard to define. Imagine an object created in an online game or Second Life that is used as a channel of communication – you could talk of sending time (creation) and receipt time as in postal communication, but decoding the entire simulation to find such interaction involves logging and processing practically all activity in the game world. No doubt there are proposed methods to deal with this, but extracting metadata from content requires human interpretation. Unlike the phone system, internet services are based on indefinitely extensible software and while the law can similarly pretend to be indefinitely extensible, in practice the proliferation of software makes the task of keeping up to date and generating “ICR” metadata approximately doubles software expenditure (first to design and code, then to log or monitor, and more if reverse engineering is required).
- Another complexity is the effect of possible abuse. If you don't really want to communicate, you don't need a reply and can forge the various sender addresses, including of someone you want to incriminate (frequently referred to as a “joe job”). Approximately 90% of email SMTP sessions are spam, which can be defined fairly rigorously as unsolicited bulk email, and there are similar problems on websites that accept user-generated content including Twitter. Thus practicalities vastly increase the size of logs when held over a period of time. When the 2005 Data Retention Directive was transposed, I heard it said that “in practice we're not interested in spam; we don't need you to log that”. If only it were that simple: while spam can be defined, a lot of work goes into algorithms to decide what is unsolicited and bulk and they are often fallible. Generally speaking, service providers log bits of metadata in order to detect problems including spam and non-delivery, whereas law enforcement wants logs for the exact opposite reason. Similarly, logging packet headers (which for small samples of data used for troubleshooting is easy enough using software such as tcpdump, ethereal or wireshark) when faced with a Distributed Reflected Denial of Service attack (DRDoS) can in our experience fill up a terabyte hard drive with metadata for DNS queries or TCP SYN connections in a matter of minutes.
- It is notable that the proposals do not include repeal of the Intelligence Service Act s5, nor of RIPA sections 2 onwards, thus retaining the power to compel disclosure of an encryption key and passphrase (or other factor in multi-factor authentication). This is both highly oppressive (note the first prosecutions using this power[11]) and useless. People cannot necessarily reveal their SSL key, and perfect forward secrecy and off-the-record messaging mean it would be useless for decrypting previously captured data. Steganography solutions in future will mean that it may be impossible to even detect the existence of a message, either in random noise or a carrier message.[12] This cannot be “future-proofed” against, and the only effect of such powers (including compulsions in the draft for CSPs to decrypt) is likely to be to discourage providers from being based in the UK.[13]
- Part of the function of the review of powers was to simplify RIPA, which was deemed difficult to interpret even by those agencies created by it.[14] Without clearer language and rigorous delimitations and justifications for all defined terms, a replacement bill could well fail to be an improvement.
Definitions in the draft
- Many of the definitions are very similar to RIPA 2000, but the scope is extended since we no longer have a monopoly or oligopoly of providers. A provider now might be held to include a hotel or café, a home user who has installed wireless, or indeed the producers and suppliers of such equipment. A “private telecommunications system” (clause 193(14)) literally includes a phone, while a “telecommunications operator” (10) controls a telecommunications system, for example, uses a phone.
- The term “Internet” connection record appears to be misleading not only in that it is not correlated to any meaningful “connection”, but that it does not necessarily apply to the Internet or an internet. For example FireChat evaded censorship in Hong Kong by creating a mesh between phones using bluetooth or wireless.
- Members of the Science and Technology committee will appreciate that some definitions are specious. “Telecommunications system” is a system that “facilitates transmission of communications by any means involving the use of electrical or electro-magnetic energy”. Since human beings have not yet learned how to communicate using gravitational waves or neutrinos, this definition includes any type of communication whatsoever, from smoke signals and two cups tied with string. The transmission of sound through air to aural apparatus involves electromagnetic energy in repulsion of N2 molecules, so this literally raises the issue of the “right to whisper”. Is private communication deemed to be essentially antisocial? Did Bentham's Panopticon (before being rejected by Mill), or 1984's “telescreens” even go this far?
- “equipment means equipment that...” (Chapter 3, 149(1)) seems a pleonasm qualified only by the fact that it doesn't produce any “emissions”, that is, is a physical object with a temperature above absolute zero. It's possible a concept more like modulated emissions encoding a signal would define communication, but so does the word “communication”.
- So it is arguable whether such definitions add anything other than embedding obfuscated implications of (extra)territoriality. It may be better to detail what kinds of service are meant to be included, and what not.
Computer Network Exploitation (CNE)/intrusion/cracking
- This concerns those parts of the document that appear to serve to provide legal cover for the secret services more than the police (or other Schedule 4 body with access to CD like DWP, NHS trusts, Ofcom), although c89 extends the technique to law enforcement and armed forces. From one perspective, the admission that the UK Intelligence Services (GCHQ, SIS and MI5/SS, collectively “UKIS”) indulge in “hacking” (also known as “cracking”), potentially opening up unauthorised access to communications and credentials of thousands or millions of people, is a shock. Such behaviour, on smaller scale, has previously been assumed to be the preserve of criminal gangs, misguided teenagers in darkened bedrooms and more recently of News International and Mirror Group Newspapers.
- Various leaked documents provide convincing evidence that GCHQ has attempted to gain access to, for example, SIM card manufacturer Gemalto and ISP Belgacom, who are in all respects unremarkable service providers, with the effect of breaking the privacy of ordinary people on a very wide scale.
- The thrust of much of the report of the Parliamentary Intelligence and Security Committee (ISC)[15] was that the scale of interception was somehow minimal and proportionate although it failed to disclose even approximate figures. This claim that UKIS's privacy violations may be “proportionate” is belied by the disclosures and personal testimony of whistleblowers including Ed Snowden, William Binney and Thomas Drake, as well as independent analysis of the successful subversion of security (eg Dual EC DRBG). Thus few will believe the pseudo-denials of the intelligence services, and most are still waiting for a serious political response.
- From another perspective, cracking is less surprising. Popular awareness of life-saving work in the Second World War by the Government Code and Cipher School remains high, and GCHQ remains a well-resourced organisation with advanced skills in decryption. But the ability to decode enemy messages by brute force is largely history. It has always been the case that with care an uncrackable code can be devised between two people who have met; with public-key cryptography it is possible without meeting so long as there is some third-party validation of identity. The uptake and security of strong crypto is a continuing trend that is understandably welcomed by technology users. This leaves security agencies with the possibility of interception of the plaintext at the endpoint of a communication, planting a software or hardware bug on computers, phones, or intermediate servers. Some may find the intellectual challenge of subverting the security of a remote system, or physically interfering with equipment without detection, appealing even though it may have no immediate use or renders accessible far more information than can possibly be justified (“collateral intrusion”). On the face of it, this is still an unethical and criminal act.
- Besides the direct propagation of malware and weakened security, an indirect concern is that while UKIS is exploiting existing security vulnerabilities, it will have an interest in those vulnerabilities remaining unaddressed.
- Given the way technology is used now, access to endpoint devices gives disproportionate insight into an ordinary person's personal and political life. People share their confidences with a machine in the assumption that it is an inanimate medium of communication with friends, not one that may be actively spying on them.
- A draft code of practice or statutory order to be enacted under an investigatory powers bill would need to elaborate precise classes of “equipment interference”, perhaps defined numerically, and not only specify that they are used in the last resort, but also state under what conditions they are considered potentially proportionate.
- Many senior staff may not fully appreciate how likely software errors are, nor the severity of their consequences, especially with skilled programmers/operators, nor unintended consequences of interference when interacting with other intruders. "Clever" people don't make fewer mistakes, just bigger ones. In one accident (not even involving other parties), the NSA cut off all Syria's link to the internet, and were unable to reverse this.[16] The overall classes of techniques needs public debate in the technical community to assess proportionality. Risks can only be minimised so far; the real question is whether a system administrator would or should ever trust an external bureaucracy with root access (super-user privileges).[17] Any secondary legislation could and should prohibit certain activity outright such as :
• subversion of the Public Key Infrastructure with inappropriate certificates.
• undermining the security of any widely-used system
• introduction of systemic vulnerabilities to software
• interference in BGP routing (whether or not it might propagate)
- Clause 187 seems to extend these (new) EI and bulk EI powers to allow launching of attacks from providers' systems, as well as provide (possibly misleading) security advice without oversight (and extent MI6 and GCHQ's activities on serious crime to the UK.)
Legal and human rights issues
- While we appreciate that this enquiry focusses on technological aspects of state surveillance and interception, it would be a mistake to think that legal aspects can ever be cleanly separated from them. As Lawrence Lessig, professor of law at Harvard put it, “Code is Law”. Not only are there analogies between programming code and law (internal vocabulary and references to define behaviour, problems caused by not considering edge cases), but architecture and practice are things law must accommodate. Thus the inability of this inquiry to reliably assess cost and impact of proposals without fully-worked examples is not totally unrelated to our inability as citizens to judge the same proposals' proportionality. We can also see a attempt to specify behaviour in the definition of “request filter”, and as such it should be read as (arguably badly written and documented) pseudocode.
- The paper uses some of the same phrases as in ECHR Article 8.2: “national security”, “economic well-being” and “prevention... of crime”. However it is far from clear that they have the same meaning as in the ECHR, and Article 8.2 exceptions can themselves be criticised as over-broad. Unmentioned by the draft is the fact that Article 8.2 restricts these exceptions to the right to privacy to those “in accordance with the law and... necessary in a democratic society”.
- However, the cases in mind there are more around CSP data retention than mass (or “bulk”) application of ultra-intrusive device interference giving access to local data. It needs to be considered whether extreme measures are compatible with such principles at all.
- We have not seen the ostensible distinction between “mass” and “bulk” clarified. It is presumed that the government denies mass surveillance because coverage is not expected to be reach 100%. Whether or not access is limited by code or law, blanket metadata generation and storage on a service (one example of “bulk personal datasets”) constitutes mass surveillance, affecting all users of the service indiscriminately.
- Human rights courts have repeatedly confirmed that it is not the analysis of data that is the sole threat to privacy, but the mechanical collection and storage of data violates Article 8 (Kennedy v UK, s162; Amman v Switzerland s 69; Rotaru v Romania s43). Also the Halford case showed that all interception requires the subject to know when they might be being spied on.
- These proposals may grant powers to law enforcement previously reserved for the secret services, but there we see nothing to prohibit a class warrant from allowing UKIS (or indeed law enforcement) access to the content of the co-ordinating “request filter” to be used in data mining or fishing expeditions.
- Senior secret service staff are far from the only ones dealing with potentially life or death decisions - there's a whole area of medical ethics written about it. This doesn't just affect emergency services, but even town planners, and a common sense of proportion might could be applied. If right to life overcame all other rights and considerations, wouldn't we ban cars?
- Related to such programmes, the amount of data that can be generated by unauthorised bugging of CSP routers or servers is apparently so great that it needs to be analysed by “expert systems” (“artificial intelligence”) to “develop targets”. This raises interesting questions about whether individuals, organisations or an international complex of automated systems are the ones doing the interception/CNE. In particular use of such systems could be expected to intrude on “false positives”, flagging up what a human analyst is looking for. The analyst then uses related criteria to assess whether further intrusion on that “target” is justified, and the human being falls into a folie a deux with the machine.
- There are outstanding questions about how the bill would interact with Data Protection obligations. Is the data processor in this case the service provider, and the data controller is the Home Office (or someone else?), since they determine what personal information is to be kept. It seems possible the Home Office has not even considered CD as personal data. It clearly is: since the main purpose of “communications data”, “event data” and “entity data” is to identify a (ultimately human) sender, it is almost the very definition of personal data. How do data protection principles apply to data that is retained for a fixed period of time regardless of whether it is of any use? Is a completely different standard of data protection to be applied here than to all other personal data, and if so exactly how?
- Fuzzy or ill-defined powers will be extended in practice, much as they were under RIPA (for example, “thematic” s8(1) warrants), and are more open to abuse than well-defined foreseeable ones. We share concerns that judicial oversight in these proposals is only rubber-stamping the process, and not up to international standards. We would also strongly recommend that if these powers are used, the actual record of what they are used for and what effects they have becomes public after a period of time, thus enabling both people to know what the intrusion on their privacy has been and also a debate on whether practice is in accordance with human rights. This is the situation in other countries. While the Interception of Communications Commissioner has made welcome progress in publishing statistics, we still have not seen examples and statistics about when surveillance and interception have been strictly necessary, and when merely possibly helpful or routine.
- The social effect of government interference would suggest we're already beyond the point at which a decrease in privacy would improve public safety, at least to any significant extent. As is frequently remarked, perpetrators of truly serious crime (mass murder and injury) are almost always already identified (the exception being the lone wolf Anders Breivik, where monitoring fertiliser and weapon deliveries would have been more effective than creating social graphs). What is contemplated here does not exist in any other EU member state to my knowledge, and blocking and monitoring at this level is most developed in Iran, and particularly China, and those countries are good exemplars that interference in the internet medium does not significantly inhibit crime such as fraud, but can be an effective means of social engineering and suppressing dissent. (Ed Snowden is particularly concerned about how surveillance causes censorship to be internalised or introjected.) Indeed proper network security (and measures like BCP-38) is one of the best defences against e-crime or online crime. A Professor Glees gave evidence on the DCDB that surveillance techniques used in the UK are nothing like the Stasi on the grounds that a much smaller proportion of the nation are engaged in spying, at least on their fellow citizens.[18] This was entirely misleading. Nowadays most of us leave a massive digital footprint, capable of being analysed in an automated fashion to provide a partial picture of one's life and interests. (Will some of this data relevant to investigations be stored for longer than 12 months? What codes and safeguards apply then?) In Amsterdam from the 1900s onwards, city authorities collected detailed religious demographic data for municipal service provision. Once the Nazis occupied the Netherlands, this data was used for non-benign purposes. Government hacking activities and access to “bulk personal datasets” show that unanticipated use of technology has enabled vast expansion of powers from the days when phone tapping required the laborious attachment of crocodile clips and was reserved for the most serious cases. Internet service providers need legal clarity and accountability of state actions, and also security and full control over their equipment. Users too need security against hackers, and the confidence to use the services. The possible grounds of privacy violations need to be resolved and properly specified at the earliest opportunity to ensure that genuinely free conversation is possible.
- It would be good if more people in law enforcement could realise that (a) just because some information is sometimes available and useful in some cases at a particular period of time, it does not mean that it is or should be equally available or even meaningful in all cases; (b) nor does it mean that such information will always be equally available as technology changes; (c) communications networks and the entire value system of society do not exist just to make their jobs easier. As former Home Secretary James Callaghan noted, it is a given that senior police will always want more power no matter how much they have (or how impractical), it also being easiest to blame one's tools when something goes wrong.
Interaction with technology industry
- For certain markets, such as those promoting DPI, these kind of proposals will be very welcome and contractors who can provide a general solution stand to benefit greatly. China has both the greatest market and the greatest expertise in manufacturing DPI boxes, but European companies can also provide malware to order.[19] We are concerned about influence of marketing of such “solutions”, overstating capabilities and accuracy and understating maintenance costs of repeatedly proposed legislation.
- On the other hand, if hardware or software manufacturers might be compelled to install backdoors under new legislation, or even perceived to be, then national manufacturing would suffer and lose investment. Similarly, the UK would hardly be first choice for data storage, against somewhere like Iceland or Germany, if it were believed that detailed user logs were kept on a VPS.
- Data storage would be a significant but unknown cost in the short term. RADIUS logs had been stored under the DRD on CDs and DVDs using automated disk-changer equipment. A lower cost solution nowadays for the larger quantities of data produced by NAT (and other even more intrusive possibilities) tends to be cheap 2-3 TB 3.5” hard drives, but there are few solutions for automatically changing and archiving HDDs. Most commercial solutions are likely to involve NAS (network attached storage), which requires additional network infrastructure and being online continuously exposes large quantities of data to attack; it most likely would also be subcontracted to a larger supplier, raising additional competition concerns besides the cost-recovery subsidy of larger ISPs. If all TCP/IP headers have to be captured, the amount of data would be “silly expensive”, a large fraction (perhaps a third, depending on application protocol) of all traffic. It seems unlikely this is what is contemplated in the short term given the lower costs in impact assessments that for the Draft CD Bill, but as shown above it has so far proved impossible to get a commitment on this, and very difficult to get a clear indication of intention.
- Denmark has attempted to force gathering by ISPs of usage data of over-the-top services by sampling IP headers, with difficulty and no reported benefit. An alternative approach to minimise data might theoretically be to log (only) the start (SYN+ACK) and end (FIN/RST) of TCP sessions, with the complication that TCP also has persistent “half-open” and “half-closed” states.[20] Of course this data may not have much value since it doesn't log data flow and can be evaded by anyone with a mind to do so by using other protocols.
- Even if retention is feasible, security requirements (see clause 74(3)) may not be. Note that data erasure of media is non-trivial and subject to error: to remove data usually requires around ten slow write operations of the drive with random data, and even then some security experts (eg Glyn Wintle) suggest that overlapping magnetic domains may retain much old data that could be recovered with high-tech processes.
- “Cost recovery” is of concern to most ISP management, because the bill does not assure providers that 100% of the cost of data retention, generation and access will be reimbursed; given the fiscal climate, once the bill is passed, there are fears government would not provide enough resources for even cheapest, low-quality suppliers, nor to recruit skilled developers.
- With the caveat already stated that the newer provisions are obscure or arcane or we just haven't got to grips with them, it is also possible that ICRs are meant to include “web logs” in the original sense of logs kept on a web server that is run by a hosting company or social network service, or indeed other logs such as email or FTP. However if they are defined not to specify content (as c193(5-6) suggests), it does not include the URL path (to the right of the hostname and first single slash), and these cannot obviously be used to, for example, trace all members of a Facebook group (hostnames and IP addresses just identify Facebook). In this case it seems maybe that drawing a social graph from social media would rely on warranted cracking/CNE or “bulk personal datasets” of all Facebook users in order to get at those with a “common purpose” (c83(b)). All the same, caveats about reasonableness may not be sufficient to protect smaller providers and could well fall most heavily on them: for example, maybe it is intended to ask web and email hosts to log remote port numbers as well as remote IPv4 addresses for correlation with CGNAT logging. It would not be desirable, realistic or proportionate to put on notice and gag so many application-layer providers operating in the UK, of whom there are at least tens of thousands. Incidentally, the current draft does not recognise internet layers and so the exclusion of “content of a communication” in 193(5) means “communications data” cannot be communicated or, effectively, exist.
- Technological lock-in has not been discussed enough. While the intention now before the bill is passed appears to be not to compel uneconomic retention by smaller providers, the possibility of doing so is open and would stifle innovation and gains in security, performance and efficiency (eg a DPI box meant to gather Host & SNI headers from HTTP/1.1 might require changes to architecture of a HTTP/2 server, or prevent introduction of secure encryption between data centres). c189(4) and c188 could be read to grant unlimited powers to compel service providers to adopt particular technologies or potentially to continue to provide insecure service.
- The language used to describe the request filter has not changed since the draft CD Bill. It does however clearly involve an application programming interface to live databases with huge amounts of metadata held by CSPs.
- The CNE (cracking) and Computer Network Attack powers have already been used against known foreign service providers, and directly affect the security and performance of CSP networks, the privacy of their users, and particularly of their staff.
- A Technical Advisory Board set up under the bill seems insufficient, and certainly does not enable informed citizens or parliamentarians to assess the effects of the suggested legislation. It is the manner in which previous lack of clarity and foreseeability in existing legislation violates human rights and legal principle that has forced for example disclosure (or avowal) of the use of Telecommunications Act s94 and publication of the draft Equipment Interference Code and contributed to the striking down of the EU Data Retention Directive (DRD) as a result of the Digital Rights Ireland case. The wide undefined powers in this draft would seem to have more severe problems in that respect.
Conclusion
- The concept of an “Internet Connection Record” is perhaps the most prominent of the problematic areas in the Home Office draft, along with the many other new powers to weaken security and the draconian, impractical and unnecessary new disclosure offences particularly for CSPs in paragraphs 31, 77(2), 101, 133, 148 and 190(8). Questions of the feasibility and proportionality of generating and retaining data cannot be answered unless we have a good idea of what data is intended, in other words how an “ICR” would need to be generated. If it were to mean logging new entries in the table of established TCP connections (and just possibly UDP datagrams) in NAT boxes run by mobile operators, it means a cost of many millions, for which law enforcement gets the same likelihood of tracing an IPv4 address to a device on mobile networks as it does for a device on home broadband. However, that was supposedly why the Queen's Speech in 2013 mentioned the “problem of matching internet protocol addresses”[21] and a “gap” that supposedly the Counter-Terrorism and Security Act 2015 (CTSA) s21 closes. It may be that there was a defect in the CTSA such that it did not enable the intended powers (see footnote 5). However the impact assessment of the new draft has a surprisingly precise figure of £187.1m for “communications data” that is additional to the CTSA impacts. It could be that the intention is to capture similar (Layer 4) data (IP addresses and ports) for non-mobile consumers as was attempted in Denmark; this involves a considerable interference with ISP routing, would probably cost billions as in the CDB, has no evident benefit whatsoever (duplicates data available by other means; error-prone and partial coverage), and would be a massive privacy intrusion for all users of UK internet services. It may also be that Host and Cookie headers (Layer 7) are intended, which again is “feasible” but unnecessary and raises costs (and incidentally carbon emissions) by perhaps an order of magnitude again, probably to more than the entire police budget.
- We believe the questions asked by the Science and Technology enquiry are very important to answer. However, the task of pinning down officials to provide the necessary information for independent experts to assess has proved difficult if not impossible. As it is, the impacts of the current draft for technology providers and users alike are vague and practically unlimited. After several failures to satisfy all interest groups, it is clear that the Home Office (as advised by private corporations) is merely one interested party in a conflict, and may not have the necessary neutrality and technical expertise to draft legislation to resolve that conflict.
- Thank you for your consideration of these comments. The principal author was Cedric Knight, technical consultant for GreenNet. This response may be published and attributed, and we would welcome any further questions via our details as below.
December 2015
[1]Willetts, Peter, Non-Governmental Organizations in World Politics: The Construction of Global Governance, Routledge, 2010. p113.
[2]See Freedom of Information request by David Davis MP at http://www.theyworkforyou.com/wrans/?id=2013-01-31b.132300.h
[3]See the account at http://www.revk.uk/2015/11/home-office-ipbill.html and associated material
[4]A Question of Trust, Chapter 15, p 288. Also p 178: “I am not aware of other European or Commonwealth countries in which service providers are compelled to retain their customers’ web logs for inspection by law enforcement. I was told by law enforcement both in Canada and in the US that there would be constitutional difficulties in such a proposal.”
[5]It appears to us that the necessary clause was written the wrong way around: the information requested was for resolving an IP address to an individual, not vice versa.
[6]Does this include any Public Order Act offence? Or is limited in case law?
[7]A Question of Trust, pp 177-8.
[8]Target Detection Identifiers, 2009, https://theintercept.com/document/2015/09/25/tdi-introduction/
[9]See Privacy International's short explanatory film on metadata at https://privacyinternational.org/node/573
[10]The word “cyberspace”, used by fiction writers from 1980 onwards to imagine a Matrix-style sub-reality, may thus be doubly misleading. The prefix “cyber-” refers to cybernetics, the use in engineering of control mechanisms drawn from living organisms, and its usage is mocked in expert circles where “digital”, “electronic” or “e-” would be more respected as descriptive.
[11]Williams, C, “UK jails schizophrenic for refusal to decrypt files”, The Register, 24 Nov 2009
[12]Richard Feynman had an amusing story about Manhattan Project security; he wrote to his wife mentioning a pattern in the decimal expansion of 1/243, a series of digits which was redacted and caused suspicion. Feynman pointed out that since 1/243 could not be anything else, the number conveyed no information whatsoever. To no avail.
[13]This “technical illiteracy” in the face of serious crime has been criticised recently in for example New Scientist (https://www.newscientist.com/article/dn28445-uk-surveillance-bill-makes-a-scientific-ass-out-of-the-law/) and the NewYork Times: http://www.nytimes.com/2015/11/22/opinion/stopping-whatsapp-wont-stop-terrorists.html , http://www.nytimes.com/2015/11/18/opinion/mass-surveillance-isnt-the-answer-to-fighting-terrorism.html
[14]See Graham Smith's evidence to Anderson https://terrorismlegislationreviewer.independent.gov.uk/wp-content/uploads/2015/06/Submissions-H-Z.pdf and his blog http://cyberleagle.blogspot.co.uk/2015/08/the-coming-surveillance-debate-legal.html
[15]“Privacy and Security: A modern and transparent legal framework ”, ISC 2015
[16]See Wired magazine interview with Snowden: http://www.wired.com/2014/08/edward-snowden/ For an famous error involving UKIS see http://www.theregister.co.uk/2013/12/31/nsa_weapons_catalogue_promises_pwnage_at_the_speed_of_light
[17]The short answer is no. Experienced sysadmins do not trust their own management with root.
[18]http://www.parliament.uk/documents/joint-committees/communications-data/Oral-Evidence-Volume.pdf
[19]For example FinFisher, based in Hampshire: http://www.dawn.com/news/1177605/fishing-in-troubled-waters
[20]The famous TCP 'three-way handshake' inspires the joke: A TCP packet walks into a bar and says: “Hello, I’d like a beer.”. The barman replies “Hello, you’d like a beer.” “Yes,” replies the packet, “I’d like a beer.”
[21]See Richard Clayton (2013): https://www.lightbluetouchpaper.org/2013/05/08/traceability-in-the-queens-speech/