Minimal personal data

What this criterion guarantees: The operator knows nothing about you. Data it does not hold can be neither hacked nor sold.

Without it: Your messaging app is tied to your real identity, and that link is exposed to leaks and cross-referencing.

At a glance
• Olvid — ✅ Good
• Signal — ❌ Poor
• WhatsApp — ❌ Poor
• Telegram — ❌ Poor
• Matrix-based — ❌ Poor
• SimpleX — ✅ Good
• Threema — 🟠 Partial

To function, a communication system does not need to know a user’s name: a “random” identifier is enough. With email, for example, everything would work just as well with addresses made of randomly drawn characters (even if it would probably be less convenient).

What do we mean here by a “random” identifier? It is an identifier that protects its holder by guaranteeing two properties:

  • it is dedicated, that is, it is used only for this communication service and nothing else;
  • and it is tied to no pre-existing personal data, nor derived from a name, an email address or a phone number.

An identifier drawn at random, specifically for the service, has both properties by construction: it reveals nothing, and allows no cross-referencing.

The phone number: a rigid identifier, impossible to multiply, and recycled

From this perspective, the phone number fails twice over. It may look “random” (a string of digits with no apparent meaning), but it is neither dedicated nor unlinked: it is used to make calls, it is demanded by countless websites and services, and it is assigned by a carrier that knows the legal identity of its holder. Anyone who knows your number can therefore link you to your other accounts, and often to your real identity. Far from being an anonymous identifier, the phone number is a pivot for cross-referencing.

The phone number has one last flaw: you do not choose it, and you cannot freely get rid of it. A random identifier can be created, multiplied and abandoned at no cost; a phone number cannot. How do you keep two perfectly separate messaging accounts, one personal, one professional, when the identifier is a phone number? You need two phone lines. How do you keep your account when you change numbers, for instance when moving abroad? And what becomes of an abandoned number? It is reassigned: the SIM card canceled when leaving a country, or the work line handed back with the phone when leaving a company, end up in a stranger’s hands within a few weeks. Contacts who still write to that number are then addressing a stranger, and that new holder can re-register the messaging account that was attached to it. We explain, moreover, in the identity substitution criterion that this need to let a reassigned number change hands is precisely what prevents Signal and WhatsApp from permanently locking their users’ accounts.

A name for the contacts, not for the server

If a random identifier is enough, why do our email addresses so often contain our names? For a purely practical reason: finding a correspondent in your address book is easier with “jane.smith” than with a meaningless string of characters. But that convenience does not require exposing the name to the server. It is enough for the name to be known to the contacts, who each associate it, in their own address book, with the random identifier. The name is information meant for humans; the identifier, information meant for machines. Nothing requires the service operator to know the former.

What does the server really need to know?

In an instant messaging app, the server needs an identifier to play its role as a relay: knowing which messages to deliver when a device comes to fetch them, and which device to notify when a message arrives. As with any online service, it also necessarily sees the IP address from which requests reach it. Finally, in practice, it most often keeps one notification token per device, which we come back to below. The inventory stops there: an identifier, an IP address, a token. Nothing else is needed to operate the service, that is, to relay messages. Any additional data (name, phone number, email address, address book) is collected for other reasons: to facilitate contact discovery, or to fight fake accounts and spam. We will see in the next criterion that this need itself stems from a design choice: that of letting anyone contact anyone.

What every relay server observes: the social graph

Small as it is, this inventory is enough to build a sensitive piece of information: the social graph. To distribute messages, the server necessarily knows, for each user, the public key with which they authenticate in order to retrieve the messages addressed to them, and the IP address they connect from. Sending, on the other hand, can remain anonymous from the server’s point of view: as with a letter sent by post, nothing requires identifying whoever drops off a message, and some solutions, such as Olvid, indeed refrain from doing so. This application-level anonymity does not protect the graph, however: by correlating IP addresses and the timing of drop-offs and pick-ups, the server can reconstruct which identifier communicates with which, how often and when, whatever the quality of the encryption.

Minimizing personal data does not make this graph disappear; it determines what the operator can attach to each of its nodes. Without any personal data, the nodes remain pseudonymous: a public key, an IP address. When the application requires a phone number, as Signal or WhatsApp do, each node is, on the contrary, tied to a legal identity from the start: the graph of pseudonyms becomes a named directory of everyone’s relationships. In between, every piece of data provided, even if optional, enriches the corresponding node. Note that SimpleX is an exception by construction: identifiers there are specific to each conversation, with no global identifier linking them, which fragments the very nodes of the graph.

The case of push notification tokens

To notify the arrival of a message, most messaging apps rely on Apple’s and Google’s notification services. For that, the messaging server keeps one “push token” per device. From its point of view, this token is just one more opaque identifier, which reveals nothing about its holder. But it is opaque only to the messaging server: at Apple and Google, that same token is tied to the user’s account, and therefore to their name and email address. In late 2023, a letterarchived from US Senator Ron Wyden revealed that governments were demanding from Apple and Google the data associated with these tokens, precisely in order to link anonymous messaging users to their Apple or Google account; Apple indicates, moreover, in its guidelines for law enforcement, that the identity associated with a token could initially be obtained with a simple subpoena; since the revelation, Apple requiresarchived a court order. In other words, even if the messaging app knows nothing about you, a third party can link your messaging identity to your real identity. This is what led several messaging apps to offer, on Android, alternative notification mechanisms, without a token, relying on a permanent connection to the server: Threema with Threema Pusharchived, Signal, whose application installed without Google servicesarchived switches to this mode on its own, and Olvid, where a single settingarchived is all it takes. The price is slightly higher battery consumption; on iOS, this approach runs into the severe limitationsarchived the system imposes on permanent connections.

Less data, less risk

Every piece of data the server does not hold is data that can be neither hacked, nor sold, nor requisitioned. Data minimization is therefore not just a matter of principle: it is a direct reduction of the attack surface. It also has a welcome regulatory consequence for organizations: deploying a messaging app that collects no personal data means fewer processing operations to justify under the GDPR, and simpler compliance.

Applying this criterion

A messaging app gets a ✅Good if it can be used without providing any personal data to the operator — no name, no email address, no phone number — and offers no way to provide it with any. A name shared with contacts only, and never with the server, falls outside this scope, as explained above. Solutions where providing such data to the operator is possible but optional get a 🟠Partial. Both levels further assume that collection stops at the inventory described above: an identifier, an IP address, a notification token. A solution whose servers retain conversation metadata persistently (members, senders, timestamps), or even spread it to other servers, gets a ❌Poor, regardless of the identifiers requested at sign-up.