How to Encrypt PII in Data Pipelines While Keeping It Searchable

At Intuit, my team built a data pipeline that processed TurboTax e-filing data. This data needed to be made available for dashboards, analytics, and other downstream use cases.

One of the biggest challenges was protecting Personally Identifiable Information (PII) such as Social Security Numbers (SSNs).

SSNs are highly sensitive and must be protected at every stage of the data pipeline. At the same time, analytics teams sometimes need to perform legitimate business operations such as identifying records associated with a particular customer or performing exact-match lookups.

So how do you protect sensitive data while still allowing controlled equality searches without exposing the plaintext?

In this article, I'll explain the encryption concepts behind this problem, including:

  • Probabilistic vs. deterministic encryption
  • Equality searches on encrypted data
  • Base keys and derived keys
  • Key separation
  • Local vs. remote cryptographic operations
  • Tokenization
  • The security trade-offs involved in making encrypted data searchable

Table of Contents

Why Encrypt PII?

Regulations such as GDPR, CCPA, and HIPAA impose strict requirements on how PII is stored and processed. Beyond compliance, a data breach involving unencrypted PII can lead to severe financial and reputational damage. Encrypting PII at rest and in transit is a foundational control, but it often conflicts with the need to query and analyze the data. This is the central tension we address here.

AES: Advanced Encryption Standard

AES is a symmetric block cipher that forms the backbone of most modern encryption. It operates on fixed-size blocks (128 bits) and supports key sizes of 128, 192, or 256 bits. When used with a mode of operation like GCM or CBC, it can encrypt arbitrary data. AES is fast, well-analyzed, and widely available in hardware, making it ideal for high-throughput data pipelines.

However, AES alone does not solve the searchable encryption problem. The mode and initialization vector (IV) handling determine whether identical plaintexts produce identical ciphertexts—a property that directly affects searchability.

Probabilistic Encryption

Probabilistic encryption uses a random initialization vector (IV) or nonce for each encryption operation. As a result, encrypting the same plaintext multiple times yields different ciphertexts. This is the gold standard for security because it prevents attackers from inferring information by comparing ciphertexts.

But there's a catch: you cannot perform equality searches on probabilistically encrypted data. If every ciphertext is unique, there's no way to look up a record by its encrypted value without decrypting everything. This makes probabilistic encryption unsuitable for scenarios that require exact-match lookups.

Deterministic Encryption

Deterministic encryption always produces the same ciphertext for the same plaintext and key. This property enables equality searches: you encrypt the search term with the same key and compare it to the stored ciphertext. No decryption is needed.

The Security Trade-off

The downside is that deterministic encryption leaks information. An attacker with access to the ciphertext can see which records share the same value. If the plaintext space is small or predictable (e.g., SSNs, which have a known format and limited range), this can enable frequency analysis or even brute-force attacks. Therefore, deterministic encryption should be used only for fields that require equality searches, and ideally combined with other controls such as access restrictions and key separation.

Base Keys and Derived Keys

In a production pipeline, you don't want to use a single master key for all encryption operations. Instead, you use a base key (or master key) to derive multiple purpose-specific keys. This practice, known as key separation, limits the blast radius if one derived key is compromised.

Key Derivation with HKDF

HKDF (HMAC-based Key Derivation Function) is a standard for deriving keys from a master key. It takes a salt and context information (e.g., "deterministic-ssn-key") to produce a derived key. This allows you to have separate keys for deterministic encryption, probabilistic encryption, tokenization, and so on.

Where Does the Master Key Live?

The master key must be stored securely, typically in a hardware security module (HSM) or a key management service (KMS) such as AWS KMS, Azure Key Vault, or Google Cloud KMS. The pipeline retrieves the master key (or a derived key) at runtime, never hardcoding it or storing it in plaintext.

Local Cryptographic Operations

For performance, encryption and decryption are usually performed locally within the pipeline using keys fetched from a KMS. This avoids sending plaintext PII to a remote service, reducing exposure. The KMS is used only for key management, not for bulk encryption. This hybrid approach balances security and efficiency.

Tokenization

Tokenization replaces sensitive data with a non-sensitive surrogate (a token). The original value is stored in a secure vault, and the token is used in the pipeline. Unlike encryption, tokenization does not involve mathematical operations on the data; the mapping is stored separately. Tokenization can be used for equality searches if the token is deterministic (e.g., a hash of the value with a secret salt). However, it requires managing a token vault, which can become a bottleneck.

Choosing the Right Approach

There is no one-size-fits-all solution. The choice depends on the required search capabilities, the sensitivity of the data, and the threat model. Often, a combination is used: deterministic encryption for fields that need equality searches (like SSN), and probabilistic encryption for other PII. In all cases, key management and access controls are critical. As data privacy regulations evolve in 2026, staying informed about emerging techniques like searchable symmetric encryption (SSE) and fully homomorphic encryption (FHE) is advisable, though these are not yet practical for most production pipelines.

By understanding these concepts, you can design a pipeline that protects PII while still enabling the analytics your business needs.

via FreeCodeCamp

Related