Tokenization vs Encryption for Card Data

Every few months a product manager or a board member asks me some version of the same question: "We encrypt the card numbers, so we are PCI compliant and safe,...

Originally published onanselmfowel.com

Every few months a product manager or a board member asks me some version of the same question: "We encrypt the card numbers, so we are PCI compliant and safe, right?" The honest answer is "it depends," and the part it depends on is whether we are talking about encryption or tokenization. The two words get used interchangeably in slide decks, but they solve different problems, carry different operational burdens, and put different parts of your business in scope for an audit.

Tokenization vs Encryption for Card Data
Tokenization vs Encryption for Card Data

I have built and inherited systems that did both, sometimes badly. What follows is how I think about the choice today as a CTO responsible for card data in a regulated environment. It is not a religious argument for one over the other. It is a description of where each one earns its place, where each one quietly fails you, and how to reason about the decision before a QSA forces the conversation for you.

What the two words actually mean

Encryption is a reversible mathematical transformation. You take a primary account number, run it through an algorithm with a key, and get ciphertext. Anyone with the key and the algorithm can recover the original value. The security of the scheme rests entirely on the secrecy of the key. The ciphertext still contains the card number in a very real sense; it is just locked. If the key leaks, every value you ever encrypted is exposed retroactively.

Tokenization replaces the card number with a surrogate value that has no mathematical relationship to the original. The token is meaningless on its own. To get back to the real PAN you must look it up in a token vault, which is a separate, hardened system that stores the mapping. The token itself, if stolen, is useless to an attacker because there is no key that reverses it. The only path back to the card number is access to the vault.

That distinction sounds academic until you trace where the sensitive data physically lives. With encryption, the sensitive data is in your application database, your backups, your replication streams, and your logs if you are careless. With tokenization, the sensitive data lives in exactly one place, and everything else holds an inert reference.

Why the scope question dominates everything

The single biggest reason I lean toward tokenization for card data is PCI DSS scope reduction. The standard applies to any system that stores, processes, or transmits cardholder data, and "in scope" is expensive. Every server in scope needs hardening, file integrity monitoring, quarterly scans, access controls, and an audit trail that you can defend to an assessor. The more systems touch the PAN, the larger and more costly your compliance perimeter becomes.

Encrypting the PAN does not remove a system from scope. If the encrypted data and the means to decrypt it can both be reached from a host, that host is in scope. In practice, encrypted card data spreads. The order service needs it, the analytics pipeline copies it, the data warehouse ingests it, and suddenly half your estate is in scope because the ciphertext is reachable even if no human can read it.

The fastest way to fail a card-data audit is to be unable to draw an accurate diagram of everywhere a PAN can flow. Tokenization lets you draw a small diagram. Encryption alone usually forces you to draw a large one.

With tokenization done correctly, only the vault and the narrow path into it are in scope. The hundred other services that used to carry encrypted PANs now carry tokens, and tokens are not cardholder data. I have seen this take a compliance footprint from dozens of servers down to a handful, which is the difference between an audit that takes a week and one that consumes a quarter.

Where encryption still earns its keep

None of this means encryption is obsolete. Encryption is the right tool when you genuinely need to recover the original value at scale, inside your own trust boundary, and when the data is not a payment card. Customer addresses, bank account details for payouts, government identifiers, and the contents of the token vault itself all want strong encryption. The vault has to store real PANs somewhere, and inside it those PANs should be encrypted at rest with keys managed in a hardware security module.

Encryption is also unavoidable as a layer underneath tokenization. Transport security, encrypted backups, and column-level encryption for non-card sensitive fields are all baseline hygiene. The mistake is treating application-layer PAN encryption as a substitute for tokenization rather than as a defense-in-depth complement to it. They operate at different layers and answer different questions.

There is also a latency and dependency argument. A token vault is a network call and a runtime dependency. If you have a use case where you must transform millions of values offline with no external lookup, format-preserving encryption can be the pragmatic answer, provided you keep the keys in an HSM and accept that those systems remain in scope.

Format-preserving encryption and its traps

Format-preserving encryption deserves its own warning because it is frequently marketed as "tokenization without the vault." FPE produces ciphertext that looks like a card number, sixteen digits, so it slots into existing database columns and validation rules without schema changes. That convenience is seductive and it is exactly where teams get burned.

FPE is still encryption. There is a key, the transformation is reversible, and the value is mathematically derived from the PAN. A host that can perform FPE decryption is in scope, full stop. Teams adopt FPE to avoid changing their schema and then convince themselves they have tokenized, which gives them the schema convenience of tokens with none of the scope benefit. When the assessor arrives, the distinction is the first thing they probe.

I am not against FPE, but I insist the team be precise about what it is. If the goal is scope reduction, FPE alone does not deliver it. If the goal is preserving format inside an already-in-scope vault, FPE is a reasonable internal mechanism. Confusing those two goals has cost teams I have advised a great deal of remediation work.

Key management is the real job

Whichever path you choose, the actual engineering difficulty is key management, not the cryptographic primitive. AES is not going to fail you. Your key lifecycle will. The questions that keep me up are mundane and unglamorous: who can request a decryption, how often do keys rotate, how do you re-encrypt historical data after rotation, where do the keys live when the application boots, and who is alerted when a key is exported.

A few principles I hold to regardless of the approach:

  • Keys never live in source control, environment variables, or config files. They live in a key management service or HSM and are fetched at runtime with audited access.
  • Data-encrypting keys are themselves encrypted by a key-encrypting key, so rotation of the master key does not force re-encryption of every record.
  • Every decryption operation is logged with the identity that requested it, and those logs are immutable and monitored.
  • Rotation is a tested, scheduled procedure, not a fire drill you run for the first time after a suspected compromise.
  • The blast radius of any single leaked key is bounded by design, through key hierarchies and short rotation windows.

Tokenization shifts this burden rather than eliminating it. The vault still uses encryption internally, so you still own key management, but you own it in one concentrated, heavily monitored place instead of scattered across every service that handles payments. Concentrating the hardest problem into the smallest, best-defended footprint is the entire point.

A concrete vault lookup pattern

Here is the shape of the boundary I aim for in a .NET service. The application code only ever sees tokens. When it genuinely needs the PAN, for example to send a transaction to an acquirer, it calls the vault through a narrow, audited interface and never persists the result.

public interface ITokenVault
{
    Task<string> TokenizeAsync(string pan, CancellationToken ct);
    Task<string> DetokenizeAsync(string token, string reason, CancellationToken ct);
}

public sealed class CardPaymentService
{
    private readonly ITokenVault _vault;
    private readonly IAcquirerClient _acquirer;

    public CardPaymentService(ITokenVault vault, IAcquirerClient acquirer)
    {
        _vault = vault;
        _acquirer = acquirer;
    }

    public async Task<ChargeResult> ChargeAsync(Order order, CancellationToken ct)
    {
        // order.CardToken is an inert reference; it is safe in our database.
        // The real PAN exists only for the lifetime of this call.
        var pan = await _vault.DetokenizeAsync(
            order.CardToken,
            reason: $"charge:{order.Id}",
            ct);

        var result = await _acquirer.SubmitChargeAsync(pan, order.Amount, ct);

        // Never log, cache, or store pan. It falls out of scope here.
        return result;
    }
}

The important properties are not in the cryptography. They are that the PAN has the shortest possible lifetime, that every detokenization carries a reason that lands in an audit log, and that the application database stores only the token. Notice the service has no decryption key anywhere; it cannot reverse the token even if compromised. That is the structural advantage you simply cannot replicate by encrypting a column in your own database.

Build, buy, or use the network token

Once a team accepts tokenization, the next argument is whether to build a vault, buy one, or lean on the payment networks. Building your own vault is rarely the right call. A vault is a piece of infrastructure that must never lose data, never leak, and survive every audit for the life of the company. That is a serious, ongoing engineering commitment, and most organizations underestimate the operational tail.

For the vast majority of teams, the right answer is to let the gateway or processor tokenize at the point of capture. The card data goes from the customer's browser to the processor through a hosted field or SDK, and your systems receive a token from the very first moment. You never touch the PAN at all, which is the strongest possible scope position. Network tokenization, offered by the card schemes, goes a step further by issuing tokens that the networks themselves manage and that update automatically when a card is reissued, which materially improves authorization rates on recurring payments.

I reserve a self-hosted vault for the narrow cases where regulation, data residency, or a multi-processor strategy makes provider lock-in genuinely unacceptable. Even then I treat it as the most sensitive system we run, with its own network segment, its own change process, and a deliberately boring technology choice. Excitement in a token vault is a bug.

How I actually decide

When a team brings me this decision, I work through a short sequence rather than reaching for a default. First, is the data a payment card? If yes, tokenization is the presumption and the burden of proof is on anyone who wants to keep raw or encrypted PANs in application systems. If the data is some other sensitive field, encryption with disciplined key management is usually the right and sufficient answer.

Second, do we ever need the original value in our own systems, or only a reference plus an occasional gateway call? If we only need a reference, tokenization removes an entire category of risk for free. Third, what does the scope diagram look like under each option, and can I defend it to an assessor in five minutes? The option that produces the smaller, more honest diagram usually wins, because the smaller perimeter is cheaper to secure and far cheaper to audit, year after year.

The framing I keep coming back to is that encryption protects data you have decided to keep, while tokenization lets you stop keeping it at all. Not keeping the data is almost always the stronger position. You cannot leak what you never stored.

Anselm Fowel, CTO and fintech architect
Anselm Fowel — CTO & fintech architect

Conclusion

Tokenization and encryption are not competitors so much as tools for different layers of the same problem. Encryption secures data within a trust boundary you control and is indispensable for non-card secrets and for the internals of the vault itself. Tokenization removes card data from your systems entirely, shrinking your audit scope and your breach exposure to a single hardened component. For payment cards specifically, I default to tokenization, ideally performed by the processor at capture so the raw PAN never enters my estate, with encryption and rigorous key management underneath as defense in depth. Decide deliberately, draw the scope diagram before you write the code, and remember that the safest card number is the one you never had to store.

Chat with us