Shadow AI is already using your data. Get the complimentary Gartner® report. Read the report

PK Protect for z/OS · Data discovery, encryption, and masking for IBM Z

Find what’s sensitive in your mainframe files. Then encrypt it, mask it, or both.

PK Protect reads your z/OS data sets the way your applications do and finds the sensitive values inside them. What happens next is your call. Encrypt what has to stay exact, mask what doesn’t.

40+ yrs
On the mainframe
504M
Records scanned in a single engagement
$150M
In fines avoided

Trusted by leading organizations for over 40 years

JPMorgan ChaseTruistFiservWestern Union

Protecting 21 of the 25 largest U.S. commercial banks and 30% of the Fortune 100.

What changed

Compliance stopped being something you pass. Now it’s something you prove, again, every year.

The bar keeps moving. PCI DSS 4.0 raised it, the post-quantum deadlines have real dates attached, and the question your auditor asks has changed. It used to be whether you encrypt. Now it’s what was in the file, who could read it, and show me. Passing last year’s assessment buys you twelve months, and the standard you passed under has already been revised.

At the same time your data started going places nobody approved. Mainframe records land in a data lake, the lake feeds a model, and somewhere along that chain an employee pastes an extract into an AI tool your security team has never heard of. Shadow AI isn’t a future problem for this platform. It’s a pipeline that already exists and that nobody drew.

Both of those come back to a single question, and it’s the one the mainframe has never been able to answer.

The problem

A mainframe file is one long line of characters. Nothing in it is labeled.

Open a file on a laptop and you can usually tell what you’re looking at. Column headers, a file name, a preview. Mainframe data doesn’t work that way. A record is an unbroken run of characters, and the only thing explaining where one field ends and the next begins is a separate definition, agreed to decades ago by programs that still run every night, written by people who have since retired.

It’s a city with no map. The roads work, the traffic moves, and the people who laid them are gone. You can drive it if you already know it. Ask which street can close for repairs and nobody can tell you what else stops working.

That’s why most data security tools skip the platform, or scan the databases sitting next to it and call it coverage. A scanner with no definitions reads the record byte by byte looking for anything shaped like a Social Security number, and reports what it thinks it found. That’s optimistic. It isn’t precise.

So when an auditor asks which files contain card numbers, the honest answer at most companies is an estimate. The estimate is what ends up in the finding.

Fig. 01 What a mainframe record actually looks like Identified fieldFlagged sensitive
WHAT A SCANNER SEES MARTINEZ00J4471982110847392847562398450112 One record. No delimiters, no column headers, nothing that says which part is sensitive. WHAT PK PROTECT SEES MARTINEZ LAST-NAME Last name 00J BRANCH-CODE Branch code SENSITIVE 447198211084 ACCT-NUM Account number SENSITIVE 7392847562398450 CC-NUM Card number 112 MFT-STATUS Status The same 42 characters. The second one is the only one you can act on.
Fig. 01 — The same record, before and after PK Protect reads it.

What it does

Nobody has to document anything. The machine has been documenting itself for forty years.

Compiler listings, job control, record definitions. Your mainframe has been writing all of it down since the day it was switched on, in a format no human wants to read. PK Protect’s SchemaLink reads that material and builds the schema from your own definitions, so the scan knows where every field starts and stops before it begins.

The question stops being what might be in this file and becomes how many, and exactly where. That covers the file types the platform actually runs on: VSAM, sequential data sets, partitioned data sets (PDS), and generation data groups (GDG). Application data, flat files, reports.

Once you know what’s in there, you have three things you can do with it.

01

Find it.

A complete inventory of every sensitive value, down to the field and the data set it sits in, so the next time an auditor asks the question you answer it from a report instead of starting an archaeology project. Where a definition isn’t available, that data set is still inventoried and reported as unresolved. We would rather tell you we don’t know than hand you a confident answer we can’t stand behind.

02

Encrypt it.

IBM Pervasive Encryption protects data while it sits on the mainframe. The moment a file leaves, that protection stops, and every copy downstream lands readable: the partner’s server, the cloud bucket, the tape in the archive. PK Protect puts the encryption inside the file itself, so it travels with the data instead of ending at the door.

03

Mask it.

Sometimes the right answer isn’t guarding the real value. It’s not putting the real value in that copy at all. Masking replaces a Social Security number with a different number that looks and behaves exactly like one. Your applications keep running. Your test team keeps testing. The real number stays in production where it belongs.

Which one you need

One question tells you which protection to use.

The question that separates them is simple: does the person on the other end need the real value back?

“Yes. They need the actual data.”

Encrypt it. A partner receiving your settlement file, an archive you may have to restore from, a feed a downstream system has to process for real. The data stays exact and stays protected, and anyone holding the key can get it back.

“No. They just need data that behaves like the real thing.”

Mask it. Developers, test environments, analytics, AI training sets, and any screen that only needs the last four digits. The record keeps its structure and loses its risk, and there’s no key to lose because there’s nothing to recover.

“Both, on the same data.” That’s the normal answer. Most organizations encrypt the files that leave the platform and mask the copies that feed lower environments, from the same discovery run and under the same policy. You aren’t choosing a product. You’re choosing per data set.

Why teams mask

Three reasons to mask mainframe data. None of them are unique to the mainframe.

People expect this part to be complicated because it’s the mainframe. It isn’t. Companies mask data on z/OS for the same three reasons they mask it in a cloud warehouse or on a laptop.

Your developers need real-looking data they aren’t allowed to see.

Testing against fake data written by a script finds fake bugs. Testing against production data gets you a compliance finding. Masked data sets give your team the same record structure, the same field lengths, and the same application behavior, with none of the real values.

Your mainframe data is already feeding AI models.

It moves directly in some shops and through a data lake in most of them, and either way it left the platform before anyone classified it. Mask the card numbers and Social Security numbers on the way out. The alternative is finding out afterward, when somebody types a prompt and sees a customer record they were never supposed to see.

Most screens only need part of the number.

Your service rep needs the last four digits to verify a caller. They don’t need the other five, and there’s no upside to the full value ever reaching the application that renders that screen. Mask it at the source and the exposure never gets created.

The hard part

Masking one system is easy. Making the results match across all of them is the actual work.

Say you mask a customer’s Social Security number on the mainframe, and you mask that same customer in your cloud database. If the two systems invent different fake numbers, the two records stop matching. Reports break. Joins fail. Testing produces results nobody can trust. Your masking project just created a data problem while it was solving a privacy one.

Deterministic masking is what prevents that. The same input always produces the same output, on every platform, every time. Mask that customer on z/OS and mask them in Oracle, and both come out identical. The records still line up, and nobody spends a quarter reconciling them.

That’s the reason the mainframe can’t be a side project with its own separate masking tool.

Fig. 02 Deterministic masking across platforms One policySeparate tools
One masking policy Deterministic. The same input always produces the same output, on every platform, every time. z/OS DATA SETS 123-45-6789 481-27-3095 A separate tool 230-58-7719ORACLE 123-45-6789 481-27-3095 A separate tool 902-64-1188CLOUD STORAGE 123-45-6789 481-27-3095 A separate tool 315-70-8842ENDPOINT 123-45-6789 481-27-3095 A separate tool 677-19-4506 The same masked value in all four. Every join still works. THE FAILURE MODE · A SEPARATE TOOL PER PLATFORM ≠ ≠ ≠ Records no longer match.
Fig. 02 — The same customer, masked on four platforms, producing one consistent value.

The same input always produces the same output, whatever platform it runs on. That is what keeps a masking project from becoming a reconciliation project.

Read: Consistent Masking Needs Deterministic Masking

What covers what

Your data doesn’t stay on one platform. Neither does the policy.

Most mainframe shops run at least two of these, because the data was never going to sit still. One discovery engine and one set of classification rules sit underneath all three, so what counts as sensitive is defined once and means the same thing everywhere it lands.

What you’re protectingWhat runs it
Application data sets on z/OS: VSAM, sequential, PDS, GDGPK Protect for z/OS
DB2, other databases, data lakes, cloud storage, ERPPK Protect Data Store Manager
Laptops, servers, file shares, Microsoft 365PK Protect Endpoint Manager

Already know where your sensitive data is and just need the files protected as they leave? That’s PK Encrypt, and there’s a 30-day trial.

Start the 30-day PK Encrypt trial

Why it’s one product

Reading the code and protecting the data are sold as two markets. Your data doesn’t care.

Application analysis tools read your source and stop there. They never look inside the data. Data discovery tools scan the data and stop there. They never read the source, and they hand protection off to somebody else’s product. Buy both and you own two tools that don’t know about each other, plus the job of joining their output yourself.

PK Protect reads the application in order to understand the data, finds what’s sensitive inside it, and protects it. One platform, one policy, one vendor.

That matters most at the moment nothing else covers. Encryption applied at the data set or volume level protects data while it sits on Z, and it does that job well. Protection applied to the data itself is still in force after the file has gone to a partner, landed in a lake, or been copied into a test region.

Proof

504 million records. 422 million card numbers nobody had accounted for.

One of the largest financial institutions in the US was facing a PCI DSS 4.0 deadline and couldn’t answer a basic question: where are the card numbers. PK Protect scanned 504 million VSAM records and found 422 million card numbers and 450 million Social Security numbers. Eighty-eight percent of it was sitting exposed. Finding it and protecting it kept an estimated $150 million in penalties off the table.

Read the full story

504M
VSAM records scanned
88%
Of the sensitive data found was sitting exposed
$150M
In estimated penalties kept off the table
“

PKWARE is the backbone of our data protection strategy. With agents across all endpoints, we seamlessly encrypt and decrypt files while managing petabytes of data. For over 5 years, PKWARE has helped us meet audit requirements with confidence and trust.

Fiserv
Fiserv encrypts more than a million files a day across 50,000 desktops and 1,200 servers, under the same platform that reads your data sets.

Straight answers

Mainframe discovery, encryption, and masking, answered.

Mainframe data encryption protects data on and leaving IBM Z by encrypting the data itself rather than the channel it moves through. The protection stays attached to the file after it reaches a distributed server, a partner system, or cloud storage, and anyone holding the key can recover the original values. It’s the right choice when the data has to stay exact for whoever receives it.

Mainframe data masking replaces sensitive values inside z/OS data sets with realistic substitutes, so the data keeps its structure and stops carrying risk. A masked card number is still sixteen digits and still passes the checks your applications run, but it isn’t anyone’s actual card number. There’s no key, because there’s nothing to recover.

Encrypt when the recipient needs the real values back, and mask when they don’t. Partner feeds, archives, and downstream systems that process live transactions need encryption, because the data has to stay exact. Development, test, analytics, AI training sets, and most screens need masking, because those users need data that behaves correctly and never needed the real numbers. Most organizations do both, from the same discovery run, choosing per data set rather than per product.

You prove it with an inventory and a record of what was done to it. An auditor asking about PCI DSS 4.0 or a privacy regulation wants to know which data sets hold sensitive values, what protection was applied to each one, and when. PK Protect produces the inventory through discovery and the evidence through the protection it applies, so the answer is a report rather than an explanation from the person who has been there longest.

Yes. They serve different copies of the same data, not competing approaches to one copy. A common pattern is encrypting the production file that leaves the platform on a nightly job while masking the copy that populates the test region, both driven by one discovery scan and one policy.

Yes. PK Protect discovers, encrypts, and masks sensitive data in VSAM data sets, along with sequential data sets, partitioned data sets (PDS), and generation data groups (GDG). VSAM is the one people ask about because it holds the transactional records, but the coverage isn’t limited to it.

Yes. PK Protect uses deterministic masking, meaning the same input always produces the same output. A customer masked on z/OS and masked again in Oracle, Snowflake, or a file share comes out identical in each one, so the records still join and your reporting still works.

Yes, and that’s the mechanism rather than a limitation. SchemaLink generates the schema from the source your programs are actually compiled against, which is what makes the results exact instead of probabilistic. A definition that no longer matches what a program writes shows up as non-conforming data instead of passing silently. Data sets with no available definition are still inventoried, and reported as unresolved rather than guessed at.

Those tools read the record byte by byte and flag anything that matches a pattern. The claim is accurate and it’s also the point: without your definitions, nothing can tell the scanner that the nine digits sitting at byte 51 are a Social Security number rather than a policy number that happens to be nine digits long. You get a number you can’t act on and a false-positive count that scales with the size of your data.

Pervasive Encryption protects data while it’s on the platform, and it does that well. It stops at the edge. When a batch job writes a file and something moves it, the file lands readable on the other side, then gets copied into test environments, vendor systems, and cloud storage by people who weren’t in the original design review. PK Protect finds what’s sensitive first, then protects the data itself so it stays protected everywhere it goes next.

No. Masked values keep the original format, length, and data type, so an application reading a masked record can’t tell the difference. A sixteen-digit card number stays sixteen digits and a date stays a valid date. Encrypted files are protected at the file level and decrypt on the other side with the key, so the receiving application reads what it always read.

DB2 is covered by PK Protect Data Store Manager, which handles structured sources across databases, data lakes, cloud repositories, and packaged applications. Most mainframe organizations run both alongside each other, under one policy.

Discovery produces findings on the first scan. What takes time is deciding what to do with them, since that’s a conversation between security, compliance, and the application teams who own the data sets. The technical work is rarely the long pole.

Start here

Find out what’s actually in your data sets.

Your next assessment is already scheduled and your data is already moving into systems nobody approved. Start with the inventory. Everything after that is a decision you make per data set, and you’ll have the evidence either way.