How to Handle Personal Data (PII) in Scraped Datasets

Published:

13 minute read

Acar Diveroli
Written by: Acar Diveroli
A dark page of redacted text lines with one heavy data row in the middle, its email and phone fields masked by a blue band

You scrape 20,000 product reviews to see which complaints come up most. The plan was ratings and text. When you open the CSV, it also holds reviewer names, links to their profile pages, and in a few hundred reviews an email address or a phone number that the author typed into the text. None of it helps the analysis, but it now sits on your laptop and in last night's backup.

This post covers what counts as personal data (PII, personally identifiable information: anything that points to a specific person) in a scraped dataset, why "it was public" does not settle the question, and how to handle it: collect less, mask what slips through, pseudonymize the keys you need and delete on a schedule. It also shows why hashing an email is weaker than it looks, gives a tested Python masking script, and points to the GDPR, Türkiye's KVKK and California's CCPA.

What counts as personal data in a scraped dataset?

The GDPR's definition in Article 4(1) is broad: any information relating to an identified or identifiable natural person, directly or indirectly, with examples such as a name, an identification number, location data and an online identifier. Türkiye's Law No. 6698 on the Protection of Personal Data (KVKK) uses the same core wording in its Article 3.

In a scraping project, personal data shows up in three places:

  • Fields you selected on purpose. Author names, usernames, profile URLs, avatars, job titles on a company page, seller names on a marketplace.
  • Free text. Reviews, comments, forum posts and listing descriptions, where people write their email, phone number, street or car plate.
  • Your own logs. Request logs and error dumps often keep the full HTML of a page, including everything above.

Harder to spot: fields that identify nobody alone. A username, a city, a date and a rare product can together point to one person, so a dataset without a "name" column can still hold personal data.

Does "public" mean free to use?

Not in general, and the answer differs by law.

  • GDPR. There is no exemption for publicly available personal data. Scraped data must still have a lawful basis and follow the principles in Article 5. Article 14 sets out what you owe people whose data you obtained from somewhere other than them.
  • KVKK. Article 5 lists the conditions for processing without explicit consent. One of them is that the person has made the data public themselves. It is not a blank cheque: the principles in Article 4, such as specified purposes and proportionality, still apply.
  • CCPA. The California definition of personal information in Civil Code 1798.140 lists IP and email addresses as identifiers, then excludes "publicly available" information, defined narrowly: for example government records, or information the consumer made available to the general public.

The wider legal question, including terms of service and copyright, is covered in Is Data and Web Scraping Legal?.

How to handle personal data in a scraping pipeline

GDPR Article 25 asks for "data protection by design and by default": safeguards belong in the design, not a cleanup job at the end. For a scraper:

  1. Write down the purpose. "Find the most common delivery complaints per product" is a purpose. "Collect reviews in case they are useful" is not.
  2. List the fields that serve it. For the example above: product, date, rating, text. Author name and profile URL are not on the list.
  3. Filter at collection. Do not select what you will not use.
  4. Mask free text on ingest. Replace emails and phone numbers with placeholders before the row is written anywhere.
  5. Pseudonymize the keys you need. If you must count repeat reviewers, keep a keyed hash of the author ID instead of the ID.
  6. Keep the key apart. The pseudonymization key lives in a secrets manager or an environment variable, never in the dataset or the same repository.
  7. Secure the storage. Encrypt at rest, restrict access to the people who work on the project, and keep raw dumps out of shared drives.
  8. Delete on a schedule. Raw HTML and unmasked files get a short retention period; the cleaned dataset gets its own, tied to the purpose.
  9. Record what you did. Purpose, fields, retention and legal basis, in one short note.

Masking, pseudonymization, anonymization: the difference

These terms are not interchangeable, and the difference decides whether the law still applies to your data.

TechniqueWhat it doesCan the person be re-identified?Still personal data?
Deletion at collectionThe field is never storedNo, the data does not existNo, for that field
MaskingReplaces a value with a placeholder like [EMAIL]Not from the masked value, but maybe from other fieldsDepends on what is left in the row
PseudonymizationReplaces an identifier with a token; the link sits in separate, protected informationYes, with the key or lookup tableYes (GDPR Recital 26)
AnonymizationRemoves or generalizes data until no one can be identified by reasonably likely meansNoNo
EncryptionMakes data unreadable without a keyYes, for anyone with the keyYes
AggregationKeeps only counts, averages or groupsOnly if groups are very smallUsually not, if groups are large enough

What the GDPR says about pseudonymization

Article 4(5) defines pseudonymization as processing personal data so that it "can no longer be attributed to a specific data subject without the use of additional information", provided that information is kept separately and protected. Recital 26 then draws the line: pseudonymized data that could be attributed to a person with additional information "should be considered to be information on an identifiable natural person". Anonymous information falls outside the Regulation.

To decide whether someone is identifiable, Recital 26 asks you to consider "all the means reasonably likely to be used", including cost, time and available technology. This is why anonymizing scraped text is hard: removing the name rarely helps when the review still mentions the town, the date and the car model.

The European Data Protection Board adopted Guidelines 01/2025 on pseudonymisation in January 2025. The European Commission's "Digital Omnibus" proposal of November 2025 would change how the definition of personal data applies to pseudonymized data; according to the European Parliament's legislative train page, it had not been adopted when this post was written. Until a change is in force, treat pseudonymized data as personal data.

KVKK does not define pseudonymization. Its Article 3 defines anonymization as making data impossible to link to a person "even through matching them with other data", and Article 7 requires erasure, destruction or anonymization once the reason for processing is gone.

Why hashing an email is not anonymization

A common shortcut is to run sha256(email) and call the column anonymous. It is not, for two reasons.

First, the hash is the same every time. Anyone with a list of email addresses, such as one leaked in a breach, can hash every address on it and look for matches; in our test a two-item guess list found the "anonymous" address at once. The same leaked lists feed credential stuffing, one more reason not to keep raw emails you do not need.

Second, some identifiers have a small search space. A national phone number has a limited number of digits, so hashing every possible number in that range is feasible on ordinary hardware, and the matching token gives the number back.

What works better:

  • A keyed hash (HMAC) with a secret key stored away from the data. The result is pseudonymized, not anonymous, data.
  • A random token and a lookup table stored separately, if you need to reverse the mapping.
  • No identifier at all, if you only need counts. Aggregate first and drop the key.

Python example: masking emails and phone numbers

The script reads a scraped reviews CSV and writes a cleaned copy: it drops unneeded columns, masks emails and phone numbers in the free-text column, and replaces the author ID with a keyed hash so repeat reviewers can still be counted. Standard library only, tested with Python 3.13.

python
import csv
import hashlib
import hmac
import os
import re

# Columns we never need for the analysis: drop them on ingest
DROP_COLUMNS = {"author_name", "profile_url"}
# Column we keep only as a pseudonym, to count repeat reviewers
PSEUDONYM_COLUMN = "author_id"
# Free-text columns where people type contact details; structured columns stay untouched
TEXT_COLUMNS = {"text"}

EMAIL_RE = re.compile(r"[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}")
# Phone-like runs: optional + or (, then 9-15 digits with spaces, dots, dashes or brackets between them
PHONE_RE = re.compile(r"(?<!\w)[+(]?\d(?:[\s.()-]*\d){8,14}(?!\w)")

# The key lives outside the dataset (env var, secrets manager), never in the same file
SECRET = os.environ.get("PSEUDONYM_KEY", "").encode()


def mask_text(text: str) -> str:
    text = EMAIL_RE.sub("[EMAIL]", text)
    return PHONE_RE.sub("[PHONE]", text)


def pseudonym(value: str) -> str:
    # Keyed hash (HMAC-SHA256): without the key, nobody can rebuild the mapping by hashing guesses
    return hmac.new(SECRET, value.encode(), hashlib.sha256).hexdigest()[:16]


def clean_row(row: dict) -> dict:
    out = {}
    for col, value in row.items():
        if col in DROP_COLUMNS:
            continue
        if col == PSEUDONYM_COLUMN:
            out[col] = pseudonym(value)
        elif col in TEXT_COLUMNS:
            out[col] = mask_text(value)
        else:
            out[col] = value
    return out


def main(src: str, dst: str) -> None:
    if not SECRET:
        raise SystemExit("Set PSEUDONYM_KEY first")
    with open(src, newline="", encoding="utf-8") as f_in, \
         open(dst, "w", newline="", encoding="utf-8") as f_out:
        reader = csv.DictReader(f_in)
        fields = [c for c in reader.fieldnames if c not in DROP_COLUMNS]
        writer = csv.DictWriter(f_out, fieldnames=fields)
        writer.writeheader()
        for row in reader:
            writer.writerow(clean_row(row))


if __name__ == "__main__":
    main("reviews_raw.csv", "reviews_clean.csv")

Run it with the key set in the environment (export PSEUDONYM_KEY=... on Linux and macOS, $env:PSEUDONYM_KEY="..." in PowerShell), then python mask_pii.py. On our made-up sample, "Call me on +90 532 000 00 00 or (0212) 000-0000" became "Call me on [PHONE] or [PHONE]", emails became [EMAIL], and one author's two reviews got the same token.

Know the limits:

  • Regex over-matches. In our tests the phone pattern also masked a 13-digit ISBN, a 10-digit order number and a date written together with a time. That is why the script only touches free-text columns.
  • Regex under-matches. Numbers spelled out in words, "name at domain dot com", IBANs, addresses and names in the text are not caught; those need a named-entity model or a manual review of a sample.
  • Masking is not anonymization. The row keeps text, date and product; check whether those alone can point to a person.

Run this step where rows are first written, not as a later job. The pandas cleaning guide shows where it fits in a full cleaning pass, and saving scraped data to CSV, JSON or SQLite covers storage.

Storage security, access and retention

GDPR Article 32 asks for security "appropriate to the risk", naming pseudonymization and encryption, resilient systems, restoring data after an incident and regular testing. KVKK Article 12 sets a similar duty. For a scraper that means:

  • Encryption at rest for disks, buckets and databases holding raw scrapes.
  • Separate zones. Raw HTML and unmasked rows in one restricted place; the cleaned dataset where analysts work.
  • Least access. Only the people and service accounts that need the raw zone can read it.
  • Retention as code. A scheduled deletion job beats a policy document.
  • Backups count. A file that lives on in a year-old backup has not been deleted. Align the two retention periods.

Our data security page covers proxies in security work; a proxy changes the IP a request comes from, not what you store.

Where this comes up

  • Price and stock monitoring. Product pages rarely hold personal data, but seller names on marketplaces can. See competitor price tracking.
  • Review and sentiment analysis. Author fields and contact details in text. See market research.
  • Research and data mining. Forum posts are dense with personal data; aggregate early. See what data mining is.
  • Large crawls. A general web crawler collects whole pages, so everything on them lands in your storage unless you filter.
  • Lead lists. Collecting individuals' contact details for outreach is the highest-risk case here; check the law and site terms first.

Common mistakes

  • Scraping the whole page "to be safe" and leaving the raw dump in storage indefinitely.
  • Calling a SHA-256 of an email "anonymized".
  • Storing the pseudonymization key in the same repository or bucket as the data.
  • Masking the analysis table but not the logs, the error dumps or the debug HTML.
  • Ignoring robots.txt and site terms because the data "is public anyway". robots.txt says what the owner allows crawlers to fetch.
  • Keeping data after the project ends because nobody owns the deletion.
  • Assuming the rules of your own country apply, when the people in the dataset live elsewhere.

Decision guide

NeedRecommendation
Analysis of prices, stock or product dataDo not collect seller or user fields at all
Text analysis of reviews or commentsDrop author fields, mask contact details in text on ingest
Count repeat authors or track one account over timeKeyed hash (HMAC) of the ID, key stored separately
Share a dataset with a client or publish itAggregate, then check small groups and rare combinations
Keep raw HTML for debuggingShort retention, restricted zone, encrypted
Contact details of individuals are the purposeStop and get legal advice before collecting

Frequently asked questions

Is scraped public data exempt from the GDPR?

No. The GDPR has no general exemption for publicly available personal data. You still need a lawful basis, Article 5 still applies, and Article 14 covers the information owed to people whose data you did not collect from them.

Is an IP address personal data?

It can be. The GDPR lists online identifiers in Article 4(1), and the CCPA names "Internet Protocol address" among its examples of identifiers. In scraping it matters mostly for your own logs.

Does hashing make data anonymous?

No. A plain hash of an email or phone number can be reversed by hashing a list of guesses. A keyed hash with a secret key stored elsewhere is pseudonymization, which the GDPR still treats as personal data.

What does KVKK say about data a person made public?

Article 5 allows processing without explicit consent when the person made the data public themselves. The principles in Article 4, such as specified purposes and proportionality, still apply.

How long can I keep scraped personal data?

Only as long as the purpose needs it. GDPR Article 5 calls this storage limitation, and KVKK Article 7 requires erasure, destruction or anonymization once the reason for processing is gone.

Does using a proxy change my obligations?

No. A proxy routes requests through a different IP address. It does not change what you collect, where you store it or which laws apply.

Summary

Personal data in a scraped dataset rarely arrives because someone wanted it; it comes along with the page. The fix is mostly engineering: decide the purpose, collect only the fields that serve it, mask what slips into free text, pseudonymize keys with a secret stored elsewhere, encrypt the storage and delete on a schedule. Pseudonymized data is still personal data under the GDPR; anonymous means nobody can be identified by reasonably likely means. For the collection side, see web data scraping: Residential Proxy addresses give you country and city targeting, Rotating Proxy spreads requests, and our proxy page lists every product.

Ask ChatGPTAsk Claude