Skip to content

The scrub pipeline

shipped 1.0.0

Runs after every cwp db pull unless you pass --no-scrub. It belongs to the database import, so a tree pull neither scrubs nor needs to: it brings no personal data down. Pulling a production database onto a laptop is a GDPR processing operation; scrubbing is the default for that reason, not for convenience.

  1. Neutralise outbound mail: install the mail guard. Always, even with scrubbing switched off.

    Nothing reaches a real recipient, and where a message goes instead depends on what answers on loopback. DDEV runs Mailpit in the web container. When its SMTP port answers, the guard diverts the message there, addressed to local@cwp.invalid. It removes Cc and Bcc from the headers, keeps the original recipients as X-CWP-Original-To / -Cc / -Bcc, and prefixes the subject with [cwp]. Read the whole message, rendering, headers and attachments, with:

    ddev launch -m

    When nothing local answers, the guard blocks the message: it short-circuits wp_mail() before WordPress builds a mailer. Every failure falls back to blocking. The guard also rewrites the site’s own SMTP settings at the last possible moment, so a database carrying production SMTP credentials cannot use them from here.

    Either way one line lands in cwp/tmp/mail.log, naming which of the two happened:

    [2026-08-12 16:57:15] diverted mail to kunde@example.com — Ihre Bestellung

    The log stays one line per message on purpose. Mailpit holds the whole of a diverted message already, and a second copy of message bodies on disk is personal data this pipeline exists to reduce.

    One case the guard can only report, not prevent: a plugin that declares its own wp_mail() takes over the function the guard hooks into. cwp doctor fails on it and names the plugin. The guard writes an UNGUARDED line into the same log.

  2. Anonymise users, except the site’s own staff. Users holding one of pull.keep_roles (default administrator and editor), super admins, and pull.keep_admin_email keep their real data. Everyone else gets user-<ID>@example.invalid, a blank display name and a reset password: subscribers, customers, visitors, where the bulk personal data lives.

    cwp writes straight to the database rather than through wp_update_user(), so plugins that sync users to external services never fire.

    Keeping staff accounts is deliberate: they are what you log in with locally. Set keep_roles: [] for no exceptions, but then no local account has a password you know. cwp destroys sessions for everyone, staff included.

    Whoever wrote a post cwp manages keeps their data too, whatever role they hold, and keep_roles: [] does not reach them. An item records the login that wrote it, so an anonymised author is an author the other side cannot resolve. pull.keep_logins adds people by name. The run says how many accounts each of the three reasons covered:

    12 anonymised, 5 kept (2 by role, 1 by address, 2 by authorship)

    The tree has what that means for a repository.

  3. Truncate log and submission tables: form submissions by default, configurable via pull.truncate_tables. cwp checks existence first, so it reports a name that does not apply to your site as absent rather than failing.

  4. Delete log posts, configurable via pull.truncate_post_types.

    These are not tables. Every SNN-BRX logging feature registers a custom post type and writes entries to wp_posts, so truncating tables could never reach them. The mail log stores message bodies and the chat history stores transcripts. Deletion goes through wp_delete_post(…, true), which takes the postmeta with it and forces past the trash. A trashed log is still the full record.

    cwp init fills this list with the SNN-BRX log types when it finds that child theme installed, and leaves it empty otherwise. Which post types hold logs is a fact about the plugins a site runs, not something to guess at (B-001).

    snn_301_redirects is deliberately not in the defaults. It holds the redirect rules: site configuration rather than a log. It sits one word away from the log type in the same feature file.

  5. Delete log taxonomy terms: snn_ip_address and snn_user_agent by default, configurable via pull.truncate_taxonomies.

    Deleting the posts is not enough. SNN-BRX’s search logger stores the visitor IP and user agent as taxonomy terms rather than postmeta, and wp_delete_post() removes term relationships while leaving the terms themselves in wp_terms. Without this step the scrub deleted every search log, reported success, and left a browsable list of every visitor’s IP behind (B-011).

    snn_search_count is deliberately not in the defaults: it hangs off the same posts, and a count is not personal data.

  6. Delete transients and expired sessions.

  7. Set blog_public = 0, so the local copy is never indexed.

  8. Report what the scrub removed.

Two steps are fatal

Steps 2 and 4 abort the scrub on failure. Everything else warns and carries on.

A database that reads as scrubbed and is not is worse than an obvious error. It is the precise outcome this pipeline exists to prevent, and nobody sees it until somebody goes looking.

Re-runnable on its own with cwp scrub. Run that after importing a dump that arrived any other way.