Word's "Remove Personal Information" Doesn't.
Microsoft Told You. Nine Federal Documents Prove It.
Microsoft has, on an Office Resource Kit blog post dated June 30, 2009 and unchanged in substance since, listed many of the types of metadata that Word's Document Inspector does not remove. We took Microsoft at its word, then downloaded nine real DOCX files from six federal agencies. The bytes match the disclaimer.
Why this matters in 2026
The AI Hallucination Cases database maintained by French legal academic Damien Charlotin, one of the most-cited public trackers of the phenomenon, had logged 1,397 court filings worldwide as of May 5, 2026, where an attorney cited a case the AI invented. United States courts alone issued more than $145,000 in sanctions in the first quarter of 2026.
Two of those orders matter for what is about to follow. In Couvrette v. Wisnovsky, No. 1:21-cv-00157-CL, U.S. Magistrate Judge Mark D. Clarke of the District of Oregon sanctioned attorney Stephen Brigandi a total of $110,204.38 across two orders issued in December 2025 and March 2026, after Brigandi included fifteen non-existent cases and eight fabricated quotations across three summary-judgment briefs. When defendants flagged the fakes, Brigandi tried to fix his filings by silently deleting the bogus citations and refiling without disclosing the change.
Judge Clarke's order said it directly. "Rather than a correction, Mr. Brigandi attempted a cover-up. He failed at both."
In Bunce v. Visual Technology Innovations, No. 2:23-cv-01740 in the Eastern District of Pennsylvania, U.S. District Judge Kai N. Scott imposed a $5,000 sanction on attorney Raja Rajan on April 20, 2026, for a second misuse of AI-generated case citations. Rajan had asked one chatbot to verify another chatbot's work. Judge Scott wrote, "It will be your name and license on the line, not ChatGPT's."
Brigandi was caught the easy way: opposing counsel read both versions of the filed PDF and noticed the citations had changed between drafts. Most cover-ups will not be that visible. When the substantive change is subtle, when the surface text reads identically, when the only evidence is a comparison of who edited what and when, that evidence lives in the source document's bytes. Lawyers draft in Word, then export to PDF for filing. The Word file is where the editing trail accumulates. Word offers a feature called "Document Inspector" that promises to remove that trail before sharing. This post is about what it actually leaves behind.
What Document Inspector says it does
Open a .docx file in Microsoft 365's Word application. Select File, then Info, then Check for Issues, then Inspect Document. A dialog appears with six checkboxes. Microsoft's official support article names them, verbatim:
That is the entire surface area of the feature for Word. Six modules. The user clicks Inspect, then clicks Remove All on each category that returns hits, and Word reports an "All clear" green check.
The feature is presented as the privacy step before sharing a document outside an organization. Court-filing guides at the American Bar Association, the William & Mary Law School curriculum, and the North Carolina Central University law libguides all recommend it. Microsoft's KB describes it as the way to "remove hidden data and personal information by inspecting documents."
What the dialog does not show is what its own author once disclosed.
What Microsoft itself says it does NOT remove
The disclosure exists. It is not on the support article most users read. It is on a blog post titled "What Document Inspector doesn't catch" that lives in the archived Office Resource Kit section of learn.microsoft.com. The post is dated June 30, 2009. Microsoft Learn's metadata shows the page last touched in 2024 with no content change. Seventeen years on a Microsoft URL.
Direct quotes follow. All of these are Microsoft's own description of what their feature leaves behind, not ours.
"if you put white text on a white background … Document Inspector assumes you meant to make the text white on a white background, so it doesn't consider it hidden and it won't remove it."
"a shape that is covered by another shape isn't considered invisible, and a shape that has no fill and no outline is not considered invisible … Document Inspector does not remove the shapes."
"Document Inspector does not remove cached data from files." Specifically named: pivot table caches, sort and filter caches, "any data that's associated with an embedded object."
"Document Inspector does not remove this information from the file." Concerning database connection strings.
"Document Inspector can remove printer name and printer path information from a file, but it can't remove all of the printer-specific information from a file."
"Document Inspector doesn't remove any metadata that's in a protected or restricted file." Concerning signed or IRM-protected documents.
"Document Inspector doesn't remove any code or comments from Visual Basic for Application (VBA) modules, and Document Inspector doesn't remove any data that's associated with an ActiveX control."
"Document Inspector doesn't remove the email address because it assumes that you want someone to send the document back." Concerning send-for-review email addresses.
"Document Inspector doesn't remove email addresses that are added to the content of a document."
"Document Inspector doesn't remove hyperlinks, unless the hyperlinks are contained in some type of metadata that Document Inspector does remove."
"In general, Document Inspector does not remove any of these things from a file." Concerning file names, file paths, template names, template paths.
"Document Inspector removes the contents of field codes, but it doesn't remove the field code itself."
"Document Inspector assumes these things are part of the content … it doesn't remove them and it doesn't remove the labels or text that you add to them." Concerning SmartArt, WordArt, shapes, and quick parts.
That is Microsoft's own list. It runs to fourteen named categories of residue that the privacy feature does not touch.
The list is also incomplete. Microsoft's documentation never names the OOXML element family w:rsid, never names <w:docVars> in settings.xml, never names the contents of word/fontTable.xml. We will get to those.
The most recent public Microsoft blog announcement of a Word Document Inspector module change is from December 9, 2014, by Office program manager Steve Kraynak. It added two modules: "Embedded Documents" and "Macros, forms or ActiveX Controls." No subsequent "What's new in Word" page from 2016, 2019, 2021, or 2024 has announced a Word module change publicly. The internal WdRemoveDocInfoType enum has grown quietly since (a wdRDITaskpaneWebExtensions value at position 17 was added at some point without a heralding post), but the user-visible feature has been frozen in Microsoft's documentation for nearly twelve years while OOXML itself has continued to grow. Microsoft Information Protection sensitivity labels, AI-authored content markers in Microsoft 365, and the per-document save flags that Word now writes into settings.xml all postdate Inspector's last announced change.
What follows is what happens when you take Microsoft's own gap list, apply it to documents you can download from .gov sites today, and read what survives.
The federal corpus
The dominant pattern across the United States federal government is upload-as-authored. Document Inspector is rarely if ever invoked at the agency-publication boundary. Native DOCX files travel from a staffer's desktop, through internal review, into a Drupal CMS, and out to the public web with no scrub.
Below are nine DOCX files, all freely downloadable from six federal agencies. All nine were verified as valid OOXML zips and their SHA-256 hashes recorded. We ran them through File X-Ray's parser and through cross-references against ExifTool. The findings are reproducible.
| Field | Files retaining |
|---|---|
| Author (dc:creator) | 8 of 9 |
| Last Modified By | 9 of 9 |
| Company name | 5 of 9 |
| Original RSID Root (w:rsidRoot) | 9 of 9 |
| Editing time greater than zero | 8 of 9 |
| Revision count greater than one | 9 of 9 |
| Last printed timestamp | 5 of 9 |
| Embedded media (images) | 5 of 9 |
| Document Inspector run before publication | 0 of 9 |
Document Inspector zero. The feature exists. It is documented. It is recommended. Across this corpus, no agency ran it.
HUD: HOTMA Net Family Assets Training Script
The file at hud.gov/sites/dfiles/Housing/documents/HOTMA_Net_Family_Assets_Script-4_26_24_Final.docx is a 384 KB training script for the Housing Opportunity Through Modernization Act, distributed by HUD's Office of Multifamily Housing. The filename ends in _Final, the canonical signal of a document uploaded straight from a desktop.
The footer of every page reads, in plain text: 7475 Wisconsin Avenue, Suite 1000 | Bethesda, MD | www.EconometricaInc.com.
Econometrica, Inc. is not HUD. It is a private contractor in Bethesda, Maryland. The HUD-published file outs the contractor's name, the contractor's office address, the contractor's website, the contractor employee who drafted the file (Kurt von Tish), the HUD employee who reviewed it (Robert L Norman), and a chain of more than three thousand editing sessions. The file is presented as HUD training material. Its metadata says it is contractor work product, last touched by HUD.
There is also a chart embedded in the file. It is a JPEG at word/media/image4.jpeg. The XMP block inside that JPEG records its creation tool: Adobe Illustrator CC 2015.3 (Macintosh). The XMP metadata date stamp on the chart is 2017-03-22T16:15:41-04:00. A 2017 chart, made on a Mac in Illustrator CC 2015.3, embedded in a 2024 training script, published in 2025. The chart predates the script by seven years. Document Inspector, as Microsoft's archived blog admits, does not remove embedded image metadata.
NIH: Data Management & Sharing Plan template
The file at grants.nih.gov/sites/default/files/DMS-Plan-blank-format-page.docx is the official NIH grant submission template for data management and sharing plans. Every researcher applying for an NIH grant downloads it.
A grant submission template, in active use through 2026, was last printed on a NIH desktop on March 11, 2011 at 7:43 PM. Twelve years and two months separate the template's last print job from its current upload. The print timestamp survives every save in between. Every researcher who downloads the file, edits it, and submits it to NIH is forwarding a chain of custody that includes that single 2011 timestamp, embedded in cp:lastPrinted in every grant they submit.
The lastModifiedBy field reads Sharma, Priyanka (NIH/OD) [C]. Per a March 2004 memo from NIH's Deputy Chief Information Officer, contractor accounts on NIH's Global Address List are required to carry a [C] suffix in the form Last, First (NIH/Institute) [C], so that contractor identity is visible in every email and directory entry. The (NIH/OD) segment tags the NIH Office of the Director. The DMS Plan template was last saved by an account provisioned as a contractor in the Office of the Director, not a federal NIH employee.
NIDCR: Clinical Studies Oversight Committee Report Template
The file at nidcr.nih.gov/sites/default/files/2025-02/csoc-report-template.docx is a 278 KB template for Clinical Studies Oversight Committee monitoring reports.
Two embedded JPEG images. Header text reads <Insert Abbreviated Study Title> CSOC Report Meeting Date:< Insert Meeting Date>, placeholder syntax visible to any researcher who downloads the template. Footer reads Template Version 4.0-20140421, dating the template revision to April 21, 2014. The file is a 2025 publication carrying lineage from a template branch that was already eleven years old.
DOJ: COPS Office Budget Justification
The file at justice.gov/d9/jmd/legacy/2014/07/01/cops-justification.docx is a Department of Justice budget submission for the Community Oriented Policing Services office, originally drafted in 2014 and re-served on the modern Drupal 9 platform in 2021.
The author field is the literal string debjones. Lower-case, no space, no comma. That is not a configured Word display name. That is a Windows account login. Either the original drafter never set their Office user name, or the file was created with a generic shared account. Either way, a username from someone's Windows session is sitting in the published DOJ budget document.
CIGIE: 2024 Guide for the Payment Integrity Information Act
The file at ignet.gov/sites/default/files/files/10-22-2024 CIGIE Guide for PIIA.docx is a 437 KB cross-IG working-group document distributed by the Council of the Inspectors General on Integrity and Efficiency.
Two body hyperlinks survive in word/_rels/document.xml.rels: mailto:Judith.Oliveira@ssa.gov and mailto:IPcompliancereports@gao.gov. The first is the work email at SSA-OIG of the file's last modifier. The second is the GAO Inspector-General compliance mailbox. The document outs not only the cross-IG editing chain but the institutional email path between SSA-OIG and GAO that the working group uses for coordination.
Quick callouts on the other four
EPA, Environmental Information Document Template. Author "EPA" (institutional, not personal). Last Modified By: Zolandz, Julia. RSID registry: 139 sessions. Body rsids: 89. The cleanest file in the corpus, and it still surfaces a named EPA staffer.
CMS, CY 2026 NY Dual Special Needs Plan Model Material. Author "CMS/MMCO." Last Modified By "Julie Jones." Last printed 2023-04-24. RSID registry: 5,321 sessions, the largest in the corpus. A Medicare model contract template with 5,321 distinct editing-session signatures stored in its bytes.
HHS ACF, PREP Webinar Transcript. No author set. Header reads Mathematica Policy Research, Inc., 20250514 PREP Webinar, Edited. Mathematica is HHS's contractor for PREP evaluation. The file is hosted on prepeval.acf.hhs.gov. The header outs the contractor that produced the transcript.
ED.gov, National Certificate of Eligibility Instructions Template 2025-2026. Author "Sarah Martinez." Last Modified By "Meyertholen, Patricia." Used by every Title I-C migrant student certification annually. The header retains an OMB control number that survives every annual revision.
Cross-corpus genealogy
The nine corpus files contain 13,803 total editing-session identifiers (w:rsid values) across their settings.xml registries. Ten of those rsids appear in two different files.
Random collisions happen, but the rate has to be calculated against Word's actual rsid generator, not against the nominal 32-bit value space. Every rsid we observed in this corpus, and every rsid in our control torture file, begins with the byte 00. The high-byte constraint narrows the effective entropy to about 24 bits. Under that conservative model, the expected number of cross-document rsid collisions across nine files of these sizes is approximately four. We observed ten. The probability of seeing ten or more collisions when chance predicts about four is roughly 1.3 percent.
The largest single cross-document cluster is between the CMS Dual Special Needs Plan model material (5,321 rsids in registry) and the DOJ COPS budget justification (2,042 rsids). They share three: 009C7A02, 00B40823, and 00D242E2. Two unrelated agencies, two unrelated documents, three matching session identifiers. For two specific files of those sizes drawn at random, the probability of three matches in 24-bit-effective space is approximately 2.8 percent. The CIGIE PIIA guide and the DOJ COPS justification share two rsids. The CMS Dual SNP material and the HUD HOTMA training script share two rsids. The HUD HOTMA script and the NIDCR CSOC report share one. Each pair, individually, would not survive a strict false-discovery correction. Taken together, the corpus-wide observed rate is significantly above what chance predicts.
The mechanism is documented in the academic literature. Spennemann and Spennemann's 2023 paper "Establishing Genealogies of Born Digital Content: The Suitability of Revision Identifier (RSID) Numbers in MS Word for Forensic Enquiry" (Publications 11(3):35) examined more than four hundred journal-template DOCX files from a single publisher and demonstrated that rsids survive copy-and-paste between files, persist across forks, and can seriate document genealogy. Jeong and Lee's 2017 Digital Investigation paper on Word file revision history reached the same conclusion from a different angle, focusing on the auxiliary .tmp and .asd files Word generates during editing. Spennemann's 2024 follow-up in Forensic Science International: Digital Investigation applied the same rsid forensic technique to detect academic plagiarism and ghostwriting.
What works on a four-hundred-paper academic corpus also works on a nine-document federal corpus. Given two DOCX files with shared rsids, you can attest with statistical confidence that at some point text passed between them, or both inherited from a common ancestor file. None of the federal documents in this corpus had any reason to share rsids by design. They share rsids because someone copied text from one to another, or both sourced the same template branch.
This is the genealogy layer Document Inspector does not address.
What we put in. What survived.
To complete the picture we built a synthetic DOCX with every leak vector we could think of, ran it through Word's full Save and Document Inspector "Remove All" pipeline, and compared the bytes before and after.
Methodology caveat first. Microsoft Word, on opening a non-canonical OOXML file, silently repairs parts that do not match its expected structure. Several of our injected sections (custom XML payload, comments, people list) were repaired into nothing on the first save, before Document Inspector itself ran. We cannot make precise claims about what Inspector specifically removed. We can make precise claims about what survived all of Word's pipeline including Inspector's most aggressive setting.
A field that exists in both the pre-Inspector file and the post-Inspector file definitionally survived Inspector. That is the claim category we have evidence for.
After running Inspector with every category checked, the field count surfaced by the parser dropped from 60 to 14. Inspector's surface-level work was real. The marketing-grade summary would call this a 76 percent scrub. The 14 surviving fields are the relevant answer:
| Surface | Value | Survived? |
|---|---|---|
| w:rsidRoot | 00A1B2C3 | Yes |
| Editing-session registry | 13 sessions | Yes |
| Per-run rsids in document.xml | 5 distinct sessions | Yes |
| Document Variable: DraftingAttorney | Stephen Brigandi | Yes |
| Document Variable: MatterNumber | WIS-2024-0119 | Yes |
| Document Variable: ReviewingPartner | Tim Murphy | Yes |
| Font reference | Wisnovsky-Display | Yes |
| w14:docId | 3FE64D49 | Yes |
| Application string | Microsoft Office Word | Yes |
| Application version | 16.0000 | Yes |
Inspector cleaned dc:creator, cp:lastModifiedBy, cp:revision, the Company field, the attached-template path (rewrote to "Normal"), cp:lastPrinted, every custom property, the comment authors, the tracked-change wrappers, the hidden text run, the header content, the footer content, and the body hyperlinks.
It did not touch the editing-session genealogy. It did not touch the document variables, which sit in word/settings.xml rather than docProps/core.xml, and which Microsoft's documentation never lists as a removable category. It did not touch the font reference, even though "Wisnovsky-Display" is plainly a custom organizational font name. It did not touch the document UUID Word wrote when the file was first saved.
The three lawyer names from the document variables are still in the bytes. Anyone with a hex editor can read them. The original document RSID is still there. Anyone with another DOCX produced by the same drafting machine can match it.
There is one more behavior worth noting. After Inspector ran, dcterms:created and dcterms:modified in core.xml were rewritten to the moment of the scrub. The original creation date is gone. So is the original modification date. They are replaced with a single timestamp marking when Inspector was last run on the file. That is a minor privacy gain (you cannot tell when the document was originally drafted) and a minor privacy signal (you can tell, with a one-minute margin, when its owner last tried to scrub it).
The tools that don't see what we see
The federal corpus residue and the controlled-experiment survivors are visible to anyone with a parser that walks the OOXML tree. Most parsers do not.
| Field | Inspector | ExifTool | Tika | mat2 | LibreOffice | File X-Ray |
|---|---|---|---|---|---|---|
| dc:creator | removes | reads | reads | strips | strips | reads + strips |
| cp:lastModifiedBy | removes | reads | reads | strips | strips | reads + strips |
| Company / Manager | removes | reads | reads | strips | strips | reads + strips |
| Application + AppVersion | not listed | reads | reads | strips | preserves own | reads + strips |
| Custom properties | removes | reads | reads | removes file | preserves | reads + strips |
| customXml/ parts | removes | no | no | removes dir | preserves | reads |
| w:rsidRoot / w:rsid | NOT REMOVED | no | no | strips | NOT removed | reads + strips |
| Comments | removes | no | text only | strips ranges | strips author | reads + strips |
| Tracked-change wrappers | removes | no | text only | strips | strips author | reads + strips |
| Embedded image EXIF | NOT addressed | -ee flag only | recursive | not per-image | not stripped | recurses + strips |
| Document variables (w:docVars) | NOT REMOVED | no | no | not stripped | not stripped | reads + strips |
| Hidden text (w:vanish) | removes | no | text only | not stripped | not stripped | reads |
| vbaProject.bin | warns only | flags | flags + strings | strips file | disables macros | detects + strips |
| Sensitivity labels (MIP) | removes via customXml | no | no | removes dir | preserves | reads + strips |
| w14:docId | not removed | no | no | not stripped | not stripped | reads + strips |
mat2 (Metadata Anonymisation Toolkit 2, by Julien Voisin, the privacy tool that ships in journalism kits like Tails) is, to our knowledge, the only widely-used open-source general-purpose metadata scrubber that strips DOCX rsids by default. The commit titled "Implement rsid stripping for office files" landed in the mat2 repository at GitHub jvoisin/mat2 and removes w:rsidR, w:rsidRDefault, w:rsidP, w:rsidRPr, and the <w:rsids> listing in settings.xml. No major commercial DOCX cleaner we surveyed does this by default.
LibreOffice's "Remove personal information on saving" feature is worse than not stripping rsids. It generates new ones on save. Document Foundation bug 68183 ("FILESAVE: Disabling the creation of automatic RSID marks in automatic-styles is impossible") was filed in 2013 and remained the standing position for years; a subsequent commit added an opt-in configuration option, which a typical user is unlikely to enable. Anyone running a DOCX through stock LibreOffice with the privacy option enabled comes out the other side with a freshly-rsidded file in LibreOffice's distinct rsid pattern. The pattern itself is a fingerprint, separate from Word's.
ExifTool, Apache Tika, python-docx, Aspose.Words CleanupOptions, oletools: none surface rsids. The big-name forensic suites (Cellebrite, Magnet AXIOM, FTK, EnCase, Autopsy) publish no DOCX-specific extractor for the w:rsid family. Investigators who care about document genealogy have been using unzip -p plus grep for two decades.
What this all means
Six findings in one paragraph each.
Microsoft has been honest, in a place no one looks. The "What Document Inspector doesn't catch" page is a clean list of fourteen residue categories in Microsoft's own writing. It dates from June 30, 2009. Seventeen years on a Microsoft URL. Anyone who reads it learns most of what this post just demonstrated. Anyone who reads only the support article does not.
The feature has not changed materially in twelve years. The last public Microsoft announcement of a Document Inspector module change was December 9, 2014. OOXML has continued to grow. New element families appear in settings.xml every few years. The 2014 module list is what Word 2024's Inspector still applies. There is no public indication this is being revisited.
Federal agencies do not run it. Across nine downloaded files from six agencies, the Author field intact in eight, the lastModifiedBy intact in nine, rsidRoot intact in nine, headers and footers carrying named staffers and contractor full addresses intact in five, last-printed timestamps reaching back to 2011 in two. Document Inspector takes one click. None of these documents got it.
Even when it runs, the editing-session genealogy survives. rsids are present in nearly every Word document and Inspector does not address them. Two unrelated agency files in our corpus shared three editing-session identifiers, in a way that statistical analysis says is not random. Cross-document linkage is real. The forensic methodology has been peer-reviewed since 2023.
Document variables are an unmarked bypass. Three lawyer names in our torture experiment survived Inspector verbatim because they live in <w:docVars> inside settings.xml, not in docProps/core.xml. Microsoft's documentation never mentions docVars in any Inspector module description. Any law firm, any consultant, any agency that uses document variables for matter numbers, draft attorneys, or template fields is publishing those fields on every "scrubbed" file.
Embedded images and embedded objects pass through untouched. Microsoft says so. The HUD HOTMA training script proves it: a 2017 Adobe Illustrator chart embedded in a 2024 file. Anyone embedding photos from an iPhone is republishing the EXIF GPS coordinates. Anyone embedding an Excel chart is republishing the XLSX's own author, modifier, and rsid lineage.
These six findings are not theoretical. They are reproducible against the file URLs cited in the methodology section below.
Methodology
The federal corpus was downloaded on May 5, 2026, from the URLs listed below. SHA-256 hashes are recorded for every file. The parser used is the production parser at filexray.orygn.tech/scan. Cross-checks were performed against ExifTool 12.97 and against direct unzip -p against the OOXML zip parts.
Reproduction. Download any of the federal DOCX files from the URLs below. Drop into File X-Ray's /scan page. The fields described in this post will surface for any reader who downloads any of the source files. The torture experiment is a synthetic DOCX built by hand for the test; the parser that reads it is in the project repository.
For technical background on the metadata standards, see the Field Manual. For the previous post on document forensics in famous PDF leaks, see /leaks.
The /scan page accepts any DOCX, PDF, photo, video, audio, or other supported file. Parsing runs locally in your browser. No file leaves your device. The same parser that produced the findings in this post is the one you'll be running.
Scan a Document