ShareMaster V2 is in beta, a complete rebuild. See what is new and request access →
  1. Home
  2. Guides
  3. Find Duplicate Files

How to Find Duplicate Files in SharePoint Online, and Remove Them Safely

SharePoint Online has no duplicate file finder. Nothing in a document library, in Site settings or in the SharePoint admin centre will list the files that exist twice. What Microsoft does give you is a way to look at one library at a time (sort by name, show the file size) and an API that stores a content hash for every file, which a script can compare. For anything bigger than a single library you need a scan that reads every file's name, size and hash, groups the matches, and lets someone decide which copy stays.

The rest of this page covers what the built-in tools can and cannot show, how to decide what "duplicate" means for you, and how ShareMaster reports and removes copies across libraries and sites.

What SharePoint's own tools can show you

A library view, one library at a time

Open the library, sort by Name and add the File Size column to the view. Copies that kept the same name, or picked up a number such as Report (1).docx, end up next to each other. That works for a small library where you already suspect a problem. It does nothing for a copy that sits in another folder view, another library or another site, and with thousands of files the list is too long to read.

Search, which is built to hide duplicates

Searching for a file name feels like the obvious check, but SharePoint search removes duplicate results on purpose. Microsoft's Search REST API documentation lists a TrimDuplicates setting whose default is true. A search can tell you a name exists somewhere. It is a poor way to count how many times.

Microsoft Graph and a script

Files in SharePoint Online and OneDrive carry a content hash that Microsoft Graph can return with the file. Microsoft's hashes resource reference describes quickXorHash as a hash that can be used to tell whether a file's contents have changed. Two files with the same size and the same quickXorHash are, for any practical purpose, the same bytes. A PowerShell or Graph script that walks every library, collects size and hash, and groups the matches is the honest native answer. It is also a project: paging, throttling, libraries over the view threshold, and an output format someone can act on are all yours to build.

Native optionFindsMisses
Sorted library view with File SizeSame-name copies in one libraryCopies in other libraries or sites, renamed copies, identical content under a different name
SharePoint searchThat a name exists somewhereHow many copies exist, because duplicates are trimmed from results
Graph script on quickXorHashTrue content duplicates anywhere you script it to lookNothing in principle, but you write and run the scan, the grouping and the clean-up yourself

Where SharePoint duplicates come from

Knowing the source tells you where to look. The usual culprits: the same email attachment saved into three team sites by three people; a file copied into a project library instead of linked; a folder uploaded twice after a failed upload; a migration or copy job re-run over content that had already arrived. The first two produce identical files in different sites under the same or similar names. The last two produce copies in the same library, often with a number added to the name.

Decide what counts as a duplicate before you scan

"Duplicate" means three different things, and the difference decides what you are allowed to delete.

  • Same name. Free to check and says nothing about the contents. Two files called Minutes.docx in different project folders are almost never the same document.
  • Same name and same byte size. Still cheap, and usually right. A copy uploaded twice, or restored next to the original, normally matches on both.
  • Same contents. The only proof. It compares the bytes, through a hash, so it also catches a byte-identical Q3 Report FINAL.docx sitting beside Q3 Report.docx.

Two more decisions shape the result. Copy suffixes such as (1), - Copy and Copy of can be stripped so those names match their original. And some libraries should be left out entirely: a file in Site Assets or Site Pages may be the image or page a site is pointing at, so removing the "extra" copy can break a page that nobody thought of as a duplicate.

Tip: duplicates are often not the biggest storage cost. Version history frequently is. If the goal is space rather than tidiness, check how much version history is costing you and look for the largest files in the tenant before a duplicate hunt.

Find duplicate files with ShareMaster

In the current release: Find Duplicated Files

ShareMaster's Reports menu has a Find Duplicated Files report, part of Report Master, which is free on every plan. You choose one document library and select Run Report. It produces an Excel workbook with two sheets: every file in the library with its size, created and modified dates and authors, and a second sheet listing only the files that look duplicated.

Be clear about how it matches. The report looks for the numbered copies that appear when the same file name is uploaded or restored again, such as test.psd beside test(1).psd or test (2).psd. Files are grouped when they sit in the same folder and share a name once the bracketed number and letter case are ignored. It does not compare sizes or contents, and it does not match copies across folders, libraries or sites. It is a quick way to find "(1)" clutter in one library. It only reports; nothing is removed.

In the V2 beta: a cross-site report, then removal

The ShareMaster V2 beta goes further, and splits the job into two stages that can be run weeks apart.

  1. Generate the report. Under Reports, the Duplicate file report tile (also reachable from Optimise Storage Tools > Remove duplicate files > Generate report) scans the SharePoint sites, libraries or folders you tick, or OneDrive. Copies are matched across everything you ticked, so a file in one site can be grouped with its twin in another.
  2. Pick a match level. Same name, same name and size (the default), or same contents. For a content match, the beta asks SharePoint for the quickXorHash it already holds and downloads and hashes only the files it cannot get a hash for. Only files that share an exact byte size with another file are hashed at all, and nothing is written to your disk. If you signed in with a SharePoint cookie the stored hashes are not reachable, so every candidate file is downloaded, and the options screen tells you so.
  3. Set the filters. Optional: treat (1), - Copy and Copy of as the same name, include zero-byte files, ignore files below a size, include hidden libraries. Site Assets, Site Pages, Pages, Style Library and Form Templates are excluded by default.
  4. Read the workbook. Each set of matching files is a group, ordered so the set wasting the most space comes first, and the summary counts how many sets span more than one site. Alongside it comes a decision CSV. You can ask it to pre-mark all but one copy using a rule (keep the oldest, the newest, or the one with the shortest path); that is off by default.
  5. Remove using the report. Remove files using report loads the marked file, rechecks every marked copy against the live library and leaves alone anything that changed or moved since the scan. No set is ever emptied: if no unmarked copy is still there unchanged, the whole set is skipped. You see every copy with its own tick before anything happens.

Removal goes to the recycle bin by default, restorable for about 93 days. Permanent deletion is a separate choice behind a second dialog where you type DELETE. The report stage is free; removing files needs a paid plan or the trial.

A safe order for a duplicate clean-up

Whatever tool you use, the order matters more than the tool.

  1. Report first and change nothing. Share the workbook with the people who own the content.
  2. Start with the groups that waste the most space. A handful of large media files usually outweighs thousands of small documents.
  3. Keep the copy in the place people actually use, not necessarily the oldest. A link, a Teams tab or a page may point at one specific copy.
  4. Recycle, don't purge. If someone shouts a week later, the file is still in the recycle bin.

For deleting whole folders or libraries rather than individual copies, bulk deleting files in SharePoint covers the options. For a broader plan to cut storage, start with reducing SharePoint storage fast.

Frequently Asked Questions

Does SharePoint Online have a duplicate file finder?

No. There is no report or command for it in SharePoint Online or the admin centre. A sorted library view helps inside one library; anything wider needs a script against Microsoft Graph or a tool.

Why does SharePoint search not show all copies of a file?

Search trims duplicates from results by design. The Search REST API's TrimDuplicates setting is true unless a query turns it off, so several copies collapse into one result.

Are two files with the same name and size duplicates?

Usually. Only a content comparison proves it, either with the quickXorHash SharePoint stores for each file or by hashing the downloaded bytes.

Can I get duplicate files back after removing them?

Yes, from the recycle bin, for 93 days. A permanent delete skips the recycle bin and cannot be reversed.

Try ShareMaster free for 14 days