Remove duplicate lines

Paste your list: repeated lines disappear and only the first occurrence of each one is kept. You can also ignore upper and lower case.

0Total lines
0Unique lines
0Duplicates removed

Your list stays in your browser: it is never saved or uploaded.

When you need a list without duplicates

Duplicates creep in everywhere: email lists merged from several spreadsheets (with the risk of sending the same newsletter twice), SEO keyword lists, discount codes, warehouse SKUs, columns exported from a CSV. This tool always keeps the first occurrence of each line, so the original order of your list never changes, only the later repetitions are dropped.

Upper and lower case: when to ignore them

For email addresses it's best to turn the option on: JOHN@MAIL.COM and john@mail.com are the same mailbox, so they should count as duplicates. For product codes, passwords, or technical identifiers, on the other hand, case can be meaningful, leave the box unchecked. Either way, the line that is kept preserves its original spelling, completely untouched.

The duplicates a computer cannot see

Two lines can look identical without being so: one invisible trailing space is enough to make them count as different, and it is the commonest case when the list comes from a copy-paste or a spreadsheet. The same goes for double spaces in the middle and for a tab used instead of a space.

The other trap is characters that resemble each other: the straight apostrophe and the curly one, the hyphen and the dash, the Latin «a» and the Cyrillic one. To the eye they are the same line, to a comparison they are two different lines, and no amount of human attention tells them apart.

Case and accents: it depends on the data

For email addresses the part after the at sign is case-insensitive, so MARIO@MAIL.COM and mario@mail.com are the same mailbox and should be merged. For a password or a code, by contrast, case matters and merging would destroy the data.

On accents it is nearly always better to distinguish: «pero» and «però» are different words. The exception is lists of personal names typed by different people, where «Nicolo» and «Nicolò» are almost certainly the same person: there the choice is not technical, it is a decision about what the list contains.

Removing duplicates is not always the goal

Sometimes the real question is which lines are repeated and how often, not making them disappear: in a list of orders, a duplicate is an error to understand, not to hide. Deleting it before looking at it loses the most useful information.

It is also worth deciding which copy you keep: the first or the last. If the lines carry a date or a sequence number, the last is usually the good version; if the list is a chronological log, the first. With identical lines it makes no difference, with similar ones it does.

Nearby tools

To reorder the list before or after there is Sort lines, and to see how long it is Word counter. If the lines come from a CSV, Clean a CSV does the job on columns.