Step 1

Decide what you are trying to prove

Email validation answers two different questions that are easy to confuse. The first is whether an address is well-formed: does the text follow the rules for how an email address is written? The second is whether it is deliverable: will a message sent to it reach a person who wants it? A CSV can answer the first question completely and the second not at all.

Syntax validation catches real problems. It finds a missing @ sign, a space pasted into the middle of an address, a domain that ends in a dot, two addresses jammed into one cell, and invisible characters copied from a web page.

What syntax cannot tell you is whether the domain accepts mail, whether the mailbox exists today, whether it belongs to the person in the name column, or whether that person agreed to hear from you. Keep those questions separate from the start. The CSV email validator asks you to acknowledge exactly that before it exports anything: well-formed is not the same as deliverable.

Step 2

Find the email column and normalize what you cannot see

Start by choosing the right column. Exports name it Email, E-mail, Email Address, Work Email, or something less obvious. A header suggestion is a convenience, not a guarantee, so confirm that the column you selected actually holds addresses and not, for example, a username or a display name.

Some cells hold more than one address, separated by semicolons or commas. If each row should be one subscriber, fix that at the source; if a row can legitimately carry several addresses, check each one separately.

Next, look for characters you cannot see. Leading and trailing spaces are common after copying from a spreadsheet. Zero-width spaces, zero-width joiners, byte order marks, soft hyphens, and non-breaking spaces arrive from web pages, word processors, and chat tools. They make two identical-looking addresses compare as different and can make an import reject a row that looks fine. Trimming and removing them is a safe normalization; it does not change what the address means. Changing letter case or correcting a domain does change it, so leave those for review.

Step 3

Use the WHATWG HTML rule as the practical baseline

The most practical definition of a well-formed address is the one browsers already enforce. The WHATWG HTML Standard defines a valid email address with a short grammar: one or more characters from a fixed set, an @ sign, and one or more domain labels separated by dots. The allowed local-part characters are letters, digits, the dot, and !#$%&'*+/=?^_`{|}~-. Each domain label starts and ends with a letter or digit, may contain hyphens in between, and is limited to 63 characters.

The standard says plainly that this definition is a willful violation of RFC 5322. It describes the RFC 5322 syntax as too strict before the @ sign, too vague after it, and too lax in allowing comments, whitespace, and quoted strings in ways unfamiliar to most users. That is why the HTML rule is a good default for a mailing list: it describes the addresses that real sign-up forms accept.

Two details are worth knowing. The HTML rule accepts a single-label host such as user@localhost, which is fine for an intranet form but not for an internet mailing list, so a list validator should additionally require at least one dot in the domain. It also accepts dots anywhere in the local part, including first..last@example.com, which the next rule set treats differently.

Step 4

Know where RFC 5322 is stricter and where it is looser

RFC 5322 section 3.4.1 defines an address as a local part, an @ sign, and a domain. The local part is a dot-atom, a quoted string, or an obsolete form, and the domain is a dot-atom, a domain literal in square brackets, or an obsolete form. A dot-atom is one or more allowed characters with single dots between them, so a leading dot, a trailing dot, or two consecutive dots are not part of it.

That gives two reasonable baselines for local-part dots. Under the HTML rule alone, a..b@example.com is valid, but many sending systems reject it because it is not a dot-atom. A validator can report it as a review under the HTML baseline, or as a blocker when you choose the stricter RFC 5322 dot rule. Dots in the domain are different: an empty domain label is invalid under every rule.

RFC 5322 is also looser in two ways that matter for lists. A quoted local part such as "first last"@example.com and a domain literal such as user@[192.0.2.1] are both legal. The same section says the dot-atom form SHOULD be used and the quoted-string form SHOULD NOT be used. The HTML rule rejects both forms. Flag them for review rather than silently passing or deleting them.

Step 5

Apply the RFC 5321 length limits exactly

Length limits come from the SMTP specification. RFC 5321 section 4.5.3.1 sets them out under size limits and minimums. The maximum total length of a local part is 64 octets. The maximum total length of a domain name or number is 255 octets. The maximum total length of a reverse-path or forward-path is 256 octets, including the punctuation and element separators.

That last limit is often misquoted as the maximum address length. A path wraps the address in angle brackets, so the two brackets use two of the 256 octets. A verified erratum to RFC 3696 corrects that document's earlier figure and states that the upper limit on address length should normally be treated as 254. The validator therefore checks 64, 255, and 254, measured in octets.

Octets are bytes, not characters. An ASCII character is one octet in UTF-8, but an accented letter is two and many other characters are three or four. Measure the UTF-8 byte length, not the character count, or long internationalized addresses will slip through.

Step 6

Check the shape of the domain

After the @ sign, split the domain on dots and check each label. An empty label means the domain starts with a dot, ends with a dot, or contains two dots in a row. A label longer than 63 characters breaks the HTML rule. A label that starts or ends with a hyphen is not allowed by the HTML grammar either.

Two more checks are useful for a mailing list even though they go beyond the HTML grammar. Require at least one dot, because a single-label host is not a public mail domain. And reject a final label made only of digits, such as user@example.123, because that is not a usable top-level domain name and usually signals a mangled IP address or a typo.

Step 7

Treat non-ASCII addresses as a capability question

Addresses can contain non-ASCII characters in the local part, the domain, or both. RFC 6531 defines the SMTPUTF8 extension that makes this possible. A server that advertises SMTPUTF8 must be prepared to accept a UTF-8 string wherever RFC 5321 allows a mailbox, and internationalized domains follow IDNA rules.

The catch is that every hop must support it. RFC 6531 says a client must not transmit an internationalized address to a server that does not advertise the extension. The HTML rule is ASCII-only, so many sign-up forms and sending tools reject these addresses outright.

Step 8

Find duplicates without assuming case rules

Duplicates inflate counts and can mean the same person receives the same message twice. The right comparison depends on which half of the address you are looking at. RFC 5321 section 2.4 says mailbox domains follow DNS rules and are not case-sensitive, while the local part must be treated as case-sensitive.

So compare the domain in lowercase and the local part exactly. Two rows that match under that comparison are exact duplicates: keep the first and move the rest to the flagged list. Two rows that differ only by letter case in the local part are probably the same mailbox, because most providers ignore case, but the standard does not promise that. Flag them for a person to decide instead of merging them automatically.

Plus addressing, such as name+news@example.com, is a similar judgment call. Some providers deliver it to the same mailbox as name@example.com, but that is a provider behavior, not a syntax rule. Leave it alone unless you know the provider.

Step 9

Review role accounts and typo domains

Role-based addresses such as info@, support@, sales@, admin@, and noreply@ are perfectly valid. They often reach a shared inbox, a ticketing system, or nobody, and they are rarely tied to one person's consent.

Typo domains are the most common real-world error that syntax alone misses. gmial.com, hotmial.com, and yaho.com are well-formed, but they are probably not what the subscriber meant. A validator can compare domains against a short list of known misspellings and suggest the likely domain. It should never rewrite the address, because a wrong guess sends mail to a different domain and a different owner.

Step 10

Hand off deliverability to the sender

Once the list is well-formed, the remaining questions belong to the system that sends mail. The most reliable proof that an address works and that its owner wants your messages is a confirmation message: send an email asking the person to confirm, and add them only when they do. This is often called double opt-in.

After sending, process bounces. A hard bounce means the receiving system rejected the address; suppress it so you do not keep sending to it. A soft bounce is temporary and is usually retried by the sender. Your sending platform documents its own bounce categories and thresholds. Follow those rather than numbers from a generic guide.

Step 11

Use a repeatable local workflow

Keep the original export untouched. Load a copy, confirm the email column, choose whether to split multi-address cells, and choose the syntax baseline. Read the findings by rule, starting with blockers.

Export three files: the clean list with your original rows and a normalized email column, the flagged list with row numbers, rules, severities, and masked addresses, and a counts-only report you can share without exposing anyone's address. Fix the flagged rows at the source and run the check again.

Common questions
  • *

    Can I validate email addresses in a CSV without uploading it?

    Yes. This validator parses the CSV or one Excel sheet in the browser tab, checks every address in the chosen column against syntax and length rules, and builds clean and flagged lists locally. Nothing is sent to a server.

  • *

    Does a valid result mean the email will be delivered?

    No. Syntax validation proves only that an address is well-formed. It cannot prove that the domain accepts mail, that the mailbox exists, or that the person agreed to receive email. Confirmation messages and bounce handling in your sending platform answer those questions.

  • *

    Which syntax rule does the validator use?

    The default is the WHATWG HTML definition of a valid email address, the same rule browsers apply to an email input. An optional stricter baseline adds the RFC 5322 dot-atom rule that forbids leading, trailing, or consecutive dots in the local part.

  • *

    How long can an email address be?

    RFC 5321 sets 64 octets for the local part and 255 octets for the domain, and limits a path to 256 octets including its angle brackets. A verified erratum to RFC 3696 explains that this leaves 254 octets for the address itself.

  • *

    Does the validator check MX records or disposable domains?

    No. It performs no DNS, MX, SMTP, or disposable-domain lookup, so it never contacts a mail server or reveals your list to a third party.

  • *

    Are uppercase and lowercase duplicates the same address?

    The domain is case-insensitive. RFC 5321 says the local part must be treated as case-sensitive, even though most mailbox providers ignore case. The validator blocks exact duplicates and flags case-only duplicates for review.

Keep going