Developer Toolbox

Regex: HTML tag

The pattern finds a <, an optional / for a closing tag, a tag name (captured as group 1), then everything up to the next >. It is handy for a quick look at markup in a log or a search across files, as long as the HTML is simple.

Open in Regex Tester The pattern and every example below are filled in.

Pattern

</?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*>

Flags: g

How it works

<
The opening angle bracket.
/?
An optional slash: a closing tag.
([a-zA-Z][a-zA-Z0-9-]*)
The tag name, captured: a letter, then letters, digits or dashes (custom elements like my-element).
\b
A word boundary, so the name ends here.
[^>]*
Attributes: anything except >.
>
The closing angle bracket.

Matches

  • <div>
  • </p>
  • <img src="a.png" alt="x">
  • <my-element data-x="1">

Doesn't match

  • a < b > c
  • <>
  • <!-- note -->
  • 3 <5

In your language

JavaScript
/<\/?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*>/g

A literal; new RegExp(source, flags) builds the same from a string.

Python
re.compile(r"</?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*>", re.ASCII)

A raw string, so backslashes reach re as written. re.ASCII keeps \d to 0-9, as in JavaScript (Python matches any Unicode digit otherwise). Use re.fullmatch to test a whole string.

Java
Pattern.compile("</?([a-zA-Z][a-zA-Z0-9-]*)\\b[^>]*>")

A normal string literal, so every backslash is doubled. matcher(s).matches() tests the whole string.

Go
regexp.MustCompile(`</?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*>`)

A raw string in backticks. RE2 runs in linear time but has no lookaround and no backreferences.

PHP
preg_match('/<\/?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*>/', $input)

PCRE with / delimiters inside a single-quoted string. Without the D modifier, $ also matches before a final newline, so it is added to patterns that end in $.

C#
new Regex(@"</?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*>", RegexOptions.ECMAScript)

A verbatim string: backslashes stay, a quote is doubled. RegexOptions.ECMAScript keeps \d to 0-9, as in JavaScript. $ also matches before a final newline; to reject one, end the pattern with \z instead.

Common mistakes

  • Regex cannot parse HTML

    A > inside an attribute value (<a title="x > y">) ends the match early, and comments and <script> contents confuse it. HTML nests, and a regular expression cannot keep count of nesting.

  • To strip tags, use the DOM

    In a browser, new DOMParser().parseFromString(html, "text/html").body.textContent gives the text. On a server, use an HTML parser: cheerio in Node.js, html.parser in Python, golang.org/x/net/html in Go.

  • Never use it to make HTML safe

    Removing tags with a regex does not prevent XSS: attributes, entities and malformed markup slip through. Use a sanitizer such as DOMPurify, or escape the text instead of filtering it.