Regex: HTML tag
The pattern finds a <, an optional / for a closing tag, a tag name (captured as group 1), then everything up to the next >. It is handy for a quick look at markup in a log or a search across files, as long as the HTML is simple.
Pattern
</?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*> Flags: g
How it works
- <
- The opening angle bracket.
- /?
- An optional slash: a closing tag.
- ([a-zA-Z][a-zA-Z0-9-]*)
- The tag name, captured: a letter, then letters, digits or dashes (custom elements like
my-element). - \b
- A word boundary, so the name ends here.
- [^>]*
- Attributes: anything except
>. - >
- The closing angle bracket.
Matches
- <div>
- </p>
- <img src="a.png" alt="x">
- <my-element data-x="1">
Doesn't match
- a < b > c
- <>
- <!-- note -->
- 3 <5
In your language
- JavaScript
/<\/?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*>/gA literal;
new RegExp(source, flags)builds the same from a string.- Python
re.compile(r"</?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*>", re.ASCII)A raw string, so backslashes reach
reas written.re.ASCIIkeeps\dto0-9, as in JavaScript (Python matches any Unicode digit otherwise). Usere.fullmatchto test a whole string.- Java
Pattern.compile("</?([a-zA-Z][a-zA-Z0-9-]*)\\b[^>]*>")A normal string literal, so every backslash is doubled.
matcher(s).matches()tests the whole string.- Go
regexp.MustCompile(`</?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*>`)A raw string in backticks. RE2 runs in linear time but has no lookaround and no backreferences.
- PHP
preg_match('/<\/?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*>/', $input)PCRE with
/delimiters inside a single-quoted string. Without theDmodifier,$also matches before a final newline, so it is added to patterns that end in$.- C#
new Regex(@"</?([a-zA-Z][a-zA-Z0-9-]*)\b[^>]*>", RegexOptions.ECMAScript)A verbatim string: backslashes stay, a quote is doubled.
RegexOptions.ECMAScriptkeeps\dto0-9, as in JavaScript.$also matches before a final newline; to reject one, end the pattern with\zinstead.
Common mistakes
Regex cannot parse HTML
A
>inside an attribute value (<a title="x > y">) ends the match early, and comments and<script>contents confuse it. HTML nests, and a regular expression cannot keep count of nesting.To strip tags, use the DOM
In a browser,
new DOMParser().parseFromString(html, "text/html").body.textContentgives the text. On a server, use an HTML parser:cheerioin Node.js,html.parserin Python,golang.org/x/net/htmlin Go.Never use it to make HTML safe
Removing tags with a regex does not prevent XSS: attributes, entities and malformed markup slip through. Use a sanitizer such as DOMPurify, or escape the text instead of filtering it.