Developer Toolbox

Regex: URL

The pattern checks the shape of a whole string as a web address: http:// or https://, then a host that does not start with a dot or a slash, then one more character of any kind and a run with no whitespace up to the end. It is short enough to read at a glance; to validate input properly, a URL parser tells you more.

Open in Regex Tester The pattern and every example below are filled in.

Pattern

^https?://[^\s/$.?#].[^\s]*$

How it works

^
Start of the string.
https?://
The scheme: http, an optional s, then ://.
[^\s/$.?#]
The first character of the host: not whitespace, /, $, ., ? or #.
.
One more character of any kind, so http://a alone fails (http://a/ passes).
[^\s]*
The rest: any run of characters that are not whitespace.
$
End of the string.

Matches

  • https://example.com
  • http://localhost:3000/api?x=1
  • https://example.com/a/b#top

Doesn't match

  • example.com
  • ftp://example.com/file
  • https://
  • http://exa mple.com

In your language

JavaScript
/^https?:\/\/[^\s\/$.?#].[^\s]*$/

A literal; new RegExp(source, flags) builds the same from a string.

Python
re.compile(r"^https?://[^\s/$.?#].[^\s]*$", re.ASCII)

A raw string, so backslashes reach re as written. re.ASCII keeps \d to 0-9, as in JavaScript (Python matches any Unicode digit otherwise). Use re.fullmatch to test a whole string.

Java
Pattern.compile("^https?://[^\\s/$.?#].[^\\s]*$")

A normal string literal, so every backslash is doubled. matcher(s).matches() tests the whole string.

Go
regexp.MustCompile(`^https?://[^\s/$.?#].[^\s]*$`)

A raw string in backticks. RE2 runs in linear time but has no lookaround and no backreferences.

PHP
preg_match('/^https?:\/\/[^\s\/$.?#].[^\s]*$/D', $input)

PCRE with / delimiters inside a single-quoted string. Without the D modifier, $ also matches before a final newline, so it is added to patterns that end in $.

C#
new Regex(@"^https?://[^\s/$.?#].[^\s]*$", RegexOptions.ECMAScript)

A verbatim string: backslashes stay, a quote is doubled. RegexOptions.ECMAScript keeps \d to 0-9, as in JavaScript. $ also matches before a final newline; to reject one, end the pattern with \z instead.

Common mistakes

  • A parser knows more

    new URL(s) in JavaScript and url.Parse in Go reject bad ports and malformed hosts. Python's urllib.parse.urlparse is lenient: it accepts spaces and bad percent escapes, so check the parts it returns.

  • No check on the host

    http://a.b, http://localhost and http://-x- all pass. Check the host against what you expect when it matters, for instance before your server fetches the address.

  • Finding links in text

    The ^ and $ make the pattern test a whole string. To find links inside prose, drop them and add the g flag; a closing ) or a full stop right after a link then becomes part of the match, so trim trailing punctuation afterwards.