AdSense Placeholder
Slot: header_reference_page
A regular expression (regex) describes a pattern of text. You can use it to find, validate or replace text in editors, command-line tools and every programming language. Try any pattern live in the Regex Tester. Every example below was run in JavaScript and Python and produces the result shown.
Characters and classes
| Pattern |
Meaning |
Example |
. |
any one character except a newline |
a-c → a, -, c |
\d+ |
one or more digits (\d is a digit, \D is not) |
a12b345 → 12, 345 |
\w+ |
word characters: letters, digits and underscore (\W is the opposite) |
hi_there! ok → hi_there, ok |
a\sb |
whitespace: space, tab, newline (\S is the opposite) |
a b ab → a b |
gr[ae]y |
one character from the set |
gray grey groy → gray, grey |
[^0-9]+ |
anything not in the set |
ab12cd3 → ab, cd |
[a-f0-9]+ |
ranges inside a set |
ff09 xyz 1a → ff09, 1a |
Anchors and boundaries
| Pattern |
Meaning |
Example |
^\d+$ |
^ is the start and $ the end: the whole text must match |
12345 → 12345 |
^\d+$ |
|
12a45 → no match |
\bcat\b |
word boundary: matches cat as a whole word only |
cat concat cat. → cat, cat |
Quantifiers
| Pattern |
Meaning |
Example |
ab* |
* means zero or more of the previous item |
a ab abbb → a, ab, abbb |
ab+ |
+ means one or more |
a ab abbb → ab, abbb |
colou?r |
? means optional (zero or one) |
color colour → color, colour |
\d{3} |
{n} means exactly n times |
12 123 12345 → 123, 123 |
\d{2,3} |
{n,m} means between n and m times |
1 12 123 1234 → 12, 123, 123 |
<.+> |
greedy: takes as much as it can |
<a><b> → <a><b> |
<.+?> |
lazy: add ? to take as little as possible |
<a><b> → <a>, <b> |
Groups, alternation and lookaround
| Pattern |
Meaning |
Example |
cat|dog |
| means either side |
hotdog cat → dog, cat |
(?:ab)+ |
a group that does not capture, so a quantifier can apply to it |
ababab ab → ababab, ab |
(\w)\1 |
\1 repeats what group 1 matched |
hello book → ll, oo |
(?<year>\d{4})-\d{2} |
a named group, (?<name>...) (JavaScript and PCRE; Python writes (?P<name>...)) |
on 2026-10-03 → 2026-10 |
\d+(?=%) |
lookahead: followed by %, without consuming it |
50% 20 30% → 50, 30 |
foo(?!bar) |
negative lookahead: not followed by bar |
foobar foobaz → foo |
(?<=\$)\d+ |
lookbehind: preceded by $, without consuming it |
cost $15 or 20 → 15 |
Flags
| Flag |
Name |
Effect |
i |
Ignore case |
cat also matches Cat and CAT |
g |
Global |
Find every match, not only the first |
m |
Multiline |
^ and $ match at each line, not only the whole text |
s |
Dot matches newline |
. also matches a line break |
u |
Unicode |
Treat the pattern as Unicode, enabling \p{...} classes |
Characters to escape
These have a special meaning, so write a backslash before them to match them literally: . ^ $ * + ? ( ) [ ] { } | \. For example, 3\.14 matches 3.14 but not 3x14.
Practical patterns
| For |
Pattern |
Matches |
| Integer |
^-?\d+$ |
42, -7 |
| Decimal number |
^-?\d+(\.\d+)?$ |
3.14, -2, 10 |
| Hex colour |
^#(?:[0-9a-fA-F]{3}){1,2}$ |
#fff, #a1B2c3 |
| Email (rough check) |
^[^\s@]+@[^\s@]+\.[^\s@]+$ |
[email protected], [email protected] |
| IPv4 address |
^(?:(?:25[0-5]|2[0-4]\d|1?\d?\d)\.){3}(?:25[0-5]|2[0-4]\d|1?\d?\d)$ |
192.168.1.1, 0.0.0.0, 255.255.255.255 |
| ISO date shape |
^\d{4}-\d{2}-\d{2}$ |
2026-10-03 |
| Repeated word |
\b(\w+)\s+\1\b |
the the, is is |
These patterns check the shape of text. A real email address, for example, cannot be fully validated with a regex: send a confirmation message instead.
Pitfalls
- Greedy matching takes as much as possible; use
? after a quantifier to make it lazy.
- Catastrophic backtracking. Nested quantifiers such as
(a+)+ can take exponential time on certain inputs. Keep patterns simple when they run on untrusted text.
- Flavours differ. JavaScript, Python, PCRE (PHP, many tools) and POSIX support slightly different syntax, especially for named groups, lookbehind and Unicode.
- Do not parse HTML with regex. Use a real parser for nested structures.