> For the complete documentation index, see [llms.txt](https://cleberjamaral.gitbook.io/mind/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://cleberjamaral.gitbook.io/mind/knowledge/programming/regex.md).

# regex

regular expressions

## Metacharacters (source [freeformatter.com](https://www.freeformatter.com/regex-tester.html))

| Character | What does it do?                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| --------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| $         | Matches the **end of the input**. If in multiline mode, it also matches **before a line break character**, hence every end of line.                                                                                                                                                                                                                                                                                                                                                                                                                      |
| (?:x)     | Matches 'x' but **does NOT remember the match**. Also known as NON-capturing parenthesis.                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| (x)       | Matches 'x' and **remembers the match**. Also known as capturing parenthesis.                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| \*        | Matches the preceding character **0 or more times**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| +         | Matches the preceding character **1 or more times**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| .         | Matches **any single character except the newline character**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| ?         | <ul><li>Matches the preceding character <strong>0 or 1 time</strong>.</li><li>When used after the quantifiers \*, +, ? or {}, <strong>makes the quantifier non-greedy</strong>; it will match the minimum number of times as opposed to matching the maximum number of times.</li></ul>                                                                                                                                                                                                                                                                  |
| \[\b]     | Matches a **backspace**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| \[^abc]   | Matches **anything NOT enclosed by the brackets**. Also known as a negative character set.                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| \[abc]    | Matches **any of the enclosed characters**. Also known as a character set. You can create range of characters using the hyphen character such as A-Z (A to Z). Note that in character sets, special characters (., \*, +) do not have any special meaning.                                                                                                                                                                                                                                                                                               |
| \\        | <ul><li>Used to indicate that the <strong>next character should NOT be interpreted literally</strong>. For example, the character 'w' by itself will be interpreted as 'match the character w', but using '\w' signifies 'match an alpha-numeric character including underscore'.</li><li>Used to indicate that a <strong>metacharacter is to be interpreted literally</strong>. For example, the '.' metacharacter means 'match any single character but a new line', but if we would rather match a dot character instead, we would use '.'.</li></ul> |
| \0        | Matches a **NULL character**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| \b        | Matches a **word boundary**. Boundaries are determined when a word character is NOT followed or NOT preceeded with another word character.                                                                                                                                                                                                                                                                                                                                                                                                               |
| \B        | Matches a **NON-word boundary**. Boundaries are determined when two adjacent characters are word characters OR non-word characters.                                                                                                                                                                                                                                                                                                                                                                                                                      |
| \cX       | Matches a **control character**. X must be between A to Z inclusive.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| \d        | Matches a **digit character**. Same as \[0-9] or \[0123456789].                                                                                                                                                                                                                                                                                                                                                                                                                                                                                          |
| \D        | Matches a **NON-digit character**. Same as \[^0-9] or \[^0123456789].                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| \f        | Matches a **form feed**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| \n        | Matches a **line feed**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| \r        | Matches a **carriage return**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| \s        | Matches a **single white space character**. This includes space, tab, form feed and line feed.                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| \S        | Matches **anything OTHER than a single white space character**. Anything other than space, tab, form feed and line feed.                                                                                                                                                                                                                                                                                                                                                                                                                                 |
| \t        | Matches **a tab**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| \uhhhh    | Matches a character with the **4-digits hexadecimal code**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| \v        | Matches **a vertical tab**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| \w        | Matches **any alphanumeric character incuding underscore**. Equivalent to \[A-Za-z0-9\_].                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| \W        | Matches **anything OTHER than an alphanumeric character incuding underscore**. Equivalent to \[^A-Za-z0-9\_].                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| \x        | A back reference to the substring matched by the x parenthetical expression. x is a positive integer.                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| \xhh      | Matches a character with the **2-digits hexadecimal code**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| ^         | <ul><li>Matches the <strong>beginning of the input</strong>. If in multiline mode, it also matches <strong>after a line break character</strong>, hence every new line.</li><li>When used in a set pattern (\[^abc]), it negates the set; <strong>match anything not enclosed in the brackets</strong></li></ul>                                                                                                                                                                                                                                         |
| x(?!y)    | Matches **'x' only if 'x' is NOT followed by 'y'**. Also known as a negative lookahead.                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |
| x(?=y)    | Matches **'x' only if 'x' is followed by 'y'**. Also known as a lookahead.                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| x\|y      | Matches **'x' OR 'y'**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |
| {n,m}     | Matches the preceding character **at least n times and at most m times. n and m can be omitted if zero.**.                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| {n}       | Matches the preceding character **exactly n times**.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |

## Testers

* [pythex](https://pythex.org/): a python regular expression composer with a cheatsheet and multiline support.
* [debuggex](https://www.debuggex.com/): online regex tester for python and javascript presents the expression in a graph way describying each part and multiline support.
* [extendsclass](https://extendsclass.com/regex-tester.html): online regex tester for javascript, python, ruby, java, php and mysql with a regex visualizer.

## Source

* [Regular-Expressions-info](https://www.regular-expressions.info/): brings a comprehensive documentation about regex.
