正则表达式完全指南:从入门到精通,看完这篇就够了

正则表达式(Regular Expression)是文本处理的利器。掌握正则表达式,可以大幅提升你的文本处理效率。

1. 基础语法

元字符

元字符 说明 示例
. 匹配任意单个字符(除换行符) a.c 匹配 abc, aac, a@c
^ 匹配行首 ^Hello 匹配行首的Hello
$ 匹配行尾 world$ 匹配行尾的world
* 匹配前一个字符0次或多次 ab*c 匹配 ac, abc, abbc
+ 匹配前一个字符1次或多次 ab+c 匹配 abc, abbc
? 匹配前一个字符0次或1次 ab?c 匹配 ac, abc
{n} 匹配前一个字符恰好n次 a{3} 匹配 aaa
{n,m} 匹配前一个字符n到m次 a{2,4} 匹配 aa, aaa, aaaa
[] 字符集合 [abc] 匹配 a, b 或 c
[^] 否定字符集合 [^abc] 匹配除a,b,c外的任意字符
| 或操作 cat|dog 匹配 cat 或 dog
() 分组 (ab)+ 匹配 ab, abab

字符类

\d  # 数字 [0-9]
\D  # 非数字 [^0-9]
\w  # 单词字符 [a-zA-Z0-9_]
\W  # 非单词字符
\s  # 空白字符(空格、制表符、换行)
\S  # 非空白字符
\b  # 单词边界
\B  # 非单词边界

2. 常用正则示例

邮箱验证

^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$

手机号验证(中国)

^1[3-9]\d{9}$

身份证号

^[1-9]\d{5}(18|19|20)\d{2}(0[1-9]|1[0-2])(0[1-9]|[12]\d|3[01])\d{3}[\dXx]$

IP地址

^((25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)$

URL

^https?://[\w\-]+(\.[\w\-]+)+[/#?]?.*$

日期YYYY-MM-DD

^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])$

时间HH:MM:SS

^([01]\d|2[0-3]):[0-5]\d:[0-5]\d$

十六进制颜色

^#?([a-f0-9]{6}|[a-f0-9]{3})$

3. Python正则应用

re模块基础

import re

# 匹配
pattern = r"\d+"
text = "Order 123, Order 456"
matches = re.findall(pattern, text)
# 结果: ['123', '456']

# 替换
new_text = re.sub(r"\d+", "XXX", text)
# 结果: "Order XXX, Order XXX"

# 分割
parts = re.split(r",\s*", "a, b, c")
# 结果: ['a', 'b', 'c']

分组与提取

import re

text = "Name: Alice, Age: 30, City: Beijing"
pattern = r"Name: (\w+), Age: (\d+), City: (\w+)"

match = re.search(pattern, text)
if match:
    name, age, city = match.groups()
    print(f"Name: {name}, Age: {age}, City: {city}")

命名分组

import re

pattern = r"(?P\d{4})-(?P\d{2})-(?P\d{2})"
text = "2025-01-15"

match = re.search(pattern, text)
if match:
    print(match.group("year"))   # 2025
    print(match.group("month"))  # 01
    print(match.group("day"))    # 15

预编译正则

import re

# 预编译提升性能
pattern = re.compile(r"\d{3}-\d{4}")

phones = ["010-1234", "021-5678", "0755-9999"]
for phone in phones:
    if pattern.match(phone):
        print(f"Valid: {phone}")

4. 高级技巧

贪婪与非贪婪

import re

# 贪婪匹配(默认)
text = "
content1
content2
" re.findall(r"
.*
", text) # 结果: ['
content1
content2
'] # 非贪婪匹配 re.findall(r"
.*?
", text) # 结果: ['
content1
', '
content2
']

零宽断言

# 正向先行断言 (?=...)
re.findall(r"\d+(?=\s*USD)", "100 USD 200 EUR")
# 结果: ['100']

# 正向后行断言 (?<=...)
re.findall(r"(?<=USD\s)\d+", "100 USD 200 EUR")
# 结果: ['200']

回溯引用

# 匹配重复单词
pattern = r"\b(\w+)\s+\1\b"
text = "Hello hello world world"
re.findall(pattern, text, re.IGNORECASE)
# 结果: ['hello', 'world']

条件匹配

# (?(id/name)yes-pattern|no-pattern)
pattern = r"(\d{3})-(?(1)\d{4}|\d{7})"
re.match(pattern, "010-1234")  # 匹配
re.match(pattern, "010-1234567")  # 不匹配

5. 实用案例

提取HTML标签

import re

html = "

Hello

World
" tags = re.findall(r"<([a-z]+)>", html) # 结果: ['p', 'div']

提取Markdown链接

import re

md = "[Google](https://google.com) and [GitHub](https://github.com)"
links = re.findall(r"\[([^\]]+)\]\(([^\)]+)\)", md)
# 结果: [('Google', 'https://google.com'), ('GitHub', 'https://github.com')]

日志解析

import re

log = "2025-01-15 10:30:45 [ERROR] Database connection failed"
pattern = r"(\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}) \[(\w+)\] (.+)"

match = re.match(pattern, log)
if match:
    timestamp, level, message = match.groups()
    print(f"Time: {timestamp}, Level: {level}, Message: {message}")

6. 性能优化

使用预编译

# 不好的做法
for i in range(10000):
    re.match(r"\d+", text)

# 好的做法
pattern = re.compile(r"\d+")
for i in range(10000):
    pattern.match(text)

避免回溯

# 不好的做法(可能回溯)
pattern = r"^(a+)+$"

# 好的做法(明确限定)
pattern = r"^a{1,10}$"

7. 在线工具

  • Regex101:https://regex101.com/
  • Regexr:https://regexr.com/
  • RegExr:https://www.regexpal.com/

总结

正则表达式学习曲线较陡,但一旦掌握,将成为你文本处理的得力助手。建议:

  • 从简单开始:先掌握基础元字符
  • 多练习:使用在线工具验证正则
  • 注意性能:避免过度回溯
  • 添加注释:复杂正则添加说明
📖 推荐资源:《精通正则表达式》(Jeffrey Friedl著)是正则领域的经典之作。在线练习推荐RegexOne。

发表评论