正则表达式(Regular Expression)是文本处理的利器。掌握正则表达式,可以大幅提升你的文本处理效率。
1. 基础语法
元字符
| 元字符 | 说明 | 示例 |
|---|---|---|
| . | 匹配任意单个字符(除换行符) | a.c 匹配 abc, aac, a@c |
| ^ | 匹配行首 | ^Hello 匹配行首的Hello |
| $ | 匹配行尾 | world$ 匹配行尾的world |
| * | 匹配前一个字符0次或多次 | ab*c 匹配 ac, abc, abbc |
| + | 匹配前一个字符1次或多次 | ab+c 匹配 abc, abbc |
| ? | 匹配前一个字符0次或1次 | ab?c 匹配 ac, abc |
| {n} | 匹配前一个字符恰好n次 | a{3} 匹配 aaa |
| {n,m} | 匹配前一个字符n到m次 | a{2,4} 匹配 aa, aaa, aaaa |
| [] | 字符集合 | [abc] 匹配 a, b 或 c |
| [^] | 否定字符集合 | [^abc] 匹配除a,b,c外的任意字符 |
| | | 或操作 | cat|dog 匹配 cat 或 dog |
| () | 分组 | (ab)+ 匹配 ab, abab |
字符类
\d # 数字 [0-9]
\D # 非数字 [^0-9]
\w # 单词字符 [a-zA-Z0-9_]
\W # 非单词字符
\s # 空白字符(空格、制表符、换行)
\S # 非空白字符
\b # 单词边界
\B # 非单词边界
2. 常用正则示例
邮箱验证
^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$
手机号验证(中国)
^1[3-9]\d{9}$
身份证号
^[1-9]\d{5}(18|19|20)\d{2}(0[1-9]|1[0-2])(0[1-9]|[12]\d|3[01])\d{3}[\dXx]$
IP地址
^((25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)$
URL
^https?://[\w\-]+(\.[\w\-]+)+[/#?]?.*$
日期YYYY-MM-DD
^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])$
时间HH:MM:SS
^([01]\d|2[0-3]):[0-5]\d:[0-5]\d$
十六进制颜色
^#?([a-f0-9]{6}|[a-f0-9]{3})$
3. Python正则应用
re模块基础
import re
# 匹配
pattern = r"\d+"
text = "Order 123, Order 456"
matches = re.findall(pattern, text)
# 结果: ['123', '456']
# 替换
new_text = re.sub(r"\d+", "XXX", text)
# 结果: "Order XXX, Order XXX"
# 分割
parts = re.split(r",\s*", "a, b, c")
# 结果: ['a', 'b', 'c']
分组与提取
import re
text = "Name: Alice, Age: 30, City: Beijing"
pattern = r"Name: (\w+), Age: (\d+), City: (\w+)"
match = re.search(pattern, text)
if match:
name, age, city = match.groups()
print(f"Name: {name}, Age: {age}, City: {city}")
命名分组
import re
pattern = r"(?P\d{4})-(?P\d{2})-(?P\d{2})"
text = "2025-01-15"
match = re.search(pattern, text)
if match:
print(match.group("year")) # 2025
print(match.group("month")) # 01
print(match.group("day")) # 15
预编译正则
import re
# 预编译提升性能
pattern = re.compile(r"\d{3}-\d{4}")
phones = ["010-1234", "021-5678", "0755-9999"]
for phone in phones:
if pattern.match(phone):
print(f"Valid: {phone}")
4. 高级技巧
贪婪与非贪婪
import re
# 贪婪匹配(默认)
text = "content1content2"
re.findall(r".*", text)
# 结果: ['content1content2']
# 非贪婪匹配
re.findall(r".*?", text)
# 结果: ['content1', 'content2']
零宽断言
# 正向先行断言 (?=...)
re.findall(r"\d+(?=\s*USD)", "100 USD 200 EUR")
# 结果: ['100']
# 正向后行断言 (?<=...)
re.findall(r"(?<=USD\s)\d+", "100 USD 200 EUR")
# 结果: ['200']
回溯引用
# 匹配重复单词
pattern = r"\b(\w+)\s+\1\b"
text = "Hello hello world world"
re.findall(pattern, text, re.IGNORECASE)
# 结果: ['hello', 'world']
条件匹配
# (?(id/name)yes-pattern|no-pattern)
pattern = r"(\d{3})-(?(1)\d{4}|\d{7})"
re.match(pattern, "010-1234") # 匹配
re.match(pattern, "010-1234567") # 不匹配
5. 实用案例
提取HTML标签
import re
html = "Hello
World"
tags = re.findall(r"<([a-z]+)>", html)
# 结果: ['p', 'div']
提取Markdown链接
import re
md = "[Google](https://google.com) and [GitHub](https://github.com)"
links = re.findall(r"\[([^\]]+)\]\(([^\)]+)\)", md)
# 结果: [('Google', 'https://google.com'), ('GitHub', 'https://github.com')]
日志解析
import re
log = "2025-01-15 10:30:45 [ERROR] Database connection failed"
pattern = r"(\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}) \[(\w+)\] (.+)"
match = re.match(pattern, log)
if match:
timestamp, level, message = match.groups()
print(f"Time: {timestamp}, Level: {level}, Message: {message}")
6. 性能优化
使用预编译
# 不好的做法
for i in range(10000):
re.match(r"\d+", text)
# 好的做法
pattern = re.compile(r"\d+")
for i in range(10000):
pattern.match(text)
避免回溯
# 不好的做法(可能回溯)
pattern = r"^(a+)+$"
# 好的做法(明确限定)
pattern = r"^a{1,10}$"
7. 在线工具
- Regex101:https://regex101.com/
- Regexr:https://regexr.com/
- RegExr:https://www.regexpal.com/
总结
正则表达式学习曲线较陡,但一旦掌握,将成为你文本处理的得力助手。建议:
- 从简单开始:先掌握基础元字符
- 多练习:使用在线工具验证正则
- 注意性能:避免过度回溯
- 添加注释:复杂正则添加说明
📖 推荐资源:《精通正则表达式》(Jeffrey Friedl著)是正则领域的经典之作。在线练习推荐RegexOne。