Araneida · 笔记

用 Python 把日志统计写短一点

2026-06-05 · Python

统计日志里最常见的需求是"按某个字段数一数",以前习惯写一堆字典和 try/except,其实标准库已经做完了。

import re
from collections import Counter, defaultdict

line = re.compile(r"^(?P<ip>\S+) .*? (?P<status>\d{3}) (?P<size>\d+)$")

status = Counter()
bytes_by_ip = defaultdict(int)
for row in open("/var/log/app/access.log", errors="ignore"):
    m = line.match(row)
    if not m:
        continue
    status[m["status"]] += 1
    bytes_by_ip[m["ip"]] += int(m["size"])

print(status.most_common(5))

Counter 负责计数,defaultdict(int) 负责累加,都不用再判断键是否存在。 数据量大到内存放不下时,再把 Counter 换成逐行写盘、最后用 SQL 聚合,思路是一样的。

正则不要写得过于严格,日志格式会悄悄变化(多一个字段、少一个空格),匹配失败直接跳过比抛异常好。 统计脚本跑在日志机上,输出给别的程序用的话,尽量保持一行一个结果的格式,别打印表格。

← 返回笔记列表