生产级网页爬虫工具链:侦察 → 生成 → 审计,一条命令走到底。
市面上爬虫框架要么太重(Scrapy 要学整个框架),要么太轻(手写 requests+BS4 每次都要重复造轮子)。这套工具取中间态:脚本即产物,生成即可跑,不用学框架。
- 全自动侦察 —
recon.py一键输出页面结构、API 端点、WAF 检测、加载机制,不用手动 curl 试探 - 参数化生成 —
scaffold.py根据侦察结果自动选模板,一条命令生成完整脚本,覆盖静态/API/动态/登录 6 种场景 - 6 大件出厂自带 — 反检测、Set 去重、侧边栏排除、随机延迟、断点续爬、数据校验,生成即生产
- 反检测审计 —
anti_detection_checker.py扫描脚本的反爬漏洞,PASS/WARN/FAIL 三级判定
pip install playwright requests beautifulsoup4
playwright install chromium
npm install -g @playwright/cli@latest # 需 Node.js 18+# 1. 侦察
python scripts/recon.py --json https://example.com > recon.json
# 2. 生成脚本(自动选模板 + 填入选择器)
python scripts/scaffold.py --recon recon.json --selectors ".title,.date" --output scraper.py
# 3. 运行
python scraper.py可选:跑一遍反检测审计确认脚本没有反爬漏洞。
python scripts/anti_detection_checker.py --file scraper.py| 场景 | 引擎 | 命令 |
|---|---|---|
| 静态 HTML 页面 | requests + BS4 | --template 1 |
| JSON API(无需 session) | requests | --template 2 |
| JSON API(需浏览器 session) | Playwright + fetch | --template 2B |
| 无限滚动 / 动态加载 | Playwright 滚动抓取 | --template 3 |
| 列表 → 逐个详情页 | Playwright 详情页遍历 | --template 4 |
| 需要登录的站点 | Playwright + auth.json | --template 5 |
scaffold.py --recon recon.json 会自动选择正确模板,无需手动判断。
web-crawler-skill/
├── SKILL.md # AI 使用指南
├── scripts/
│ ├── recon.py # 页面侦察
│ ├── scaffold.py # 脚本生成器
│ └── anti_detection_checker.py # 反检测审计
└── references/
├── code_templates.md # 模板选用速查
├── anti_crawl_strategies.md # 反爬对抗策略
├── auth-and-sessions.md # 登录态管理
├── testing_playbook.md # 页面诊断手册
└── common_pitfalls.md # 常见陷阱与解法