First Commit
ci / go (push) Waiting to run
ci / go-db (agent) (push) Waiting to run
ci / go-db (config) (push) Waiting to run
ci / go-db (db) (push) Waiting to run
ci / go-db (evidence) (push) Waiting to run
ci / go-db (llmrec) (push) Waiting to run
ci / go-db (server) (push) Waiting to run
detections / detections (push) Waiting to run
web / web (push) Waiting to run
docs / links (push) Canceled after 0s
ci / go (push) Waiting to run
ci / go-db (agent) (push) Waiting to run
ci / go-db (config) (push) Waiting to run
ci / go-db (db) (push) Waiting to run
ci / go-db (evidence) (push) Waiting to run
ci / go-db (llmrec) (push) Waiting to run
ci / go-db (server) (push) Waiting to run
detections / detections (push) Waiting to run
web / web (push) Waiting to run
docs / links (push) Canceled after 0s
This commit is contained in:
@@ -0,0 +1,425 @@
|
||||
---
|
||||
name: api-recon
|
||||
description: 收集网站API接口时调用该skill。
|
||||
---
|
||||
|
||||
# API Recon(前端接口侦察)
|
||||
|
||||
在**已授权**前提下,尽可能完整地发现:**后端 API**(路径、方法、参数、响应体)、**前端路由**、**UI 功能触发点**(Tab、弹窗、表格操作等)。
|
||||
|
||||
---
|
||||
|
||||
## 边界与禁止(Agent 必读 · 违反即越界)
|
||||
|
||||
本 skill **仅做 API / 参数面侦察**,不是漏洞挖掘或渗透利用阶段。
|
||||
|
||||
### 任务边界
|
||||
|
||||
| 范围 | 允许 | 禁止 |
|
||||
|---|---|---|
|
||||
| **目标** | 枚举 path、method、参数、路由、UI 触发点 | SQLi/XSS/越权/爆破/fuzz 漏洞、改包攻击、破坏性操作 |
|
||||
| **鉴权** | Hook + stub/mock 绕过**客户端**登录门 | 向用户索要或猜测账号密码;尝试真实登录表单提交 |
|
||||
| **运行时** | 无凭据下 hook 接口,用 mock 响应让 SPA 进入登录后壳层 | 依赖真实后端会话才能继续的流程 |
|
||||
|
||||
### 无凭据动态分析(Phase 3 默认)
|
||||
|
||||
1. 通过 `preload.js` / `runtime_harvest.js` **拦截并 stub** 登录、权限、菜单等 bootstrap 接口;
|
||||
2. 对业务查询接口返回 **结构正确、业务码成功、数据可为空** 的 mock body;
|
||||
3. 使前端在无后端或 401 环境下仍能渲染登录后页面,从而触发更多 XHR/fetch/WebSocket;
|
||||
4. **空数据、空白表格、占位 UI 均属预期**——勿为此转向真实登录或漏洞测试。
|
||||
|
||||
**一句话**:用 mock 撑开前端路由与组件挂载,**只录 outbound 请求**;后端返回什么不重要,重要的是前端**还会发哪些接口**。
|
||||
|
||||
### 流程硬禁止
|
||||
|
||||
| 禁止 | 替代做法 |
|
||||
|---|---|
|
||||
| Phase 1 完成前 grep/curl/Read 主 entry `index-*.js` 提取 API path | 跑 `OUTDIR/harvest_static.py` |
|
||||
| 手写 `extract_apis.py` 等替代 harvest 的脚本 | 改 `OUTDIR/harvest_static.py` 后重跑 |
|
||||
| 同一 grep/命令失败 ≥2 次仍重复 | 换策略:读 tool_logs、改 harvest、查 reference |
|
||||
| 跳过门禁 A/B,直接跑 `scripts/` 原版 | 复制到 OUTDIR 并按目标改 |
|
||||
| 真实用户名/密码、OTP、OAuth 等鉴权 | stub/mock(见上文) |
|
||||
| 以「拿真实数据」为由跳过 stub,做越权/注入测试 | 只录 outbound,属 recon 边界 |
|
||||
| 删除、导出敏感数据、批量写等不可逆操作 | coverage 点击亦同 |
|
||||
| 未完成 runtime + 动态枚举,声称已获全部页面和接口 | 见「完成定义」或标注局限 |
|
||||
| 未完成参数触发矩阵 + diff,声称已掌握全部参数 | Phase 3b 矩阵 + Phase 5 diff |
|
||||
| 用单一 runtime 样本推断必填/可选 | 多样本 diff 或校验规则/错误反推 |
|
||||
|
||||
---
|
||||
|
||||
## 两层模型 + 运行模式
|
||||
|
||||
| 层 | 产出 | 上限 |
|
||||
|---|---|---|
|
||||
| **静态**(JS bundle) | 全量 endpoint 路径、路由草案、组包点字段候选 | 无 HTTP 方法;参数须 Phase 1b;漏掉运行时拼接 URL |
|
||||
| **运行时**(活会话) | 方法 + body + 响应 + 动态 URL + WS/SSE;多样本 diff 补全参数 | 页面须实际渲染才会发请求;单样本不足以定必填/可选 |
|
||||
|
||||
| 运行模式 | 引擎 | 适用 |
|
||||
|---|---|---|
|
||||
| **depth** | `runtime_harvest.js`(Puppeteer) | API 清单、METHOD/params/响应体、WS/SSE、可复现批量跑 |
|
||||
| **coverage** | browser + `preload.js` | 点 Tab/弹窗/表格,功能点覆盖更深 |
|
||||
| **both** | 先 depth 再 coverage | 最完整,耗时最长 |
|
||||
|
||||
**参数方法论**(无通用脚本):path 用 harvest/正则;参数用 **锚点扩窗 + UI 绑定链 + 多样本 diff + 错误反推**(grep 配方见 [reference.md](reference.md) J 节)。
|
||||
|
||||
---
|
||||
|
||||
## 完成定义
|
||||
|
||||
全部满足方可声称 recon 完成:
|
||||
|
||||
- [ ] **静态**:Phase 1 harvest 产出 `api_static.txt`、`routes.txt`、`js/`
|
||||
- [ ] **运行时**:至少 depth 或 coverage 之一;coverage/both 须 **Hook 生效 + 动态枚举环**
|
||||
- [ ] **进壳**:访问业务 path 时非 `/login`(注意 hash 路由)
|
||||
- [ ] **参数**:coverage/both 完成参数触发矩阵 + `param_samples.json`;Phase 5 合并 `params_merged.json`
|
||||
- [ ] **深度**(若模块页空白):Phase 4 权限树还原并重跑,直至出现 **module 级 API**(非仅 locale/bootstrap)
|
||||
- [ ] **交付**:Phase 5 产出齐全(见 Phase 5 产出表);`insert_assets` 写入服务与端点资产
|
||||
|
||||
---
|
||||
|
||||
## 脚本与门禁
|
||||
|
||||
`scripts/` 仅为参考模板,**禁止**直接跑原版并当最终结果。
|
||||
|
||||
**规则**:先读 → 按目标改 → 写入 `OUTDIR`(如 `recon/`)→ 记 `CHANGES.md`;不匹配则按方法论重写,只借结构。
|
||||
|
||||
| 门禁 | 何时 | 参考脚本 → OUTDIR 副本 | 常见必改项 |
|
||||
|---|---|---|---|
|
||||
| **A(静态)** | Phase 0 后、**第一次**跑 harvest/spider 前 | `harvest_static.py` / `spider_mpa.py` | **多数站点默认 regex 可直接跑**;仅 manifest/方言不匹配时改 endpoint 正则、webpack/Vite `publicPath`、MPA exclude/cookie |
|
||||
| **B(运行时)** | Phase 2 后、跑 depth/coverage 前 | `runtime_harvest.js` / `preload.js` + `config.json` | Cookie/localStorage 键、neutralize 成功值、stubs、login 正则、api 前缀、hash/history |
|
||||
|
||||
**SPA 强制顺序**(不可交换;Phase 编号优先于「先探索再脚本」):
|
||||
|
||||
| 步骤 | 必须 | 禁止 |
|
||||
|---|---|---|
|
||||
| Phase 0 完成后 | 下一条 Bash = `python3 OUTDIR/harvest_static.py <URL> OUTDIR` | curl/grep/Read 主 entry `index-*.js`(通常 >500KB) |
|
||||
| 门禁 A | 复制脚本 → 按需小改 → **立刻运行** | 先手工提取 API 再决定是否 harvest |
|
||||
| Phase 1 完成前 | `wc -l` 校验产出;404 改 harvest 重试 | 手写 extract 脚本;对未下载 URL 反复 grep |
|
||||
| Phase 1b 起 | grep 仅 `OUTDIR/js/*.js` | 用主 bundle 代替 harvest |
|
||||
|
||||
- ✅ 复制 `harvest_static.py` → (可选)改 regex → **立即运行**
|
||||
- ❌ curl 主 bundle → grep 多次 → 写临时 extract → 最后才 harvest
|
||||
- **MPA**:Phase 0 后下一条 Bash = `python3 OUTDIR/spider_mpa.py ...`
|
||||
|
||||
---
|
||||
|
||||
## 工具与输出约束
|
||||
|
||||
| 约束 | 说明 |
|
||||
|---|---|
|
||||
| 大文件 | >100KB 的 `index-*.js` **禁止** Read/grep 进上下文;用 OUTDIR 脚本批处理 |
|
||||
| grep 输出 | 必须 `\| head -20` 或 `-m 5`;对话只保留 path 摘要,勿贴 bundle 片段 |
|
||||
| 校验 | 用 `wc -l`、`ls \| wc -l`;勿 Read 整目录 |
|
||||
| regex 初探 | 可选、≤1 次、仅 ≤50KB 小 chunk 或 HTML;正式静态以 harvest 为准 |
|
||||
| reference | 配方/模板/排障见 [reference.md](reference.md),勿重复 inline 全文 |
|
||||
|
||||
---
|
||||
|
||||
## 执行路线图
|
||||
|
||||
```
|
||||
Phase 0 分类 + OUTDIR
|
||||
→ 门禁 A → Phase 1 harvest(★ 立刻运行 ★)
|
||||
→ Phase 1b 参数逆向
|
||||
→ Phase 2 鉴权三道门 → config.json
|
||||
→ 门禁 B → Phase 3 运行时 + 参数矩阵
|
||||
→ Phase 4 权限树(必要时)→ 重跑 Phase 3
|
||||
→ Phase 5 合并报告 + insert_assets批量插入所有发现的服务、端点api资产,无论如何插入时不允许漏掉已发现的资产
|
||||
```
|
||||
|
||||
按序勾选;**前一项未完成不得进入下一 Phase**。
|
||||
|
||||
1. [ ] **Phase 0**:初探 SPA/MPA;创建 `OUTDIR` → [Phase 0](#phase-0--分类)
|
||||
2. [ ] **门禁 A + Phase 1**:复制脚本 → **立刻** harvest → `wc -l` 校验 → [Phase 1](#phase-1--静态)
|
||||
3. [ ] **Phase 1b**:锚点扩窗 + 绑定层 → `param_candidates.json` → [Phase 1b](#phase-1b--参数逆向)
|
||||
4. [ ] **Phase 2**:鉴权三道门 → `config.json` → [Phase 2](#phase-2--鉴权三道门)
|
||||
5. [ ] **门禁 B**:调整 runtime 脚本 → [Phase 3](#phase-3--运行时)
|
||||
6. [ ] **Phase 3**:depth / coverage / both;确认进壳;参数触发矩阵 → `param_samples.json`
|
||||
7. [ ] **Phase 4**(若需要):权限树 → patch stubs → 重跑 Phase 3 → [Phase 4](#phase-4--权限树还原)
|
||||
8. [ ] **Phase 5**:合并产出 + 报告 + `insert_assets` → [Phase 5](#phase-5--合并与报告)
|
||||
|
||||
---
|
||||
|
||||
## Phase 0 — 分类
|
||||
|
||||
拉取入口 HTML,**创建 `OUTDIR`**(勿改 skill 内 `scripts/`):
|
||||
|
||||
- **SPA**:空壳 + `<div id=app>` + chunk → Phase 1–5
|
||||
- **MPA**:SSR + `<form>`、无 endpoint bundle → 门禁 A 后:
|
||||
|
||||
```bash
|
||||
python3 recon/spider_mpa.py <BASE_URL> <OUTDIR> [--cookie "session=..."] [--max 300] [--depth 5] [--exclude "logout|delete|destroy"]
|
||||
```
|
||||
|
||||
产出 `forms.txt`、`links.txt`、`api_inline.txt`。SPA 若 forms ≈ 0 → 切 Phase 1。
|
||||
|
||||
---
|
||||
|
||||
## Phase 1 — 静态
|
||||
|
||||
遵守 [脚本与门禁](#脚本与门禁) · [工具与输出约束](#工具与输出约束)。
|
||||
|
||||
```bash
|
||||
python3 recon/harvest_static.py <BASE_URL> <OUTDIR>
|
||||
```
|
||||
|
||||
harvest:解析 HTML script → webpack/Vite manifest → 下载全部 lazy chunk → 产出 `js/`、`api_static.txt`、`routes.txt`、`chunkmap.txt`。
|
||||
|
||||
```bash
|
||||
wc -l OUTDIR/api_static.txt OUTDIR/routes.txt
|
||||
ls OUTDIR/js | wc -l
|
||||
```
|
||||
|
||||
- chunk 数 vs manifest:404 须改 harvest 重试,勿手工 curl 逐个 chunk
|
||||
- `api_static.txt` 过少 → 放宽 OUTDIR 内 endpoint 正则后重跑(见 reference)
|
||||
|
||||
### Phase 1b — 参数逆向
|
||||
|
||||
path 来自 Phase 1;参数字段须单独 recon。grep 规则见 [工具与输出约束](#工具与输出约束)。
|
||||
|
||||
**完成标准**:重要接口能答——字段名、传输位置、类型推断、是否必填、样本值、置信度。
|
||||
|
||||
#### 1b.0 — 传输形态
|
||||
|
||||
| 形态 | 参数在哪 | 静态优先看 |
|
||||
|---|---|---|
|
||||
| REST JSON | body + query | path 锚点旁 `(params\|data\|body)\s*:\s*\{` |
|
||||
| GraphQL | `variables` | gql 模板、`$page: Int` |
|
||||
| 传统 form | urlencoded | `<form>`、`FormData` |
|
||||
| 文件上传 | multipart | `FormData.append` |
|
||||
| 路径参数 | `/user/:id` | 路由表 + `useParams` / `$route.params` |
|
||||
| 加密/签名 | 包进 `sign`/`data` | Hook 加密函数入参(reference D 节) |
|
||||
|
||||
产出:每接口标注 `transport: query|json|form|graphql|encrypted`。
|
||||
|
||||
#### 1b.1 — 锚点扩窗
|
||||
|
||||
以已知 path 为锚,扩窗口找组包对象:
|
||||
|
||||
```bash
|
||||
grep -n '"/api/user/list"' OUTDIR/js/*.js | head -20
|
||||
grep -rhoaE '.{0,120}("/api[^"]+").{0,200}' OUTDIR/js/*.js | head -20
|
||||
grep -rhoaE '(params|data|body|payload)\s*:\s*\{' OUTDIR/js/*.js | head -20
|
||||
```
|
||||
|
||||
| 包装层 | 参数线索 |
|
||||
|---|---|
|
||||
| axios 实例 | `data` / `params` |
|
||||
| 统一 request | 拦截器注入全局字段 |
|
||||
| OpenAPI 客户端 | 生成 method 签名 |
|
||||
| React Query / SWR | hook 第二参数 |
|
||||
| Vue composable | composable 入参 |
|
||||
|
||||
类型残留:`yup`/`zod`/rules、`Form.Item name=`、内嵌 Swagger。
|
||||
|
||||
→ `param_candidates.json`:`{ path, fields[], source: "static-callsite", confidence }`
|
||||
|
||||
#### 1b.2 — 绑定层
|
||||
|
||||
```
|
||||
Form field → onFinish/handleSubmit → transform → API payload
|
||||
```
|
||||
|
||||
| 绑定源 | 手法 |
|
||||
|---|---|
|
||||
| 表单 submit | 跟 submit → transform → API |
|
||||
| 表格搜索 | `getFieldsValue()` → `params` |
|
||||
| 路由 | `:id` / `?tab=` |
|
||||
| 拦截器 | 全局 `tenantId`、分页、sign |
|
||||
| 枚举 select | `options` → API 枚举值 |
|
||||
|
||||
DevTools call stack 从 `fetch`/`XHR.send` 往上追组包函数。
|
||||
|
||||
#### 1b.3 — 组包三问(≠ Phase 2 鉴权三门)
|
||||
|
||||
| 问 | 要答什么 |
|
||||
|---|---|
|
||||
| **组装** | payload 在哪 build、transform 痕迹 |
|
||||
| **校验** | required、pattern、enum |
|
||||
| **传输** | path / query / body / multipart / 头 |
|
||||
|
||||
拦截器门(Phase 2)顺带读全局注入字段(Authorization、`X-Tenant-Id`、sign)。
|
||||
|
||||
#### 1b.4 — 与 Phase 3 衔接
|
||||
|
||||
候选字段来自静态/绑定层;**必填/可选/条件依赖**须 Phase 3 参数矩阵 + diff + Phase 5 错误反推。
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 — 鉴权三道门
|
||||
|
||||
在 `OUTDIR/js/` grep(带 `head`),写入 `config.json`(配方见 reference):
|
||||
|
||||
| 门 | 问题 | 关键词 |
|
||||
|---|---|---|
|
||||
| **渲染门** | 如何判断已登录? | `isLogin`、`getToken`、Cookie/localStorage |
|
||||
| **拦截器门** | 什么触发跳 `/login`? | `response_code`、`errno`、axios interceptor |
|
||||
| **内容门** | 菜单/权限从哪来? | `menu`、`permission`、`role`、`acl`、`routes` |
|
||||
|
||||
禁止把 localStorage 键名当凭据——须从 chunk/请求链确认。
|
||||
|
||||
**出口 = 门禁 B**:结论落到 `config.json`,并改 `OUTDIR/runtime_harvest.js` / `preload.js`。
|
||||
|
||||
### Phase 2b — API 观察(可选)
|
||||
|
||||
用 OUTDIR 内 `preload.js` 确认会话键名、Authorization、嵌套 API URL:
|
||||
|
||||
| 配置 | 产出 |
|
||||
|---|---|
|
||||
| `recordDetail: true` | `__API_RECON_DETAIL__` |
|
||||
| `observe.xhrHeaders: true` | headers 观察 |
|
||||
| `extractUrlsFromResponse: true` | 响应内子 API |
|
||||
| `observe.storageReads/cookieReads: true` | 回填 config |
|
||||
| `neutralizeVueRouter: true` | `__API_RECON_ROUTES__` |
|
||||
|
||||
coverage 每轮导出:`__API_RECON_LOG__`、`__API_RECON_DETAIL__`、`__API_RECON_ROUTES__`、`__API_RECON_OBSERVE__`。
|
||||
|
||||
---
|
||||
|
||||
## Phase 3 — 运行时
|
||||
|
||||
须已过门禁 B;遵守 [边界与禁止](#边界与禁止agent-必读--违反即越界) · 无凭据 mock 策略。
|
||||
|
||||
`config.json` 设置 `"runtimeMode": "depth" | "coverage" | "both"`(模板见 reference)。
|
||||
|
||||
### Hook 与 stub(depth + coverage 共用)
|
||||
|
||||
| 层 | 范围 | 目的 |
|
||||
|---|---|---|
|
||||
| L1 精确 | auth/权限/bootstrap stub | 过首屏鉴权 |
|
||||
| L2 负向修正 | 所有 JSON 响应 | 未登录码 → 成功 |
|
||||
| L3 兜底 | 未命中 L1 的 `/api` 等 | 空成功体,撑开 UI |
|
||||
|
||||
- **depth**:fake auth + `forward` 改业务码 + `stubs`;遍历 `routes`(hash/history);产出 `runtime_api.json`
|
||||
- **coverage**:**document-start** 注入 `preload.js`(CDP `addScriptToEvaluateOnNewDocument` 或 Userscript)
|
||||
|
||||
验证:`window.__API_RECON_PRELOAD__` 存在;业务 path 不回 `/login`。
|
||||
|
||||
```bash
|
||||
cd recon && npm install
|
||||
node runtime_harvest.js config.json
|
||||
```
|
||||
|
||||
### 3b — coverage 动态枚举(必做)
|
||||
|
||||
1. 主导航/侧栏 — 每项点击,等网络 1–3s
|
||||
2. Tab — `role=tab`、`.ant-tabs-tab`
|
||||
3. 表格 — 首行查看/编辑/详情
|
||||
4. 工具栏 — 导出、筛选、新建(**避免不可逆删除**)
|
||||
5. 每进模块 — 合并 API/路由
|
||||
6. SPA — 对 `routes.txt` 未覆盖 path 受控 `pushState`(MPA 禁止)
|
||||
|
||||
**参数触发矩阵**(必做):每模块按操作类型各录一次,**diff 多样本**:
|
||||
|
||||
| 操作 | 通常多出的参数 |
|
||||
|---|---|
|
||||
| 列表首屏 | 分页 + 默认筛选 |
|
||||
| 点搜索 | keyword、filter |
|
||||
| 高级筛选 | 更多 optional |
|
||||
| 新建/编辑 | 完整 entity |
|
||||
| 批量/导出/排序 | `ids[]`、`exportType`、`sortField` |
|
||||
|
||||
**stub 下 outbound body/headers 仍真实**——以请求为准。录制 → `scan_raw.json`、`param_samples.json`、`api_detail.json`。
|
||||
|
||||
- **Vue**:`neutralizeVueRouter: true` + document-start preload
|
||||
- **React**:`routes.txt` + 侧栏点击 + `pushState`
|
||||
- **both**:先 3a depth,再 3b coverage
|
||||
|
||||
---
|
||||
|
||||
## Phase 4 — 权限树还原
|
||||
|
||||
**触发**:模块页空白 / 每路由仅 bootstrap(如 locale)→ 内容门未过。
|
||||
|
||||
| 现象 | 含义 |
|
||||
|---|---|
|
||||
| 进壳成功 | 渲染门 + 拦截器门已过 |
|
||||
| 侧栏缺项/点击空白 | stub shape 或权限码不全 |
|
||||
| 每路由 API 相同且极少 | `v-if permission` 未通过 |
|
||||
| `routes.txt` 远少于 bundle | 须从 auth 模块补全 |
|
||||
|
||||
```bash
|
||||
grep -rhoaE '"/api[^"]*(permission|perm|role|menu|acl)[^"]*"' OUTDIR/js/*.js | sort -u | head -30
|
||||
grep -rhoaE 'userRouteAuth|getResultTree|routeMap|routeLink|menuList|authList' OUTDIR/js/*.js | head -20
|
||||
```
|
||||
|
||||
典型链:`role_permissions`(flat codes)+ `permissions/all`(tree)→ `getResultTree` → `userRouteAuth[CODE].url`。
|
||||
|
||||
```bash
|
||||
python3 recon/extract_route_map.py recon/js recon/
|
||||
python3 recon/build_perm_tree.py recon/js recon/ --config recon/config.json
|
||||
```
|
||||
|
||||
中间产出:`route_map.json`、`userRouteAuth.json`、`permissions_tree.json`、`*_stub.json`、`perm_codes_all.txt`。
|
||||
|
||||
stub 检查:外层 `response_code` 与拦截器门一致;flat codes 与 tree 对齐;`routes` 覆盖 `route_map` 全部 link。
|
||||
|
||||
更新 `config.json` 后**重跑 Phase 3**。大型 SPA 可调 `waitUntil`、`routeTimeout`、`perRouteMs`(见 reference A3/I 节)。
|
||||
|
||||
---
|
||||
|
||||
## Phase 5 — 合并与报告
|
||||
|
||||
### 产出表
|
||||
|
||||
| 文件 | 阶段 | 内容 |
|
||||
|---|---|---|
|
||||
| `js/`、`api_static.txt`、`routes.txt`、`chunkmap.txt` | 1 | 静态 bundle 与 path |
|
||||
| `param_candidates.json` | 1b | 静态参数字段候选 |
|
||||
| `config.json` | 2 | 三道门 + runtime 配置 |
|
||||
| `runtime_api.json` | 3a | depth 详细录制(含 WS/SSE) |
|
||||
| `param_samples.json`、`scan_raw.json`、`api_detail.json` | 3b | 多样本、点击日志、detail |
|
||||
| `route_map.json` 等 | 4 | 权限树中间文件(若执行) |
|
||||
| `params_merged.json` | 5 | 合并参数字段 + 置信度 |
|
||||
| `api_merged.txt` | 5 | `METHOD /path [params] [static\|runtime\|both]` |
|
||||
| `site_map.json` | 5 | 路由、API、params、功能点、局限 |
|
||||
| **insert_assets** | 5 | 将所有服务、端点资产写入资产库 |
|
||||
|
||||
### 5b — 参数合并
|
||||
|
||||
从 `param_samples.json` diff,**无通用合并脚本**。置信度规则见 reference J7(高/中/低/待触发)。
|
||||
|
||||
### 5c — 错误反推
|
||||
|
||||
授权范围内可发不完整请求读 400(**属参数 recon,非漏洞测试**):`field 'x' is required`、枚举错误等。注意 `data` 包装、`variables`、加密前 `bizData`。
|
||||
|
||||
报告须注明:runtimeMode、静态/运行时 API 数、参数置信度、未覆盖模块、相对参考脚本的 `CHANGES.md` 摘要。
|
||||
|
||||
`site_map.json` 建议结构:
|
||||
|
||||
```json
|
||||
{
|
||||
"site": "https://example.com",
|
||||
"runtimeMode": "both",
|
||||
"appType": "vue-spa",
|
||||
"routeGuardStrategy": ["nav-neutralize", "L1-auth", "L2-patch", "forward"],
|
||||
"apisFromStatic": [],
|
||||
"apisFromRuntime": [],
|
||||
"apis": [],
|
||||
"params": [{ "method": "POST", "path": "/api/user/list", "transport": "json", "fields": [] }],
|
||||
"frontendRoutes": [],
|
||||
"routesVerifiedByClick": [],
|
||||
"featuresTriggered": [],
|
||||
"limitations": ""
|
||||
}
|
||||
```
|
||||
|
||||
更多字段与 grep 配方见 [reference.md](reference.md)。
|
||||
|
||||
---
|
||||
|
||||
## 通用说明
|
||||
|
||||
- **框架无关**:webpack/Vite/Angular lazy load 方法相同
|
||||
- **传输**:REST/JSON、GraphQL、WebSocket、SSE;gRPC-web 不在范围
|
||||
- **SSR**:客户端 fetch 可录;RSC/Server Actions 不完全可枚举
|
||||
- **盲区**:JSVMP、WASM、HMAC/mTLS 强校验 → 静态 + 标注局限
|
||||
- **参数盲区**:条件联动、hidden params、WASM 组包 → 「待触发」/「不可达」
|
||||
- **静态是安全网**:runtime 被挡时静态仍能枚举 endpoint
|
||||
|
||||
---
|
||||
|
||||
## 附加资源
|
||||
|
||||
- Grep 配方、`config.json` 模板、排障、Hook、参数逆向 J 节、site_map 模板:**[reference.md](reference.md)**
|
||||
- 参考脚本路径见 [脚本与门禁](#脚本与门禁) 表
|
||||
@@ -0,0 +1,500 @@
|
||||
# api-recon — 参考手册
|
||||
|
||||
Grep 配方、`config.json` 模板与排障。所有 grep 针对 `js/` 目录执行。bundle 单行时可先 `js-beautify` 或 `sed 's/}/}\n/g'`,通常带上下文窗口的 raw grep 即可。
|
||||
|
||||
## 脚本说明
|
||||
|
||||
`scripts/` 内所有文件均为**参考模板**,执行前必须按目标站点调整。典型改动点:
|
||||
|
||||
| 脚本 | 常见需调整项 |
|
||||
|---|---|
|
||||
| `harvest_static.py` | endpoint 正则、webpack/Vite manifest 解析、微前端 publicPath、重试/并发 |
|
||||
| `runtime_harvest.js` | neutralize 字段名与成功值、stub 匹配规则与 body 结构、routes 来源、WS 录制、`waitUntil`/`routeTimeout`/`proxy` |
|
||||
| `preload.js` | `loginPathRe`、L1 stubs、`neutralize.fields`、`apiPattern`、是否启用 L3、`recordDetail`、`observe.*`、`neutralizeVueRouter` |
|
||||
| `spider_mpa.py` | `--exclude` 破坏性链接、cookie、depth/max、同域过滤 |
|
||||
| `extract_route_map.py` | `routeMap` / `routeLink` 正则、KEY 命名模式 |
|
||||
| `build_perm_tree.py` | `userRouteAuth` 解析、`ROOTS`/`PREFIX_PARENT` 层级启发式、stub 外层字段名 |
|
||||
| `config.json` | 以上全部站点专属参数的统一入口 |
|
||||
|
||||
调整后的文件建议放在任务工作目录(如 `recon/`),报告中注明相对参考脚本的具体改动。
|
||||
|
||||
---
|
||||
|
||||
## A. 逆向三道门
|
||||
|
||||
### A1. 渲染门 — 「如何判断已登录?」
|
||||
|
||||
```bash
|
||||
grep -rhoaE '.{0,40}(isLogin|isAuthenticated|loggedIn|hasLogin|requireAuth)\b.{0,80}' js | head
|
||||
grep -rhoaE 'function (getUser|getToken|getAuth)[0-9]?\([^)]*\)\{.{0,200}' js | head
|
||||
grep -rhoaE '(localStorage|sessionStorage)\.getItem\("[^"]+"\)' js | sort -u
|
||||
grep -rhoaE '(Cookies?|cookie)\.(get|load)\("[^"]+"\)' js | sort -u
|
||||
grep -rhoaE '\batob\(|JSON\.parse\(|jwt|decode' js | head
|
||||
```
|
||||
|
||||
找链路 `isLogin = f(getUser())` → `getUser = decode(storage.read(KEY))`,确定 **存储键**、**容器**(Cookie vs localStorage)、**编码**:
|
||||
|
||||
| 编码 | config 伪造方式 |
|
||||
|---|---|
|
||||
| 明文字符串 / `"1"` / token | `"value": "anything-truthy"` |
|
||||
| `JSON.parse(x)` | `"value": "json:{\"id\":1,\"username\":\"admin\"}"` |
|
||||
| `JSON.parse(atob(x))` | `"value": "b64json:{\"id\":1,\"username\":\"admin\"}"` |
|
||||
| JWT | 无签名/`alg:none` JWT,或 bundle 内密钥签名 |
|
||||
| 加密(SM2/AES/RSA) | 找硬编码密钥;渲染门仅需可解码 blob 时可 forge;否则静态兜底 |
|
||||
|
||||
→ 写入 `cookies` / `localStorage`。
|
||||
|
||||
### A2. 拦截器门 — 「什么触发跳 /login?」
|
||||
|
||||
```bash
|
||||
grep -rhoaE '.{0,60}(interceptors\.response|axios|request\.use).{0,120}' js | head
|
||||
grep -rhoaE '.{0,40}(response_code|errcode|errno|\bcode\b|\bret\b|\bstatus\b)\s*[=!]==?\s*[\-0-9]{1,4}.{0,60}' js | head -20
|
||||
grep -rhoaE '.{0,40}(未登录|请重新登录|登录已过期|unauthorized|登录失效|授权|token.{0,10}invalid).{0,40}' js | head
|
||||
grep -rhoaE '.{0,30}(location\.href|router\.(push|replace)|navigate)\([^)]*login[^)]*\)' js | head
|
||||
```
|
||||
|
||||
确定:**字段名**、**成功值**(通常 `0` 或 `200`)、**触发跳转的失败值**。用 junk session 验证:
|
||||
|
||||
```bash
|
||||
curl -sk -X POST -H 'Cookie: <fakekey>=junk' https://target/api/<protected> -d '{}' -H 'Content-Type: application/json'
|
||||
```
|
||||
|
||||
→ 写入 `neutralize.fields` + `neutralize.success`。
|
||||
|
||||
### A3. 内容门 — 「菜单/权限从哪来?」
|
||||
|
||||
```bash
|
||||
grep -rhoaE '"/api[^"]*(permission|perm|role|menu|acl|resource|nav)[^"]*"' js | sort -u
|
||||
grep -rhoaE '.{0,30}(menus|permissions|menuList|routeList|authList|role_permissions)\b.{0,120}' js | head
|
||||
grep -rhoaE 'userRouteAuth|getResultTree|routeMap|routeLink|hasPermission|checkAuth' js | head
|
||||
grep -rhoaE '([A-Z_][A-Z0-9_]*):\{name:"[^"]*",link:"/[^"]+"\}' js | head
|
||||
```
|
||||
|
||||
**两层数据**(常见企业后台):
|
||||
|
||||
| API | 典型 payload | 消费方 |
|
||||
|---|---|---|
|
||||
| `.../role_permissions` | `{ permissions: string[], role_type }` | 路由守卫、按钮级 ACL |
|
||||
| `.../permissions/all` | `tree[{ code, position, children }]` | 侧栏菜单渲染 |
|
||||
| bundle 内 `userRouteAuth` | `{ CODE: { url, name? } }` | code → 前端 path |
|
||||
| bundle 内 `routeMap` | `{ KEY: { name, link } }` | 别名解析(webpack `o.DASHBOARD`) |
|
||||
|
||||
读消费方代码确认:`getResultTree(tree, permissions)` 如何过滤、`v-if` / `hasAuth(code)` 检查哪个字段。
|
||||
|
||||
**手工 forge**(小站点):构建 permissive payload → `stubs`。
|
||||
|
||||
**完整权限树还原**(大站点,侧栏/子模块仍空白):见 **I 节**。
|
||||
|
||||
---
|
||||
|
||||
## B. config.json 模板
|
||||
|
||||
```json
|
||||
{
|
||||
"baseUrl": "https://target/",
|
||||
"runtimeMode": "both",
|
||||
"chromium": "/usr/bin/chromium",
|
||||
|
||||
"cookies": [
|
||||
{ "name": "auth", "value": "b64json:{\"id\":1,\"username\":\"admin\",\"role\":\"admin\",\"func\":{},\"permissions\":[\"*\"]}" }
|
||||
],
|
||||
"localStorage": { "token": "faketoken", "isLogin": "1" },
|
||||
|
||||
"neutralize": {
|
||||
"fields": ["response_code", "code", "errno", "ret", "status"],
|
||||
"success": 0,
|
||||
"flags": { "success": true, "message": "ok" }
|
||||
},
|
||||
"forward": true,
|
||||
"loginUrlPattern": "/login",
|
||||
"apiPattern": "/api/|/rest/|/graphql",
|
||||
|
||||
"mockTier": "L1+L2",
|
||||
"recordDetail": true,
|
||||
"observe": {
|
||||
"storageReads": false,
|
||||
"cookieReads": false,
|
||||
"xhrHeaders": true
|
||||
},
|
||||
"neutralizeVueRouter": true,
|
||||
"stubs": [
|
||||
{
|
||||
"match": "permissions/all|/menu|role_permissions",
|
||||
"body": {
|
||||
"response_code": 0, "code": 0,
|
||||
"data": {
|
||||
"permissions": ["*"],
|
||||
"menus": [
|
||||
{ "name": "dashboard", "path": "/dashboard", "show": true, "children": [] },
|
||||
{ "name": "alert", "path": "/alert", "show": true, "children": [] }
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
],
|
||||
|
||||
"explore": {
|
||||
"clickTabs": true,
|
||||
"clickTables": true,
|
||||
"pushStateFallback": true,
|
||||
"maxMenuItems": 50
|
||||
},
|
||||
|
||||
"routes": ["/dashboard", "/alert", "/asset", "/device", "/report", "/config", "/system"],
|
||||
"waitMs": 1500, "perRouteMs": 900, "headless": true,
|
||||
"waitUntil": "domcontentloaded",
|
||||
"routeTimeout": 12000,
|
||||
"proxy": "",
|
||||
|
||||
"captureResponses": true, "recordWs": true, "respMax": 600
|
||||
}
|
||||
```
|
||||
|
||||
字段说明:
|
||||
- `runtimeMode`:`depth`(Puppeteer)、`coverage`(browser MCP)、`both`
|
||||
- `cookies[].value` 前缀:`b64json:` → base64(JSON);`json:` → 原始 JSON;无前缀 → 字面量
|
||||
- `forward: true` 转发真实请求并改写码字段;`false` 完全离线 stub
|
||||
- `mockTier`:coverage 模式 preload 启用层级,如 `L1+L2`、`L1+L2+L3`
|
||||
- `routes` 来自 `routes.txt`;forge 菜单后 harness 自动追加 `<a href>`
|
||||
- `captureResponses` / `recordWs` 仅 depth 模式有效
|
||||
- `waitUntil`:大型 SPA 用 `domcontentloaded`,避免 `networkidle2` 挂起
|
||||
- `routeTimeout`:单路由 `page.goto` 超时(毫秒)
|
||||
- `proxy`:Puppeteer `--proxy-server`;也可设 `HTTP_PROXY` / `HTTPS_PROXY`
|
||||
|
||||
### B1. 双 stub 模板(role_permissions + permissions/all)
|
||||
|
||||
```json
|
||||
"stubs": [
|
||||
{
|
||||
"match": "role_permissions",
|
||||
"body": {
|
||||
"response_code": 0,
|
||||
"data": {
|
||||
"permissions": ["MONITOR", "MONITOR_ALERT", "THREAT", "ASSETS_RISK"],
|
||||
"role_type": "SUPER_ADMIN"
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"match": "permissions/all",
|
||||
"body": {
|
||||
"response_code": 0,
|
||||
"data": [
|
||||
{
|
||||
"code": "MONITOR",
|
||||
"position": 1,
|
||||
"children": [
|
||||
{ "code": "MONITOR_ALERT", "position": 1, "children": [] }
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
外层字段名(`response_code` / `code` / `data`)须与 A2 拦截器门一致;`permissions` 须覆盖 tree 中所有 leaf code。
|
||||
|
||||
---
|
||||
|
||||
## C. coverage 模式:preload 配置
|
||||
|
||||
编辑 `scripts/preload.js` 顶部 `CONFIG` 对象,或通过 CDP 注入前替换:
|
||||
|
||||
```javascript
|
||||
const CONFIG = {
|
||||
loginPathRe: /\/(login|signin)(\/|$|\?)/i,
|
||||
mockTier: 'L1+L2',
|
||||
forward: true,
|
||||
recordDetail: true,
|
||||
extractUrlsFromResponse: true,
|
||||
neutralizeVueRouter: true,
|
||||
observe: { storageReads: false, cookieReads: false, xhrHeaders: true },
|
||||
neutralize: { fields: ['response_code', 'code'], success: 0 },
|
||||
stubs: [ /* 同 config.json stubs */ ],
|
||||
apiPattern: /\/(api|apis|v\d+|dev|internal|graphql)\//i,
|
||||
};
|
||||
```
|
||||
|
||||
验证:`window.__API_RECON_PRELOAD__ === true` 且 pathname 稳定。
|
||||
|
||||
导出录制结果:
|
||||
|
||||
```javascript
|
||||
JSON.stringify({
|
||||
apis: [...window.__API_RECON_LOG__],
|
||||
detail: window.__API_RECON_DETAIL__,
|
||||
routes: [...(window.__API_RECON_ROUTES__ || [])],
|
||||
observe: window.__API_RECON_OBSERVE__,
|
||||
}, null, 2)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## D. preload / runtime Hook 能力
|
||||
|
||||
preload(coverage)与 runtime_harvest(depth)内置的浏览器 Hook 能力及覆盖范围:
|
||||
|
||||
| Hook 能力 | 对 API 发现的价值 | 覆盖 |
|
||||
|---|---|---|
|
||||
| Hook fetch / XHR.open | 录请求 URL/方法 | ✅ `recordDetail` + `__API_RECON_LOG__` |
|
||||
| Hook XHR.setRequestHeader | 发现 Authorization 等头 | ✅ `observe.xhrHeaders` |
|
||||
| Hook localStorage/cookie 读 | 确认会话键名 | ⚠️ 可选 `observe.storageReads/cookieReads` |
|
||||
| Vue 获取路由 | 补全 frontendRoutes | ✅ `__API_RECON_ROUTES__`(已加载路由) |
|
||||
| Vue 路由守卫中和 / 登录跳转阻断 | 撑开模块触发 API | ✅ `neutralizeVueRouter` + 原生跳转中和 |
|
||||
| React 获取路由 | 补路由 | ⚠️ 静态 + 点击;无专用 Hook |
|
||||
| 页面跳转阻断(登录 path) | 留页分析 | ⚠️ 仅阻断登录 path,避免挡业务导航 |
|
||||
| Hook 加密库(CryptoJS/SM 等) | 加密参数 → 明文 API body | ❌ 须手工 Hook 加密函数入参;结论写 config |
|
||||
| 反调试 bypass | 否则 runtime 录不到 API | ❌ 须手工处理;静态仍可用 |
|
||||
|
||||
---
|
||||
|
||||
## E. Endpoint 提取正则(静态过少时)
|
||||
|
||||
在 `harvest_static.py` 的 `extract_endpoints` 放宽,或手动:
|
||||
|
||||
```bash
|
||||
grep -rhoaE '"/[a-z][A-Za-z0-9_/\-]{3,}"' js | sort -u
|
||||
grep -rhoaE '/api/[a-zA-Z0-9_./-]+' js | sort -u
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## F. 排障
|
||||
|
||||
| 现象 | 原因 → 处理 |
|
||||
|---|---|
|
||||
| 静态 API 很少 | endpoint 方言不匹配 → 放宽正则(D 节) |
|
||||
| chunk 数 ≪ manifest | CSS-only 或未部署 chunk;404 已重试 |
|
||||
| runtime 仍显示登录页 | 渲染门错误 → 复查 A1:键名、容器、编码、domain |
|
||||
| 进壳但模块空白 | 内容门 → forge 菜单(A3);`routes` path 可能不对 |
|
||||
| 每路由只有 bootstrap/locale | 权限码不全 → I 节权限树还原;检查 `role_permissions` + `permissions/all` 双 stub |
|
||||
| 侧栏有项但子页空白 | tree 缺 intermediate 节点或 code 与 `userRouteAuth` 不一致 |
|
||||
| 每个 API 都跳登录 | 拦截器门 → 确认 `neutralize`;嵌套字段需扩展 walk 逻辑 |
|
||||
| WS 帧为 0 | 需用户交互后才 subscribe;加长 `perRouteMs` |
|
||||
| 响应体空 | 仅 `forward: true` 时有真实响应 |
|
||||
| Chromium 缺失 | 安装 chromium 或设置 `config.chromium` / `CHROMIUM` |
|
||||
| Mock 很多仍回登录 | Hook 太晚或缺 `location.href` setter → document-start + preload |
|
||||
| 列表全空 | L3 空数组正常;继续点 Tab/设置/详情 |
|
||||
| 误把 Redux action 当路由 | 过滤含 get/set/change/clear/toggle/upload 的内部 path |
|
||||
| Vue 仍跳登录 | preload 非 document-start → 改注入时机;或 `neutralizeVueRouter: false` 时手动清守卫 |
|
||||
| 响应里有 URL 但未进 log | 开 `extractUrlsFromResponse`;或从 `__API_RECON_DETAIL__` 人工提取 |
|
||||
| 不知 Authorization 头名 | 开 `observe.xhrHeaders` 或 DevTools 查看请求头 |
|
||||
| runtime 极慢 / 超时 | 改 `waitUntil: domcontentloaded`;降 `routeTimeout`;勿用 `networkidle2` |
|
||||
| 代理连接失败 | 检查 `proxy` / 环境变量;Puppeteer 与 curl 代理端口一致 |
|
||||
|
||||
---
|
||||
|
||||
## G. hardened 目标
|
||||
|
||||
服务端逐步校验会话(不可 forge 的签名 cookie、服务端渲染且不可 stub 的菜单)时,runtime 会在 shell 处卡住。预期行为:
|
||||
|
||||
- **静态足够做 endpoint 枚举** — 模块 path 在代码里
|
||||
- 若授权允许,用**真实会话**跑同一 harness:`forward: true`、无需 neutralize,捕获真实 methods/params/responses
|
||||
|
||||
---
|
||||
|
||||
## H. 单次任务清单
|
||||
|
||||
1. 确认授权范围
|
||||
2. **阅读** `scripts/harvest_static.py` → 按目标调整 → 运行 → 审 `api_static.txt`、`routes.txt`
|
||||
3. **Phase 1b**:path 锚点扩窗 + 绑定层 → `param_candidates.json`(J 节)
|
||||
4. 逆向 A1/A2/A3 → 写站点专属 `config.json`
|
||||
5. **阅读并调整** `runtime_harvest.js` / `preload.js` 后再执行
|
||||
6. `runtimeMode=depth`:`npm install` → 运行调整后的 harvest 脚本
|
||||
7. `runtimeMode=coverage/both`:document-start 注入调整后的 preload → browser MCP 动态枚举 + **参数触发矩阵**
|
||||
8. 模块不渲染 → **I 节权限树还原** → patch stubs → 重跑
|
||||
9. 参数多样本 diff + 错误反推 → `params_merged.json`
|
||||
10. 合并 → `site_map.json` + `api_merged.txt`,诚实标注覆盖、缺口及脚本改动点
|
||||
|
||||
---
|
||||
|
||||
## I. 权限树还原(Phase 4 深化)
|
||||
|
||||
当 forge 简单 `menus: [{ path, show: true }]` 无效、子模块仍不 mount 时使用。
|
||||
|
||||
### I1. 定位 auth 模块
|
||||
|
||||
```bash
|
||||
grep -l 'userRouteAuth' js/*.js
|
||||
grep -l 'routeMap\|routeLink' js/*.js
|
||||
grep -rhoaE 'getResultTree|role_permissions|permissions/all' js | head
|
||||
```
|
||||
|
||||
记录:**权限 API path**、**响应字段名**、**消费 chunk 文件名**。
|
||||
|
||||
### I2. 提取 routeMap
|
||||
|
||||
```bash
|
||||
python3 scripts/extract_route_map.py recon/js recon/
|
||||
# 产出 recon/route_map.json
|
||||
```
|
||||
|
||||
若 `[!] no routeMap pattern found`:放宽 `extract_route_map.py` 中正则,或手工 grep:
|
||||
|
||||
```bash
|
||||
grep -rhoaE '([A-Z_][A-Z0-9_]*):\{name:"[^"]*",link:"/[^"]+"\}' js | head -20
|
||||
```
|
||||
|
||||
### I3. 构建权限树 + stub
|
||||
|
||||
```bash
|
||||
python3 scripts/build_perm_tree.py recon/js recon/ --config recon/config.json
|
||||
```
|
||||
|
||||
脚本逻辑:
|
||||
1. 解析 `userRouteAuth={MONITOR:{url:...},...}`(含 webpack 别名 `He=o.DASHBOARD`)
|
||||
2. 用 `route_map.json` 解析 alias → 真实 path
|
||||
3. 按 code 前缀推断 parent(`MONITOR_ALERT` → `MONITOR`)
|
||||
4. 输出 `permissions_tree.json`、`permissions_all_stub.json`、`role_permissions_stub.json`
|
||||
5. `--config` 时自动写入 `config.json` 的 `stubs` 与扩展 `routes`
|
||||
|
||||
**按目标调整**(在脚本顶部):
|
||||
- `DEFAULT_ROOTS`:顶级模块 code 列表
|
||||
- `DEFAULT_PREFIX_PARENT`:`PREFIX_` → parent 映射
|
||||
- `DEFAULT_EXTRA_PARENT`:非前缀关系的 orphan 节点
|
||||
|
||||
### I4. 校验 stub 一致性
|
||||
|
||||
```bash
|
||||
# permissions 数量应 ≈ userRouteAuth 条目数
|
||||
wc -l recon/perm_codes_all.txt
|
||||
# routes 应覆盖 route_map 全部 link
|
||||
python3 -c "import json; m=json.load(open('recon/route_map.json')); r=set(json.load(open('recon/config.json'))['routes']); print('missing', [v['link'] for v in m.values() if v['link'] not in r])"
|
||||
```
|
||||
|
||||
### I5. 重跑 runtime 并对比
|
||||
|
||||
```bash
|
||||
node recon/runtime_harvest.js recon/config.json
|
||||
# 对比 forge 前后 runtime_api.json 条数;检查 /attack、/asset 等是否出现模块 API
|
||||
```
|
||||
|
||||
| forge 前 | forge 后(成功) |
|
||||
|---|---|
|
||||
| 每路由相同 3–5 条 bootstrap | 不同路由触发不同 module API |
|
||||
| 仅 `/api/locale/language` | 出现 `/api/web/...` 模块 endpoint |
|
||||
| `routes.txt` 个位路由 | `routes` 80–110+ 来自 route_map |
|
||||
|
||||
### I6. 仍失败时
|
||||
|
||||
- **coverage 模式**:点击侧栏 + Tab,权限 gating 可能在交互后才请求
|
||||
- **stub 字段**:对比真实 API(curl + 真实 session)与 stub 的 nesting
|
||||
- **额外守卫**:grep `hasPermission|checkRole|func.` 等按钮级检查,扩展 `role_permissions.permissions`
|
||||
- **静态兜底**:模块 API path 仍在 `api_static.txt`,runtime 仅补 METHOD/body;参数保留 `param_candidates.json` + 已录样本
|
||||
|
||||
---
|
||||
|
||||
## J. 参数逆向(Phase 1b / 5b / 5c)
|
||||
|
||||
**方法论,非通用脚本。** 找 path 用正则;找参数用锚点扩窗 + UI 绑定链 + 多样本 diff + 错误反推。
|
||||
|
||||
### J1. 锚点扩窗 — 从 path 找组包对象
|
||||
|
||||
```bash
|
||||
# 以 Phase 1 已知 path 为锚
|
||||
grep -n '"/api/user/list"' js/*.js
|
||||
grep -rhoaE '.{0,120}("/api[^"]+").{0,200}' js | head
|
||||
grep -rhoaE '(params|data|body|payload)\s*:\s*\{' js | head
|
||||
grep -rhoaE '(get|post|put|delete|patch)\([^,]+,\s*\{' js | head
|
||||
```
|
||||
|
||||
### J2. 包装层与传输形态
|
||||
|
||||
```bash
|
||||
# axios / 统一 request
|
||||
grep -rhoaE '(axios|request)\.(get|post|put|delete|patch)\(' js | head
|
||||
grep -rhoaE 'interceptors\.(request|response)' js | head
|
||||
|
||||
# GraphQL
|
||||
grep -rhoaE '(query|mutation)\s+\w+|gql`|graphql\(' js | head
|
||||
grep -rhoaE '\$[a-zA-Z_]+\s*:\s*(Int|String|Boolean|\[)' js | head
|
||||
|
||||
# FormData / multipart
|
||||
grep -rhoaE 'FormData|\.append\(' js | head
|
||||
|
||||
# 路径参数
|
||||
grep -rhoaE 'path:\s*"/[^"]*:[^"]+"' js | head
|
||||
grep -rhoaE 'useParams|route\.params|\$route\.params' js | head
|
||||
```
|
||||
|
||||
### J3. 校验门 — 必填 / 格式 / 枚举
|
||||
|
||||
```bash
|
||||
grep -rhoaE '(required|message|pattern|enum|validator)\s*:' js | head
|
||||
grep -rhoaE 'yup\.|zod\.|async-validator|Form\.Item|a-form-item|el-form-item' js | head
|
||||
grep -rhoaE 'rules\s*:\s*\[|name:\s*["\'][a-zA-Z_]+["\']' js | head
|
||||
grep -rhoaE 'label.*value|options\s*:\s*\[' js | head
|
||||
```
|
||||
|
||||
### J4. 绑定层 — 表单 → API
|
||||
|
||||
```bash
|
||||
grep -rhoaE 'onFinish|handleSubmit|getFieldsValue|validateFields' js | head
|
||||
grep -rhoaE '(pick|omit|transform|dayjs|moment)\(' js | head
|
||||
```
|
||||
|
||||
runtime 补位:DevTools → Network → 请求 → **发起程序**(call stack)从 `fetch`/`send` 往上追组包函数。
|
||||
|
||||
### J5. 加密参数
|
||||
|
||||
```bash
|
||||
grep -rhoaE 'encrypt|decrypt|sign|CryptoJS|sm2|sm3|sm4|RSA|AES' js | head
|
||||
```
|
||||
|
||||
**勿在密文上猜字段** — Hook 加密函数**入参**,在加密前录 plaintext payload;结论写 `config.json` / `param_candidates.json`。
|
||||
|
||||
### J6. 参数触发矩阵(Phase 3 必做)
|
||||
|
||||
对每模块按操作各录一次,diff 请求 body/query:
|
||||
|
||||
| 操作 | 关注 |
|
||||
|---|---|
|
||||
| 列表首屏 | 分页默认值 |
|
||||
| 搜索 | keyword、filters |
|
||||
| 高级筛选 | optional 字段 |
|
||||
| 新建/编辑 | 完整 entity |
|
||||
| 批量/导出 | `ids[]`、`exportType` |
|
||||
| 排序/翻页 | `sortField`、`order` |
|
||||
|
||||
产出 `param_samples.json`:`[{ "path", "method", "action": "search", "body", "query", "headers" }]`
|
||||
|
||||
### J7. 置信度规则
|
||||
|
||||
| 置信度 | 条件 |
|
||||
|---|---|
|
||||
| **高** | 静态 callsite + runtime ≥2 样本一致 |
|
||||
| **中** | 仅静态,或仅 1 次 runtime |
|
||||
| **低** | 响应/错误反推,未二次验证 |
|
||||
| **待触发** | 静态已知字段,UI/权限未跑到 |
|
||||
|
||||
### J8. 场景快配
|
||||
|
||||
| 场景 | 顺序 |
|
||||
|---|---|
|
||||
| REST 列表页 | J1 组包对象 → J6 四次 diff → J3 rules |
|
||||
| 新建/编辑表单 | J3 Form name → J4 submit 链 → runtime 提交 + 故意留空看 400 |
|
||||
| GraphQL | J2 variables 声明 → runtime 各 operation 录 variables |
|
||||
| 加密 body | J5 Hook 入参 → 加密前字段即真实 params |
|
||||
|
||||
### J9. 与 api-recon 阶段映射
|
||||
|
||||
| api-recon | 参数 recon |
|
||||
|---|---|
|
||||
| Phase 1 静态 | J1 锚点扩窗 |
|
||||
| Phase 2 A2 拦截器 | 全局注入字段(tenantId、sign) |
|
||||
| Phase 3 runtime | J6 触发矩阵 + `param_samples.json` |
|
||||
| Phase 4 权限树 | 不同模块表单不同 → 权限够才触发全字段 |
|
||||
| Phase 5 合并 | `params_merged.json` + 置信度;勿单样本定必填 |
|
||||
|
||||
### J10. 排障
|
||||
|
||||
| 现象 | 处理 |
|
||||
|---|---|
|
||||
| 静态有字段名 runtime 从未出现 | 标注「待触发」;补权限树 / 点高级筛选 / 联动 select 各 option |
|
||||
| 同 path 不同 body 形状 | 正常 — 按 `action` 分条记录,勿强行合并 schema |
|
||||
| stub 响应假但想看 params | **看 outbound 请求** body/headers,勿从 stub 响应反推 |
|
||||
| 400 报 nested field | 注意外层包装 `data`/`bizData`/`variables` |
|
||||
| GraphQL 只见 operation 名 | 展开 `variables` JSON;静态找 `$var: Type` |
|
||||
|
||||
---
|
||||
@@ -0,0 +1,210 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
build_perm_tree.py <JSDIR> <OUTDIR> [--config recon/config.json]
|
||||
|
||||
Rebuild a permissive permission/menu stub from frontend auth modules.
|
||||
|
||||
Typical consumer chain (TDP / enterprise admin SPAs):
|
||||
POST /api/web/user/role_permissions -> { permissions: [...], role_type }
|
||||
POST /api/web/permissions/all -> tree[{ code, position, children }]
|
||||
getResultTree(tree, permissions) -> menu include lists
|
||||
userRouteAuth[code].url -> frontend route path
|
||||
|
||||
This script:
|
||||
1. Locates userRouteAuth={MONITOR:{url:...},...} in js/
|
||||
2. Resolves webpack alias refs (He=o.DASHBOARD) via route_map.json
|
||||
3. Infers hierarchy from code prefixes (MONITOR_ -> MONITOR)
|
||||
4. Writes permissions_tree.json, permissions_all_stub.json,
|
||||
role_permissions_stub.json, userRouteAuth.json
|
||||
5. Optionally patches config.json stubs (permissions/all + role_permissions)
|
||||
|
||||
Adjust ROOTS / PREFIX_PARENT / EXTRA_PARENT per target if heuristics miss nodes.
|
||||
"""
|
||||
import re, json, os, sys, glob, argparse
|
||||
|
||||
DEFAULT_ROOTS = [
|
||||
'MONITOR', 'THREAT', 'ASSETS_RISK', 'INVESTIGATION', 'MANAGEMENT',
|
||||
'AGENT_EVIDENCE', 'MDR', 'PLATFORM',
|
||||
]
|
||||
DEFAULT_PREFIX_PARENT = {
|
||||
'MONITOR_': 'MONITOR', 'THREAT_': 'THREAT', 'ASSETS_RISK_': 'ASSETS_RISK',
|
||||
'INVESTIGATION_': 'INVESTIGATION', 'MANAGEMENT_': 'MANAGEMENT', 'MDR_': 'MDR',
|
||||
'PLATFORM_': 'PLATFORM', 'CUSTOM_': 'PLATFORM', 'FALSE_': 'PLATFORM',
|
||||
}
|
||||
DEFAULT_EXTRA_PARENT = {
|
||||
'LINKAGE_DISPOSAL': 'MANAGEMENT',
|
||||
'ASSETS_RISK_API': 'ASSETS_RISK',
|
||||
'ASSETS_RISK_WEAK_PWD': 'ASSETS_RISK',
|
||||
}
|
||||
|
||||
|
||||
def find_auth_file(jsdir):
|
||||
best_fp, best_len = None, 0
|
||||
for fp in glob.glob(os.path.join(jsdir, '*.js')):
|
||||
try:
|
||||
txt = open(fp, encoding='utf-8', errors='ignore').read()
|
||||
except Exception:
|
||||
continue
|
||||
if 'userRouteAuth' not in txt:
|
||||
continue
|
||||
m = re.search(r'userRouteAuth=\{MONITOR:', txt)
|
||||
if m and len(txt) > best_len:
|
||||
best_fp, best_len = fp, len(m.group(0))
|
||||
elif 'userRouteAuth' in txt and best_fp is None:
|
||||
best_fp = fp
|
||||
return best_fp
|
||||
|
||||
|
||||
def parse_aliases(chunk):
|
||||
aliases = {}
|
||||
for m in re.finditer(r'([a-zA-Z_$][\w$]*)=o\.([A-Z_0-9]+)\b', chunk[:8000]):
|
||||
aliases[m.group(1)] = m.group(2)
|
||||
return aliases
|
||||
|
||||
|
||||
def resolve_ref(token, aliases, route_map):
|
||||
token = token.strip()
|
||||
if token.startswith('"'):
|
||||
return json.loads(token)
|
||||
if token in aliases:
|
||||
key = aliases[token]
|
||||
return route_map.get(key, {}).get('link', key)
|
||||
return token
|
||||
|
||||
|
||||
def parse_user_route_auth(txt, route_map):
|
||||
m = re.search(r'(?:t\.)?userRouteAuth=(\{MONITOR:.*?\})\},', txt, re.S)
|
||||
if not m:
|
||||
m = re.search(r'userRouteAuth=(\{[A-Z_0-9]+:\{url:', txt)
|
||||
if not m:
|
||||
return None, {}
|
||||
# greedy fallback — trim at next webpack module
|
||||
body = m.group(1)
|
||||
end = body.rfind('}')
|
||||
obj_src = body[: end + 1] if end > 0 else body
|
||||
else:
|
||||
obj_src = m.group(1)
|
||||
|
||||
chunk_start = txt.find('userRouteAuth')
|
||||
chunk = txt[chunk_start:chunk_start + 35000]
|
||||
aliases = parse_aliases(chunk)
|
||||
|
||||
entries = {}
|
||||
for em in re.finditer(
|
||||
r'([A-Z_0-9]+):\{url:([^,}]+)(?:,control:(\[.*?\]|[^,}]+))?\}', obj_src
|
||||
):
|
||||
key = em.group(1)
|
||||
url = resolve_ref(em.group(2).strip(), aliases, route_map)
|
||||
controls = []
|
||||
if em.group(3):
|
||||
raw = em.group(3).strip()
|
||||
vars_ = re.findall(r'([A-Za-z_$][\w$]*)', raw) if raw.startswith('[') else [raw]
|
||||
controls = [aliases.get(v, v) for v in vars_]
|
||||
entries[key] = {'code': key, 'url': url, 'control': controls}
|
||||
return obj_src, entries
|
||||
|
||||
|
||||
def parent_of(code, roots, prefix_parent, extra_parent):
|
||||
if code in extra_parent:
|
||||
return extra_parent[code]
|
||||
if code in roots:
|
||||
return None
|
||||
for pref, par in prefix_parent.items():
|
||||
if code.startswith(pref):
|
||||
return par
|
||||
return None
|
||||
|
||||
|
||||
def build_tree(entries, roots, prefix_parent, extra_parent):
|
||||
children_of = {k: [] for k in entries}
|
||||
for code in entries:
|
||||
p = parent_of(code, roots, prefix_parent, extra_parent)
|
||||
if p:
|
||||
children_of.setdefault(p, []).append(code)
|
||||
|
||||
def make_node(code):
|
||||
node = {
|
||||
'code': code,
|
||||
'position': 'top' if code in roots else 'left',
|
||||
'name': code,
|
||||
}
|
||||
kids = sorted(children_of.get(code, []))
|
||||
if kids:
|
||||
node['children'] = [make_node(c) for c in kids]
|
||||
return node
|
||||
|
||||
tree = [make_node(r) for r in roots if r in entries or children_of.get(r)]
|
||||
orphans = [c for c in entries if parent_of(c, roots, prefix_parent, extra_parent) is None and c not in roots]
|
||||
for code in sorted(orphans):
|
||||
tree.append(make_node(code))
|
||||
return tree
|
||||
|
||||
|
||||
def flat_codes(nodes):
|
||||
out = []
|
||||
for n in nodes:
|
||||
out.append(n['code'])
|
||||
out.extend(flat_codes(n.get('children', [])))
|
||||
return out
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser(description='Build permission tree stubs from JS auth modules')
|
||||
ap.add_argument('jsdir')
|
||||
ap.add_argument('outdir')
|
||||
ap.add_argument('--config', help='patch stubs into config.json')
|
||||
ap.add_argument('--role', default='SUPER_ADMIN', help='role_type in role_permissions stub')
|
||||
args = ap.parse_args()
|
||||
os.makedirs(args.outdir, exist_ok=True)
|
||||
|
||||
route_map_path = os.path.join(args.outdir, 'route_map.json')
|
||||
if not os.path.exists(route_map_path):
|
||||
print('[*] route_map.json missing — run extract_route_map.py first')
|
||||
route_map = {}
|
||||
else:
|
||||
route_map = json.load(open(route_map_path, encoding='utf-8'))
|
||||
|
||||
auth_fp = find_auth_file(args.jsdir)
|
||||
if not auth_fp:
|
||||
print('[!] userRouteAuth module not found in js/')
|
||||
sys.exit(1)
|
||||
print(f'[*] auth module: {os.path.basename(auth_fp)}')
|
||||
txt = open(auth_fp, encoding='utf-8', errors='ignore').read()
|
||||
_, entries = parse_user_route_auth(txt, route_map)
|
||||
if not entries:
|
||||
print('[!] failed to parse userRouteAuth object — adjust regex in script')
|
||||
sys.exit(1)
|
||||
|
||||
tree = build_tree(entries, DEFAULT_ROOTS, DEFAULT_PREFIX_PARENT, DEFAULT_EXTRA_PARENT)
|
||||
all_codes = sorted(set(flat_codes(tree) + [c for e in entries.values() for c in e.get('control', [])]))
|
||||
|
||||
perm_all = {'response_code': 0, 'verbose_msg': 'ok', 'data': tree}
|
||||
role_perm = {
|
||||
'response_code': 0,
|
||||
'verbose_msg': 'ok',
|
||||
'data': {'role_type': args.role, 'permissions': all_codes},
|
||||
}
|
||||
|
||||
json.dump(entries, open(os.path.join(args.outdir, 'userRouteAuth.json'), 'w', encoding='utf-8'), ensure_ascii=False, indent=2)
|
||||
json.dump(tree, open(os.path.join(args.outdir, 'permissions_tree.json'), 'w', encoding='utf-8'), ensure_ascii=False, indent=2)
|
||||
open(os.path.join(args.outdir, 'perm_codes_all.txt'), 'w', encoding='utf-8').write('\n'.join(all_codes))
|
||||
json.dump(perm_all, open(os.path.join(args.outdir, 'permissions_all_stub.json'), 'w', encoding='utf-8'), ensure_ascii=False, indent=2)
|
||||
json.dump(role_perm, open(os.path.join(args.outdir, 'role_permissions_stub.json'), 'w', encoding='utf-8'), ensure_ascii=False, indent=2)
|
||||
|
||||
print(f'[+] entries={len(entries)} codes={len(all_codes)} tree_roots={len(tree)}')
|
||||
|
||||
cfg_path = args.config or os.path.join(args.outdir, 'config.json')
|
||||
if os.path.exists(cfg_path):
|
||||
cfg = json.load(open(cfg_path, encoding='utf-8'))
|
||||
stubs = [s for s in cfg.get('stubs', []) if not re.search(r'permissions/all|role_permissions', s.get('match', ''))]
|
||||
stubs = [
|
||||
{'match': 'permissions/all', 'body': perm_all},
|
||||
{'match': 'role_permissions', 'body': role_perm},
|
||||
] + stubs
|
||||
cfg['stubs'] = stubs
|
||||
json.dump(cfg, open(cfg_path, 'w', encoding='utf-8'), ensure_ascii=False, indent=2)
|
||||
print(f'[+] patched stubs -> {cfg_path}')
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,42 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
extract_route_map.py <JSDIR> <OUTDIR>
|
||||
|
||||
Scan downloaded JS chunks for routeMap / routeLink style objects:
|
||||
KEY:{name:"...",link:"/path"}
|
||||
|
||||
Writes route_map.json — used by build_perm_tree.py to resolve alias refs.
|
||||
"""
|
||||
import re, json, os, sys, glob
|
||||
|
||||
def main():
|
||||
if len(sys.argv) < 3:
|
||||
print("usage: extract_route_map.py <JSDIR> <OUTDIR>")
|
||||
sys.exit(1)
|
||||
jsdir, outdir = sys.argv[1], sys.argv[2]
|
||||
os.makedirs(outdir, exist_ok=True)
|
||||
|
||||
best = {}
|
||||
best_file = None
|
||||
pat = re.compile(r'([A-Z_][A-Z0-9_]*):\{name:"([^"]*)",link:"([^"]+)"')
|
||||
|
||||
for fp in glob.glob(os.path.join(jsdir, '*.js')):
|
||||
try:
|
||||
txt = open(fp, encoding='utf-8', errors='ignore').read()
|
||||
except Exception:
|
||||
continue
|
||||
hits = pat.findall(txt)
|
||||
if len(hits) > len(best):
|
||||
best = {k: {'name': n, 'link': l} for k, n, l in hits}
|
||||
best_file = fp
|
||||
|
||||
if not best:
|
||||
print('[!] no routeMap pattern found — widen regex or grep manually')
|
||||
sys.exit(1)
|
||||
|
||||
out = os.path.join(outdir, 'route_map.json')
|
||||
json.dump(best, open(out, 'w', encoding='utf-8'), ensure_ascii=False, indent=2)
|
||||
print(f'[+] {len(best)} routes from {os.path.basename(best_file)} -> {out}')
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,207 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
harvest_static.py <BASE_URL> <OUTDIR>
|
||||
|
||||
参考模板 — 非通用成品。执行前须按目标站点调整,常见改动:
|
||||
- extract_endpoints 正则(endpoint 方言)
|
||||
- webpack/Vite manifest 解析逻辑
|
||||
- 微前端 publicPath、重试策略
|
||||
|
||||
Static SPA bundle harvester. Framework-agnostic; tuned for webpack + Vite.
|
||||
1. Fetch entry HTML, collect script/module references.
|
||||
2. Parse the runtime chunk manifest(s) and download EVERY chunk (not just the
|
||||
ones in HTML), looping until no new chunk ids appear. Handles multiple
|
||||
micro-frontend runtimes, each with its own publicPath.
|
||||
3. Extract API endpoints and route paths from all downloaded JS.
|
||||
|
||||
Stdlib only. TLS verification is disabled (recon against self-signed/internal hosts).
|
||||
"""
|
||||
import sys, os, re, ssl, json, urllib.request, urllib.parse
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
|
||||
CTX = ssl.create_default_context(); CTX.check_hostname = False; CTX.verify_mode = ssl.CERT_NONE
|
||||
UA = "Mozilla/5.0 (spa-api-recon)"
|
||||
|
||||
def fetch(url, binary=False):
|
||||
try:
|
||||
req = urllib.request.Request(url, headers={"User-Agent": UA})
|
||||
with urllib.request.urlopen(req, context=CTX, timeout=25) as r:
|
||||
data = r.read()
|
||||
return (data if binary else data.decode("utf-8", "ignore")), r.status
|
||||
except Exception as e:
|
||||
code = getattr(e, "code", 0)
|
||||
return "", code
|
||||
|
||||
def origin_of(u):
|
||||
p = urllib.parse.urlparse(u)
|
||||
return f"{p.scheme}://{p.netloc}"
|
||||
|
||||
# ---- chunk manifest parsing -------------------------------------------------
|
||||
def matched_block(s, end_brace_idx):
|
||||
"""Walk back from a '}' index to its matching '{' and return the object body."""
|
||||
depth = 0; j = end_brace_idx
|
||||
while j >= 0:
|
||||
if s[j] == '}': depth += 1
|
||||
elif s[j] == '{':
|
||||
depth -= 1
|
||||
if depth == 0: return s[j:end_brace_idx+1]
|
||||
j -= 1
|
||||
return ""
|
||||
|
||||
def parse_chunk_maps(js_text):
|
||||
"""Return list of (publicPath_hint, {id: hash}) for every `}[x]+".js"` map
|
||||
(webpack __webpack_require__.u) found in the text."""
|
||||
maps = []
|
||||
# publicPath hints in this file: .p="..." or publicPath="..."
|
||||
pubs = re.findall(r'(?:\.p|publicPath)\s*=\s*"([^"]*)"', js_text)
|
||||
for m in re.finditer(r'\}\[[A-Za-z_$]\]\s*\+\s*"\.js"', js_text):
|
||||
body = matched_block(js_text, m.start())
|
||||
pairs = re.findall(r'(\d+):"([0-9a-fA-F]{6,16})"', body)
|
||||
if pairs:
|
||||
maps.append((pubs, dict(pairs)))
|
||||
# Vite style: __vite__mapDeps map of "assets/xx.js"
|
||||
for fn in re.findall(r'"(assets/[^"]+\.js)"', js_text):
|
||||
maps.append((["/"], {"__vite__": fn}))
|
||||
return maps
|
||||
|
||||
def chunk_urls(base, js_text):
|
||||
"""Yield absolute chunk URLs reconstructable from this file's manifest(s)."""
|
||||
origin = origin_of(base)
|
||||
out = set()
|
||||
for pubs, idmap in parse_chunk_maps(js_text):
|
||||
paths = pubs or ["/assets/", "/"]
|
||||
for cid, h in idmap.items():
|
||||
if cid == "__vite__":
|
||||
fn = h # already "assets/xx.js"
|
||||
for pp in paths:
|
||||
out.add(urllib.parse.urljoin(origin + "/", fn))
|
||||
continue
|
||||
for pp in paths:
|
||||
if pp.startswith("http"): bareorigin = ""; prefix = pp
|
||||
else: bareorigin = origin; prefix = pp if pp.startswith("/") else "/"+pp
|
||||
if not prefix.endswith("/"): prefix += "/"
|
||||
out.add(f"{bareorigin}{prefix}{cid}.{h}.js")
|
||||
return out
|
||||
|
||||
# ---- endpoint / route extraction -------------------------------------------
|
||||
API_HINT = re.compile(r'/(?:api|rest|service|services|gateway|graphql|v\d|web|admin|backend|open)\b', re.I)
|
||||
ASSET_EXT = re.compile(r'\.(js|css|png|jpe?g|svg|gif|woff2?|ttf|ico|map|json|mp4|webp)(\?|$)', re.I)
|
||||
|
||||
def extract_endpoints(js_text):
|
||||
paths = set()
|
||||
# quoted ('/...'), double-quoted, and backtick template paths
|
||||
for m in re.findall(r'''["'`](/[A-Za-z0-9_\-./{}$:]+)["'`]''', js_text):
|
||||
paths.add(m)
|
||||
# concatenation heads: "/api/x/" + var
|
||||
for m in re.findall(r'''["'](/[A-Za-z0-9_\-./]+/)["']\s*\+''', js_text):
|
||||
paths.add(m)
|
||||
api, other = set(), set()
|
||||
for p in paths:
|
||||
if ASSET_EXT.search(p): continue
|
||||
if p.count('/') < 2 and not API_HINT.search(p): continue
|
||||
(api if API_HINT.search(p) else other).add(p)
|
||||
return api, other
|
||||
|
||||
def extract_routes(js_text):
|
||||
r = set()
|
||||
for key in ('path', 'to', 'redirect', 'href'):
|
||||
for m in re.findall(key + r'''\s*:\s*["'](/[A-Za-z0-9_\-/:]*)["']''', js_text):
|
||||
if not ASSET_EXT.search(m) and not API_HINT.search(m):
|
||||
r.add(m)
|
||||
return r
|
||||
|
||||
# ---- main -------------------------------------------------------------------
|
||||
def main():
|
||||
if len(sys.argv) < 3:
|
||||
print("usage: harvest_static.py <BASE_URL> <OUTDIR>"); sys.exit(1)
|
||||
base, outdir = sys.argv[1], sys.argv[2]
|
||||
if not base.startswith("http"): base = "https://" + base
|
||||
jsdir = os.path.join(outdir, "js"); os.makedirs(jsdir, exist_ok=True)
|
||||
|
||||
print(f"[*] entry: {base}")
|
||||
html, status = fetch(base)
|
||||
open(os.path.join(outdir, "index.html"), "w").write(html)
|
||||
origin = origin_of(base)
|
||||
|
||||
# initial scripts from HTML
|
||||
srcs = set(re.findall(r'<script[^>]+src="([^"]+\.js[^"]*)"', html))
|
||||
srcs |= set(re.findall(r'(?:src|href)="([^"]*\.js)"', html))
|
||||
seed = set()
|
||||
for s in srcs:
|
||||
seed.add(s if s.startswith("http") else urllib.parse.urljoin(base, s))
|
||||
print(f"[*] {len(seed)} scripts referenced in HTML")
|
||||
|
||||
have = {} # url -> local path
|
||||
def dl(url):
|
||||
fn = os.path.basename(urllib.parse.urlparse(url).path)
|
||||
if not fn.endswith(".js"): return None
|
||||
lp = os.path.join(jsdir, fn)
|
||||
if url in have: return have[url]
|
||||
data, st = fetch(url, binary=True)
|
||||
if st == 200 and data and not data[:15].lstrip().startswith(b"<"):
|
||||
open(lp, "wb").write(data); have[url] = lp; return lp
|
||||
return None
|
||||
|
||||
with ThreadPoolExecutor(max_workers=20) as ex:
|
||||
list(ex.map(dl, seed))
|
||||
|
||||
# iteratively expand via chunk manifests (chunks reference more chunks)
|
||||
seen_urls = set(have.keys()); frontier = list(have.values())
|
||||
rounds = 0
|
||||
while frontier and rounds < 6:
|
||||
rounds += 1
|
||||
new_urls = set()
|
||||
for lp in frontier:
|
||||
try: txt = open(lp, encoding="utf-8", errors="ignore").read()
|
||||
except: continue
|
||||
for cu in chunk_urls(base, txt):
|
||||
if cu not in seen_urls: new_urls.add(cu)
|
||||
seen_urls |= new_urls
|
||||
if not new_urls: break
|
||||
print(f"[*] round {rounds}: {len(new_urls)} new chunk urls from manifest")
|
||||
before = set(have.values())
|
||||
with ThreadPoolExecutor(max_workers=24) as ex:
|
||||
list(ex.map(dl, new_urls))
|
||||
frontier = [p for p in have.values() if p not in before]
|
||||
|
||||
# retry-once any manifest chunk that 404'd (transient failures are real)
|
||||
all_manifest = set()
|
||||
for lp in list(have.values()):
|
||||
try: all_manifest |= chunk_urls(base, open(lp, encoding="utf-8", errors="ignore").read())
|
||||
except: pass
|
||||
missing = [u for u in all_manifest if u not in have]
|
||||
if missing:
|
||||
with ThreadPoolExecutor(max_workers=24) as ex:
|
||||
list(ex.map(dl, missing))
|
||||
still = [u for u in all_manifest if u not in have]
|
||||
print(f"[*] manifest chunks: {len(all_manifest)} | downloaded {len(have)} | "
|
||||
f"unreachable {len(still)} (CSS-only / undeployed)")
|
||||
|
||||
print(f"[+] total JS downloaded: {len(have)}")
|
||||
|
||||
# extract from everything
|
||||
api, other, routes = set(), set(), set()
|
||||
for lp in have.values():
|
||||
try: txt = open(lp, encoding="utf-8", errors="ignore").read()
|
||||
except: continue
|
||||
a, o = extract_endpoints(txt); api |= a; other |= o
|
||||
routes |= extract_routes(txt)
|
||||
|
||||
def dump(name, items):
|
||||
path = os.path.join(outdir, name)
|
||||
open(path, "w").write("\n".join(sorted(items)))
|
||||
return path
|
||||
dump("api_static.txt", api)
|
||||
dump("paths_other.txt", other)
|
||||
dump("routes.txt", routes)
|
||||
dump("chunkmap.txt", sorted(os.path.basename(u) for u in all_manifest))
|
||||
|
||||
print(f"[+] api endpoints: {len(api)} (api_static.txt)")
|
||||
print(f"[+] other paths: {len(other)} (paths_other.txt)")
|
||||
print(f"[+] route paths: {len(routes)} (routes.txt)")
|
||||
print(f"[+] outdir: {outdir}")
|
||||
print("\n[next] reverse the 3 gate facts (see reference.md), fill config.json, "
|
||||
"then: node runtime_harvest.js config.json")
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
+1011
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"name": "api-recon-runtime",
|
||||
"version": "1.0.0",
|
||||
"private": true,
|
||||
"description": "Headless runtime API harvester for the api-recon skill",
|
||||
"dependencies": {
|
||||
"puppeteer-core": "^23.11.1"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,348 @@
|
||||
/**
|
||||
* api-recon coverage 模式预加载脚本 — 参考模板
|
||||
*
|
||||
* ⚠ 非通用成品:须按目标站点调整后再注入。
|
||||
* 常见改动:loginPathRe、stubs、neutralize.fields/success、apiPattern、mockTier、forward
|
||||
*
|
||||
* document-start 注入(CDP addScriptToEvaluateOnNewDocument 或 userscript)
|
||||
* CONFIG 字段应与 recon/config.json 保持一致
|
||||
*/
|
||||
(function () {
|
||||
'use strict';
|
||||
|
||||
const CONFIG = {
|
||||
loginPathRe: /\/(login|signin)(\/|$|\?)/i,
|
||||
mockTier: 'L1+L2',
|
||||
forward: true,
|
||||
neutralize: {
|
||||
fields: ['response_code', 'code', 'errno', 'ret', 'status'],
|
||||
success: 0,
|
||||
flags: { success: true, message: 'ok' },
|
||||
},
|
||||
stubs: [],
|
||||
apiPattern: /\/(api|apis|v\d+|dev|internal|graphql)\//i,
|
||||
apiFallbackRe: /^\/(api|apis|v\d+|dev|internal)\//i,
|
||||
// API 发现增强:fetch/XHR Hook、请求头与响应录制
|
||||
recordDetail: true, // 详细录制 method/url/headers/body/响应
|
||||
respMax: 600,
|
||||
extractUrlsFromResponse: true, // 从 JSON 响应里抠嵌套 URL
|
||||
neutralizeVueRouter: true, // Vue beforeEach/push 登录跳转中和
|
||||
observe: {
|
||||
storageReads: false, // 观察 localStorage.getItem(辅助确认会话键名)
|
||||
cookieReads: false, // 观察 document.cookie 读取
|
||||
xhrHeaders: true, // 录制 XHR setRequestHeader
|
||||
},
|
||||
};
|
||||
|
||||
window.__API_RECON_PRELOAD__ = true;
|
||||
window.__API_RECON_LOG__ = window.__API_RECON_LOG__ || new Set();
|
||||
window.__API_RECON_DETAIL__ = window.__API_RECON_DETAIL__ || [];
|
||||
window.__API_RECON_ROUTES__ = window.__API_RECON_ROUTES__ || new Set();
|
||||
window.__API_RECON_OBSERVE__ = window.__API_RECON_OBSERVE__ || { storage: [], headers: [] };
|
||||
|
||||
function trunc(s, n) {
|
||||
n = n || CONFIG.respMax || 600;
|
||||
s = String(s == null ? '' : s);
|
||||
return s.length > n ? s.slice(0, n) + '…' : s;
|
||||
}
|
||||
|
||||
function extractApiUrlsFromText(text) {
|
||||
if (!CONFIG.extractUrlsFromResponse || !text) return;
|
||||
const re = /["'](\/(?:api|apis|v\d+|dev|internal|graphql)[^"'`\\s]*)["'`]/gi;
|
||||
let m;
|
||||
while ((m = re.exec(text))) {
|
||||
const p = m[1].split('?')[0];
|
||||
window.__API_RECON_LOG__.add('GET ' + p);
|
||||
}
|
||||
const broad = /\/api\/[a-zA-Z0-9_./-]+/g;
|
||||
while ((m = broad.exec(text))) {
|
||||
const p = m[0].split('?')[0];
|
||||
if (!/\.(svg|png|jpg|gif|ico)$/i.test(p)) window.__API_RECON_LOG__.add('GET ' + p);
|
||||
}
|
||||
}
|
||||
|
||||
function recordApi(url, method) {
|
||||
const p = String(url || '').split('?')[0];
|
||||
if (CONFIG.apiPattern.test(p)) {
|
||||
window.__API_RECON_LOG__.add((method || 'GET').toUpperCase() + ' ' + p);
|
||||
}
|
||||
}
|
||||
|
||||
function recordApiDetail(entry) {
|
||||
recordApi(entry.url, entry.method);
|
||||
if (!CONFIG.recordDetail) return;
|
||||
window.__API_RECON_DETAIL__.push(entry);
|
||||
if (entry.responseBody) extractApiUrlsFromText(entry.responseBody);
|
||||
}
|
||||
|
||||
function blocked(url) {
|
||||
return url && CONFIG.loginPathRe.test(String(url));
|
||||
}
|
||||
|
||||
// --- 跳转中和 ---
|
||||
(function neutralizeNativeNavigation() {
|
||||
const rawAssign = Location.prototype.assign;
|
||||
const rawReplace = Location.prototype.replace;
|
||||
Location.prototype.assign = function (url) {
|
||||
if (blocked(url)) return;
|
||||
return rawAssign.call(this, url);
|
||||
};
|
||||
Location.prototype.replace = function (url) {
|
||||
if (blocked(url)) return;
|
||||
return rawReplace.call(this, url);
|
||||
};
|
||||
|
||||
const rawPush = history.pushState;
|
||||
const rawRep = history.replaceState;
|
||||
history.pushState = function (s, t, url) {
|
||||
if (blocked(url)) return;
|
||||
return rawPush.apply(this, arguments);
|
||||
};
|
||||
history.replaceState = function (s, t, url) {
|
||||
if (blocked(url)) return;
|
||||
return rawRep.apply(this, arguments);
|
||||
};
|
||||
|
||||
const hrefDesc = Object.getOwnPropertyDescriptor(Location.prototype, 'href');
|
||||
if (hrefDesc && hrefDesc.set) {
|
||||
const nativeSet = hrefDesc.set;
|
||||
Object.defineProperty(Location.prototype, 'href', {
|
||||
configurable: true,
|
||||
enumerable: hrefDesc.enumerable,
|
||||
get: hrefDesc.get,
|
||||
set(url) {
|
||||
if (blocked(url)) return;
|
||||
return nativeSet.call(this, url);
|
||||
},
|
||||
});
|
||||
}
|
||||
window.close = function () {};
|
||||
})();
|
||||
|
||||
// --- Vue Router 登录跳转中和 ---
|
||||
function neutralizeVueRouter() {
|
||||
if (!CONFIG.neutralizeVueRouter) return;
|
||||
try {
|
||||
const el = document.querySelector('[data-v-app]');
|
||||
const app = (el && el.__vue_app__) || window.__VUE__;
|
||||
if (!app) return;
|
||||
const router = app.config && app.config.globalProperties && app.config.globalProperties.$router;
|
||||
if (!router) return;
|
||||
router.beforeEach(function (to, from, next) { next(); });
|
||||
if (router.beforeResolve) router.beforeResolve(function (to, from, next) { next(); });
|
||||
function wrapNav(fn) {
|
||||
return function (loc) {
|
||||
const path = typeof loc === 'string' ? loc : (loc && (loc.path || loc.fullPath)) || '';
|
||||
if (blocked(path)) return Promise.resolve();
|
||||
return fn.apply(this, arguments);
|
||||
};
|
||||
}
|
||||
router.push = wrapNav(router.push.bind(router));
|
||||
router.replace = wrapNav(router.replace.bind(router));
|
||||
if (router.getRoutes) {
|
||||
router.getRoutes().forEach(function (r) {
|
||||
if (r.path) window.__API_RECON_ROUTES__.add(r.path);
|
||||
});
|
||||
}
|
||||
} catch (e) { /* ignore */ }
|
||||
}
|
||||
|
||||
function runPostLoadHooks() {
|
||||
neutralizeVueRouter();
|
||||
// 业务层跳转函数(goPage / navigateTo 等)
|
||||
['goPage', 'navigateTo', 'jumpTo', 'redirectTo'].forEach(function (name) {
|
||||
if (typeof window[name] !== 'function' || window[name].__apiReconWrapped) return;
|
||||
const raw = window[name];
|
||||
window[name] = function () {
|
||||
const arg = arguments[0];
|
||||
const path = typeof arg === 'string' ? arg : (arg && arg.path) || '';
|
||||
if (blocked(path)) return;
|
||||
return raw.apply(this, arguments);
|
||||
};
|
||||
window[name].__apiReconWrapped = true;
|
||||
});
|
||||
}
|
||||
document.addEventListener('DOMContentLoaded', runPostLoadHooks);
|
||||
window.addEventListener('load', runPostLoadHooks);
|
||||
|
||||
// --- 观察 Hook:辅助发现会话键名与请求头(可选,默认关 storage/cookie)---
|
||||
if (CONFIG.observe && CONFIG.observe.storageReads) {
|
||||
const rawGet = Storage.prototype.getItem;
|
||||
Storage.prototype.getItem = function (key) {
|
||||
window.__API_RECON_OBSERVE__.storage.push({ type: 'getItem', key: key, at: location.pathname });
|
||||
return rawGet.apply(this, arguments);
|
||||
};
|
||||
}
|
||||
if (CONFIG.observe && CONFIG.observe.cookieReads) {
|
||||
const cookieDesc = Object.getOwnPropertyDescriptor(Document.prototype, 'cookie');
|
||||
if (cookieDesc && cookieDesc.get) {
|
||||
const nativeGet = cookieDesc.get;
|
||||
Object.defineProperty(document, 'cookie', {
|
||||
configurable: true,
|
||||
get: function () {
|
||||
window.__API_RECON_OBSERVE__.storage.push({ type: 'cookieRead', at: location.pathname });
|
||||
return nativeGet.call(this);
|
||||
},
|
||||
set: cookieDesc.set,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
// --- Mock 辅助 ---
|
||||
const NEGATIVE_RE = /未登录|未授权|授权|not\s*login|unauthorized|forbidden/i;
|
||||
const tier = CONFIG.mockTier || 'L1+L2';
|
||||
|
||||
function patchJsonBody(text) {
|
||||
if (!tier.includes('L2')) return text;
|
||||
try {
|
||||
const j = JSON.parse(text);
|
||||
if (j && typeof j === 'object') {
|
||||
const msg = String(j.message || j.msg || '');
|
||||
for (const f of CONFIG.neutralize.fields) {
|
||||
if (f in j && j[f] !== CONFIG.neutralize.success && NEGATIVE_RE.test(msg)) {
|
||||
j[f] = CONFIG.neutralize.success;
|
||||
}
|
||||
}
|
||||
Object.assign(j, CONFIG.neutralize.flags || {});
|
||||
if (j.data == null) j.data = {};
|
||||
return JSON.stringify(j);
|
||||
}
|
||||
} catch (e) { /* ignore */ }
|
||||
return text;
|
||||
}
|
||||
|
||||
function lookupPrecise(url) {
|
||||
if (!tier.includes('L1')) return null;
|
||||
const p = String(url || '');
|
||||
for (const s of CONFIG.stubs) {
|
||||
if (s._re ? s._re.test(p) : new RegExp(s.match).test(p)) return s.body;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
function fallbackBody(url) {
|
||||
const p = String(url).split('?')[0];
|
||||
if (/\/(list|search|query|page|all|tree|nodes|options)/i.test(p)) {
|
||||
return { response_code: 0, data: [] };
|
||||
}
|
||||
if (/\/(get|detail|info|config|status)/i.test(p)) {
|
||||
return { response_code: 0, data: {} };
|
||||
}
|
||||
return { response_code: 0, data: {}, success: true };
|
||||
}
|
||||
|
||||
function matchMock(url) {
|
||||
const precise = lookupPrecise(url);
|
||||
if (precise) return precise;
|
||||
if (tier.includes('L3') && CONFIG.apiFallbackRe.test(String(url || '').split('?')[0])) {
|
||||
return fallbackBody(url);
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
CONFIG.stubs.forEach(function (s) {
|
||||
if (s.match && !s._re) s._re = new RegExp(s.match);
|
||||
});
|
||||
|
||||
// --- fetch Hook ---
|
||||
const rawFetch = window.fetch;
|
||||
window.fetch = async function (input, init) {
|
||||
const url = typeof input === 'string' ? input : (input && input.url) || '';
|
||||
const method = ((init && init.method) || 'GET').toUpperCase();
|
||||
const reqHeaders = (init && init.headers) || {};
|
||||
const reqBody = init && init.body ? trunc(init.body) : null;
|
||||
|
||||
const mock = matchMock(url);
|
||||
if (mock && !CONFIG.forward) {
|
||||
recordApiDetail({ method: method, url: url, reqHeaders: reqHeaders, reqBody: reqBody, status: 200, responseBody: JSON.stringify(mock), source: 'mock' });
|
||||
return new Response(JSON.stringify(mock), {
|
||||
status: 200,
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
});
|
||||
}
|
||||
|
||||
const resp = await rawFetch.apply(this, arguments);
|
||||
const clone = resp.clone();
|
||||
let text = '';
|
||||
try { text = await clone.text(); } catch (e) { /* ignore */ }
|
||||
|
||||
recordApiDetail({
|
||||
method: method, url: url, reqHeaders: reqHeaders, reqBody: reqBody,
|
||||
status: resp.status, responseBody: trunc(text), source: 'fetch',
|
||||
});
|
||||
|
||||
if (!tier.includes('L2') && !mock) return resp;
|
||||
|
||||
const patched = patchJsonBody(text);
|
||||
if (patched !== text) {
|
||||
return new Response(patched, {
|
||||
status: resp.status,
|
||||
statusText: resp.statusText,
|
||||
headers: resp.headers,
|
||||
});
|
||||
}
|
||||
return resp;
|
||||
};
|
||||
|
||||
// --- XHR Hook ---
|
||||
const rawOpen = XMLHttpRequest.prototype.open;
|
||||
const rawSend = XMLHttpRequest.prototype.send;
|
||||
const rawSetHeader = XMLHttpRequest.prototype.setRequestHeader;
|
||||
|
||||
XMLHttpRequest.prototype.open = function (method, url) {
|
||||
this.__apiReconUrl = url;
|
||||
this.__apiReconMethod = method;
|
||||
this.__apiReconHeaders = {};
|
||||
return rawOpen.apply(this, arguments);
|
||||
};
|
||||
|
||||
if (CONFIG.observe && CONFIG.observe.xhrHeaders) {
|
||||
XMLHttpRequest.prototype.setRequestHeader = function (name, value) {
|
||||
if (!this.__apiReconHeaders) this.__apiReconHeaders = {};
|
||||
this.__apiReconHeaders[name] = value;
|
||||
window.__API_RECON_OBSERVE__.headers.push({ name: name, url: this.__apiReconUrl });
|
||||
return rawSetHeader.apply(this, arguments);
|
||||
};
|
||||
}
|
||||
|
||||
XMLHttpRequest.prototype.send = function (body) {
|
||||
const xhr = this;
|
||||
const url = xhr.__apiReconUrl || '';
|
||||
const method = (xhr.__apiReconMethod || 'GET').toUpperCase();
|
||||
const mock = matchMock(url);
|
||||
|
||||
if (mock && !CONFIG.forward) {
|
||||
const bodyStr = JSON.stringify(mock);
|
||||
recordApiDetail({ method: method, url: url, reqHeaders: xhr.__apiReconHeaders, reqBody: trunc(body), status: 200, responseBody: bodyStr, source: 'mock' });
|
||||
setTimeout(function () {
|
||||
Object.defineProperty(xhr, 'readyState', { configurable: true, get: function () { return 4; } });
|
||||
Object.defineProperty(xhr, 'status', { configurable: true, get: function () { return 200; } });
|
||||
Object.defineProperty(xhr, 'responseText', { configurable: true, get: function () { return bodyStr; } });
|
||||
xhr.onreadystatechange && xhr.onreadystatechange();
|
||||
xhr.onload && xhr.onload();
|
||||
}, 0);
|
||||
return;
|
||||
}
|
||||
|
||||
const orig = xhr.onreadystatechange;
|
||||
xhr.onreadystatechange = function () {
|
||||
if (xhr.readyState === 4) {
|
||||
recordApiDetail({
|
||||
method: method, url: url, reqHeaders: xhr.__apiReconHeaders,
|
||||
reqBody: trunc(body), status: xhr.status,
|
||||
responseBody: trunc(xhr.responseText), source: 'xhr',
|
||||
});
|
||||
if (tier.includes('L2') && xhr.responseText) {
|
||||
try {
|
||||
const patched = patchJsonBody(xhr.responseText);
|
||||
if (patched !== xhr.responseText) {
|
||||
Object.defineProperty(xhr, 'responseText', { configurable: true, get: function () { return patched; } });
|
||||
}
|
||||
} catch (e) { /* ignore */ }
|
||||
}
|
||||
}
|
||||
orig && orig.apply(this, arguments);
|
||||
};
|
||||
return rawSend.apply(this, arguments);
|
||||
};
|
||||
})();
|
||||
@@ -0,0 +1,228 @@
|
||||
#!/usr/bin/env node
|
||||
/*
|
||||
* runtime_harvest.js <config.json>
|
||||
*
|
||||
* 参考模板 — 非通用成品。执行前须按目标站点调整 config.json 及脚本内逻辑:
|
||||
* cookies/localStorage、neutralize 字段、stubs 结构、loginUrlPattern、apiPattern
|
||||
*
|
||||
* Drives a headless browser through an authorized SPA to capture the live API
|
||||
* surface (method + url + body) by defeating three client-side gates:
|
||||
* 1. render gate -> inject fake auth state (cookies / localStorage)
|
||||
* 2. interceptor gate-> rewrite the "unauthorized" code field to success
|
||||
* 3. content gate -> stub the menu/permission endpoint with a full-feature payload
|
||||
*
|
||||
* Requires puppeteer-core + a system Chromium. ( cd scripts && npm install )
|
||||
*
|
||||
* config.json schema (all fields optional except baseUrl):
|
||||
* {
|
||||
* "baseUrl": "https://target/",
|
||||
* "chromium": "/usr/bin/chromium", // or env CHROMIUM
|
||||
* "cookies": [{"name":"auth","value":"b64json:{\"id\":1,\"username\":\"admin\"}"}],
|
||||
* "localStorage": {"token":"x","isLogin":"1"},
|
||||
* "neutralize": { "fields":["response_code","code","errno","ret"], "success":0,
|
||||
* "flags":{"success":true,"message":"ok"} },
|
||||
* "forward": true, // forward real req then rewrite code; false = offline stub
|
||||
* "loginUrlPattern": "/login", // navigations matching this are suppressed as a fallback
|
||||
* "stubs": [ {"match":"permission|menu|role", "body": { ...full-feature menu... }} ],
|
||||
* "routes": ["/dashboard","/device", ...], // from routes.txt or the forged menu
|
||||
* "apiPattern": "/api/|/rest/|/graphql", // what counts as an API call to record
|
||||
* "proxy": "http://127.0.0.1:8080", // optional; also HTTP_PROXY / HTTPS_PROXY
|
||||
* "waitUntil": "domcontentloaded", // prefer over networkidle2 for large SPAs
|
||||
* "routeTimeout": 12000,
|
||||
* "waitMs": 1200, "perRouteMs": 900, "headless": true
|
||||
* }
|
||||
*/
|
||||
process.env.NODE_TLS_REJECT_UNAUTHORIZED = '0';
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
|
||||
let puppeteer;
|
||||
try { puppeteer = require('puppeteer-core'); }
|
||||
catch (e) { console.error("[!] run `npm install` in the scripts/ dir first (needs puppeteer-core)"); process.exit(1); }
|
||||
|
||||
function loadCfg(p) {
|
||||
const c = JSON.parse(fs.readFileSync(p, 'utf8'));
|
||||
c.baseUrl || (() => { throw new Error("config.baseUrl required"); })();
|
||||
c.chromium = c.chromium || process.env.CHROMIUM || '/usr/bin/chromium';
|
||||
c.neutralize = c.neutralize || { fields: ['response_code', 'code', 'errno', 'ret', 'status'], success: 0, flags: { success: true } };
|
||||
c.forward = c.forward !== false;
|
||||
c.apiPattern = new RegExp(c.apiPattern || '/api/|/rest/|/service/|/graphql|/gateway/');
|
||||
c.loginUrlPattern = c.loginUrlPattern || '/login';
|
||||
c.routes = c.routes || ['/'];
|
||||
c.waitMs = c.waitMs || 1200; c.perRouteMs = c.perRouteMs || 900;
|
||||
c.headless = c.headless !== false;
|
||||
c.waitUntil = c.waitUntil || 'domcontentloaded';
|
||||
c.routeTimeout = c.routeTimeout || 12000;
|
||||
c.captureResponses = c.captureResponses !== false; // record response body samples (forward mode)
|
||||
c.recordWs = c.recordWs !== false; // record WebSocket frames + SSE endpoints
|
||||
c.respMax = c.respMax || 600; // truncation length for captured bodies
|
||||
c.proxy = c.proxy || process.env.HTTP_PROXY || process.env.HTTPS_PROXY || '';
|
||||
(c.stubs || []).forEach(s => s._re = new RegExp(s.match));
|
||||
return c;
|
||||
}
|
||||
|
||||
function cookieValue(v) {
|
||||
if (typeof v === 'string' && v.startsWith('b64json:'))
|
||||
return Buffer.from(v.slice(8)).toString('base64');
|
||||
if (typeof v === 'string' && v.startsWith('json:'))
|
||||
return v.slice(5);
|
||||
return v;
|
||||
}
|
||||
|
||||
function neutralize(txt, n) {
|
||||
try {
|
||||
const j = JSON.parse(txt);
|
||||
if (j && typeof j === 'object') {
|
||||
for (const f of n.fields) if (f in j) j[f] = n.success;
|
||||
Object.assign(j, n.flags || {});
|
||||
return JSON.stringify(j);
|
||||
}
|
||||
} catch (e) {}
|
||||
return txt;
|
||||
}
|
||||
|
||||
(async () => {
|
||||
const cfg = loadCfg(process.argv[2] || 'config.json');
|
||||
const origin = new URL(cfg.baseUrl).origin;
|
||||
const host = new URL(cfg.baseUrl).hostname;
|
||||
const rec = []; // {m,u,b,resp,ct}
|
||||
const chunks = new Set();
|
||||
const ws = []; // {url,dir,data} WebSocket frames
|
||||
const sse = new Set(); // SSE (text/event-stream) endpoints
|
||||
|
||||
if (cfg.proxy) {
|
||||
process.env.HTTP_PROXY = cfg.proxy;
|
||||
process.env.HTTPS_PROXY = cfg.proxy;
|
||||
}
|
||||
const launchArgs = ['--no-sandbox', '--disable-dev-shm-usage', '--ignore-certificate-errors'];
|
||||
if (cfg.proxy) launchArgs.push(`--proxy-server=${cfg.proxy}`);
|
||||
const browser = await puppeteer.launch({
|
||||
executablePath: cfg.chromium, headless: cfg.headless ? 'new' : false,
|
||||
args: launchArgs
|
||||
});
|
||||
const page = await browser.newPage();
|
||||
|
||||
// ---- WebSocket frame capture via CDP (fetch/XHR interception can't see WS) ----
|
||||
if (cfg.recordWs) {
|
||||
try {
|
||||
const cdp = await page.target().createCDPSession();
|
||||
await cdp.send('Network.enable');
|
||||
const wsUrl = {}; // requestId -> url
|
||||
cdp.on('Network.webSocketCreated', e => { wsUrl[e.requestId] = e.url; });
|
||||
const onFrame = dir => e => {
|
||||
const p = e.response && e.response.payloadData;
|
||||
if (p != null) ws.push({ url: (wsUrl[e.requestId] || '').replace(origin, ''), dir, data: String(p).slice(0, cfg.respMax) });
|
||||
};
|
||||
cdp.on('Network.webSocketFrameSent', onFrame('send'));
|
||||
cdp.on('Network.webSocketFrameReceived', onFrame('recv'));
|
||||
} catch (e) { console.log('[ws] CDP capture unavailable:', e.message); }
|
||||
}
|
||||
|
||||
// inject localStorage on every document
|
||||
if (cfg.localStorage) {
|
||||
await page.evaluateOnNewDocument((kv) => {
|
||||
try { for (const k in kv) localStorage.setItem(k, kv[k]); } catch (e) {}
|
||||
}, cfg.localStorage);
|
||||
}
|
||||
// fallback: block full-page navigations to the login url
|
||||
await page.evaluateOnNewDocument((pat) => {
|
||||
const bad = u => { try { return String(u).indexOf(pat) >= 0; } catch (e) { return false; } };
|
||||
try {
|
||||
const d = Object.getOwnPropertyDescriptor(Location.prototype, 'href');
|
||||
Object.defineProperty(Location.prototype, 'href', { configurable: true,
|
||||
get() { return d.get.call(this); },
|
||||
set(v) { if (bad(v)) return; return d.set.call(this, v); } });
|
||||
const a = Location.prototype.assign, r = Location.prototype.replace;
|
||||
Location.prototype.assign = function (v) { if (bad(v)) return; return a.call(this, v); };
|
||||
Location.prototype.replace = function (v) { if (bad(v)) return; return r.call(this, v); };
|
||||
} catch (e) {}
|
||||
}, cfg.loginUrlPattern);
|
||||
|
||||
await page.setRequestInterception(true);
|
||||
page.on('request', async (req) => {
|
||||
const u = req.url(), m = req.method();
|
||||
if (/\.js(\?|$)/.test(u) && /\/assets\/|\/static\/|\/js\//.test(u)) chunks.add(u.split('/').pop());
|
||||
|
||||
// suppress fallback login navigations (after the app has bootstrapped)
|
||||
if (req.isNavigationRequest() && req.frame() === page.mainFrame()
|
||||
&& u.includes(cfg.loginUrlPattern) && rec.length > 3) {
|
||||
return req.respond({ status: 204, body: '' });
|
||||
}
|
||||
if (!cfg.apiPattern.test(u)) return req.continue();
|
||||
|
||||
const entry = { m, u: u.replace(origin, '').split('?')[0], full: u, b: req.postData() ? req.postData().slice(0, 400) : null };
|
||||
rec.push(entry);
|
||||
|
||||
// explicit stubs (menu / permission forgery) win
|
||||
const stub = (cfg.stubs || []).find(s => s._re.test(u));
|
||||
if (stub) return req.respond({ status: 200, contentType: 'application/json', body: JSON.stringify(stub.body) });
|
||||
|
||||
// SSE: forwarding a text/event-stream would hang the handler — record + short-circuit
|
||||
const accept = (req.headers().accept || '');
|
||||
if (/text\/event-stream/.test(accept)) { sse.add(entry.u); return req.respond({ status: 200, contentType: 'application/json', body: '{}' }); }
|
||||
|
||||
if (!cfg.forward) {
|
||||
return req.respond({ status: 200, contentType: 'application/json',
|
||||
body: neutralize('{"data":{},"list":[],"total":0}', cfg.neutralize) });
|
||||
}
|
||||
// forward real request, then rewrite the unauthorized code field
|
||||
try {
|
||||
const headers = Object.assign({}, req.headers());
|
||||
if (cfg.cookies) headers.cookie = cfg.cookies.map(c => `${c.name}=${cookieValue(c.value)}`).join('; ');
|
||||
const r = await fetch(u, { method: m, headers, body: (m !== 'GET' && m !== 'HEAD') ? req.postData() : undefined });
|
||||
const t = await r.text();
|
||||
if (cfg.captureResponses) { entry.resp = t.slice(0, cfg.respMax); entry.ct = r.headers.get('content-type') || ''; }
|
||||
req.respond({ status: 200, contentType: 'application/json', body: neutralize(t, cfg.neutralize) });
|
||||
} catch (e) {
|
||||
req.respond({ status: 200, contentType: 'application/json', body: '{"response_code":0,"code":0,"data":{}}' });
|
||||
}
|
||||
});
|
||||
|
||||
// set auth cookies
|
||||
if (cfg.cookies) {
|
||||
for (const c of cfg.cookies)
|
||||
await page.setCookie({ name: c.name, value: cookieValue(c.value), domain: host, path: c.path || '/' });
|
||||
}
|
||||
|
||||
// boot
|
||||
await page.goto(cfg.baseUrl, { waitUntil: cfg.waitUntil, timeout: 45000 }).catch(e => console.log('[goto]', e.message));
|
||||
await new Promise(r => setTimeout(r, cfg.waitMs));
|
||||
const shell = await page.evaluate(() => ({
|
||||
url: location.href,
|
||||
loginForm: !!document.querySelector('input[type=password]'),
|
||||
links: [...document.querySelectorAll('a[href^="/"]')].map(a => a.getAttribute('href'))
|
||||
}));
|
||||
console.log(`[*] after boot: url=${shell.url} loginForm=${shell.loginForm}`);
|
||||
if (shell.loginForm) console.log('[!] still on login — recheck render-gate facts (cookies/localStorage) in config');
|
||||
|
||||
// discovered menu links extend the route list
|
||||
const routes = [...new Set([...cfg.routes, ...shell.links.filter(h => h && h.length > 1)])];
|
||||
const perRoute = {};
|
||||
for (const rt of routes) {
|
||||
const before = rec.length;
|
||||
try { await page.goto(origin + rt, { waitUntil: cfg.waitUntil, timeout: cfg.routeTimeout }); } catch (e) {}
|
||||
await new Promise(r => setTimeout(r, cfg.perRouteMs));
|
||||
const fresh = [...new Set(rec.slice(before).map(r => r.m + ' ' + r.u))];
|
||||
perRoute[rt] = fresh;
|
||||
}
|
||||
|
||||
const uniq = [...new Set(rec.map(r => r.m + ' ' + r.u))].sort();
|
||||
const wsUniq = [...new Set(ws.map(f => f.url))].filter(Boolean).sort();
|
||||
const outdir = path.dirname(process.argv[2] || '.');
|
||||
fs.writeFileSync(path.join(outdir, 'runtime_api.json'),
|
||||
JSON.stringify({ shell, uniq, perRoute, rec, ws, wsEndpoints: wsUniq, sse: [...sse] }, null, 2));
|
||||
|
||||
// merge with static
|
||||
let merged = new Set(uniq.map(x => x.split(' ')[1]));
|
||||
const staticFile = path.join(outdir, 'api_static.txt');
|
||||
if (fs.existsSync(staticFile)) fs.readFileSync(staticFile, 'utf8').split('\n').filter(Boolean).forEach(p => merged.add(p));
|
||||
fs.writeFileSync(path.join(outdir, 'api_merged.txt'), [...merged].sort().join('\n'));
|
||||
|
||||
console.log(`[+] runtime endpoints (with method): ${uniq.length}`);
|
||||
console.log(`[+] chunks seen at runtime: ${chunks.size}`);
|
||||
if (cfg.recordWs) console.log(`[+] websocket endpoints: ${wsUniq.length} (${ws.length} frames) | SSE endpoints: ${sse.size}`);
|
||||
if (cfg.captureResponses) console.log(`[+] response bodies captured for ${rec.filter(r => r.resp != null).length}/${rec.length} calls`);
|
||||
console.log(`[+] merged unique paths: ${merged.size} -> api_merged.txt`);
|
||||
console.log(`[+] full detail -> runtime_api.json (per-route + bodies + ws frames + sse)`);
|
||||
await browser.close();
|
||||
})();
|
||||
@@ -0,0 +1,123 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
spider_mpa.py <BASE_URL> <OUTDIR> [--cookie "k=v; k2=v2"] [--max 200] [--depth 4]
|
||||
|
||||
参考模板 — 非通用成品。执行前须按目标调整 --exclude、cookie、depth/max 等同域策略。
|
||||
|
||||
Fallback for NON-SPA targets (traditional server-rendered MPAs: Django/Rails/PHP/
|
||||
JSP, classic admin panels). When there is no JS endpoint bundle, the API surface
|
||||
lives in HTML <form action>, <a href>, and inline-JS ajax urls. This BFS-crawls
|
||||
the same origin and extracts:
|
||||
- forms.txt : METHOD action [param1, param2, ...] (the real "endpoints")
|
||||
- links.txt : every same-origin URL reached
|
||||
- api_inline.txt : url-ish strings found in inline <script> / onclick (fetch/ajax)
|
||||
|
||||
Stdlib only. Same-origin, bounded, polite. Provide --cookie for an authed crawl.
|
||||
"""
|
||||
import sys, os, re, ssl, argparse, urllib.request, urllib.parse, time
|
||||
from html.parser import HTMLParser
|
||||
from collections import deque
|
||||
|
||||
CTX = ssl.create_default_context(); CTX.check_hostname = False; CTX.verify_mode = ssl.CERT_NONE
|
||||
UA = "Mozilla/5.0 (spa-api-recon spider)"
|
||||
SKIP_EXT = re.compile(r'\.(css|png|jpe?g|gif|svg|ico|woff2?|ttf|pdf|zip|mp4|webp|js)(\?|$)', re.I)
|
||||
API_HINT = re.compile(r'/(?:api|rest|service|ajax|action|do|rpc|graphql|v\d|admin|backend)\b', re.I)
|
||||
|
||||
def fetch(url, cookie):
|
||||
h = {"User-Agent": UA}
|
||||
if cookie: h["Cookie"] = cookie
|
||||
try:
|
||||
with urllib.request.urlopen(urllib.request.Request(url, headers=h), context=CTX, timeout=20) as r:
|
||||
ct = r.headers.get("Content-Type", "")
|
||||
if "html" not in ct and "xml" not in ct: return "", ct
|
||||
return r.read().decode("utf-8", "ignore"), ct
|
||||
except Exception:
|
||||
return "", ""
|
||||
|
||||
class Page(HTMLParser):
|
||||
def __init__(self):
|
||||
super().__init__()
|
||||
self.links, self.scripts, self.inline = [], [], []
|
||||
self.forms, self._cur = [], None
|
||||
self._in_script = False
|
||||
def handle_starttag(self, tag, attrs):
|
||||
a = dict(attrs)
|
||||
if tag == "a" and a.get("href"): self.links.append(a["href"])
|
||||
elif tag == "script":
|
||||
self._in_script = True
|
||||
if a.get("src"): self.scripts.append(a["src"])
|
||||
elif tag == "form":
|
||||
self._cur = {"action": a.get("action", ""), "method": (a.get("method") or "GET").upper(), "params": []}
|
||||
elif tag in ("input", "select", "textarea", "button") and self._cur is not None:
|
||||
n = a.get("name")
|
||||
if n: self._cur["params"].append(n)
|
||||
# ajax-ish handlers
|
||||
for v in a.values():
|
||||
for m in re.findall(r'''["'](/[^"']{2,120})["']''', v or ""):
|
||||
if API_HINT.search(m): self.inline.append(m)
|
||||
def handle_endtag(self, tag):
|
||||
if tag == "script": self._in_script = False
|
||||
elif tag == "form" and self._cur is not None:
|
||||
self.forms.append(self._cur); self._cur = None
|
||||
def handle_data(self, data):
|
||||
if self._in_script and data:
|
||||
for m in re.findall(r'''["'`](/[A-Za-z0-9_\-./{}$:?=&]{2,120})["'`]''', data):
|
||||
if API_HINT.search(m): self.inline.append(m)
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("base"); ap.add_argument("outdir")
|
||||
ap.add_argument("--cookie", default=""); ap.add_argument("--max", type=int, default=200)
|
||||
ap.add_argument("--depth", type=int, default=4); ap.add_argument("--delay", type=float, default=0.1)
|
||||
ap.add_argument("--include", default=""); ap.add_argument("--exclude", default="logout|signout|delete|remove|destroy")
|
||||
args = ap.parse_args()
|
||||
base = args.base if args.base.startswith("http") else "https://" + args.base
|
||||
os.makedirs(args.outdir, exist_ok=True)
|
||||
origin = "{0.scheme}://{0.netloc}".format(urllib.parse.urlparse(base))
|
||||
inc = re.compile(args.include) if args.include else None
|
||||
exc = re.compile(args.exclude, re.I) if args.exclude else None
|
||||
|
||||
seen, forms, inline, links = set(), {}, set(), set()
|
||||
q = deque([(base, 0)]); seen.add(base.split("#")[0])
|
||||
n = 0
|
||||
while q and n < args.max:
|
||||
url, d = q.popleft(); n += 1
|
||||
html, ct = fetch(url, args.cookie)
|
||||
if args.delay: time.sleep(args.delay)
|
||||
if not html: continue
|
||||
p = Page()
|
||||
try: p.feed(html)
|
||||
except Exception: pass
|
||||
links.add(url.replace(origin, "") or "/")
|
||||
for f in p.forms:
|
||||
act = urllib.parse.urljoin(url, f["action"] or url)
|
||||
key = f["method"] + " " + act.replace(origin, "")
|
||||
forms.setdefault(key, set()).update(f["params"])
|
||||
for s in p.inline: inline.add(s)
|
||||
if d < args.depth:
|
||||
for href in p.links:
|
||||
if href.startswith(("mailto:", "tel:", "javascript:", "#")): continue
|
||||
nxt = urllib.parse.urljoin(url, href).split("#")[0]
|
||||
if not nxt.startswith(origin): continue
|
||||
if SKIP_EXT.search(nxt): continue
|
||||
if exc and exc.search(nxt): continue
|
||||
if inc and not inc.search(nxt): continue
|
||||
if nxt not in seen:
|
||||
seen.add(nxt); q.append((nxt, d + 1))
|
||||
|
||||
with open(os.path.join(args.outdir, "forms.txt"), "w") as fh:
|
||||
for k in sorted(forms):
|
||||
ps = ", ".join(sorted(forms[k]))
|
||||
fh.write(f"{k} [{ps}]\n")
|
||||
open(os.path.join(args.outdir, "links.txt"), "w").write("\n".join(sorted(links)))
|
||||
open(os.path.join(args.outdir, "api_inline.txt"), "w").write("\n".join(sorted(inline)))
|
||||
|
||||
print(f"[+] crawled {n} pages (cap {args.max}, depth {args.depth})")
|
||||
print(f"[+] forms (endpoints): {len(forms)} -> forms.txt")
|
||||
print(f"[+] inline ajax urls: {len(inline)} -> api_inline.txt")
|
||||
print(f"[+] pages reached: {len(links)} -> links.txt")
|
||||
if not args.cookie:
|
||||
print("[i] no --cookie: only public pages crawled. Pass an authed session cookie for the full surface.")
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,420 @@
|
||||
---
|
||||
name: playwright-cli
|
||||
description: Automate browser interactions, test web pages and work with Playwright tests.
|
||||
allowed-tools: Bash(playwright-cli:*) Bash(npx:*) Bash(npm:*)
|
||||
---
|
||||
|
||||
# Browser Automation with playwright-cli
|
||||
|
||||
## Quick start
|
||||
|
||||
```bash
|
||||
# open new browser
|
||||
playwright-cli open
|
||||
# navigate to a page
|
||||
playwright-cli goto https://playwright.dev
|
||||
# interact with the page using refs from the snapshot
|
||||
playwright-cli click e15
|
||||
playwright-cli type "page.click"
|
||||
playwright-cli press Enter
|
||||
# take a screenshot (rarely used, as snapshot is more common)
|
||||
playwright-cli screenshot
|
||||
# close the browser
|
||||
playwright-cli close
|
||||
```
|
||||
|
||||
## Commands
|
||||
|
||||
### Core
|
||||
|
||||
```bash
|
||||
playwright-cli open
|
||||
# open and navigate right away
|
||||
playwright-cli open https://example.com/
|
||||
playwright-cli goto https://playwright.dev
|
||||
playwright-cli type "search query"
|
||||
playwright-cli click e3
|
||||
playwright-cli dblclick e7
|
||||
# --submit presses Enter after filling the element
|
||||
playwright-cli fill e5 "user@example.com" --submit
|
||||
playwright-cli drag e2 e8
|
||||
# drop files or data onto an element (from outside the page)
|
||||
playwright-cli drop e4 --path=./image.png
|
||||
playwright-cli drop e4 --data="text/plain=hello world"
|
||||
playwright-cli hover e4
|
||||
playwright-cli select e9 "option-value"
|
||||
playwright-cli upload ./document.pdf
|
||||
playwright-cli check e12
|
||||
playwright-cli uncheck e12
|
||||
playwright-cli snapshot
|
||||
# search the snapshot for text or a regexp, returns matching nodes with surrounding context
|
||||
playwright-cli find "Sign in"
|
||||
playwright-cli find --regex "Sign (in|up)"
|
||||
# wrap the regexp in slashes to add flags, e.g. /i for case-insensitive
|
||||
playwright-cli find --regex "/sign (in|up)/i"
|
||||
playwright-cli eval "document.title"
|
||||
playwright-cli eval "el => el.textContent" e5
|
||||
# get element id, class, or any attribute not visible in the snapshot
|
||||
playwright-cli eval "el => el.id" e5
|
||||
playwright-cli eval "el => el.getAttribute('data-testid')" e5
|
||||
playwright-cli dialog-accept
|
||||
playwright-cli dialog-accept "confirmation text"
|
||||
playwright-cli dialog-dismiss
|
||||
playwright-cli resize 1920 1080
|
||||
playwright-cli close
|
||||
```
|
||||
|
||||
### Navigation
|
||||
|
||||
```bash
|
||||
playwright-cli go-back
|
||||
playwright-cli go-forward
|
||||
playwright-cli reload
|
||||
```
|
||||
|
||||
### Keyboard
|
||||
|
||||
```bash
|
||||
playwright-cli press Enter
|
||||
playwright-cli press ArrowDown
|
||||
playwright-cli keydown Shift
|
||||
playwright-cli keyup Shift
|
||||
```
|
||||
|
||||
### Mouse
|
||||
|
||||
```bash
|
||||
playwright-cli mousemove 150 300
|
||||
playwright-cli mousedown
|
||||
playwright-cli mousedown right
|
||||
playwright-cli mouseup
|
||||
playwright-cli mouseup right
|
||||
playwright-cli mousewheel 0 100
|
||||
```
|
||||
|
||||
### Save as
|
||||
|
||||
```bash
|
||||
playwright-cli screenshot
|
||||
playwright-cli screenshot e5
|
||||
playwright-cli screenshot --filename=page.png
|
||||
playwright-cli screenshot --hires
|
||||
playwright-cli pdf --filename=page.pdf
|
||||
```
|
||||
|
||||
### Tabs
|
||||
|
||||
```bash
|
||||
playwright-cli tab-list
|
||||
playwright-cli tab-new
|
||||
playwright-cli tab-new https://example.com/page
|
||||
playwright-cli tab-close
|
||||
playwright-cli tab-close 2
|
||||
playwright-cli tab-select 0
|
||||
```
|
||||
|
||||
### Storage
|
||||
|
||||
```bash
|
||||
playwright-cli state-save
|
||||
playwright-cli state-save auth.json
|
||||
playwright-cli state-load auth.json
|
||||
|
||||
# Cookies
|
||||
playwright-cli cookie-list
|
||||
playwright-cli cookie-list --domain=example.com
|
||||
playwright-cli cookie-get session_id
|
||||
playwright-cli cookie-set session_id abc123
|
||||
playwright-cli cookie-set session_id abc123 --domain=example.com --httpOnly --secure
|
||||
playwright-cli cookie-delete session_id
|
||||
playwright-cli cookie-clear
|
||||
|
||||
# LocalStorage
|
||||
playwright-cli localstorage-list
|
||||
playwright-cli localstorage-get theme
|
||||
playwright-cli localstorage-set theme dark
|
||||
playwright-cli localstorage-delete theme
|
||||
playwright-cli localstorage-clear
|
||||
|
||||
# SessionStorage
|
||||
playwright-cli sessionstorage-list
|
||||
playwright-cli sessionstorage-get step
|
||||
playwright-cli sessionstorage-set step 3
|
||||
playwright-cli sessionstorage-delete step
|
||||
playwright-cli sessionstorage-clear
|
||||
```
|
||||
|
||||
### Network
|
||||
|
||||
```bash
|
||||
playwright-cli route "**/*.jpg" --status=404
|
||||
playwright-cli route "https://api.example.com/**" --body='{"mock": true}'
|
||||
playwright-cli route-list
|
||||
playwright-cli unroute "**/*.jpg"
|
||||
playwright-cli unroute
|
||||
```
|
||||
|
||||
### DevTools
|
||||
|
||||
```bash
|
||||
playwright-cli console
|
||||
playwright-cli console warning
|
||||
playwright-cli requests
|
||||
playwright-cli request 5
|
||||
playwright-cli run-code "async page => await page.context().grantPermissions(['geolocation'])"
|
||||
playwright-cli run-code --filename=script.js
|
||||
playwright-cli tracing-start
|
||||
playwright-cli tracing-stop
|
||||
playwright-cli video-start video.webm
|
||||
playwright-cli video-chapter "Chapter Title" --description="Details" --duration=2000
|
||||
playwright-cli video-stop
|
||||
|
||||
# annotate each subsequent action (click, type, ...) with a callout naming the action and highlighting the target
|
||||
playwright-cli video-show-actions --duration=600 --position=top-right
|
||||
playwright-cli video-hide-actions
|
||||
|
||||
# launch the dashboard for UI review / design feedback — user annotates the page, you receive the annotated screenshot, snapshot, and notes
|
||||
playwright-cli show --annotate
|
||||
|
||||
# generate a Playwright locator for an element from its ref or selector
|
||||
playwright-cli generate-locator e5 --raw
|
||||
|
||||
# show a persistent highlight overlay for an element, optionally with a custom style
|
||||
playwright-cli highlight e5
|
||||
playwright-cli highlight e5 --style="outline: 3px dashed red"
|
||||
# hide a single element highlight, or all page highlights when no target is given
|
||||
playwright-cli highlight e5 --hide
|
||||
playwright-cli highlight --hide
|
||||
```
|
||||
|
||||
## Raw output
|
||||
|
||||
The global `--raw` option strips page status, generated code, and snapshot sections from the output, returning only the result value. Use it to pipe command output into other tools. Commands that don't produce output return nothing.
|
||||
|
||||
```bash
|
||||
playwright-cli --raw eval "JSON.stringify(performance.timing)" | jq '.loadEventEnd - .navigationStart'
|
||||
playwright-cli --raw eval "JSON.stringify([...document.querySelectorAll('a')].map(a => a.href))" > links.json
|
||||
playwright-cli --raw snapshot > before.yml
|
||||
playwright-cli click e5
|
||||
playwright-cli --raw snapshot > after.yml
|
||||
diff before.yml after.yml
|
||||
TOKEN=$(playwright-cli --raw cookie-get session_id)
|
||||
playwright-cli --raw localstorage-get theme
|
||||
```
|
||||
|
||||
For structured output wrapping every reply as JSON, pass --json
|
||||
```bash
|
||||
playwright-cli list --json
|
||||
```
|
||||
|
||||
## Open parameters
|
||||
```bash
|
||||
# Use specific browser when creating session
|
||||
playwright-cli open --browser=chrome
|
||||
playwright-cli open --browser=firefox
|
||||
playwright-cli open --browser=webkit
|
||||
playwright-cli open --browser=msedge
|
||||
|
||||
# Emulate a generic mobile device (Pixel 10 for Chromium, iPhone 17 for WebKit).
|
||||
# Prefer this when a mobile layout is acceptable: mobile pages are usually
|
||||
# lighter, so snapshots are smaller and cheaper.
|
||||
playwright-cli open --mobile
|
||||
playwright-cli open --device="iPhone 15"
|
||||
|
||||
# Use persistent profile (by default profile is in-memory)
|
||||
playwright-cli open --persistent
|
||||
# Use persistent profile with custom directory
|
||||
playwright-cli open --profile=/path/to/profile
|
||||
|
||||
# Connect to browser via Playwright Extension
|
||||
playwright-cli attach --extension=chrome
|
||||
|
||||
# Connect to a running Chrome or Edge by channel name
|
||||
playwright-cli attach --cdp=chrome
|
||||
playwright-cli attach --cdp=msedge
|
||||
|
||||
# Connect to a running browser via CDP endpoint
|
||||
playwright-cli attach --cdp=http://localhost:9222
|
||||
|
||||
# Start with config file
|
||||
playwright-cli open --config=my-config.json
|
||||
|
||||
# Close the browser
|
||||
playwright-cli close
|
||||
# Detach from an attached browser (leaves the external browser running)
|
||||
playwright-cli -s=msedge detach
|
||||
# Delete user data for the default session
|
||||
playwright-cli delete-data
|
||||
```
|
||||
|
||||
## URLs with `&` on Windows
|
||||
|
||||
On Windows, `cmd.exe` and PowerShell treat `&` as a command separator, so URLs with multiple query parameters get truncated before `playwright-cli` runs. Escape `&` with `^&` in `cmd.exe`, or use `--%` in PowerShell:
|
||||
|
||||
```batch
|
||||
playwright-cli goto "https://example.com/?a=1^&b=2"
|
||||
```
|
||||
|
||||
```powershell
|
||||
playwright-cli --% goto "https://example.com/?a=1&b=2"
|
||||
```
|
||||
|
||||
## Snapshots
|
||||
|
||||
After each command, playwright-cli provides a snapshot of the current browser state.
|
||||
|
||||
```bash
|
||||
> playwright-cli goto https://example.com
|
||||
### Page
|
||||
- Page URL: https://example.com/
|
||||
- Page Title: Example Domain
|
||||
### Snapshot
|
||||
.playwright-cli/page-<timestamp>.yml
|
||||
```
|
||||
|
||||
You can also take a snapshot on demand using `playwright-cli snapshot` command. All the options below can be combined as needed.
|
||||
|
||||
```bash
|
||||
# default - save to a file with timestamp-based name
|
||||
playwright-cli snapshot
|
||||
|
||||
# save to file, use when snapshot is a part of the workflow result
|
||||
playwright-cli snapshot --filename=after-click.yaml
|
||||
|
||||
# snapshot an element instead of the whole page
|
||||
playwright-cli snapshot "#main"
|
||||
|
||||
# limit snapshot depth for efficiency, take a partial snapshot afterwards
|
||||
playwright-cli snapshot --depth=4
|
||||
playwright-cli snapshot e34
|
||||
|
||||
# include each element's bounding box as [box=x,y,width,height]
|
||||
playwright-cli snapshot --boxes
|
||||
|
||||
# search a large snapshot instead of capturing it all — returns matching nodes
|
||||
# with 3 lines of context around each match (like grep -C)
|
||||
playwright-cli find "Add to cart"
|
||||
playwright-cli find --regex "\\$[0-9]+\\.[0-9]{2}"
|
||||
```
|
||||
|
||||
## Targeting elements
|
||||
|
||||
By default, use refs from the snapshot to interact with page elements.
|
||||
|
||||
```bash
|
||||
# get snapshot with refs
|
||||
playwright-cli snapshot
|
||||
|
||||
# interact using a ref
|
||||
playwright-cli click e15
|
||||
```
|
||||
|
||||
You can also use css selectors or Playwright locators.
|
||||
|
||||
```bash
|
||||
# css selector
|
||||
playwright-cli click "#main > button.submit"
|
||||
|
||||
# role locator
|
||||
playwright-cli click "getByRole('button', { name: 'Submit' })"
|
||||
|
||||
# test id
|
||||
playwright-cli click "getByTestId('submit-button')"
|
||||
```
|
||||
|
||||
## Browser Sessions
|
||||
|
||||
```bash
|
||||
# create new browser session named "mysession" with persistent profile
|
||||
playwright-cli -s=mysession open example.com --persistent
|
||||
# same with manually specified profile directory (use when requested explicitly)
|
||||
playwright-cli -s=mysession open example.com --profile=/path/to/profile
|
||||
playwright-cli -s=mysession click e6
|
||||
playwright-cli -s=mysession close # stop a named browser
|
||||
playwright-cli -s=mysession delete-data # delete user data for persistent session
|
||||
|
||||
playwright-cli list
|
||||
# Close all browsers
|
||||
playwright-cli close-all
|
||||
# Forcefully kill all browser processes
|
||||
playwright-cli kill-all
|
||||
```
|
||||
|
||||
## Installation
|
||||
|
||||
If global `playwright-cli` command is not available, try a local version via `npx playwright cli`:
|
||||
|
||||
```bash
|
||||
npx --no-install playwright --version
|
||||
```
|
||||
|
||||
When local version is available, use `npx playwright cli` in all commands. Otherwise, install `playwright-cli` as a global command:
|
||||
|
||||
```bash
|
||||
npm install -g @playwright/cli@latest
|
||||
```
|
||||
|
||||
## Example: Form submission
|
||||
|
||||
```bash
|
||||
playwright-cli open https://example.com/form
|
||||
playwright-cli snapshot
|
||||
|
||||
playwright-cli fill e1 "user@example.com"
|
||||
playwright-cli fill e2 "password123"
|
||||
playwright-cli click e3
|
||||
playwright-cli snapshot
|
||||
playwright-cli close
|
||||
```
|
||||
|
||||
## Example: Multi-tab workflow
|
||||
|
||||
```bash
|
||||
playwright-cli open https://example.com
|
||||
playwright-cli tab-new https://example.com/other
|
||||
playwright-cli tab-list
|
||||
playwright-cli tab-select 0
|
||||
playwright-cli snapshot
|
||||
playwright-cli close
|
||||
```
|
||||
|
||||
## Example: Debugging with DevTools
|
||||
|
||||
```bash
|
||||
playwright-cli open https://example.com
|
||||
playwright-cli click e4
|
||||
playwright-cli fill e7 "test"
|
||||
playwright-cli console
|
||||
playwright-cli requests
|
||||
playwright-cli close
|
||||
```
|
||||
|
||||
```bash
|
||||
playwright-cli open https://example.com
|
||||
playwright-cli tracing-start
|
||||
playwright-cli click e4
|
||||
playwright-cli fill e7 "test"
|
||||
playwright-cli tracing-stop
|
||||
playwright-cli close
|
||||
```
|
||||
|
||||
## Example: Interactive session
|
||||
|
||||
Ask the user for UI review or design feedback. The user draws boxes on the live page and types comments; you receive the annotated screenshot, the snapshot of the marked region, and the user's notes. Use this whenever the user asks for "UI review", "design feedback", or to "ask the user what they think / want / mean":
|
||||
|
||||
```bash
|
||||
playwright-cli open https://example.com
|
||||
playwright-cli show --annotate
|
||||
```
|
||||
|
||||
## Specific tasks
|
||||
|
||||
* **Running and Debugging Playwright tests** [references/playwright-tests.md](references/playwright-tests.md)
|
||||
* **Request mocking** [references/request-mocking.md](references/request-mocking.md)
|
||||
* **Running Playwright code** [references/running-code.md](references/running-code.md)
|
||||
* **Browser session management** [references/session-management.md](references/session-management.md)
|
||||
* **Storage state (cookies, localStorage)** [references/storage-state.md](references/storage-state.md)
|
||||
* **Test generation (plan / generate / heal)** [references/test-generation.md](references/test-generation.md)
|
||||
* **Tracing** [references/tracing.md](references/tracing.md)
|
||||
* **Video recording** [references/video-recording.md](references/video-recording.md)
|
||||
* **Inspecting element attributes** [references/element-attributes.md](references/element-attributes.md)
|
||||
@@ -0,0 +1,23 @@
|
||||
# Inspecting Element Attributes
|
||||
|
||||
When the snapshot doesn't show an element's `id`, `class`, `data-*` attributes, or other DOM properties, use `eval` to inspect them.
|
||||
|
||||
## Examples
|
||||
|
||||
```bash
|
||||
playwright-cli snapshot
|
||||
# snapshot shows a button as e7 but doesn't reveal its id or data attributes
|
||||
|
||||
# get the element's id
|
||||
playwright-cli eval "el => el.id" e7
|
||||
|
||||
# get all CSS classes
|
||||
playwright-cli eval "el => el.className" e7
|
||||
|
||||
# get a specific attribute
|
||||
playwright-cli eval "el => el.getAttribute('data-testid')" e7
|
||||
playwright-cli eval "el => el.getAttribute('aria-label')" e7
|
||||
|
||||
# get a computed style property
|
||||
playwright-cli eval "el => getComputedStyle(el).display" e7
|
||||
```
|
||||
@@ -0,0 +1,39 @@
|
||||
# Running Playwright Tests
|
||||
|
||||
To run Playwright tests, use the `npx playwright test` command, or a package manager script. To avoid opening the interactive html report, use `PLAYWRIGHT_HTML_OPEN=never` environment variable.
|
||||
|
||||
```bash
|
||||
# Run all tests
|
||||
PLAYWRIGHT_HTML_OPEN=never npx playwright test
|
||||
|
||||
# Run all tests through a custom npm script
|
||||
PLAYWRIGHT_HTML_OPEN=never npm run special-test-command
|
||||
```
|
||||
|
||||
# Debugging Playwright Tests
|
||||
|
||||
To debug a failing Playwright test, run it with `--debug=cli` option. This command will pause the test at the start and print the debugging instructions.
|
||||
|
||||
**IMPORTANT**: run the command in the background and check the output until "Debugging Instructions" is printed. Make sure to stop the command after you have finished.
|
||||
|
||||
Once instructions containing a session name are printed, use `playwright-cli` to attach the session and explore the page.
|
||||
|
||||
```bash
|
||||
# Run the test
|
||||
PLAYWRIGHT_HTML_OPEN=never npx playwright test --debug=cli
|
||||
# ...
|
||||
# ... debugging instructions for "tw-abcdef" session ...
|
||||
# ...
|
||||
|
||||
# Attach to the test
|
||||
playwright-cli attach tw-abcdef
|
||||
```
|
||||
|
||||
Keep the test running in the background while you explore and look for a fix.
|
||||
The test is paused at the start, so you should step over or pause at a particular location
|
||||
where the problem is most likely to be.
|
||||
|
||||
Every action you perform with `playwright-cli` generates corresponding Playwright TypeScript code.
|
||||
This code appears in the output and can be copied directly into the test. Most of the time, a specific locator or an expectation should be updated, but it could also be a bug in the app. Use your judgement.
|
||||
|
||||
After fixing the test, stop the background test run. Rerun to check that test passes.
|
||||
@@ -0,0 +1,87 @@
|
||||
# Request Mocking
|
||||
|
||||
Intercept, mock, modify, and block network requests.
|
||||
|
||||
## CLI Route Commands
|
||||
|
||||
```bash
|
||||
# Mock with custom status
|
||||
playwright-cli route "**/*.jpg" --status=404
|
||||
|
||||
# Mock with JSON body
|
||||
playwright-cli route "**/api/users" --body='[{"id":1,"name":"Alice"}]' --content-type=application/json
|
||||
|
||||
# Mock with custom headers
|
||||
playwright-cli route "**/api/data" --body='{"ok":true}' --header="X-Custom: value"
|
||||
|
||||
# Remove headers from requests
|
||||
playwright-cli route "**/*" --remove-header=cookie,authorization
|
||||
|
||||
# List active routes
|
||||
playwright-cli route-list
|
||||
|
||||
# Remove a route or all routes
|
||||
playwright-cli unroute "**/*.jpg"
|
||||
playwright-cli unroute
|
||||
```
|
||||
|
||||
## URL Patterns
|
||||
|
||||
```
|
||||
**/api/users - Exact path match
|
||||
**/api/*/details - Wildcard in path
|
||||
**/*.{png,jpg,jpeg} - Match file extensions
|
||||
**/search?q=* - Match query parameters
|
||||
```
|
||||
|
||||
## Advanced Mocking with run-code
|
||||
|
||||
For conditional responses, request body inspection, response modification, or delays:
|
||||
|
||||
### Conditional Response Based on Request
|
||||
|
||||
```bash
|
||||
playwright-cli run-code "async page => {
|
||||
await page.route('**/api/login', route => {
|
||||
const body = route.request().postDataJSON();
|
||||
if (body.username === 'admin') {
|
||||
route.fulfill({ body: JSON.stringify({ token: 'mock-token' }) });
|
||||
} else {
|
||||
route.fulfill({ status: 401, body: JSON.stringify({ error: 'Invalid' }) });
|
||||
}
|
||||
});
|
||||
}"
|
||||
```
|
||||
|
||||
### Modify Real Response
|
||||
|
||||
```bash
|
||||
playwright-cli run-code "async page => {
|
||||
await page.route('**/api/user', async route => {
|
||||
const response = await route.fetch();
|
||||
const json = await response.json();
|
||||
json.isPremium = true;
|
||||
await route.fulfill({ response, json });
|
||||
});
|
||||
}"
|
||||
```
|
||||
|
||||
### Simulate Network Failures
|
||||
|
||||
```bash
|
||||
playwright-cli run-code "async page => {
|
||||
await page.route('**/api/offline', route => route.abort('internetdisconnected'));
|
||||
}"
|
||||
# Options: connectionrefused, timedout, connectionreset, internetdisconnected
|
||||
```
|
||||
|
||||
### Delayed Response
|
||||
|
||||
```bash
|
||||
playwright-cli run-code "async page => {
|
||||
await page.route('**/api/slow', async route => {
|
||||
await new Promise(r => setTimeout(r, 3000));
|
||||
route.fulfill({ body: JSON.stringify({ data: 'loaded' }) });
|
||||
});
|
||||
}"
|
||||
```
|
||||
@@ -0,0 +1,241 @@
|
||||
# Running Custom Playwright Code
|
||||
|
||||
Use `run-code` to execute arbitrary Playwright code for advanced scenarios not covered by CLI commands.
|
||||
|
||||
## Syntax
|
||||
|
||||
```bash
|
||||
playwright-cli run-code "async page => {
|
||||
// Your Playwright code here
|
||||
// Access page.context() for browser context operations
|
||||
}"
|
||||
```
|
||||
|
||||
You can also load the function from a file:
|
||||
|
||||
```bash
|
||||
playwright-cli run-code --filename=./my-script.js
|
||||
```
|
||||
|
||||
|
||||
The code must be a single function expression, it is wrapped in `(...)` and evaluated.
|
||||
import/export/require syntax is not supported.
|
||||
|
||||
## Geolocation
|
||||
|
||||
```bash
|
||||
# Grant geolocation permission and set location
|
||||
playwright-cli run-code "async page => {
|
||||
await page.context().grantPermissions(['geolocation']);
|
||||
await page.context().setGeolocation({ latitude: 37.7749, longitude: -122.4194 });
|
||||
}"
|
||||
|
||||
# Set location to London
|
||||
playwright-cli run-code "async page => {
|
||||
await page.context().grantPermissions(['geolocation']);
|
||||
await page.context().setGeolocation({ latitude: 51.5074, longitude: -0.1278 });
|
||||
}"
|
||||
|
||||
# Clear geolocation override
|
||||
playwright-cli run-code "async page => {
|
||||
await page.context().clearPermissions();
|
||||
}"
|
||||
```
|
||||
|
||||
## Permissions
|
||||
|
||||
```bash
|
||||
# Grant multiple permissions
|
||||
playwright-cli run-code "async page => {
|
||||
await page.context().grantPermissions([
|
||||
'geolocation',
|
||||
'notifications',
|
||||
'camera',
|
||||
'microphone'
|
||||
]);
|
||||
}"
|
||||
|
||||
# Grant permissions for specific origin
|
||||
playwright-cli run-code "async page => {
|
||||
await page.context().grantPermissions(['clipboard-read'], {
|
||||
origin: 'https://example.com'
|
||||
});
|
||||
}"
|
||||
```
|
||||
|
||||
## Media Emulation
|
||||
|
||||
```bash
|
||||
# Emulate dark color scheme
|
||||
playwright-cli run-code "async page => {
|
||||
await page.emulateMedia({ colorScheme: 'dark' });
|
||||
}"
|
||||
|
||||
# Emulate light color scheme
|
||||
playwright-cli run-code "async page => {
|
||||
await page.emulateMedia({ colorScheme: 'light' });
|
||||
}"
|
||||
|
||||
# Emulate reduced motion
|
||||
playwright-cli run-code "async page => {
|
||||
await page.emulateMedia({ reducedMotion: 'reduce' });
|
||||
}"
|
||||
|
||||
# Emulate print media
|
||||
playwright-cli run-code "async page => {
|
||||
await page.emulateMedia({ media: 'print' });
|
||||
}"
|
||||
```
|
||||
|
||||
## Wait Strategies
|
||||
|
||||
```bash
|
||||
# Wait for network idle
|
||||
playwright-cli run-code "async page => {
|
||||
await page.waitForLoadState('networkidle');
|
||||
}"
|
||||
|
||||
# Wait for specific element
|
||||
playwright-cli run-code "async page => {
|
||||
await page.locator('.loading').waitFor({ state: 'hidden' });
|
||||
}"
|
||||
|
||||
# Wait for function to return true
|
||||
playwright-cli run-code "async page => {
|
||||
await page.waitForFunction(() => window.appReady === true);
|
||||
}"
|
||||
|
||||
# Wait with timeout
|
||||
playwright-cli run-code "async page => {
|
||||
await page.locator('.result').waitFor({ timeout: 10000 });
|
||||
}"
|
||||
```
|
||||
|
||||
## Frames and Iframes
|
||||
|
||||
```bash
|
||||
# Work with iframe
|
||||
playwright-cli run-code "async page => {
|
||||
const frame = page.locator('iframe#my-iframe').contentFrame();
|
||||
await frame.locator('button').click();
|
||||
}"
|
||||
|
||||
# Get all frames
|
||||
playwright-cli run-code "async page => {
|
||||
const frames = page.frames();
|
||||
return frames.map(f => f.url());
|
||||
}"
|
||||
```
|
||||
|
||||
## File Downloads
|
||||
|
||||
```bash
|
||||
# Handle file download
|
||||
playwright-cli run-code "async page => {
|
||||
const downloadPromise = page.waitForEvent('download');
|
||||
await page.getByRole('link', { name: 'Download' }).click();
|
||||
const download = await downloadPromise;
|
||||
await download.saveAs('./downloaded-file.pdf');
|
||||
return download.suggestedFilename();
|
||||
}"
|
||||
```
|
||||
|
||||
## Clipboard
|
||||
|
||||
```bash
|
||||
# Read clipboard (requires permission)
|
||||
playwright-cli run-code "async page => {
|
||||
await page.context().grantPermissions(['clipboard-read']);
|
||||
return await page.evaluate(() => navigator.clipboard.readText());
|
||||
}"
|
||||
|
||||
# Write to clipboard
|
||||
playwright-cli run-code "async page => {
|
||||
await page.evaluate(text => navigator.clipboard.writeText(text), 'Hello clipboard!');
|
||||
}"
|
||||
```
|
||||
|
||||
## Page Information
|
||||
|
||||
```bash
|
||||
# Get page title
|
||||
playwright-cli run-code "async page => {
|
||||
return await page.title();
|
||||
}"
|
||||
|
||||
# Get current URL
|
||||
playwright-cli run-code "async page => {
|
||||
return page.url();
|
||||
}"
|
||||
|
||||
# Get page content
|
||||
playwright-cli run-code "async page => {
|
||||
return await page.content();
|
||||
}"
|
||||
|
||||
# Get viewport size
|
||||
playwright-cli run-code "async page => {
|
||||
return page.viewportSize();
|
||||
}"
|
||||
```
|
||||
|
||||
## JavaScript Execution
|
||||
|
||||
```bash
|
||||
# Execute JavaScript and return result
|
||||
playwright-cli run-code "async page => {
|
||||
return await page.evaluate(() => {
|
||||
return {
|
||||
userAgent: navigator.userAgent,
|
||||
language: navigator.language,
|
||||
cookiesEnabled: navigator.cookieEnabled
|
||||
};
|
||||
});
|
||||
}"
|
||||
|
||||
# Pass arguments to evaluate
|
||||
playwright-cli run-code "async page => {
|
||||
const multiplier = 5;
|
||||
return await page.evaluate(m => document.querySelectorAll('li').length * m, multiplier);
|
||||
}"
|
||||
```
|
||||
|
||||
## Error Handling
|
||||
|
||||
```bash
|
||||
# Try-catch in run-code
|
||||
playwright-cli run-code "async page => {
|
||||
try {
|
||||
await page.getByRole('button', { name: 'Submit' }).click({ timeout: 1000 });
|
||||
return 'clicked';
|
||||
} catch (e) {
|
||||
return 'element not found';
|
||||
}
|
||||
}"
|
||||
```
|
||||
|
||||
## Complex Workflows
|
||||
|
||||
```bash
|
||||
# Login and save state
|
||||
playwright-cli run-code "async page => {
|
||||
await page.goto('https://example.com/login');
|
||||
await page.getByRole('textbox', { name: 'Email' }).fill('user@example.com');
|
||||
await page.getByRole('textbox', { name: 'Password' }).fill('secret');
|
||||
await page.getByRole('button', { name: 'Sign in' }).click();
|
||||
await page.waitForURL('**/dashboard');
|
||||
await page.context().storageState({ path: 'auth.json' });
|
||||
return 'Login successful';
|
||||
}"
|
||||
|
||||
# Scrape data from multiple pages
|
||||
playwright-cli run-code "async page => {
|
||||
const results = [];
|
||||
for (let i = 1; i <= 3; i++) {
|
||||
await page.goto(\`https://example.com/page/\${i}\`);
|
||||
const items = await page.locator('.item').allTextContents();
|
||||
results.push(...items);
|
||||
}
|
||||
return results;
|
||||
}"
|
||||
```
|
||||
@@ -0,0 +1,225 @@
|
||||
# Browser Session Management
|
||||
|
||||
Run multiple isolated browser sessions concurrently with state persistence.
|
||||
|
||||
## Named Browser Sessions
|
||||
|
||||
Use `-s` flag to isolate browser contexts:
|
||||
|
||||
```bash
|
||||
# Browser 1: Authentication flow
|
||||
playwright-cli -s=auth open https://app.example.com/login
|
||||
|
||||
# Browser 2: Public browsing (separate cookies, storage)
|
||||
playwright-cli -s=public open https://example.com
|
||||
|
||||
# Commands are isolated by browser session
|
||||
playwright-cli -s=auth fill e1 "user@example.com"
|
||||
playwright-cli -s=public snapshot
|
||||
```
|
||||
|
||||
## Browser Session Isolation Properties
|
||||
|
||||
Each browser session has independent:
|
||||
- Cookies
|
||||
- LocalStorage / SessionStorage
|
||||
- IndexedDB
|
||||
- Cache
|
||||
- Browsing history
|
||||
- Open tabs
|
||||
|
||||
## Browser Session Commands
|
||||
|
||||
```bash
|
||||
# List all browser sessions
|
||||
playwright-cli list
|
||||
|
||||
# Stop a browser session (close the browser)
|
||||
playwright-cli close # stop the default browser
|
||||
playwright-cli -s=mysession close # stop a named browser
|
||||
|
||||
# Stop all browser sessions
|
||||
playwright-cli close-all
|
||||
|
||||
# Forcefully kill all daemon processes (for stale/zombie processes)
|
||||
playwright-cli kill-all
|
||||
|
||||
# Delete browser session user data (profile directory)
|
||||
playwright-cli delete-data # delete default browser data
|
||||
playwright-cli -s=mysession delete-data # delete named browser data
|
||||
```
|
||||
|
||||
## Environment Variable
|
||||
|
||||
Set a default browser session name via environment variable:
|
||||
|
||||
```bash
|
||||
export PLAYWRIGHT_CLI_SESSION="mysession"
|
||||
playwright-cli open example.com # Uses "mysession" automatically
|
||||
```
|
||||
|
||||
## Common Patterns
|
||||
|
||||
### Concurrent Scraping
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
# Scrape multiple sites concurrently
|
||||
|
||||
# Start all browsers
|
||||
playwright-cli -s=site1 open https://site1.com &
|
||||
playwright-cli -s=site2 open https://site2.com &
|
||||
playwright-cli -s=site3 open https://site3.com &
|
||||
wait
|
||||
|
||||
# Take snapshots from each
|
||||
playwright-cli -s=site1 snapshot
|
||||
playwright-cli -s=site2 snapshot
|
||||
playwright-cli -s=site3 snapshot
|
||||
|
||||
# Cleanup
|
||||
playwright-cli close-all
|
||||
```
|
||||
|
||||
### A/B Testing Sessions
|
||||
|
||||
```bash
|
||||
# Test different user experiences
|
||||
playwright-cli -s=variant-a open "https://app.com?variant=a"
|
||||
playwright-cli -s=variant-b open "https://app.com?variant=b"
|
||||
|
||||
# Compare
|
||||
playwright-cli -s=variant-a screenshot
|
||||
playwright-cli -s=variant-b screenshot
|
||||
```
|
||||
|
||||
### Persistent Profile
|
||||
|
||||
By default, browser profile is kept in memory only. Use `--persistent` flag on `open` to persist the browser profile to disk:
|
||||
|
||||
```bash
|
||||
# Use persistent profile (auto-generated location)
|
||||
playwright-cli open https://example.com --persistent
|
||||
|
||||
# Use persistent profile with custom directory
|
||||
playwright-cli open https://example.com --profile=/path/to/profile
|
||||
```
|
||||
|
||||
## Attaching to a Running Browser
|
||||
|
||||
Use `attach` to connect to a browser that is already running, instead of launching a new one.
|
||||
|
||||
### Attach by channel name
|
||||
|
||||
Connect to a running Chrome or Edge instance by its channel name. The browser must have remote debugging enabled — navigate to `chrome://inspect/#remote-debugging` in the target browser and check "Allow remote debugging for this browser instance".
|
||||
|
||||
```bash
|
||||
# Attach to Chrome
|
||||
playwright-cli attach --cdp=chrome
|
||||
|
||||
# Attach to Chrome Canary
|
||||
playwright-cli attach --cdp=chrome-canary
|
||||
|
||||
# Attach to Microsoft Edge
|
||||
playwright-cli attach --cdp=msedge
|
||||
|
||||
# Attach to Edge Dev
|
||||
playwright-cli attach --cdp=msedge-dev
|
||||
```
|
||||
|
||||
Supported channels: `chrome`, `chrome-beta`, `chrome-dev`, `chrome-canary`, `msedge`, `msedge-beta`, `msedge-dev`, `msedge-canary`.
|
||||
|
||||
When `--session` is not provided, the session is named after the channel (e.g. `--cdp=msedge` creates a session called `msedge`), so parallel attaches to Chrome and Edge don't collide on `default`. Pass `--session=<name>` to override.
|
||||
|
||||
### Attach via CDP endpoint
|
||||
|
||||
Connect to a browser that exposes a Chrome DevTools Protocol endpoint:
|
||||
|
||||
```bash
|
||||
playwright-cli attach --cdp=http://localhost:9222
|
||||
```
|
||||
|
||||
### Attach via browser extension
|
||||
|
||||
Connect to a browser with the Playwright extension installed:
|
||||
|
||||
```bash
|
||||
playwright-cli attach --extension
|
||||
```
|
||||
|
||||
### Detach
|
||||
|
||||
Tear down an attached session without affecting the external browser:
|
||||
|
||||
```bash
|
||||
# Detach the default attached session
|
||||
playwright-cli detach
|
||||
|
||||
# Detach a specific attached session
|
||||
playwright-cli -s=msedge detach
|
||||
```
|
||||
|
||||
`detach` only works on sessions created via `attach`. For sessions created via `open`, use `close`.
|
||||
|
||||
## Default Browser Session
|
||||
|
||||
When `-s` is omitted, commands use the default browser session:
|
||||
|
||||
```bash
|
||||
# These use the same default browser session
|
||||
playwright-cli open https://example.com
|
||||
playwright-cli snapshot
|
||||
playwright-cli close # Stops default browser
|
||||
```
|
||||
|
||||
## Browser Session Configuration
|
||||
|
||||
Configure a browser session with specific settings when opening:
|
||||
|
||||
```bash
|
||||
# Open with config file
|
||||
playwright-cli open https://example.com --config=.playwright/my-cli.json
|
||||
|
||||
# Open with specific browser
|
||||
playwright-cli open https://example.com --browser=firefox
|
||||
|
||||
# Open in headed mode
|
||||
playwright-cli open https://example.com --headed
|
||||
|
||||
# Open with persistent profile
|
||||
playwright-cli open https://example.com --persistent
|
||||
```
|
||||
|
||||
## Best Practices
|
||||
|
||||
### 1. Name Browser Sessions Semantically
|
||||
|
||||
```bash
|
||||
# GOOD: Clear purpose
|
||||
playwright-cli -s=github-auth open https://github.com
|
||||
playwright-cli -s=docs-scrape open https://docs.example.com
|
||||
|
||||
# AVOID: Generic names
|
||||
playwright-cli -s=s1 open https://github.com
|
||||
```
|
||||
|
||||
### 2. Always Clean Up
|
||||
|
||||
```bash
|
||||
# Stop browsers when done
|
||||
playwright-cli -s=auth close
|
||||
playwright-cli -s=scrape close
|
||||
|
||||
# Or stop all at once
|
||||
playwright-cli close-all
|
||||
|
||||
# If browsers become unresponsive or zombie processes remain
|
||||
playwright-cli kill-all
|
||||
```
|
||||
|
||||
### 3. Delete Stale Browser Data
|
||||
|
||||
```bash
|
||||
# Remove old browser data to free disk space
|
||||
playwright-cli -s=oldsession delete-data
|
||||
```
|
||||
@@ -0,0 +1,275 @@
|
||||
# Storage Management
|
||||
|
||||
Manage cookies, localStorage, sessionStorage, and browser storage state.
|
||||
|
||||
## Storage State
|
||||
|
||||
Save and restore complete browser state including cookies and storage.
|
||||
|
||||
### Save Storage State
|
||||
|
||||
```bash
|
||||
# Save to auto-generated filename (storage-state-{timestamp}.json)
|
||||
playwright-cli state-save
|
||||
|
||||
# Save to specific filename
|
||||
playwright-cli state-save my-auth-state.json
|
||||
```
|
||||
|
||||
### Restore Storage State
|
||||
|
||||
```bash
|
||||
# Load storage state from file
|
||||
playwright-cli state-load my-auth-state.json
|
||||
|
||||
# Reload page to apply cookies
|
||||
playwright-cli open https://example.com
|
||||
```
|
||||
|
||||
### Storage State File Format
|
||||
|
||||
The saved file contains:
|
||||
|
||||
```json
|
||||
{
|
||||
"cookies": [
|
||||
{
|
||||
"name": "session_id",
|
||||
"value": "abc123",
|
||||
"domain": "example.com",
|
||||
"path": "/",
|
||||
"expires": 1893456000,
|
||||
"httpOnly": true,
|
||||
"secure": true,
|
||||
"sameSite": "Lax"
|
||||
}
|
||||
],
|
||||
"origins": [
|
||||
{
|
||||
"origin": "https://example.com",
|
||||
"localStorage": [
|
||||
{ "name": "theme", "value": "dark" },
|
||||
{ "name": "user_id", "value": "12345" }
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Cookies
|
||||
|
||||
### List All Cookies
|
||||
|
||||
```bash
|
||||
playwright-cli cookie-list
|
||||
```
|
||||
|
||||
### Filter Cookies by Domain
|
||||
|
||||
```bash
|
||||
playwright-cli cookie-list --domain=example.com
|
||||
```
|
||||
|
||||
### Filter Cookies by Path
|
||||
|
||||
```bash
|
||||
playwright-cli cookie-list --path=/api
|
||||
```
|
||||
|
||||
### Get Specific Cookie
|
||||
|
||||
```bash
|
||||
playwright-cli cookie-get session_id
|
||||
```
|
||||
|
||||
### Set a Cookie
|
||||
|
||||
```bash
|
||||
# Basic cookie
|
||||
playwright-cli cookie-set session abc123
|
||||
|
||||
# Cookie with options
|
||||
playwright-cli cookie-set session abc123 --domain=example.com --path=/ --httpOnly --secure --sameSite=Lax
|
||||
|
||||
# Cookie with expiration (Unix timestamp)
|
||||
playwright-cli cookie-set remember_me token123 --expires=1893456000
|
||||
```
|
||||
|
||||
### Delete a Cookie
|
||||
|
||||
```bash
|
||||
playwright-cli cookie-delete session_id
|
||||
```
|
||||
|
||||
### Clear All Cookies
|
||||
|
||||
```bash
|
||||
playwright-cli cookie-clear
|
||||
```
|
||||
|
||||
### Advanced: Multiple Cookies or Custom Options
|
||||
|
||||
For complex scenarios like adding multiple cookies at once, use `run-code`:
|
||||
|
||||
```bash
|
||||
playwright-cli run-code "async page => {
|
||||
await page.context().addCookies([
|
||||
{ name: 'session_id', value: 'sess_abc123', domain: 'example.com', path: '/', httpOnly: true },
|
||||
{ name: 'preferences', value: JSON.stringify({ theme: 'dark' }), domain: 'example.com', path: '/' }
|
||||
]);
|
||||
}"
|
||||
```
|
||||
|
||||
## Local Storage
|
||||
|
||||
### List All localStorage Items
|
||||
|
||||
```bash
|
||||
playwright-cli localstorage-list
|
||||
```
|
||||
|
||||
### Get Single Value
|
||||
|
||||
```bash
|
||||
playwright-cli localstorage-get token
|
||||
```
|
||||
|
||||
### Set Value
|
||||
|
||||
```bash
|
||||
playwright-cli localstorage-set theme dark
|
||||
```
|
||||
|
||||
### Set JSON Value
|
||||
|
||||
```bash
|
||||
playwright-cli localstorage-set user_settings '{"theme":"dark","language":"en"}'
|
||||
```
|
||||
|
||||
### Delete Single Item
|
||||
|
||||
```bash
|
||||
playwright-cli localstorage-delete token
|
||||
```
|
||||
|
||||
### Clear All localStorage
|
||||
|
||||
```bash
|
||||
playwright-cli localstorage-clear
|
||||
```
|
||||
|
||||
### Advanced: Multiple Operations
|
||||
|
||||
For complex scenarios like setting multiple values at once, use `run-code`:
|
||||
|
||||
```bash
|
||||
playwright-cli run-code "async page => {
|
||||
await page.evaluate(() => {
|
||||
localStorage.setItem('token', 'jwt_abc123');
|
||||
localStorage.setItem('user_id', '12345');
|
||||
localStorage.setItem('expires_at', Date.now() + 3600000);
|
||||
});
|
||||
}"
|
||||
```
|
||||
|
||||
## Session Storage
|
||||
|
||||
### List All sessionStorage Items
|
||||
|
||||
```bash
|
||||
playwright-cli sessionstorage-list
|
||||
```
|
||||
|
||||
### Get Single Value
|
||||
|
||||
```bash
|
||||
playwright-cli sessionstorage-get form_data
|
||||
```
|
||||
|
||||
### Set Value
|
||||
|
||||
```bash
|
||||
playwright-cli sessionstorage-set step 3
|
||||
```
|
||||
|
||||
### Delete Single Item
|
||||
|
||||
```bash
|
||||
playwright-cli sessionstorage-delete step
|
||||
```
|
||||
|
||||
### Clear sessionStorage
|
||||
|
||||
```bash
|
||||
playwright-cli sessionstorage-clear
|
||||
```
|
||||
|
||||
## IndexedDB
|
||||
|
||||
### List Databases
|
||||
|
||||
```bash
|
||||
playwright-cli run-code "async page => {
|
||||
return await page.evaluate(async () => {
|
||||
const databases = await indexedDB.databases();
|
||||
return databases;
|
||||
});
|
||||
}"
|
||||
```
|
||||
|
||||
### Delete Database
|
||||
|
||||
```bash
|
||||
playwright-cli run-code "async page => {
|
||||
await page.evaluate(() => {
|
||||
indexedDB.deleteDatabase('myDatabase');
|
||||
});
|
||||
}"
|
||||
```
|
||||
|
||||
## Common Patterns
|
||||
|
||||
### Authentication State Reuse
|
||||
|
||||
```bash
|
||||
# Step 1: Login and save state
|
||||
playwright-cli open https://app.example.com/login
|
||||
playwright-cli snapshot
|
||||
playwright-cli fill e1 "user@example.com"
|
||||
playwright-cli fill e2 "password123"
|
||||
playwright-cli click e3
|
||||
|
||||
# Save the authenticated state
|
||||
playwright-cli state-save auth.json
|
||||
|
||||
# Step 2: Later, restore state and skip login
|
||||
playwright-cli state-load auth.json
|
||||
playwright-cli open https://app.example.com/dashboard
|
||||
# Already logged in!
|
||||
```
|
||||
|
||||
### Save and Restore Roundtrip
|
||||
|
||||
```bash
|
||||
# Set up authentication state
|
||||
playwright-cli open https://example.com
|
||||
playwright-cli eval "() => { document.cookie = 'session=abc123'; localStorage.setItem('user', 'john'); }"
|
||||
|
||||
# Save state to file
|
||||
playwright-cli state-save my-session.json
|
||||
|
||||
# ... later, in a new session ...
|
||||
|
||||
# Restore state
|
||||
playwright-cli state-load my-session.json
|
||||
playwright-cli open https://example.com
|
||||
# Cookies and localStorage are restored!
|
||||
```
|
||||
|
||||
## Security Notes
|
||||
|
||||
- Never commit storage state files containing auth tokens
|
||||
- Add `*.auth-state.json` to `.gitignore`
|
||||
- Delete state files after automation completes
|
||||
- Use environment variables for sensitive data
|
||||
- By default, sessions run in-memory mode which is safer for sensitive operations
|
||||
@@ -0,0 +1,433 @@
|
||||
# Test generation (plan → generate → heal)
|
||||
|
||||
End-to-end workflow for authoring and maintaining Playwright tests with `playwright-cli`. Every `playwright-cli` action emits the equivalent Playwright TypeScript, and that generated code is the raw material for every test. The sections below can be used independently:
|
||||
|
||||
- **How generation works** — the core mechanic everything else relies on: actions become TypeScript, plus how to add assertions.
|
||||
- **Plan** — explore the app, produce a spec file describing what to test.
|
||||
- **Generate** — turn a spec into Playwright test files. Update the spec if it's vague or stale.
|
||||
- **Heal** — diagnose failing tests, fix the code, reconcile the spec with reality.
|
||||
|
||||
Plan / generate / heal lean on the same mechanic: run `npx playwright test --debug=cli` in the background, then `playwright-cli attach tw-XXXX` to drive the paused page interactively. See [playwright-tests.md](playwright-tests.md) for the debug/attach mechanics.
|
||||
|
||||
---
|
||||
|
||||
## 0. How generation works
|
||||
|
||||
Every action you perform with `playwright-cli` generates corresponding Playwright TypeScript code. This code appears in the output and can be copied directly into your test files.
|
||||
|
||||
```bash
|
||||
# Start a session
|
||||
playwright-cli open https://example.com/login
|
||||
|
||||
# Take a snapshot to see elements
|
||||
playwright-cli snapshot
|
||||
# Output shows: e1 [textbox "Email"], e2 [textbox "Password"], e3 [button "Sign In"]
|
||||
|
||||
# Fill form fields - generates code automatically
|
||||
playwright-cli fill e1 "user@example.com"
|
||||
# Ran Playwright code:
|
||||
# await page.getByRole('textbox', { name: 'Email' }).fill('user@example.com');
|
||||
|
||||
playwright-cli fill e2 "password123"
|
||||
# Ran Playwright code:
|
||||
# await page.getByRole('textbox', { name: 'Password' }).fill('password123');
|
||||
|
||||
playwright-cli click e3
|
||||
# Ran Playwright code:
|
||||
# await page.getByRole('button', { name: 'Sign In' }).click();
|
||||
```
|
||||
|
||||
### Building a test file
|
||||
|
||||
Collect the generated code into a Playwright test:
|
||||
|
||||
```typescript
|
||||
import { test, expect } from '@playwright/test';
|
||||
|
||||
test('login flow', async ({ page }) => {
|
||||
// Generated code from playwright-cli session:
|
||||
await page.goto('https://example.com/login');
|
||||
await page.getByRole('textbox', { name: 'Email' }).fill('user@example.com');
|
||||
await page.getByRole('textbox', { name: 'Password' }).fill('password123');
|
||||
await page.getByRole('button', { name: 'Sign In' }).click();
|
||||
|
||||
// Add assertions
|
||||
await expect(page).toHaveURL(/.*dashboard/);
|
||||
});
|
||||
```
|
||||
|
||||
### Use semantic locators
|
||||
|
||||
The generated code uses role-based locators when possible, which are more resilient:
|
||||
|
||||
```typescript
|
||||
// Generated (good - semantic)
|
||||
await page.getByRole('button', { name: 'Submit' }).click();
|
||||
|
||||
// Avoid (fragile - CSS selectors)
|
||||
await page.locator('#submit-btn').click();
|
||||
```
|
||||
|
||||
### Explore before recording
|
||||
|
||||
Take snapshots to understand the page structure before recording actions:
|
||||
|
||||
```bash
|
||||
playwright-cli open https://example.com
|
||||
playwright-cli snapshot
|
||||
# Review the element structure
|
||||
playwright-cli click e5
|
||||
```
|
||||
|
||||
### Add assertions manually
|
||||
|
||||
Generated code captures actions but not assertions. Add expectations in your test using one of the recommended matchers:
|
||||
|
||||
- `toBeVisible()` — element is rendered and visible
|
||||
- `toHaveText(text)` — element text content matches
|
||||
- `toHaveValue(value) / toBeEmpty()` — input/select value matches
|
||||
- `toBeChecked() / toBeUnchecked()` — checkbox state matches
|
||||
- `toMatchAriaSnapshot(snapshot)` — page (or locator) matches a partial accessibility snapshot
|
||||
|
||||
Use `playwright-cli generate-locator <target>` to produce the locator expression for the assertion, and the snapshot/eval commands to capture the expected value.
|
||||
|
||||
When asserting text content, make sure that generated locator does not contain text from the element itself. `getByTestId()` or `getByLabel()` usually work well with asserting text. When locator is text-based, prefer `toBeVisible()` instead.
|
||||
|
||||
Snapshot to be matched does not have to contain all the information - only capture what's necessary for the assertion. You can use regular expressions for unstable values.
|
||||
|
||||
```bash
|
||||
# Get a stable locator for an element ref to use in the assertion
|
||||
playwright-cli --raw generate-locator e5
|
||||
# getByRole('button', { name: 'Submit' })
|
||||
|
||||
# Capture expected text content for toHaveText
|
||||
playwright-cli --raw eval "el => el.textContent" e5
|
||||
|
||||
# Capture expected input value for toHaveValue/toBeEmpty
|
||||
playwright-cli --raw eval "el => el.value" e5
|
||||
|
||||
# Capture expected aria snapshot for toMatchAriaSnapshot/toBeChecked
|
||||
# (whole page, or use a ref to scope to a region)
|
||||
playwright-cli --raw snapshot
|
||||
playwright-cli --raw snapshot e5
|
||||
```
|
||||
|
||||
```typescript
|
||||
// Generated action
|
||||
await page.getByRole('button', { name: 'Submit' }).click();
|
||||
|
||||
// Manual assertions using the outputs above:
|
||||
await expect(page.getByRole('alert', { name: 'Success' })).toBeVisible();
|
||||
await expect(page.getByTestId('main-header')).toHaveText('Welcome, user');
|
||||
await expect(page.getByRole('textbox', { name: 'Email' })).toHaveValue('user@example.com');
|
||||
await expect(page.getByRole('checkbox', { name: 'Enable notifications' })).toBeChecked();
|
||||
|
||||
// toMatchAriaSnapshot on the whole page, finds a matching region
|
||||
await expect(page).toMatchAriaSnapshot(`
|
||||
- heading "Welcome, user"
|
||||
- link /\\d+ new messages?/
|
||||
- button "Sign out"
|
||||
`);
|
||||
|
||||
// toMatchAriaSnapshot scoped to a region
|
||||
await expect(page.getByRole('navigation')).toMatchAriaSnapshot(`
|
||||
- link "Home"
|
||||
- link /\\d+ new messages?/
|
||||
- link "Profile"
|
||||
`);
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 1. Planning
|
||||
|
||||
Goal: produce a spec file (e.g. `specs/<feature>.plan.md`) that enumerates the scenarios to test. **Always** write the spec to a file.
|
||||
|
||||
### 1.1 Prerequisite: workspace
|
||||
|
||||
Check the workspace has Playwright installed before anything else:
|
||||
|
||||
```bash
|
||||
# Either of these confirms a workspace:
|
||||
test -f playwright.config.ts || test -f playwright.config.js
|
||||
npx --no-install playwright --version
|
||||
```
|
||||
|
||||
If there is no Playwright install, bootstrap one and let the user pick the defaults:
|
||||
|
||||
```bash
|
||||
npm init playwright@latest
|
||||
```
|
||||
|
||||
### 1.2 Prerequisite: seed test
|
||||
|
||||
A **seed test** is a minimal test that lands the page in the state every scenario starts from: navigation to the app, any required login, feature flags, etc. Scenarios assume a fresh start *after* the seed. `--debug=cli` pauses *inside* this test, so the seed is where every planning and generation session begins.
|
||||
|
||||
Minimum viable seed:
|
||||
|
||||
```ts
|
||||
// tests/seed.spec.ts
|
||||
import { test } from '@playwright/test';
|
||||
|
||||
test('seed', async ({ page }) => {
|
||||
await page.goto('https://example.com/');
|
||||
});
|
||||
```
|
||||
|
||||
Preferred — push navigation into a fixture so scenario tests reuse it:
|
||||
|
||||
```ts
|
||||
// tests/fixtures.ts
|
||||
import { test as baseTest } from '@playwright/test';
|
||||
export { expect } from '@playwright/test';
|
||||
|
||||
export const test = baseTest.extend({
|
||||
page: async ({ page }, use) => {
|
||||
await page.goto('https://example.com/');
|
||||
await use(page);
|
||||
},
|
||||
});
|
||||
```
|
||||
|
||||
```ts
|
||||
// tests/seed.spec.ts
|
||||
import { test } from './fixtures';
|
||||
|
||||
test('seed', async ({ page }) => {
|
||||
// Fixture already navigates. This empty body tells agents where to start.
|
||||
});
|
||||
```
|
||||
|
||||
If no seed exists, create one that at least navigates to the app.
|
||||
|
||||
### 1.3 Explore the app
|
||||
|
||||
Launch the app via the seed in the background and attach:
|
||||
|
||||
```bash
|
||||
PLAYWRIGHT_HTML_OPEN=never npx playwright test tests/seed.spec.ts --debug=cli
|
||||
# wait for "Debugging Instructions" and the session name tw-XXXX
|
||||
playwright-cli attach tw-XXXX
|
||||
```
|
||||
|
||||
Resume so the seed runs, then probe the app:
|
||||
|
||||
```bash
|
||||
playwright-cli resume # resume so that seed test runs fully
|
||||
playwright-cli snapshot # inventory of interactive elements
|
||||
playwright-cli click e5 # follow a flow
|
||||
playwright-cli eval "location.href" # read URL / state
|
||||
playwright-cli show --annotate # ask the user to point at something
|
||||
```
|
||||
|
||||
Map out:
|
||||
|
||||
- Interactive surfaces (forms, buttons, lists, filters, modals).
|
||||
- Primary user journeys end-to-end.
|
||||
- Edge cases: empty states, validation errors, very long input, boundary values.
|
||||
- Persistence: reload, local/session storage, URL fragments.
|
||||
- Navigation: which controls change the URL, back/forward behaviour.
|
||||
|
||||
**Important**: Do not just open the app url with playwright-cli, always go through the test to capture any custom setup done there.
|
||||
**Important**: Stop the background test when done exploring.
|
||||
|
||||
### 1.4 Write the spec file
|
||||
|
||||
Save under `specs/<feature>.plan.md`. Use this structure:
|
||||
|
||||
```markdown
|
||||
# <Feature> Test Plan
|
||||
|
||||
## Application Overview
|
||||
|
||||
<One paragraph describing what the feature does and why it matters.>
|
||||
|
||||
## Test Scenarios
|
||||
|
||||
### 1. <Group Name>
|
||||
|
||||
**Seed:** `tests/seed.spec.ts`
|
||||
|
||||
#### 1.1. <kebab-case-scenario-name>
|
||||
|
||||
**File:** `tests/<group>/<kebab-case-scenario-name>.spec.ts`
|
||||
|
||||
**Steps:**
|
||||
1. <Concrete user step>
|
||||
- expect: <observable outcome>
|
||||
- expect: <another observable outcome>
|
||||
2. <Next step>
|
||||
- expect: <outcome>
|
||||
|
||||
#### 1.2. <next-scenario>
|
||||
...
|
||||
|
||||
### 2. <Next Group>
|
||||
|
||||
**Seed:** `tests/seed.spec.ts`
|
||||
...
|
||||
```
|
||||
|
||||
Guidelines:
|
||||
|
||||
- Each scenario is independent and starts from the seed's fresh state — never chain scenarios.
|
||||
- Scenario names are kebab-case and match the test file name (`should-add-single-todo` → `should-add-single-todo.spec.ts`).
|
||||
- Cover happy path, edge cases, validation, negative flows, persistence.
|
||||
- Write steps at the user level ("Type 'Buy milk' into the input"), not the API level ("call `fill`").
|
||||
- Put observable outcomes in `- expect:` bullets; each becomes an assertion during generation.
|
||||
|
||||
---
|
||||
|
||||
## 2. Generate
|
||||
|
||||
Goal: take a spec file and produce Playwright test files. Optionally update the spec if it has drifted.
|
||||
|
||||
### 2.1 Inputs
|
||||
|
||||
- **Spec file**, e.g. `specs/basic-operations.plan.md`.
|
||||
- **Target**: either a single scenario (e.g. `1.2`), a whole group (`1`), or all.
|
||||
- **Seed file**, read from the `**Seed:**` line of the scenario's group.
|
||||
|
||||
### 2.2 Generate one scenario
|
||||
|
||||
For each target scenario, in sequence (never in parallel — scenarios share the seed session):
|
||||
|
||||
```bash
|
||||
PLAYWRIGHT_HTML_OPEN=never npx playwright test <seed-file> --debug=cli # background
|
||||
playwright-cli attach tw-XXXX
|
||||
# resume
|
||||
```
|
||||
|
||||
**Do not** just open the app url with playwright-cli, always go through the test to capture any custom setup done there.
|
||||
|
||||
Walk the scenario's `Steps:` one by one with `playwright-cli`, treating the spec as the plan and the live app as the source of truth. If a step is vague ("click the button" — which button?), references an element that no longer exists, or contradicts the app's actual behaviour, use your judgement: update the spec to match what the app really does, then keep going. Editing the spec mid-generation is expected.
|
||||
|
||||
Every action prints the equivalent Playwright TypeScript (see [How generation works](#0-how-generation-works)):
|
||||
|
||||
```bash
|
||||
playwright-cli snapshot # find refs
|
||||
playwright-cli fill e3 "John Doe" # -> page.getByRole('textbox', {...}).fill(...)
|
||||
playwright-cli press Enter
|
||||
playwright-cli click e7
|
||||
```
|
||||
|
||||
For each `- expect:` bullet, add an explicit assertion. See [How generation works](#0-how-generation-works) for details.
|
||||
|
||||
Collect the generated code and write the test file at the path given in the spec:
|
||||
|
||||
```ts
|
||||
// spec: specs/basic-operations.plan.md
|
||||
// seed: tests/seed.spec.ts
|
||||
import { test, expect } from './fixtures'; // or '@playwright/test' if no fixtures file
|
||||
|
||||
test.describe('Signing in and out', () => {
|
||||
test('should sign in', async ({ page }) => {
|
||||
// 1. Navigate to the application
|
||||
// (handled by the seed fixture)
|
||||
|
||||
// 2. Type 'John Doe' into the username field
|
||||
await page.getByRole('textbox', { name: 'username' }).fill('John Doe');
|
||||
|
||||
// 3. Type password
|
||||
await page.getByRole('textbox', { name: 'password' }).fill('TestPassword');
|
||||
|
||||
// 4. Press Enter to submit
|
||||
await page.getByRole('textbox', { name: 'password' }).press('Enter');
|
||||
|
||||
await expect(page.getByRole('heading')).toContainText('Welcome, John Doe!');
|
||||
});
|
||||
});
|
||||
```
|
||||
|
||||
Rules:
|
||||
|
||||
- **One test per file.** File path, describe name, and test name come verbatim from the spec (minus the ordinal).
|
||||
- Prefix each numbered step with a `// N. <step text>` comment before its actions.
|
||||
- Use the describe group name verbatim from the spec (no `1.` ordinal).
|
||||
- Import from `./fixtures` if the project has one; otherwise `@playwright/test`.
|
||||
- **Important**: close the CLI session and stop the background test before moving to the next scenario.
|
||||
|
||||
### 2.3 Generate multiple scenarios
|
||||
|
||||
Loop 2.2 over the targeted scenarios one at a time, restarting the seed between each so every test starts from a clean page. This is safe to parallelise due to unique generated session names - just make sure each test run is stopped.
|
||||
|
||||
### 2.4 Run generated tests
|
||||
|
||||
After generation, run the new tests once:
|
||||
|
||||
```bash
|
||||
PLAYWRIGHT_HTML_OPEN=never npx playwright test tests/<group>/<scenario>.spec.ts
|
||||
```
|
||||
|
||||
Any failure goes to Section 3.
|
||||
|
||||
---
|
||||
|
||||
## 3. Heal
|
||||
|
||||
Goal: fix failing tests, and update the spec if the app's intended behaviour changed.
|
||||
|
||||
### 3.1 Find failing tests
|
||||
|
||||
```bash
|
||||
PLAYWRIGHT_HTML_OPEN=never npx playwright test
|
||||
```
|
||||
|
||||
Record the list of failing `<file>:<line>` entries and process them one at a time. Do not attempt parallel fixes — shared state and the single CLI session make that fragile.
|
||||
|
||||
### 3.2 Debug one failure
|
||||
|
||||
Run the single failing test in debug mode in the background, then attach:
|
||||
|
||||
```bash
|
||||
PLAYWRIGHT_HTML_OPEN=never npx playwright test tests/<group>/<scenario>.spec.ts:<line> --debug=cli
|
||||
# wait for "Debugging Instructions" and the tw-XXXX session name
|
||||
playwright-cli attach tw-XXXX
|
||||
```
|
||||
|
||||
The test is paused at the start. Step forward or run to until just before the failing action or assertion, then diagnose:
|
||||
|
||||
```bash
|
||||
playwright-cli snapshot # did the element change / move / rename?
|
||||
playwright-cli console # app-side errors?
|
||||
playwright-cli requests # failed request? wrong payload?
|
||||
playwright-cli show --annotate # ask the user to point somewhere
|
||||
```
|
||||
|
||||
Common causes: selector drift, new wrapper element, label/ARIA rename, timing (transition, async load), assertion text updated in the app, test data leaking between runs.
|
||||
|
||||
Rehearse the corrected interaction with `playwright-cli` — the generated code in the output is what you paste back into the test.
|
||||
|
||||
### 3.3 Apply the fix
|
||||
|
||||
Edit the test file: update the locator, assertion, step order, or inputs to match the corrected behaviour. Stop the background debug run. Rerun the single test to confirm green.
|
||||
|
||||
Never skip hooks or add sleeps as a fix. Never use `networkidle`.
|
||||
|
||||
### 3.4 Reconcile with the spec
|
||||
|
||||
Open the spec referenced by the `// spec:` header in the test file and locate the scenario that matches the test.
|
||||
|
||||
- **Fix was purely technical** (locator drift, better assertion shape) and the spec's user-level behaviour still matches the app → leave the spec alone.
|
||||
- **Fix changed user-visible steps, inputs, order, or expected outcomes** that the spec describes → update the spec to match reality. Keep the scenario id and file path stable; only the step / expect lines change.
|
||||
- **Unclear whether the app change is intentional** (spec is stale) **or a regression** (test was right, app is wrong) → **stop and ask the user**. Provide:
|
||||
- the scenario id (e.g. `2.3`),
|
||||
- the spec lines that no longer match,
|
||||
- the observed app behaviour (quote a snapshot excerpt or a concrete outcome).
|
||||
|
||||
Only after the user answers, either update the spec (intentional change) or file/flag the test as covering a bug (regression).
|
||||
|
||||
### 3.5 Iteration and giving up
|
||||
|
||||
- Fix failures one at a time; rerun after each.
|
||||
- If after thorough investigation you are confident the test is correct but the app is wrong *and* the user has confirmed it's a bug: mark the test `test.fixme(...)` with a comment pointing at the user's decision or issue link. Never silently skip.
|
||||
|
||||
---
|
||||
|
||||
## Cross-references
|
||||
|
||||
| For... | See |
|
||||
|---|---|
|
||||
| `--debug=cli` / attach mechanics | [playwright-tests.md](playwright-tests.md) |
|
||||
| Mocking requests during exploration/generation | [request-mocking.md](request-mocking.md) |
|
||||
| Managing the CLI browser session | [session-management.md](session-management.md) |
|
||||
@@ -0,0 +1,139 @@
|
||||
# Tracing
|
||||
|
||||
Capture detailed execution traces for debugging and analysis. Traces include DOM snapshots, screenshots, network activity, and console logs.
|
||||
|
||||
## Basic Usage
|
||||
|
||||
```bash
|
||||
# Start trace recording
|
||||
playwright-cli tracing-start
|
||||
|
||||
# Perform actions
|
||||
playwright-cli open https://example.com
|
||||
playwright-cli click e1
|
||||
playwright-cli fill e2 "test"
|
||||
|
||||
# Stop trace recording
|
||||
playwright-cli tracing-stop
|
||||
```
|
||||
|
||||
## Trace Output Files
|
||||
|
||||
When you start tracing, Playwright creates a `traces/` directory with several files:
|
||||
|
||||
### `trace-{timestamp}.trace`
|
||||
|
||||
**Action log** - The main trace file containing:
|
||||
- Every action performed (clicks, fills, navigations)
|
||||
- DOM snapshots before and after each action
|
||||
- Screenshots at each step
|
||||
- Timing information
|
||||
- Console messages
|
||||
- Source locations
|
||||
|
||||
### `trace-{timestamp}.network`
|
||||
|
||||
**Network log** - Complete network activity:
|
||||
- All HTTP requests and responses
|
||||
- Request headers and bodies
|
||||
- Response headers and bodies
|
||||
- Timing (DNS, connect, TLS, TTFB, download)
|
||||
- Resource sizes
|
||||
- Failed requests and errors
|
||||
|
||||
### `resources/`
|
||||
|
||||
**Resources directory** - Cached resources:
|
||||
- Images, fonts, stylesheets, scripts
|
||||
- Response bodies for replay
|
||||
- Assets needed to reconstruct page state
|
||||
|
||||
## What Traces Capture
|
||||
|
||||
| Category | Details |
|
||||
|----------|---------|
|
||||
| **Actions** | Clicks, fills, hovers, keyboard input, navigations |
|
||||
| **DOM** | Full DOM snapshot before/after each action |
|
||||
| **Screenshots** | Visual state at each step |
|
||||
| **Network** | All requests, responses, headers, bodies, timing |
|
||||
| **Console** | All console.log, warn, error messages |
|
||||
| **Timing** | Precise timing for each operation |
|
||||
|
||||
## Use Cases
|
||||
|
||||
### Debugging Failed Actions
|
||||
|
||||
```bash
|
||||
playwright-cli tracing-start
|
||||
playwright-cli open https://app.example.com
|
||||
|
||||
# This click fails - why?
|
||||
playwright-cli click e5
|
||||
|
||||
playwright-cli tracing-stop
|
||||
# Open trace to see DOM state when click was attempted
|
||||
```
|
||||
|
||||
### Analyzing Performance
|
||||
|
||||
```bash
|
||||
playwright-cli tracing-start
|
||||
playwright-cli open https://slow-site.com
|
||||
playwright-cli tracing-stop
|
||||
|
||||
# View network waterfall to identify slow resources
|
||||
```
|
||||
|
||||
### Capturing Evidence
|
||||
|
||||
```bash
|
||||
# Record a complete user flow for documentation
|
||||
playwright-cli tracing-start
|
||||
|
||||
playwright-cli open https://app.example.com/checkout
|
||||
playwright-cli fill e1 "4111111111111111"
|
||||
playwright-cli fill e2 "12/25"
|
||||
playwright-cli fill e3 "123"
|
||||
playwright-cli click e4
|
||||
|
||||
playwright-cli tracing-stop
|
||||
# Trace shows exact sequence of events
|
||||
```
|
||||
|
||||
## Trace vs Video vs Screenshot
|
||||
|
||||
| Feature | Trace | Video | Screenshot |
|
||||
|---------|-------|-------|------------|
|
||||
| **Format** | .trace file | .webm video | .png/.jpeg image |
|
||||
| **DOM inspection** | Yes | No | No |
|
||||
| **Network details** | Yes | No | No |
|
||||
| **Step-by-step replay** | Yes | Continuous | Single frame |
|
||||
| **File size** | Medium | Large | Small |
|
||||
| **Best for** | Debugging | Demos | Quick capture |
|
||||
|
||||
## Best Practices
|
||||
|
||||
### 1. Start Tracing Before the Problem
|
||||
|
||||
```bash
|
||||
# Trace the entire flow, not just the failing step
|
||||
playwright-cli tracing-start
|
||||
playwright-cli open https://example.com
|
||||
# ... all steps leading to the issue ...
|
||||
playwright-cli tracing-stop
|
||||
```
|
||||
|
||||
### 2. Clean Up Old Traces
|
||||
|
||||
Traces can consume significant disk space:
|
||||
|
||||
```bash
|
||||
# Remove traces older than 7 days
|
||||
find .playwright-cli/traces -mtime +7 -delete
|
||||
```
|
||||
|
||||
## Limitations
|
||||
|
||||
- Traces add overhead to automation
|
||||
- Large traces can consume significant disk space
|
||||
- Some dynamic content may not replay perfectly
|
||||
@@ -0,0 +1,143 @@
|
||||
# Video Recording
|
||||
|
||||
Capture browser automation sessions as video for debugging, documentation, or verification. Produces WebM (VP8/VP9 codec).
|
||||
|
||||
## Basic Recording
|
||||
|
||||
```bash
|
||||
# Open browser first
|
||||
playwright-cli open
|
||||
|
||||
# Start recording
|
||||
playwright-cli video-start demo.webm
|
||||
|
||||
# Add a chapter marker for section transitions
|
||||
playwright-cli video-chapter "Getting Started" --description="Opening the homepage" --duration=2000
|
||||
|
||||
# Navigate and perform actions
|
||||
playwright-cli goto https://example.com
|
||||
playwright-cli snapshot
|
||||
playwright-cli click e1
|
||||
|
||||
# Add another chapter
|
||||
playwright-cli video-chapter "Filling Form" --description="Entering test data" --duration=2000
|
||||
playwright-cli fill e2 "test input"
|
||||
|
||||
# Stop and save
|
||||
playwright-cli video-stop
|
||||
```
|
||||
|
||||
## Best Practices
|
||||
|
||||
### 1. Use Descriptive Filenames
|
||||
|
||||
```bash
|
||||
# Include context in filename
|
||||
playwright-cli video-start recordings/login-flow-2024-01-15.webm
|
||||
playwright-cli video-start recordings/checkout-test-run-42.webm
|
||||
```
|
||||
|
||||
### 2. Record entire hero scripts.
|
||||
|
||||
When recording a video for the user or as a proof of work, it is best to create a code snippet and execute it with run-code.
|
||||
It allows inserting appropriate pauses between the actions and annotating the video. There are new Playwright APIs for that.
|
||||
|
||||
1) Perform scenario using CLI and take note of all locators and actions. You'll need those locators to request their bounding boxes for highlight.
|
||||
2) Create a file with the intended script for video (below). Use pressSequentially w/ delay for nice typing, make reasonable pauses.
|
||||
3) Use playwright-cli run-code --filename your-script.js
|
||||
|
||||
**Important**: Overlays are `pointer-events: none` — they do not interfere with page interactions. You can safely keep sticky overlays visible while clicking, filling, or performing any actions on the page.
|
||||
|
||||
```js
|
||||
async page => {
|
||||
await page.screencast.start({ path: 'video.webm', size: { width: 1280, height: 800 } });
|
||||
await page.goto('https://demo.playwright.dev/todomvc');
|
||||
|
||||
// Show a chapter card — blurs the page and shows a dialog.
|
||||
// Blocks until duration expires, then auto-removes.
|
||||
// Use this for simple use cases, but always feel free to hand-craft your own beautiful
|
||||
// overlay via await page.screencast.showOverlay().
|
||||
await page.screencast.showChapter('Adding Todo Items', {
|
||||
description: 'We will add several items to the todo list.',
|
||||
duration: 2000,
|
||||
});
|
||||
|
||||
// Perform action
|
||||
await page.getByRole('textbox', { name: 'What needs to be done?' }).pressSequentially('Walk the dog', { delay: 60 });
|
||||
await page.getByRole('textbox', { name: 'What needs to be done?' }).press('Enter');
|
||||
await page.waitForTimeout(1000);
|
||||
|
||||
// Show next chapter
|
||||
await page.screencast.showChapter('Verifying Results', {
|
||||
description: 'Checking the item appeared in the list.',
|
||||
duration: 2000,
|
||||
});
|
||||
|
||||
// Add a sticky annotation that stays while you perform actions.
|
||||
// Overlays are pointer-events: none, so they won't block clicks.
|
||||
const annotation = await page.screencast.showOverlay(`
|
||||
<div style="position: absolute; top: 8px; right: 8px;
|
||||
padding: 6px 12px; background: rgba(0,0,0,0.7);
|
||||
border-radius: 8px; font-size: 13px; color: white;">
|
||||
✓ Item added successfully
|
||||
</div>
|
||||
`);
|
||||
|
||||
// Perform more actions while the annotation is visible
|
||||
await page.getByRole('textbox', { name: 'What needs to be done?' }).pressSequentially('Buy groceries', { delay: 60 });
|
||||
await page.getByRole('textbox', { name: 'What needs to be done?' }).press('Enter');
|
||||
await page.waitForTimeout(1500);
|
||||
|
||||
// Remove the annotation when done
|
||||
await annotation.dispose();
|
||||
|
||||
// You can also highlight relevant locators and provide contextual annotations.
|
||||
const bounds = await page.getByText('Walk the dog').boundingBox();
|
||||
await page.screencast.showOverlay(`
|
||||
<div style="position: absolute;
|
||||
top: ${bounds.y}px;
|
||||
left: ${bounds.x}px;
|
||||
width: ${bounds.width}px;
|
||||
height: ${bounds.height}px;
|
||||
border: 1px solid red;">
|
||||
</div>
|
||||
<div style="position: absolute;
|
||||
top: ${bounds.y + bounds.height + 5}px;
|
||||
left: ${bounds.x + bounds.width / 2}px;
|
||||
transform: translateX(-50%);
|
||||
padding: 6px;
|
||||
background: #808080;
|
||||
border-radius: 10px;
|
||||
font-size: 14px;
|
||||
color: white;">Check it out, it is right above this text
|
||||
</div>
|
||||
`, { duration: 2000 });
|
||||
|
||||
await page.screencast.stop();
|
||||
}
|
||||
```
|
||||
|
||||
Embrace creativity, overlays are powerful.
|
||||
|
||||
### Overlay API Summary
|
||||
|
||||
| Method | Use Case |
|
||||
|--------|----------|
|
||||
| `page.screencast.showChapter(title, { description?, duration?, styleSheet? })` | Full-screen chapter card with blurred backdrop — ideal for section transitions |
|
||||
| `page.screencast.showOverlay(html, { duration? })` | Custom HTML overlay — use for callouts, labels, highlights |
|
||||
| `disposable.dispose()` | Remove a sticky overlay added without duration |
|
||||
| `page.screencast.hideOverlays()` / `page.screencast.showOverlays()` | Temporarily hide/show all overlays |
|
||||
|
||||
## Tracing vs Video
|
||||
|
||||
| Feature | Video | Tracing |
|
||||
|---------|-------|---------|
|
||||
| Output | WebM file | Trace file (viewable in Trace Viewer) |
|
||||
| Shows | Visual recording | DOM snapshots, network, console, actions |
|
||||
| Use case | Demos, documentation | Debugging, analysis |
|
||||
| Size | Larger | Smaller |
|
||||
|
||||
## Limitations
|
||||
|
||||
- Recording adds slight overhead to automation
|
||||
- Large recordings can consume significant disk space
|
||||
@@ -0,0 +1,366 @@
|
||||
---
|
||||
|
||||
## name: scopesentry-mcp
|
||||
description: 通过 ScopeSentry MCP 管理安全扫描平台(项目、任务、模板、资产、节点)。在用户提到 ScopeSentry、MCP、API Key、扫描任务、资产查询时使用。
|
||||
|
||||
# ScopeSentry MCP 使用指南
|
||||
|
||||
面向**已部署 ScopeSentry 实例**的用户。通过 Cursor(或其他 MCP 客户端)连接平台,无需本地源码。
|
||||
|
||||
## 1. 准备工作
|
||||
|
||||
### 1.1 确认服务可访问
|
||||
|
||||
- 默认 Web 界面:`http://<主机>`
|
||||
- MCP 端点:`http://<主机>/mcp`(若前面有反向代理或前端代理,以实际 `/mcp` 地址为准)
|
||||
|
||||
### 1.2 创建 API Key
|
||||
|
||||
1. 浏览器登录 ScopeSentry Web 界面
|
||||
2. 进入 **API Key** 管理页创建密钥(或通过管理员提供的接口创建)
|
||||
3. 保存返回的 `ssk_...` 字符串(**仅显示一次**)
|
||||
|
||||
### 1.3 配置 Cursor MCP
|
||||
|
||||
Cursor → Settings → MCP → 添加服务器:
|
||||
|
||||
```json
|
||||
{
|
||||
"mcpServers": {
|
||||
"scopesentry": {
|
||||
"url": "http://<你的主机>:8082/mcp",
|
||||
"headers": {
|
||||
"X-API-Key": "ssk_你的密钥"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
也可使用:`Authorization: Bearer ssk_你的密钥`
|
||||
|
||||
配置完成后重启 MCP 或重载 Cursor,确认工具列表中出现 `list_projects`、`list_assets` 等。
|
||||
|
||||
---
|
||||
|
||||
## 2. 工具一览
|
||||
|
||||
|
||||
| 工具 | 用途 |
|
||||
| ---------------------- | ----------------- |
|
||||
| `list_projects` | 按标签分组的项目树(含项目 ID) |
|
||||
| `list_projects_data` | 分页项目列表,可按名称搜索 |
|
||||
| `get_project` | 项目详情 |
|
||||
| `create_project` | 新建项目 |
|
||||
| `list_tasks` | 扫描任务列表 |
|
||||
| `get_task` | 任务详情 |
|
||||
| `list_scan_templates` | 扫描模板列表 |
|
||||
| `get_scan_template` | 模板详情 |
|
||||
| `list_plugin_modules` | 扫描流水线模块名 |
|
||||
| `list_plugins` | 可用插件(含 hash、默认参数) |
|
||||
| `create_scan_template` | 创建扫描模板 |
|
||||
| `create_scan_task` | 创建扫描任务 |
|
||||
| `list_assets` | 查询各类资产(分页列表) |
|
||||
| `count_assets` | 统计资产数量(`/api/assets/common/total`) |
|
||||
| `get_asset_detail` | 资产或漏洞详情 |
|
||||
| `add_asset_tag` | 为资产添加标签 |
|
||||
| `list_nodes` | 扫描节点列表 |
|
||||
|
||||
|
||||
各工具参数以 MCP 工具描述(schema)为准;`list_assets` / `count_assets` 的 search、filter 语法一致,查询资产前可先阅读 `list_assets` description。
|
||||
|
||||
需要知道「共多少条」时用 `count_assets`(对应 Web 分页总数接口),不必为了数总数反复翻页 `list_assets`。
|
||||
|
||||
---
|
||||
|
||||
## 3. 常用工作流
|
||||
|
||||
### 3.1 按项目查资产
|
||||
|
||||
当用户或上下文**已有项目条件**时,优先带上 `filter.project` 缩小范围,避免跨项目数据过多导致响应变慢。若无明确项目,可不强制加项目筛选。
|
||||
|
||||
1. `list_projects` 或 `list_projects_data` 获取目标项目的 **ObjectID**(`id` / `children[].value`)
|
||||
2. `list_assets` 传入 `filter.project`(**必须是 ID,不能写项目中文名**)
|
||||
|
||||
```json
|
||||
{
|
||||
"asset_type": "asset",
|
||||
"pageIndex": 1,
|
||||
"pageSize": 20,
|
||||
"search": "domain=^example.com",
|
||||
"filter": {
|
||||
"project": ["<项目ObjectID>"]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 3.2 创建扫描任务
|
||||
|
||||
1. `list_nodes` 获取在线节点名称
|
||||
2. `list_scan_templates` 或 `create_scan_template` 获取模板 **ObjectID**
|
||||
3. `create_scan_task`:`name`、`node` 必填,`template` 填模板 ID(不能填模板名)
|
||||
|
||||
**目标来源 `targetSource`(与 Web 端一致):**
|
||||
|
||||
| targetSource | 说明 | 必填参数 |
|
||||
| --- | --- | --- |
|
||||
| `general` | 直接输入目标 | `target` |
|
||||
| `project` | 从项目读取目标 | `project`(项目 ObjectID 数组) |
|
||||
| `asset` | 从 Web 资产库搜索 | `search`;可选 `project`、`filter`、`targetNumber` |
|
||||
| `RootDomain` | 从根域名库搜索 | `search`;可选 `project`、`filter`、`targetNumber` |
|
||||
| `subdomain` | 从子域名库搜索 | `search`;可选 `project`、`filter`、`targetNumber` |
|
||||
| `UrlScan` | 从 URL 扫描结果搜索 | `search`;可选 `project`、`filter`、`targetNumber` |
|
||||
| `*Source`(如 `subdomainSource`) | 从资产页「选中/搜索」创建 | `targetTp=search` 时用 `search`;`targetTp=select` 时用 `targetIds` |
|
||||
|
||||
**示例 — 直接扫根域名:**
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "example-子域名收集",
|
||||
"node": ["node-1"],
|
||||
"template": "<模板ObjectID>",
|
||||
"targetSource": "general",
|
||||
"target": "example.com\nfoo.com",
|
||||
"project": ["<项目ObjectID>"]
|
||||
}
|
||||
```
|
||||
|
||||
**示例 — 从子域名库续扫(按上一任务名筛选):**
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "example-端口与漏洞",
|
||||
"node": ["node-1"],
|
||||
"template": "<后续模块模板ObjectID>",
|
||||
"targetSource": "subdomain",
|
||||
"search": "task==\"example-子域名收集\"",
|
||||
"project": ["<项目ObjectID>"]
|
||||
}
|
||||
```
|
||||
|
||||
### 3.3 根域名完整信息收集(推荐两阶段)
|
||||
|
||||
当输入为**根域名**且要进行**完整信息收集**时,建议分两次扫描,不要一次跑全流水线。
|
||||
|
||||
**原因:** 分布式任务以**单个目标**为单位分发。根域名作为目标时,某节点分到该根域名后,在该节点上扫出的子域名也会继续在该节点执行后续模块,容易造成负载不均、速度慢、易出错。
|
||||
|
||||
**最佳实践:**
|
||||
|
||||
1. **第一阶段 — 仅子域名收集**
|
||||
- `targetSource`: `general`
|
||||
- `target`: 所有根域名(多行)
|
||||
- 模板:仅启用 `SubdomainScan`、`SubdomainSecurity`(子域名扫描 + 子域名接管)
|
||||
- 用 `get_task` 等待任务完成
|
||||
|
||||
2. **第二阶段 — 后续模块**
|
||||
- `targetSource`: `subdomain`
|
||||
- `search`: `task=="<第一阶段任务名称>"`(精确匹配任务名)
|
||||
- 可选 `project` 缩小范围
|
||||
- 模板:端口扫描、资产测绘、漏洞扫描等(可不含 SubdomainScan)
|
||||
- 子域名作为独立目标分发到各节点,并行效率更高
|
||||
|
||||
也可在 Web 界面「子域名」资产页按任务名筛选后,使用「从子域名创建任务」,效果相同。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
A[根域名列表] --> B[阶段1: general + SubdomainScan]
|
||||
B --> C[子域名入库]
|
||||
C --> D[阶段2: subdomain + task==阶段1任务名]
|
||||
D --> E[端口/资产/漏洞等模块]
|
||||
```
|
||||
|
||||
### 3.4 创建扫描模板
|
||||
|
||||
1. `list_plugin_modules` → 模块名列表
|
||||
2. `list_plugins`(可按 `module` 过滤)→ 各插件 `hash` 与默认 `parameter`
|
||||
3. `create_scan_template`:用 `modules` 指定「模块 → 插件 hash 数组」
|
||||
|
||||
---
|
||||
|
||||
## 4. 资产查询(`list_assets` / `count_assets`)
|
||||
|
||||
`count_assets` 与 `list_assets` 使用相同的 `asset_type`、`search`、`filter`,返回 `{ "total": N }`,对应 Web 端 `/api/assets/common/total`。
|
||||
|
||||
```json
|
||||
{
|
||||
"asset_type": "subdomain",
|
||||
"search": "task==\"某任务名\"",
|
||||
"filter": {"project": ["<项目ObjectID>"]}
|
||||
}
|
||||
```
|
||||
|
||||
**性能建议(`list_assets` / `count_assets` 通用):** 有项目条件时优先用 `filter.project` 缩小范围;`search` 中对已建索引字段尽量用 `==` 全等或 `^` 前缀匹配(见 [4.3](#43-search-搜索表达式)),避免大面积 `=` 模糊查询拖慢响应。无项目上下文时不强制加项目筛选。
|
||||
|
||||
支持 `filter.project` 的类型见 [4.4](#44-filter-精确过滤) 表格。
|
||||
|
||||
### 4.1 资产类型 `asset_type`
|
||||
|
||||
`asset`、`RootDomain`、`subdomain`、`app`、`mp`、`UrlScan`、`SensitiveResult`、`DirScanResult`、`crawler`、`vulnerability`、`PageMonitoring`、`IPAsset`、`SubdomainTakerResult`
|
||||
|
||||
别名示例:`web`→asset、`vuln`→vulnerability、`ip`→IPAsset、`url`→UrlScan
|
||||
|
||||
### 4.2 参数说明
|
||||
|
||||
|
||||
| 参数 | 说明 |
|
||||
| ------------------------ | --------------------------------------- |
|
||||
| `pageIndex` / `pageSize` | 分页,默认 1 / 20 |
|
||||
| `search` | 搜索表达式(见下节) |
|
||||
| `filter` | 精确过滤 JSON(见下节) |
|
||||
| `sort` | 仅 UrlScan、DirScanResult 支持按 `length` 排序 |
|
||||
| `sid` | 仅 SensitiveResult:敏感规则名称 |
|
||||
|
||||
|
||||
`search` 与 `filter` **可同时使用**。
|
||||
|
||||
### 4.3 search 搜索表达式
|
||||
|
||||
自定义 DSL(**不是 SQL**):
|
||||
|
||||
|
||||
| 运算符 | 含义 | 索引 | 示例 |
|
||||
| ---- | ---- | ---- | --------------------------- |
|
||||
| `=` | 模糊匹配(regex) | 不走索引 | `domain=example` |
|
||||
| `==` | 精确匹配(全等) | **走索引** | `port==443` |
|
||||
| `!=` | 排除 | — | `port!="80"` |
|
||||
| `&&` | 与 | — | `domain==example.com && port==443` |
|
||||
| `||` | 或 | — | `title=admin || body=login` |
|
||||
|
||||
|
||||
**索引与运算符:** `domain`、`ip`、`port`、`title` 等字段已建索引,但仅 **`==` 全等** 或 **值以 `^` 开头的前缀匹配**(如 `domain=^example.com`)能走索引;**`=` 会转为 regex 模糊匹配,无法使用索引**,数据量大时易变慢。
|
||||
|
||||
**所有类型通用 search 字段:** `tag`、`task`(任务名称)、`rootDomain`
|
||||
|
||||
**project 不能写在 search 里**(无效或与 `&&` 组合时报错)。筛项目请用 `filter.project`。
|
||||
|
||||
**各类型常用 search 字段:**
|
||||
|
||||
|
||||
| asset_type | 字段 |
|
||||
| -------------------- | ----------------------------------------------------------------------------------- |
|
||||
| asset | domain, ip, port, service, app, title, statuscode, icon, banner, type, body, header |
|
||||
| RootDomain | domain, icp, company |
|
||||
| subdomain | domain, ip, type, value |
|
||||
| app | name, icp, company, category, description, url, apk |
|
||||
| mp | name, icp, company, category, description, url |
|
||||
| UrlScan | url, input, source, resultId, type |
|
||||
| SensitiveResult | url, sname, body, info, md5 |
|
||||
| DirScanResult | url, statuscode, redirect, length |
|
||||
| vulnerability | url, vulname, matched, request, response, level |
|
||||
| crawler | url, method, body, resultId |
|
||||
| PageMonitoring | url, hash, diff, response |
|
||||
| IPAsset | ip, domain, port, service, webServer, app |
|
||||
| SubdomainTakerResult | domain, value, type, response |
|
||||
|
||||
|
||||
**search 示例:**
|
||||
|
||||
- `domain==www.example.com && port==443`(全等,走索引)
|
||||
- `domain=^example.com`(前缀匹配,走索引)
|
||||
- `ip==192.168.1.1`
|
||||
- `task=="某任务名"`
|
||||
- `level==high`(vulnerability)
|
||||
- `statuscode==200`(DirScanResult)
|
||||
|
||||
需模糊包含时再用 `=`,如 `title=admin`(不走索引,宜配合项目等条件缩小范围)。
|
||||
|
||||
### 4.4 filter 精确过滤
|
||||
|
||||
JSON 对象:同 key 多个值为 **OR**,不同 key 为 **AND**。
|
||||
|
||||
**有项目条件时优先用 `project`:** 若用户或上下文已明确项目,且 asset_type 支持 `project`,应带上以缩小范围;无项目信息时不强制。
|
||||
|
||||
|
||||
| filter key | 含义 | 取值说明 |
|
||||
| ------------ | -------- | -------------------------------------------------------- |
|
||||
| `project` | 所属项目 | **ObjectID**,用 `list_projects` / `list_projects_data` 获取 |
|
||||
| `task` | 来源任务 | **任务名称**,用 `list_tasks` 的 `name` |
|
||||
| `port` | 端口 | 如 `"443"` |
|
||||
| `service` | 服务/协议 | 如 `"https"` |
|
||||
| `app` | 应用指纹 | 如 `"Nginx"` |
|
||||
| `icon` | 图标 hash | |
|
||||
| `statuscode` | HTTP 状态码 | 主要用于 asset |
|
||||
| `status` | 状态 | UrlScan/DirScan HTTP 码;漏洞/敏感信息处理状态 |
|
||||
| `level` | 漏洞等级 | critical / high / medium / low / info |
|
||||
| `type` | 类型 | 如子域名记录类型 A、CNAME |
|
||||
| `color` | 敏感规则颜色 | SensitiveResult |
|
||||
| `sname` | 敏感规则名 | SensitiveResult |
|
||||
| `tags` | 标签 | |
|
||||
|
||||
|
||||
**各类型可用 filter key:**
|
||||
|
||||
|
||||
| asset_type | filter key |
|
||||
| ------------------------------------- | --------------------------------------------------------------- |
|
||||
| asset | project, port, service, app, icon, statuscode, type, task, tags |
|
||||
| RootDomain | project, tags |
|
||||
| subdomain | project, type, task, tags |
|
||||
| app / mp | project, tags |
|
||||
| UrlScan | status, tags |
|
||||
| DirScanResult | status, tags |
|
||||
| SensitiveResult | status, color, sname, tags |
|
||||
| crawler | project, task, tags |
|
||||
| vulnerability | project, level, status, task, tags |
|
||||
| PageMonitoring / SubdomainTakerResult | tags |
|
||||
| IPAsset | project, port, service, app |
|
||||
|
||||
|
||||
**filter 示例:**
|
||||
|
||||
```json
|
||||
{"project": ["<项目ObjectID>"], "port": ["443"]}
|
||||
```
|
||||
|
||||
**组合查询示例:**
|
||||
|
||||
```json
|
||||
{
|
||||
"asset_type": "asset",
|
||||
"search": "domain=^baidu && port==443",
|
||||
"filter": {"project": ["<项目ObjectID>"]},
|
||||
"pageIndex": 1,
|
||||
"pageSize": 10
|
||||
}
|
||||
```
|
||||
|
||||
**注意:**
|
||||
|
||||
- 有项目条件时优先带 `filter.project`(支持时);无项目上下文可不强制
|
||||
- `filter.project` 勿填项目显示名称
|
||||
- 已知值用 `==`,前缀用 `^`;避免对大表滥用 `=` 模糊匹配
|
||||
- UrlScan 的 HTTP 状态用 `filter.status`;DirScanResult 可在 search 中用 `statuscode==200`
|
||||
- SensitiveResult 按规则名:`search` 用 `sname=规则名`,或 `filter.sname`
|
||||
|
||||
### 4.5 排序 sort
|
||||
|
||||
仅 **UrlScan**、**DirScanResult** 支持:
|
||||
|
||||
```json
|
||||
{"length": "ascending"}
|
||||
```
|
||||
|
||||
其他类型忽略 `sort`,按时间默认排序。
|
||||
|
||||
---
|
||||
|
||||
## 5. 扫描模板模块名
|
||||
|
||||
`TargetHandler`、`SubdomainScan`、`SubdomainSecurity`、`PortScanPreparation`、`PortScan`、`PortFingerprint`、`AssetMapping`、`AssetHandle`、`URLScan`、`WebCrawler`、`URLSecurity`、`DirScan`、`VulnerabilityScan`、`PassiveScan`
|
||||
|
||||
---
|
||||
|
||||
## 6. 故障排查
|
||||
|
||||
|
||||
| 现象 | 处理 |
|
||||
| --------- | -------------------------------------------------- |
|
||||
| MCP 无工具 | 检查 URL、API Key、ScopeSentry 是否运行 |
|
||||
| 401 / 403 | 重新创建或更换 API Key |
|
||||
| 资产查不到 | 确认 `filter.project` 为 ObjectID;勿在 search 写 project |
|
||||
| 模板/任务创建失败 | `template` 必须是模板 ObjectID;`node` 填在线节点名 |
|
||||
| 查询很慢/卡住 | 有项目时加 `filter.project`;search 对已索引字段改用 `==` 或 `^` 前缀,少用 `=`;缩小 `pageSize` |
|
||||
|
||||
|
||||
---
|
||||
|
||||
Reference in New Issue
Block a user