First Commit
ci / go (push) Waiting to run
ci / go-db (agent) (push) Waiting to run
ci / go-db (config) (push) Waiting to run
ci / go-db (db) (push) Waiting to run
ci / go-db (evidence) (push) Waiting to run
ci / go-db (llmrec) (push) Waiting to run
ci / go-db (server) (push) Waiting to run
detections / detections (push) Waiting to run
web / web (push) Waiting to run
docs / links (push) Canceled after 0s

This commit is contained in:
dela
2026-10-09 08:38:16 +08:00
commit 0335d572de
756 changed files with 201663 additions and 0 deletions
+455
View File
@@ -0,0 +1,455 @@
# ARTEX 탐지 규칙 테스트
한국어 · [English](README.md)
[`../`](../) 아래의 탐지 규칙이 실제로 발화하는지, 그리고 그에 못지않게 중요한, 양성(benign)
트래픽에는 침묵하는지를 재현 가능하게 증명하는 회귀 테스트입니다. 돌려 볼 수 없는 탐지 규칙은
주장에 지나지 않습니다. 이 테스트들은 규칙 파일과 방어 가이드에 적힌 주장을 검토자가 소스에서
다시 돌려 볼 수 있는 것으로 바꿉니다.
이진 패킷 캡처는 저장소에 넣지 않습니다. 캡처는 **매 실행마다 결정론적으로 생성**했다가 끝난 뒤
지우므로, 테스트는 불투명한 고정 파일이 아니라 읽을 수 있는 소스로 배포되며 저장소를 불리지
않습니다.
## 모든 스위트를 한 번에 실행: [`run-all.sh`](run-all.sh)
[`run-all.sh`](run-all.sh) 는 아래 여덟 스위트를 CI 와 같은 순서로 한 명령에 전부 돌리므로, 여덟 개
`run.sh` 스크립트를 손으로 하나씩 호출하지 않아도 됩니다. 앞 스위트가 실패해도 각 스위트는 끝까지
돌고, 스크립트는 마지막에 스위트마다 PASS/FAIL 한 줄 요약을 출력하며, 하나라도 실패하면 0 이 아닌
코드로 종료합니다.
스위트를 돌리기 전에 하네스 자기 점검([`check-harness-sync.sh`](check-harness-sync.sh))을 먼저 실행합니다.
이 점검은 위의 스위트 목록, [CI](../../.github/workflows/detections.yml) 의 스위트별 스텝, 디스크의 스위트
디렉터리 이 셋이 서로 다른 스위트나 다른 순서를 가리키면 실행을 실패로 끝냅니다. 이것은 여덟 스위트가
스스로 보지 못하는 유일한 공백입니다. 세 곳 중 한 곳에만 배선된 스위트(예: `run-all.sh` 항목 없이 CI
스텝만 추가하거나, 어느 쪽에도 넣지 않은 디렉터리)는 스위트별 테스트를 모두 통과하면서도, 로컬에서
초록이던 `run-all.sh` 가 더는 초록 CI 를 뜻하지 않게 만듭니다. 이 점검은 아홉째 스위트가 아니라 게이트라서
아래 요약에는 나타나지 않으므로, 탐지 스위트는 여덟 그대로입니다.
```sh
detections/tests/run-all.sh
```
예상 출력(축약):
```
===== detection suites summary =====
PASS sigma
PASS sigma_match
PASS sigma_lint
PASS sigma_backends
PASS suricata
PASS attack
PASS indicators
PASS misp
RESULT: PASS
```
실패가 하나라도 있으면 0 이 아닌 코드로 종료하므로 pre-commit 훅에 그대로 넣을 수 있습니다. 바로 쓸
수 있는 예시가 저장소 최상위 [`.pre-commit-config.yaml`](../../.pre-commit-config.yaml) 에 있습니다.
`pip install pre-commit && pre-commit install` 로 설치하면, 탐지 규칙이나 그 규칙이 고정한 상류 소스
파일을 건드리는 커밋에서 러너가 발화합니다. CI 와 같은 범위입니다. 개별 스위트가 인식하는 이미지·버전
재정의(`PYTHON_IMAGE`, `SIGMA_CLI_VERSION`, `SIGMAHQ_VALIDATORS_VERSION`, `SURICATA_IMAGE`)는 러너가
그대로 물려받으므로, 그중 어느 것을 export 해도 모든 스위트에 한꺼번에 적용됩니다.
## Suricata: [`suricata/`](suricata/)
[`suricata/run.sh`](suricata/run.sh) 는 [`../suricata/artex.rules`](../suricata/artex.rules) 의
네트워크 규칙을 종단으로 돌려 다섯 가지 속성을 단언합니다:
- **유효성**: 규칙 파일 전체가 `suricata -T --init-errors-fatal` 로 적재되므로, 아래 어떤 캡처도
건드리지 않는 규칙이라도 파싱·초기화에 실패하면 잡아냅니다. 그냥 `suricata -r` 는 그런 규칙을
건너뛰고도 0 으로 종료하므로, 이 적재 검사는 Sigma 스위트의 `sigma check` 유효성 단언에 해당하는
Suricata 쪽 장치입니다.
- **존재성(보강 프로버)**: sid `1000001` 이 보강 프로브마다 정확히 한 번 발화합니다.
- **속도**: 소스당 300 초에 30 요청이라는 `detection_filter` 임계를 넘으면 sid `1000002` 가
발화합니다.
- **존재성(WebFetch)**: sid `1000003` 이 norma WebFetch 요청마다 정확히 한 번 발화하고, 같은 캡처에서
보강 프로버 sid 는 침묵합니다. 두 네트워크 시그니처가 각자 발화할 뿐 아니라 서로 특이적임을
확인합니다.
- **특이성**: 다른 것은 같고 User-Agent 만 양성(benign) 브라우저로 바꾼 캡처는 ARTEX 경보를
**하나도** 내지 않습니다.
[`suricata/gen_pcap.py`](suricata/gen_pcap.py) 는 [scapy](https://scapy.net) 로 캡처를 만듭니다.
고정된 한 소스에서 나오는 N 개의 독립적인 평문 HTTP 요청/응답 흐름을, 각각 지정한 User-Agent 를
실어, 고정된 기준 타임스탬프에서 1 초 간격으로 배치합니다. 파일을 쓰기만 할 뿐, 패킷을 보내거나
네트워크를 건드리지 않습니다.
### 실행
Docker 만 있으면 됩니다. scapy 와 Suricata 모두 컨테이너에서 돕니다.
```sh
detections/tests/suricata/run.sh
```
예상 출력(축약):
```
PASS ruleset loads with zero parse/init errors (suricata -T)
PASS sid 1000001 presence: one alert per probe (got 35, want eq 35)
PASS sid 1000002 velocity: fires past 30-in-300s (got 5, want ge 1)
PASS sid 1000003 presence: one alert per WebFetch request (got 8, want eq 8)
PASS enrich sids stay silent on norma traffic (specificity) (got 0, want eq 0)
PASS benign browser UA produces no ARTEX alerts (got 0, want eq 0)
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료하므로 CI 나 pre-commit 훅에 그대로 넣을 수
있습니다. 내부에 미러를 두었다면 `SURICATA_IMAGE` / `PYTHON_IMAGE` 로 이미지를 재정의하십시오.
### 속도 경보 수를 정확한 값이 아니라 하한으로 단언하는 이유
`run.sh` 는 존재성 경보 수(`1000001 == 35`·`1000003 == 8`)와 양성 경보 수(`== 0`)를 정확히 단언합니다.
이들은 엔진 버전과 무관하기 때문입니다. 일치하는 요청마다 경보 하나, 다른 User-Agent 에는 불일치입니다. 속도
규칙의 경보 수는 특정 Suricata 릴리스가 경계에서 `detection_filter` 임계를 어떻게 처리하느냐에 달려
있으므로, 테스트는 `>= 1` 로 단언하고 기준값은 따로 기록합니다. **Suricata 8.0.7** 에서는 기준
실행이 sid `1000002` 에 경보 **5** 개를 냅니다(300 초에 30 임계를 넘긴 뒤의 31–35 번째 흐름).
## Sigma: [`sigma/`](sigma/)
[`sigma/run.sh`](sigma/run.sh) 는 [`../sigma/`](../sigma/) 아래의 Sigma 규칙을 구조적으로, 그리고
[sigma-cli](https://github.com/SigmaHQ/sigma-cli)(pySigma) 로 컴파일해 검증하며, 다섯 가지 속성을
단언합니다:
- **유효성**: `sigma check` 가 트리 전체에서 오류 0, 조건 오류 0, 이슈 0 을 보고합니다.
- **컴파일**: `sigma convert -t splunk` 가 트리 전체를 오류 없이 백엔드 질의 언어로 변환합니다.
- **지표 보존**: 각 원자 지표 문자열(`artex-enrich/1.0`, `artex-selfupdate`, 가드 마커, 그리고 기록용
프록시 CA 파일명 `mitmproxy-ca-cert.pem`)이 컴파일된 질의에 그대로 남아 있으므로, 규칙이 자신이
기반한 문자열을 조용히 잃을 수 없습니다.
- **상관 규칙 컴파일**: [`../sigma/correlation/`](../sigma/correlation/) 의 행동 규칙이 버려지지
않고 `event_count` / `value_count` 집계를 내보냅니다.
- **상관 규칙이 실제로 작동함**: 상관 규칙 하나만 *단독으로* 변환하면 실패합니다. 그 규칙이 원자
기반 규칙을 `id` 로 참조하기 때문이며, 이 참조는 장식이 아니라 강제됩니다. 이는 위 Suricata
특이성 단언에 해당하는 Sigma 쪽 장치입니다.
이는 [`../README.ko.md`](../README.ko.md) 에 설명한 구조 + 컴파일 검증을, 실행 가능하고 단언하는 형태로 만든
것입니다. 아래의 짝 스위트 [`sigma_match/`](sigma_match/) 가 원자 규칙과 상관 규칙 양쪽에 *매칭* 절반을
더합니다. 대표적인 악성 이벤트(또는 타임라인)가 각 규칙을 발화시키고 정상 이벤트는 발화시키지 않음을
확인하므로, 이제 Sigma 규칙도 Suricata 규칙처럼 재현 가능한 검증 테스트와 재현 가능한 매칭 테스트를 함께
갖습니다. (엉성하게 손으로 짠 매처가 규칙을
깎아내릴 수 있다는 기존 우려는, 파싱을 전부 pySigma 에 위임해 해소했습니다. 신뢰 모델은 다음 절에서 설명합니다.)
### 실행
Docker 만 있으면 됩니다. sigma-cli 와 splunk 백엔드가 컨테이너에서 돌고 저장소에는 아무것도 쓰지
않습니다.
```sh
detections/tests/sigma/run.sh
```
예상 출력(축약):
```
PASS sigma check: 0 errors, 0 condition errors, 0 issues
PASS whole tree converts to splunk (exit 0)
PASS indicator present: artex-enrich/1.0
PASS correlation rule fails to convert alone — it requires its atomic base rule
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료하므로 CI 나 pre-commit 훅에 그대로 넣을 수
있습니다. sigma-cli 는 기준 버전(`3.1.0`)으로 고정돼 있습니다. 내부에 미러를 두었다면
`SIGMA_CLI_VERSION` 으로 버전을, `PYTHON_IMAGE` 로 이미지를 재정의하십시오.
## Sigma 실시간 이벤트 매칭: [`sigma_match/`](sigma_match/)
[`sigma_match/run.sh`](sigma_match/run.sh) 는 [`../sigma/`](../sigma/) 아래의 Sigma 규칙이, 원자 규칙과
[`../sigma/correlation/`](../sigma/correlation/) 의 상관 규칙을 모두 포함해, 매칭되는 이벤트에 실제로 *발화*하고
정상 이벤트에는 침묵함을 증명합니다. Suricata 스위트가 네트워크 규칙에 주는 "돌려 볼 수 없는 탐지 규칙은
주장일 뿐"이라는 보증을, 호스트·로그 계층 규칙으로 확장한 것입니다. 원자 규칙 셋과 상관 규칙 셋, 모두 여섯
속성을 단언합니다:
- **규칙·샘플 짝짓기**: 모든 원자 규칙에는 [`events/<이름>.json`](sigma_match/events/) 샘플 파일이 있고,
모든 샘플 파일은 규칙으로 되짚어집니다. 샘플 없이 추가한 규칙은 검증을 못 받고 넘어가는 대신 여기서
실패합니다.
- **참 양성(true positive)**: 각 규칙이 자신의 악성 샘플 이벤트를 전부 매칭합니다.
- **참 음성(true negative)**: 각 규칙이 자신의 정상 샘플 이벤트를 하나도 매칭하지 않습니다. 예를 들어
`.mitmproxy/` 아래의 단독 `mitmproxy-ca-cert.pem` 은 기록용 프록시 규칙을 발화시키지 **않습니다**. 그
규칙의 `|all` 수식자가 ARTEX 가 쓰는 `_ca/` 디렉터리까지 함께 요구하기 때문이며, 이 판별을 증명하는 것이
바로 매칭 테스트입니다.
- **상관 규칙·타임라인 짝짓기**: 모든 상관 규칙에는 [`events/correlation/<이름>.json`](sigma_match/events/correlation/)
타임라인 파일이 있고, 모든 타임라인은 규칙으로 되짚어집니다. 타임라인의 각 이벤트는 상대 초를 담은 `ts`
필드를 지닙니다.
- **상관 규칙의 참 양성**: 임계를 시간 창 안에서 한 그룹이 채우는 양성 타임라인에 각 규칙이 발화합니다.
예를 들어 한 출처(`c-ip`)에서 10분 안에 서로 다른 20개 호스트로 퍼지는 요청이 수집 팬아웃 규칙을
발화시킵니다.
- **상관 규칙의 참 음성**: 임계 미달, 임계는 채웠지만 시간 창을 벗어난 경우, 그룹이 갈린 경우,
시간 상관에서 한쪽 레그가 빠진 경우에는 침묵합니다. 특히 요청량은 많아도 폭(서로 다른 호스트 수)이 작은
버스트는 팬아웃 규칙을 발화시키지 **않습니다**. 폭이 신호이지 양이 신호가 아니며, 이 판별을 증명하는 것이
바로 매칭 테스트입니다.
신뢰 모델은 이렇습니다. 손으로 짠 코드가 아니라 pySigma 가 각 규칙을 파싱합니다. 원자 규칙은 수식자와
조건을 트리로 컴파일하고(`|contains` → 와일드카드 값, `|all` → AND, `1 of selection_*` → OR), 상관 규칙은
집계 명세(유형·group-by·시간 창·임계 조건·참조하는 원자 규칙)로 컴파일합니다. [`check.py`](sigma_match/check.py)
는 그 트리와 명세를 따라 걷을 뿐이고, 상관 규칙이 어느 이벤트를 먹는지는 원자 규칙과 똑같은 매처로
판정하므로 권위 있는 Sigma 로직은 pySigma 안에 남습니다. 명시적으로 지원하지 않는 구문을 만나면 조용히
통과시키지 않고 예외를 던집니다(fail-closed). 범위와 한계는 스크립트 머리말에 밝혀 둡니다. 상관 규칙의
시간 창은 표준 슬라이딩 윈도(매칭 이벤트마다 `timespan` 길이의 창을 잡는) 해석이며, 실제 SIEM 의 윈도
방식은 다를 수 있습니다. 매칭은 **대소문자를 무시**하고(`sigma/` 스위트가 겨냥하는 splunk 백엔드의
기본값이며, 파괴 명령 규칙의 오탐 주석 자체가 이를 전제합니다), 키워드 매칭은 전문 부분 문자열 검색입니다.
이것은 규칙의 필드·값·조건·집계 로직에 대한 회귀 테스트이지, 필드 정규화가 다를 수 있는 각자의 SIEM 에서
검증하는 일을 대신하지는 않습니다.
### 실행
Docker 만 있으면 됩니다. pySigma 가 컨테이너에서 돌고 저장소에는 아무것도 쓰지 않습니다.
```sh
detections/tests/sigma_match/run.sh
```
예상 출력(축약):
```
PASS rule/sample pairing: 5 atomic rules, 5 event files, no orphans
PASS artex_enrich_user_agent: 1/1 positive events matched
PASS artex_recording_proxy_ca: 2/2 benign events correctly not matched
PASS rule/timeline pairing: 4 correlation rules, 4 timeline files, no orphans
PASS artex_enrich_fanout: fired — 20 distinct hosts from one source within the 10-minute window
PASS artex_enrich_fanout: quiet — high volume, low breadth: 25 requests from one source but only 4 distinct hosts
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료하므로 CI 나 pre-commit 훅에 그대로 넣을 수
있습니다. pySigma 는 기준 버전(`2.0.0`)으로 고정돼 있습니다. 내부에 미러를 두었다면 `PYSIGMA_VERSION`
으로 버전을, `PYTHON_IMAGE` 로 이미지를 재정의하십시오.
## Sigma 백엔드 이식성: [`sigma_backends/`](sigma_backends/)
[`sigma_backends/run.sh`](sigma_backends/run.sh) 는 규칙이 Sigma 테스트가 돌려 보는 단일 Splunk
예시를 넘어서도 변환됨을 증명하고, [`../README.ko.md`](../README.ko.md) 의 백엔드별 지원 표를 정직하게
유지합니다. Sigma 상관 규칙 변환은 백엔드에 따라 다르므로, README 는 어떤 `-t` 대상이 트리 전체를
받고 어떤 대상이 원자 규칙만 받는지 방어자에게 알려 줍니다. 다시 돌려 봐야만 믿을 수 있는
주장입니다. 두 가지 속성을 단언하는데, 둘 다 긍정형이라 실제 회귀가 있을 때만 실패합니다:
- **상관 규칙의 이식성**: 트리 전체(원자 + 상관)가 Splunk, Elasticsearch `eql` 대상, Grafana
`loki` 에서 변환되며, 보강 지표가 각 질의에 그대로 살아남습니다. 상관 규칙이 Splunk 전용이 아님을
보여 줍니다.
- **원자 전용 폴백 동작**: 다섯 개의 원자 규칙은 `lucene` 과 Microsoft `kusto` 백엔드에서도
변환됩니다. 이 백엔드들은 고정된 버전에서 Sigma 상관 규칙 변환을 지원하지 않으므로, 해당 백엔드를
쓰는 방어자는 원자 규칙을 배포하고 상관 윈도우는 그 백엔드 고유 기능으로 표현할 수 있습니다.
"백엔드 X 는 상관 규칙을 처리하지 못한다"는 부정형은 일부러 단언하지 않습니다. 그렇게 하면 백엔드가
*개선되는* 것이 빨간 빌드가 되기 때문입니다. 정직한 한계는 README 에 적어 두었고, 이 테스트의 명령이
그것을 재현합니다. [`sigma_backends/check.sh`](sigma_backends/check.sh) 는 컨테이너 안쪽 절반입니다.
고정된 sigma-cli 와 네 백엔드를 설치하고, 읽기 전용으로 마운트한 규칙 트리를 읽습니다.
### 실행
Docker 만 있으면 됩니다. sigma-cli 와 백엔드들이 컨테이너에서 돌고 저장소에는 아무것도 쓰지
않습니다.
```sh
detections/tests/sigma_backends/run.sh
```
예상 출력(축약):
```
PASS whole tree (atomic + correlation) converts on 'eql', enrich indicator survives
PASS five atomic rules convert on 'kusto', enrich indicator survives
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료합니다. sigma-cli 는 고정돼 있고(`3.1.0`,
`SIGMA_CLI_VERSION` 으로 재정의), 백엔드 플러그인은 호환되는 최신 버전으로 설치됩니다. 그래서 이
스위트는 상류 백엔드 릴리스에 가장 민감합니다. 지원을 떨어뜨린 플러그인은 빌드를 빨갛게 만들고, 이는
고정 버전과 README 표를 함께 갱신하라는 신호입니다.
## SigmaHQ 관례 린트: [`sigma_lint/`](sigma_lint/)
[`sigma_lint/run.sh`](sigma_lint/run.sh) 는 README 와 `CONTRIBUTING.md` 의 "`sigma check` 를 깨끗이
통과한다"는 약속이 pySigma 의 핵심 검사뿐 아니라 SigmaHQ 의 관례까지 포함하도록 만듭니다. 그냥
`sigma check` 는 `pySigma-validators-sigmahq` 플러그인을 적재하지 않으므로, 제목 대소문자·필드명 분류
체계·로그소스 분류 체계·참조 링크 관례가 검사되지 않고 지나갑니다. 이 스위트는 그 플러그인을 설치하고,
[`sigma_lint/validators.yml`](sigma_lint/validators.yml) 에 문서화한 기준선에 맞춰 전체 검사 집합을
돌립니다. 두 가지 속성을 단언합니다:
- **문서화한 기준선이 깨끗함**: `validators.yml` 과 함께 `sigma check` 를 돌리면 오류 0, 이슈 0 을
보고합니다.
- **전체 집합이 살아 있고, 문서화한 제외만 남음**: 제외 없이 모든 SigmaHQ 검증기를 돌려도 이슈가
보고되며, 그 각각은 `validators.yml` 이 일부러 끄는 네 검사 중 하나입니다(그 외에는 없음). 이것은
공허 방지 가드입니다. 플러그인이 적재에 실패했다면 전체 실행이 아무것도 보고하지 않아 첫 번째 속성이
잘못된 이유로 통과할 것이므로, 알려진 제외 항목이 반드시 나타나도록 요구합니다.
네 제외 항목은 SigmaHQ 의 모노레포 파일 정리 방식(로그소스 접두어가 붙은 파일명과 `correlation_`
파일명)과 분류 체계(일반 `application` 로그소스, 제품명 없는 `process_creation`), 그리고 브랜치 대
영구링크(permalink) 참조 관례를 담습니다. 어느 것도 자기 저장소의 살아 있는 문서를 참조하는 작고
독립적인 규칙 집합에는 맞지 않습니다. 각 제외 항목은 그 근거를 `validators.yml` 안에 함께 적어
두었습니다. *나머지* 모든 SigmaHQ 검사는 강제되므로, 새 관례 이슈를 들인 규칙(대소문자가 틀린 제목,
분류 체계를 벗어난 필드명)은 빌드를 빨갛게 만듭니다. `pySigma-validators-sigmahq` 는 고정돼 있고
(`0.21.0`, `SIGMAHQ_VALIDATORS_VERSION` 으로 재정의), 버전을 올리면 새 관례가 드러날 수 있는데, 이는
규칙이나 문서화한 기준선을 갱신하라는 신호입니다.
### 실행
Docker 만 있으면 됩니다. sigma-cli 와 검증기 플러그인이 컨테이너에서 돌고 저장소에는 아무것도 쓰지
않습니다.
```sh
detections/tests/sigma_lint/run.sh
```
예상 출력(축약):
```
PASS sigma check with the documented baseline: 0 errors, 0 issues
PASS every reported issue is one of the four documented exclusions
RESULT: PASS
```
## ATT&CK 레이어: [`attack/`](attack/)
[`attack/run.sh`](attack/run.sh) 는 [`../attack/artex_navigator_layer.json`](../attack/artex_navigator_layer.json)
의 [ATT&CK 커버리지 레이어](../attack/)가 커버한다고 주장하는 규칙과 어긋나지 않는지 확인합니다. 규칙
집합에서 어긋난 커버리지 레이어는 없느니만 못하므로, 이 테스트는 "이 규칙들이 이 ATT&CK 기법들을
커버한다"를 검토자가 소스에서 다시 돌려 볼 수 있는 것으로 바꿉니다. 단언하는 것:
- **유효한 레이어**: 파일이 JSON 으로 파싱되고 필수 ATT&CK Navigator v4.x 필드를 지니며, 모든
항목에 올바른 형식의 기법 ID 와 유효한 ATT&CK 전술이 있습니다.
- **양방향 일치**: 점수가 매겨진 기법이 Sigma 규칙의 `attack.*` 기법 태그와 *정확히* 일치합니다.
레이어에 빠진 규칙 기법도, 규칙에 없는 레이어 기법도 없습니다. 전술도 같은 방식으로 일치합니다.
- **근거 있음**: 점수가 매겨진 모든 기법의 주석이 실재하는 규칙 파일을 가리키므로, 레이어가 이름이
바뀌거나 삭제된 규칙을 인용할 수 없습니다.
이것은 발화 테스트가 아니라 일관성 검사입니다. 탐지 백엔드가 필요 없고 Python 표준 라이브러리만 있으면
되므로, Sigma·Suricata 테스트와 달리 버전에 의존하는 경보 수가 없습니다. [`attack/check.py`](attack/check.py)
는 컨테이너 안쪽 절반입니다. 읽기 전용으로 마운트한 탐지 트리를 읽고 아무것도 쓰지 않습니다.
### 실행
Docker 만 있으면 됩니다. 검사가 Python 컨테이너에서 돌고 저장소에는 아무것도 쓰지 않습니다.
```sh
detections/tests/attack/run.sh
```
예상 출력(축약):
```
PASS scored techniques match the rule set exactly (8: T1059, T1105, T1485, T1489, T1557, T1561.002, T1592, T1595)
PASS scored tactics match the rule set exactly (collection, command-and-control, credential-access, execution, impact, reconnaissance)
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료하므로 CI 나 pre-commit 훅에 그대로 넣을 수
있습니다. 내부에 미러를 두었다면 `PYTHON_IMAGE` 로 이미지를 재정의하십시오.
## 지표 근거(source-of-truth): [`indicators/`](indicators/)
[`indicators/run.sh`](indicators/run.sh) 는 위 세 테스트가 하지 못하는 한 가지를 증명합니다. 각 규칙이
고정한 지표가 여전히 ARTEX 자신의 소스가 실제로 내보내는 문자열인지입니다. Sigma 테스트는 지표가
규칙→질의 *컴파일*을 거쳐 살아남음을 증명하고, ATT&CK 테스트는 레이어가 규칙 태그와 일치함을
증명하며, Suricata 테스트는 네트워크 규칙이 생성한 캡처에서 *발화*함을 증명합니다. 어느 것도 지표가
유래했다고 주장하는 소스 파일을 되짚어 보지는 않습니다. 이들이 모두 놓치는 부패는, 프로버 User-Agent
를 `artex-enrich/2.0` 으로 올리거나 가드 마커를 다시 쓰는 상류 재동기화입니다. 그래도 규칙은 모두
컴파일되고, 레이어는 여전히 일치하고, pcap 테스트도 여전히 발화합니다. 그런데 배포된 규칙은 실제
ARTEX 트래픽에 조용히 매칭을 멈춥니다. 각 지표에 대해 양방향으로 단언합니다:
- **소스가 여전히 내보냄**: 값이 그것을 만들어 내는 상류 소스 파일에 존재합니다(`enrich/enrich.go`
의 `artex-enrich/1.0`, `selfupdate/` 의 `artex-selfupdate`, `guard/guard.go` 의 가드 마커). 값이
없다는 것은 규칙이 아직 따라잡지 못한 상류 변경을 뜻합니다.
- **규칙이 여전히 고정함**: 값이 그것을 기반으로 세운 규칙에 존재하므로, 규칙 편집이 지표를 소스에서
조용히 떼어 놓을 수 없습니다. Suricata 규칙은 `startswith` 접두어로 확인하는데, 이는 그 규칙이
실제로 와이어를 매칭하는 방식과 같습니다.
- **차단 목록 대응**: 파괴적 명령 토큰(`rm -rf`, `mkfs`, `DROP DATABASE`, `FLUSHALL`)이 ARTEX
가드의 차단 목록(`db/db.go`)과 그것을 반영한 헌팅 규칙 양쪽에 나타납니다. 이것들은 고유 지문이
아니라 일반 헌팅 단서이므로, 테스트는 규칙이 실제로 주장하는 대응 관계만 단언합니다.
- **공개 목록이 근거를 유지함**: 방어자가 가져다 쓰는 산출물인 기계가 읽는 지표 목록
[`detections/indicators/artex_indicators.csv`](../indicators/artex_indicators.csv) 을 행 단위로 다시
읽습니다. 모든 값은 인용한 소스 파일에 여전히 존재하고 인용한 규칙에 고정돼 있어야 하며, 테스트가
근거를 확인한 모든 지문은 이 목록에 나타나야 합니다. 그래서 공개된 CSV 는 자신이 유래했다고 주장하는
소스에서 어느 방향으로도 조용히 어긋날 수 없습니다.
- **두 관문이 고정된 각 소스에서 발화함**: 테스트가 읽는 모든 상류 소스는 그것을 돌리는 두 관문에
포함됩니다. CI 워크플로의 `push`·`pull_request` paths 필터([`.github/workflows/detections.yml`](../../.github/workflows/detections.yml))와
로컬 pre-commit 훅의 `files` 정규식([`.pre-commit-config.yaml`](../../.pre-commit-config.yaml))입니다.
필요한 집합은 지표 자체에서 파생되므로, 새 소스를 고정하면서(예전의 `cmd/artex/main.go` 포트가
그랬듯) *두* 관문에 모두 배선하지 않으면 여기서 실패합니다. 그러지 않으면 그 소스만 건드린 변경이
그 소스를 빠뜨린 관문에서 테스트를 건너뜁니다. CI 에서는 머지 게이트를 초록으로 통과하고, 훅에서는
"CI 와 같은 소스 범위"라고 약속해 놓고도 로컬에서 끝내 잡히지 않습니다.
이는 [`../README.ko.md`](../README.ko.md) 의 약속("여기 모든 지표는 추정이 아니라 이 저장소 소스에서 확인한
문자열에 근거한다")과 CONTRIBUTING 의 첫 번째 기여 계약을, 검토자가 다시 돌려 볼 수 있는 가드로
바꿉니다. ATT&CK 테스트처럼 탐지 백엔드가 필요 없고 Python 표준 라이브러리만 있으면 됩니다.
[`indicators/check.py`](indicators/check.py) 는 규칙 트리, 공개 지표 목록, 고정된 소스 패키지, 그리고
그것을 발화시키는 두 관문(CI 워크플로와 pre-commit 설정)을 읽기 전용으로 마운트해 읽고, 아무것도 쓰지
않습니다.
### 실행
Docker 만 있으면 됩니다. 검사가 Python 컨테이너에서 돌고 저장소에는 아무것도 쓰지 않습니다.
```sh
detections/tests/indicators/run.sh
```
예상 출력(축약):
```
PASS enrichment prober User-Agent: 'artex-enrich/1.0' emitted by enrich/enrich.go
PASS detections/sigma/artex_enrich_user_agent.yml pins 'artex-enrich/1.0'
PASS 'FLUSHALL' present in both db/db.go and detections/sigma/destructive_command_hunting.yml
PASS enrich-user-agent: 'artex-enrich/1.0' grounded in enrich/enrich.go
PASS tested fingerprint 'artex-enrich/1.0' is published in the list
PASS .github/workflows/detections.yml push paths covers cmd/artex/main.go
PASS .pre-commit-config.yaml files covers cmd/artex/main.go
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료하므로 CI 나 pre-commit 훅에 그대로 넣을 수
있습니다. 내부에 미러를 두었다면 `PYTHON_IMAGE` 로 이미지를 재정의하십시오.
## MISP 내보내기 일관성: [`misp/`](misp/)
[`misp/run.sh`](misp/run.sh) 는 지표의 두 번째 공개 형태, 바로 가져올 수 있는 MISP 이벤트
[`detections/indicators/artex_indicators.misp.json`](../indicators/artex_indicators.misp.json) 를
다룹니다. 위 지표 테스트가 CSV 를 소스에 근거하게 유지한다면, 이 테스트는 방어자가 실제로 위협
인텔리전스 플랫폼에 적재하는 산출물인 MISP 이벤트가 그 CSV 에서 어긋나지 않게 유지합니다. 단언하는 것:
- **정말로 MISP 임**: 이벤트가 [pymisp](https://github.com/MISP/PyMISP) 로 적재되는데, 그 객체
모델은 `type` 이 진짜 MISP 타입이 아닌 속성을 거부합니다. 그럴듯해 보여도 유효하지 않은 타입은
여기서 실패하므로, "유효한 MISP"는 그냥 주장하는 것이 아니라 MISP 서버가 쓰는 라이브러리로
증명됩니다.
- **CSV 와 행 단위 동기화**: 모든 CSV 행이 의도한 타입·카테고리를 가진 MISP 속성 정확히 하나로
대응되고(`http.user-agent` → `user-agent`, 가드 마커 `string` → `pattern-in-file`, `port` →
`port`, `ip-dst|port` → 합성 `ip|port` 값을 가진 `ip-dst|port`, 탐색 스키마 `other` → `other`), CSV 행 없이 남는 MISP 속성이
하나도 없습니다. 이벤트는 CSV 와 함께 손으로 유지하므로, `artex_indicators.misp.json` 을 같은
커밋에서 맞춰 갱신하지 않은 채 CSV 행을 추가·삭제·타입 변경하면 실패합니다.
- **`to_ids` 가 `rule` 열을 반영함**: 규칙이 뒷받침하는 지표는 `to_ids: true` 이고, 규칙이 없는
호스트 포렌식 행은 `disable_correlation: true` 와 함께 `to_ids: false` 입니다. CSV 가 함의하는
것과 다르게 플래그를 뒤집으면 실패하므로, MISP 이벤트는 어떤 지문이 실행 가능한지를 조용히
부풀리거나 줄여 주장할 수 없습니다.
- **가드 마커가 바이트 단위로 보존되고** `detections/**` 가 CI paths 필터에 있어, CSV 나 이벤트를
바꾸면 이 스위트가 발화합니다.
위의 순수 표준 라이브러리 테스트들과 달리, 이 스위트는 컨테이너 안에 고정된 `pymisp` 를
설치합니다(호스트에는 아무것도 설치하지 않음). [`misp/check.py`](misp/check.py) 는 CSV, MISP 이벤트,
CI 워크플로를 읽기 전용으로 마운트해 읽고, 아무것도 쓰지 않습니다.
### 실행
Docker 만 있으면 됩니다. pymisp 가 컨테이너에 설치되고 저장소에는 아무것도 쓰지 않습니다.
```sh
detections/tests/misp/run.sh
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료합니다. 내부에 미러를 두었다면
`PYTHON_IMAGE` 로 이미지를, `PYMISP_VERSION` 으로 고정된 라이브러리를 재정의하십시오.
## 기여
새 탐지 규칙은 그것이 발화함을 보여 주는 테스트가 있을 때 더 강합니다. 테스트는 자기 입력을 결정론적으로
생성하고, 엔진 버전과 무관한 속성은 정확히 단언하며(그보다 무른 속성은 기준값을 기록한 하한으로), 공격
안내로 읽힐 수 있는 내용은 피해야 합니다. [`../../CONTRIBUTING.md`](../../CONTRIBUTING.md) 와
[`../README.ko.md`](../README.ko.md) 의 규칙 색인을 보십시오.
여덟 스위트는 모두 `detections/` 를 건드리는 모든 push 나 pull request 에서 CI 로 돕니다
([`../../.github/workflows/detections.yml`](../../.github/workflows/detections.yml) 참조). 그리고 지표
테스트는 그것이 고정한 상류 소스 파일(`enrich/`, `selfupdate/`, `guard/`, `db/`, `cmd/artex/main.go`)이
바뀔 때도 돕니다. 그래서 지표를 떨어뜨리거나, ATT&CK 레이어에서 어긋나거나, 문서화한 백엔드에서
변환이 멈추거나, SigmaHQ 관례를 깨거나, 소스와 동기화가 어긋나거나, MISP 이벤트가 CSV 에서 어긋나게
두거나, 워크플로가 아직 감시하지 않는 새 소스를 고정하는 규칙 변경은 머지되기 전에 빌드를 빨갛게
만듭니다.
+456
View File
@@ -0,0 +1,456 @@
# ARTEX detection rule tests
English · [한국어](README.ko.md)
> 한국어: 이 디렉터리는 [`../`](../)의 탐지 규칙이 실제로 발화하는지를 재현 가능하게 증명하는
> 회귀 테스트입니다. 바이너리 캡처를 저장소에 넣지 않고, 패킷 캡처를 매번 결정론적으로 생성한 뒤
> [Suricata](https://suricata.io)로 직접 돌려 경보 수를 확인합니다. 모든 테스트는 자신이 소유하거나
> 서면 허가를 받은 시스템을 지키는 **방어·탐지 목적에만** 쓰십시오. 한국어 전체 문서는
> **[README.ko.md](README.ko.md)** 를 보십시오.
Reproducible regression tests that prove the rules under [`../`](../) actually fire — and, just as
important, stay silent on benign traffic. A detection rule you cannot run is a claim; these tests turn the
claims in the rule files and the defense guide into something a reviewer can re-run from source.
No binary packet capture is committed. The capture is **synthesized deterministically on every run** and
removed afterwards, so the test ships as readable source, not as an opaque fixture, and never bloats the
repository.
## Run every suite at once — [`run-all.sh`](run-all.sh)
[`run-all.sh`](run-all.sh) runs all eight suites below in one command, in the same order as CI, so you do not
have to invoke the eight `run.sh` scripts by hand. Each suite runs to completion even if an earlier one fails,
the script prints a one-line PASS/FAIL summary per suite at the end, and it exits non-zero if any suite failed.
Before the suites, it runs a harness self-check ([`check-harness-sync.sh`](check-harness-sync.sh)) that fails
the run if this suite list, the per-suite steps in [CI](../../.github/workflows/detections.yml), and the suite
directories on disk ever name different suites or a different order. That is the one gap the eight suites
cannot see on their own: a suite wired into only one of the three (a new CI step with no `run-all.sh` entry, or
a directory never added to either) would otherwise pass every per-suite test while a green local `run-all.sh`
quietly stopped meaning a green CI. The check is a gate, not a ninth suite: it stays out of the summary
below, so the eight detection suites stay eight.
```sh
detections/tests/run-all.sh
```
Expected output (abridged):
```
===== detection suites summary =====
PASS sigma
PASS sigma_match
PASS sigma_lint
PASS sigma_backends
PASS suricata
PASS attack
PASS indicators
PASS misp
RESULT: PASS
```
Because it exits non-zero on any failure, it drops straight into a pre-commit hook. A ready-to-use example
lives in [`.pre-commit-config.yaml`](../../.pre-commit-config.yaml) at the repository root: install it with
`pip install pre-commit && pre-commit install`, and the runner then fires on commits that touch the detection
rules or the upstream source files they pin — the same scope as CI. The image and version overrides the
individual suites honour (`PYTHON_IMAGE`, `SIGMA_CLI_VERSION`, `SIGMAHQ_VALIDATORS_VERSION`, `SURICATA_IMAGE`)
are inherited by the runner, so exporting any of them applies to every suite at once.
## Suricata — [`suricata/`](suricata/)
[`suricata/run.sh`](suricata/run.sh) exercises the network rules in
[`../suricata/artex.rules`](../suricata/artex.rules) end to end and asserts five properties:
- **Valid** — the whole rules file loads under `suricata -T --init-errors-fatal`, so a rule that fails to
parse or initialise is caught even when no capture below exercises it. Plain `suricata -r` skips such a
rule and still exits 0, so this load check is the Suricata analogue of the Sigma suite's `sigma check`
validity assertion.
- **Presence (enrich)** — sid `1000001` fires exactly once per enrichment probe.
- **Velocity** — sid `1000002` fires once the `detection_filter` rate of 30 requests in 300 s per source is
crossed.
- **Presence (WebFetch)** — sid `1000003` fires exactly once per norma WebFetch request, and the enrich sids
stay silent on that same capture — so the two network signatures are mutually specific, not just each
present.
- **Specificity** — an identical capture whose only change is a benign browser User-Agent produces **zero**
ARTEX alerts.
[`suricata/gen_pcap.py`](suricata/gen_pcap.py) builds the capture with [scapy](https://scapy.net): N
independent plaintext HTTP request/response flows from one fixed source, each carrying a chosen
User-Agent, at a fixed base timestamp spaced one second apart. It only writes a file — it never sends a
packet or touches a network.
### Run it
Needs only Docker; scapy and Suricata both run in containers.
```sh
detections/tests/suricata/run.sh
```
Expected output (abridged):
```
PASS ruleset loads with zero parse/init errors (suricata -T)
PASS sid 1000001 presence: one alert per probe (got 35, want eq 35)
PASS sid 1000002 velocity: fires past 30-in-300s (got 5, want ge 1)
PASS sid 1000003 presence: one alert per WebFetch request (got 8, want eq 8)
PASS enrich sids stay silent on norma traffic (specificity) (got 0, want eq 0)
PASS benign browser UA produces no ARTEX alerts (got 0, want eq 0)
RESULT: PASS
```
The script exits non-zero if any assertion fails, so it drops straight into CI or a pre-commit hook.
Override the images with `SURICATA_IMAGE` / `PYTHON_IMAGE` if you mirror them internally.
### Why the velocity count is a floor, not an exact match
`run.sh` asserts the presence counts (`1000001 == 35`, `1000003 == 8`) and the benign count (`== 0`) exactly,
because those are engine-version-independent: one alert per matching request, and no match on a different
User-Agent. The
velocity rule's count depends on how a given Suricata release resolves the `detection_filter` threshold at
the boundary, so the test asserts `>= 1` and records the reference value separately. On **Suricata 8.0.7**
the reference run produces **5** alerts on sid `1000002` (flows 31–35, after the 30-in-300 s threshold is
crossed).
## Sigma — [`sigma/`](sigma/)
[`sigma/run.sh`](sigma/run.sh) validates the Sigma rules under [`../sigma/`](../sigma/) structurally and by
compilation with [sigma-cli](https://github.com/SigmaHQ/sigma-cli) (pySigma), and asserts five properties:
- **Valid** — `sigma check` reports 0 errors, 0 condition errors, and 0 issues over the whole tree.
- **Compiles** — `sigma convert -t splunk` turns the whole tree into a backend query language without error.
- **Indicators survive** — each atomic indicator string (`artex-enrich/1.0`, `artex-selfupdate`, the guard
marker, and the recording-proxy CA filename `mitmproxy-ca-cert.pem`) is still present in the compiled query,
so a rule cannot silently lose the string it is built on.
- **Correlations compile** — the behaviour rules in [`../sigma/correlation/`](../sigma/correlation/) emit their
`event_count` / `value_count` aggregations rather than being dropped.
- **Correlations are load-bearing** — converting one correlation rule *alone* fails, because it references its
atomic base rule by `id`; the reference is enforced, not decorative. This is the Sigma analogue of the
Suricata specificity assertion above.
This is the structural + compilation validation documented in [`../README.md`](../README.md), made executable
and assertive. The companion [`sigma_match/`](sigma_match/) suite below adds the *matching* half for both the
atomic and the correlation rules — a representative malicious event (or timeline) fires each rule and a benign
one does not — so the Sigma rules now get both a reproducible validation test and a reproducible matching test,
the way the Suricata rule does. (The
earlier concern that a weak hand-written matcher would undercut the rules is addressed by delegating all parsing
to pySigma; see the trust model in the next section.)
### Run it
Needs only Docker; sigma-cli and the splunk backend run in a container and nothing is written to the repo.
```sh
detections/tests/sigma/run.sh
```
Expected output (abridged):
```
PASS sigma check: 0 errors, 0 condition errors, 0 issues
PASS whole tree converts to splunk (exit 0)
PASS indicator present: artex-enrich/1.0
PASS correlation rule fails to convert alone — it requires its atomic base rule
RESULT: PASS
```
The script exits non-zero if any assertion fails, so it drops straight into CI or a pre-commit hook. sigma-cli
is pinned to a reference version (`3.1.0`); override it with `SIGMA_CLI_VERSION`, or the image with
`PYTHON_IMAGE`, if you mirror them internally.
## Sigma live event-matching — [`sigma_match/`](sigma_match/)
[`sigma_match/run.sh`](sigma_match/run.sh) proves the Sigma rules under [`../sigma/`](../sigma/) — both the
atomic rules and the correlation rules under [`../sigma/correlation/`](../sigma/correlation/) — actually *fire*
on a matching event (or timeline) and stay quiet on a benign one — the "a detection you cannot run is only a
claim" guarantee the Suricata suite gives the network rule, extended here to the host/log-layer rules. It
asserts six properties, three for the atomic rules and three for the correlations:
- **Rule/sample pairing** — every atomic rule has an [`events/<name>.json`](sigma_match/events/) sample file and
every sample file maps back to a rule, so a rule added without samples fails here rather than going untested.
- **True positives** — each rule matches every one of its malicious sample events.
- **True negatives** — each rule matches none of its benign sample events. For example, a standalone
`mitmproxy-ca-cert.pem` under `.mitmproxy/` does **not** trip the recording-proxy rule, because its `|all`
modifier also requires the `_ca/` directory ARTEX writes — the matching test is what proves that discrimination.
- **Correlation rule/timeline pairing** — every correlation rule has an
[`events/correlation/<name>.json`](sigma_match/events/correlation/) timeline file and every timeline maps back
to a rule. Each timeline event carries a `ts` field in relative seconds.
- **Correlation true positives** — each rule fires on a positive timeline where the threshold is met inside the
window within one group. For example, requests from one source (`c-ip`) fanning out to 20 distinct hosts
within 10 minutes trip the enrichment fan-out rule.
- **Correlation true negatives** — each rule stays quiet when the threshold is not met, when it is met but the
events are spread beyond the window, when they are split across groups, or when a temporal rule is missing a
leg. In particular a high-volume, low-breadth burst does **not** trip the fan-out rule: breadth, not volume, is
the signal, and the matching test is what proves that discrimination.
The trust model is that pySigma — not hand-written code — parses each rule: an atomic rule into a condition tree
(`|contains` → a wildcard value, `|all` → an AND, `1 of selection_*` → an OR), and a correlation rule into its
aggregation spec (type, group-by, timespan, threshold, and the resolved references to the atomic base rules).
[`check.py`](sigma_match/check.py) only walks that tree and spec, deciding which events feed a referenced rule
with the very same atomic matcher, so the authoritative Sigma logic stays in pySigma; it raises rather than
passing on any construct it does not explicitly support (fail-closed). Scope and limits are stated in the script
header: the correlation window is the standard sliding-window interpretation (a `timespan`-second window
anchored at each matching event) and a real SIEM's windowing may differ; matching is **case-insensitive** (the
splunk-backend default the `sigma/` suite targets, which the destructive rule's own false-positive note
assumes); and keyword matching is a full-text substring search. It is a regression test for the rules'
field/value/condition/aggregation logic, not a substitute for validating in your own SIEM, whose field
normalization may differ.
### Run it
Needs only Docker; pySigma runs in a container and nothing is written to the repo.
```sh
detections/tests/sigma_match/run.sh
```
Expected output (abridged):
```
PASS rule/sample pairing: 5 atomic rules, 5 event files, no orphans
PASS artex_enrich_user_agent: 1/1 positive events matched
PASS artex_recording_proxy_ca: 2/2 benign events correctly not matched
PASS rule/timeline pairing: 4 correlation rules, 4 timeline files, no orphans
PASS artex_enrich_fanout: fired — 20 distinct hosts from one source within the 10-minute window
PASS artex_enrich_fanout: quiet — high volume, low breadth: 25 requests from one source but only 4 distinct hosts
RESULT: PASS
```
The script exits non-zero if any assertion fails, so it drops straight into CI or a pre-commit hook. pySigma is
pinned to a reference version (`2.0.0`); override it with `PYSIGMA_VERSION`, or the image with `PYTHON_IMAGE`,
if you mirror them internally.
## Sigma backend portability — [`sigma_backends/`](sigma_backends/)
[`sigma_backends/run.sh`](sigma_backends/run.sh) proves the rules convert beyond the single Splunk example the
Sigma test exercises, and keeps the per-backend support matrix in [`../README.md`](../README.md) honest. Sigma
correlation conversion is backend-dependent, so the README tells a defender which `-t` targets take the whole
tree and which take only the atomic rules — a claim that is only trustworthy if it is re-run. It asserts two
properties, both positive so the test fails only on a real regression:
- **Correlations are portable** — the whole tree (atomic + correlation) converts on Splunk, the Elasticsearch
`eql` target, and Grafana `loki`, with the enrich indicator surviving into each query. This shows the
correlation rules are not Splunk-only.
- **Atomic-only fallback works** — the five atomic rules still convert on `lucene` and the Microsoft `kusto`
backend, which do not support Sigma correlation conversion at the pinned versions, so a defender on those
backends can deploy the atomic rules and express the correlation window natively.
It deliberately does not assert the negative "backend X cannot do correlations": that would turn a backend
*improving* into a red build. The honest limitation lives in the README, reproduced by this test's commands.
[`sigma_backends/check.sh`](sigma_backends/check.sh) is the in-container half; it installs the pinned sigma-cli
plus four backends and reads the rule tree mounted read-only.
### Run it
Needs only Docker; sigma-cli and the backends run in a container and nothing is written to the repo.
```sh
detections/tests/sigma_backends/run.sh
```
Expected output (abridged):
```
PASS whole tree (atomic + correlation) converts on 'eql', enrich indicator survives
PASS five atomic rules convert on 'kusto', enrich indicator survives
RESULT: PASS
```
The script exits non-zero if any assertion fails. sigma-cli is pinned (`3.1.0`, override with
`SIGMA_CLI_VERSION`); the backend plugins install at their latest compatible version, so this suite is the one
most sensitive to an upstream backend release — a plugin that drops support turns the build red, which is the
signal to update the pin and the README matrix together.
## SigmaHQ convention lint — [`sigma_lint/`](sigma_lint/)
[`sigma_lint/run.sh`](sigma_lint/run.sh) makes the "passes `sigma check` cleanly" promise in the README and
`CONTRIBUTING.md` cover SigmaHQ's conventions, not just pySigma's core checks. Plain `sigma check` does not
load the `pySigma-validators-sigmahq` plugin, so title casing, field-name taxonomy, logsource taxonomy, and
reference-link conventions go unchecked. This suite installs that plugin and runs the full set against the
documented baseline in [`sigma_lint/validators.yml`](sigma_lint/validators.yml). It asserts two properties:
- **The documented baseline is clean** — `sigma check` with `validators.yml` reports 0 errors and 0 issues.
- **The full set is live, and only the documented exclusions remain** — running every SigmaHQ validator with
no exclusions still reports issues, and each one is among the four checks `validators.yml` deliberately
disables (nothing else). This is the anti-vacuity guard: if the plugin failed to load, the full run would
report nothing and the first property would pass for the wrong reason, so the known exclusions are required
to appear.
The four exclusions encode SigmaHQ's monorepo filing scheme (logsource-prefixed and `correlation_` filenames)
and taxonomy (a generic `application` logsource and a product-less `process_creation`), plus the
branch-vs-permalink reference convention — none of which fit a small standalone rule set that references its
own living docs. Each exclusion carries its rationale inline in `validators.yml`. Because every *other*
SigmaHQ check is enforced, a rule that picks up a new convention issue — a mis-cased title, an off-taxonomy
field name — turns the build red. `pySigma-validators-sigmahq` is pinned (`0.21.0`, override with
`SIGMAHQ_VALIDATORS_VERSION`); bumping it may surface new conventions, which is the signal to update the rules
or the documented baseline.
### Run it
Needs only Docker; sigma-cli and the validator plugin run in a container and nothing is written to the repo.
```sh
detections/tests/sigma_lint/run.sh
```
Expected output (abridged):
```
PASS sigma check with the documented baseline: 0 errors, 0 issues
PASS every reported issue is one of the four documented exclusions
RESULT: PASS
```
## ATT&CK layer — [`attack/`](attack/)
[`attack/run.sh`](attack/run.sh) checks that the [ATT&CK coverage layer](../attack/) in
[`../attack/artex_navigator_layer.json`](../attack/artex_navigator_layer.json) stays consistent with the
rules it claims to cover. A coverage layer that drifts from its rule set is worse than none, so this turns
"these rules cover these ATT&CK techniques" into something a reviewer can re-run from source. It asserts:
- **Valid layer** — the file parses as JSON and carries the required ATT&CK Navigator v4.x fields, with a
well-formed technique ID and a valid ATT&CK tactic on every entry.
- **Bidirectional match** — the scored techniques are *exactly* the `attack.*` technique tags on the Sigma
rules: no rule technique missing from the layer, no layer technique absent from the rules. The tactics
match the same way.
- **Grounded** — every scored technique's comment names a rule file that exists, so the layer cannot cite a
rule that was renamed or removed.
This is a consistency check, not a firing test: it needs no detection backend, only the Python standard
library, so unlike the Sigma and Suricata tests it carries no version-dependent counts. [`attack/check.py`](attack/check.py)
is the in-container half; it reads the detections tree mounted read-only and writes nothing.
### Run it
Needs only Docker; the check runs in a Python container and nothing is written to the repo.
```sh
detections/tests/attack/run.sh
```
Expected output (abridged):
```
PASS scored techniques match the rule set exactly (8: T1059, T1105, T1485, T1489, T1557, T1561.002, T1592, T1595)
PASS scored tactics match the rule set exactly (collection, command-and-control, credential-access, execution, impact, reconnaissance)
RESULT: PASS
```
The script exits non-zero if any assertion fails, so it drops straight into CI or a pre-commit hook.
Override the image with `PYTHON_IMAGE` if you mirror it internally.
## Indicator source-of-truth — [`indicators/`](indicators/)
[`indicators/run.sh`](indicators/run.sh) proves the one thing the three tests above do not: that each rule's
pinned indicator is still the string ARTEX's own source actually emits. The Sigma test proves an indicator
survives rule→query *compilation*; the ATT&CK test proves the layer matches the rules' tags; the Suricata
test proves the network rule *fires* on a synthesized capture. None of them look back at the source file the
indicator claims to come from. The rot they all miss is an upstream re-sync that bumps the prober User-Agent
to `artex-enrich/2.0` or rewrites the guard marker: every rule still compiles, the layer still matches, the
pcap test still fires — and the deployed rule silently stops matching real ARTEX traffic. It asserts, for
each indicator, bidirectionally:
- **Source still emits it** — the value is present in the upstream source file(s) that produce it
(`artex-enrich/1.0` in `enrich/enrich.go`, `artex-selfupdate` in `selfupdate/`, the guard marker in
`guard/guard.go`). A missing value means an upstream change the rule has not caught up with.
- **Rule still pins it** — the value is present in the rule built on it, so a rule edit cannot quietly move
the indicator away from its source. The Suricata rule is checked by its `startswith` prefix, matching how
it actually matches the wire.
- **Deny-list correspondence** — the destructive-command tokens (`rm -rf`, `mkfs`, `DROP DATABASE`,
`FLUSHALL`) appear both in ARTEX's guard deny-list (`db/db.go`) and in the hunting rule that mirrors it.
These are generic hunting leads, not unique fingerprints, so the test asserts only the correspondence the
rule actually claims.
- **Published list stays grounded** — the machine-readable indicator list
[`detections/indicators/artex_indicators.csv`](../indicators/artex_indicators.csv), the artifact a
defender imports, is re-read row by row: every value must still be present in the source file(s) it cites
and pinned in the rule(s) it cites, and every fingerprint the test grounds must appear in the list. So the
published CSV cannot silently drift from the source it claims to come from, in either direction.
- **Both gates fire on each pinned source** — every upstream source the test reads is covered by the two
gates that run it: the CI workflow's `push` and `pull_request` paths filter
([`.github/workflows/detections.yml`](../../.github/workflows/detections.yml)) and the local pre-commit
hook's `files` regex ([`.pre-commit-config.yaml`](../../.pre-commit-config.yaml)). The required set is
derived from the indicators themselves, so pinning a new source (as the `cmd/artex/main.go` ports once were)
without wiring it into *both* gates fails here — otherwise a change touching only that source skips the test
on whichever gate omits it: on CI it passes the merge gate green, on the hook it is never caught locally
even though the hook promises "the same source scope as CI".
This turns [`../README.md`](../README.md)'s promise — "every indicator here is grounded in a string verified
in this repository's source, not inferred" — and CONTRIBUTING's first contribution contract into a guard a
reviewer can re-run. Like the ATT&CK test it needs no detection backend, only the Python standard library;
[`indicators/check.py`](indicators/check.py) reads the rule tree, the published indicator list, the pinned
source packages, and the two gates that fire it (the CI workflow and the pre-commit config) mounted read-only
and writes nothing.
### Run it
Needs only Docker; the check runs in a Python container and nothing is written to the repo.
```sh
detections/tests/indicators/run.sh
```
Expected output (abridged):
```
PASS enrichment prober User-Agent: 'artex-enrich/1.0' emitted by enrich/enrich.go
PASS detections/sigma/artex_enrich_user_agent.yml pins 'artex-enrich/1.0'
PASS 'FLUSHALL' present in both db/db.go and detections/sigma/destructive_command_hunting.yml
PASS enrich-user-agent: 'artex-enrich/1.0' grounded in enrich/enrich.go
PASS tested fingerprint 'artex-enrich/1.0' is published in the list
PASS .github/workflows/detections.yml push paths covers cmd/artex/main.go
PASS .pre-commit-config.yaml files covers cmd/artex/main.go
RESULT: PASS
```
The script exits non-zero if any assertion fails, so it drops straight into CI or a pre-commit hook.
Override the image with `PYTHON_IMAGE` if you mirror it internally.
## MISP export consistency — [`misp/`](misp/)
[`misp/run.sh`](misp/run.sh) covers the second published form of the indicators — the ready-to-import MISP
event [`detections/indicators/artex_indicators.misp.json`](../indicators/artex_indicators.misp.json). The
indicator test above keeps the CSV grounded in the source; this test keeps the MISP event, the artifact a
defender actually loads into a threat-intelligence platform, from drifting away from that CSV. It asserts:
- **It is really MISP** — the event loads under [pymisp](https://github.com/MISP/PyMISP), whose object model
rejects any attribute whose `type` is not a genuine MISP type. A plausible-looking but invalid type fails
here, so "valid MISP" is proven by the library a MISP server uses, not asserted.
- **Row-for-row sync with the CSV** — every CSV row maps to exactly one MISP attribute with the intended type
and category (`http.user-agent` → `user-agent`, the guard marker `string` → `pattern-in-file`, `port` →
`port`, `ip-dst|port` → `ip-dst|port` with the composite `ip|port` value, and the exploration-schema `other` → `other`), and no MISP attribute is left
without a CSV row. The event is hand-maintained alongside the CSV, so adding, removing, or retyping a CSV
row without updating `artex_indicators.misp.json` to match in the same commit fails.
- **`to_ids` mirrors the `rule` column** — a rule-backed indicator is `to_ids: true`; a host-forensic row
with no rule is `to_ids: false` with `disable_correlation: true`. Flipping a flag away from what the CSV
implies fails, so the MISP event cannot quietly over- or under-claim which fingerprints are actionable.
- **The guard marker survives byte-for-byte** and `detections/**` is in the CI paths filter, so a change to
the CSV or the event triggers this suite.
Unlike the pure-standard-library tests above, this suite installs a pinned `pymisp` inside its container
(nothing is installed on the host); [`misp/check.py`](misp/check.py) reads the CSV, the MISP event, and the
CI workflow mounted read-only and writes nothing.
### Run it
Needs only Docker; pymisp is installed in the container and nothing is written to the repo.
```sh
detections/tests/misp/run.sh
```
The script exits non-zero if any assertion fails. Override the image with `PYTHON_IMAGE` and the pinned
library with `PYMISP_VERSION` if you mirror them internally.
## Contributing
A new detection rule is stronger with a test that shows it firing. Tests should synthesize their own input
deterministically, assert engine-version-independent properties exactly (and softer ones as floors with a
recorded reference), and avoid any content that reads as attack guidance. See
[`../../CONTRIBUTING.en.md`](../../CONTRIBUTING.en.md) and the rule indexes in [`../README.md`](../README.md).
All eight suites run in CI (see [`../../.github/workflows/detections.yml`](../../.github/workflows/detections.yml))
on every push or pull request that touches `detections/` — and the indicator test also runs when the upstream
source files it pins (`enrich/`, `selfupdate/`, `guard/`, `db/`, `cmd/artex/main.go`) change — so a rule
change that drops an indicator, drifts from the ATT&CK layer, stops converting on a documented backend,
breaks a SigmaHQ convention, falls out of sync with the source, lets the MISP event drift from the CSV, or
pins a new source the workflow does not yet watch turns the build red before it can merge.
+191
View File
@@ -0,0 +1,191 @@
#!/usr/bin/env python3
#
# In-container half of the ARTEX ATT&CK coverage-layer test. run.sh launches this
# inside a Python container with the detections tree mounted read-only at
# /detections. It proves that the ATT&CK Navigator layer in
# detections/attack/artex_navigator_layer.json stays consistent with the rules it
# claims to cover, so the layer cannot silently drift from the Sigma rule set:
#
# 1. the layer is valid JSON with the required Navigator v4.x fields
# 2. every technique entry has a well-formed ID and a valid ATT&CK tactic
# 3. the scored techniques are EXACTLY the attack.* techniques tagged on the
# rules (bidirectional: no rule technique missing from the layer, no layer
# technique absent from the rules)
# 4. the scored tactics are exactly the attack.* tactics tagged on the rules
# 5. every scored technique's comment grounds it in a rule file that exists
# 6. any score-less entry is a display-only parent of a scored sub-technique
# 7. scores stay within the gradient bounds
#
# Pure standard library (the slim image already ships python3); nothing is
# installed and nothing is written to the repo. Exits non-zero on any failure.
import glob
import json
import os
import re
import sys
DET = "/detections"
LAYER = os.path.join(DET, "attack", "artex_navigator_layer.json")
SIGMA = os.path.join(DET, "sigma")
# ATT&CK Enterprise tactic shortnames (the Navigator "tactic" field uses these).
VALID_TACTICS = {
"reconnaissance", "resource-development", "initial-access", "execution",
"persistence", "privilege-escalation", "defense-evasion", "credential-access",
"discovery", "lateral-movement", "collection", "command-and-control",
"exfiltration", "impact",
}
TECHNIQUE_RE = re.compile(r"^T\d{4}(\.\d{3})?$")
TAG_RE = re.compile(r"attack\.(t\d{4}(?:\.\d{3})?)", re.IGNORECASE)
TACTIC_TAG_RE = re.compile(r"attack\.([a-z][a-z-]+)")
fail = 0
def note(s):
print(" " + s)
def ok(s):
note("PASS " + s)
def bad(s):
global fail
note("FAIL " + s)
fail = 1
def rule_tags():
"""Techniques and tactics tagged across every Sigma rule file."""
techs, tactics = set(), set()
files = sorted(glob.glob(os.path.join(SIGMA, "**", "*.yml"), recursive=True))
for path in files:
with open(path, encoding="utf-8") as fh:
for line in fh:
m = TAG_RE.search(line)
if m:
techs.add(m.group(1).upper())
continue
t = TACTIC_TAG_RE.search(line)
if t and t.group(1) in VALID_TACTICS:
tactics.add(t.group(1))
return techs, tactics, files
print("== 1/7 layer parses as JSON with the required Navigator fields ==")
try:
with open(LAYER, encoding="utf-8") as fh:
layer = json.load(fh)
ok("artex_navigator_layer.json is valid JSON")
except Exception as exc: # noqa: BLE001
print(" FAIL cannot parse layer: %s" % exc)
print("RESULT: FAIL")
sys.exit(1)
for key in ("name", "versions", "domain", "techniques", "gradient"):
if key in layer:
ok("top-level key present: %s" % key)
else:
bad("top-level key missing: %s" % key)
for vkey in ("attack", "navigator", "layer"):
if vkey in layer.get("versions", {}):
ok("versions.%s present (%s)" % (vkey, layer["versions"][vkey]))
else:
bad("versions.%s missing" % vkey)
if layer.get("domain") == "enterprise-attack":
ok("domain is enterprise-attack")
else:
bad("domain is not enterprise-attack: %r" % layer.get("domain"))
techniques = layer.get("techniques", [])
scored = [t for t in techniques if "score" in t]
helpers = [t for t in techniques if "score" not in t]
print("== 2/7 every technique entry has a valid ID and tactic ==")
for t in techniques:
tid = t.get("techniqueID", "")
if TECHNIQUE_RE.match(tid):
ok("well-formed techniqueID: %s" % tid)
else:
bad("malformed techniqueID: %r" % tid)
tac = t.get("tactic", "")
if tac in VALID_TACTICS:
ok("valid tactic for %s: %s" % (tid, tac))
else:
bad("invalid tactic for %s: %r" % (tid, tac))
rule_techs, rule_tactics, rule_files = rule_tags()
layer_scored_ids = {t["techniqueID"] for t in scored}
layer_scored_tactics = {t["tactic"] for t in scored}
print("== 3/7 scored techniques == techniques tagged on the rules (bidirectional) ==")
if not rule_techs:
bad("found no attack.* technique tags in %s" % SIGMA)
missing_in_layer = rule_techs - layer_scored_ids
extra_in_layer = layer_scored_ids - rule_techs
if not missing_in_layer and not extra_in_layer:
ok("scored techniques match the rule set exactly (%d: %s)"
% (len(rule_techs), ", ".join(sorted(rule_techs))))
else:
if missing_in_layer:
bad("rule techniques missing from the layer: %s"
% ", ".join(sorted(missing_in_layer)))
if extra_in_layer:
bad("layer techniques not tagged on any rule: %s"
% ", ".join(sorted(extra_in_layer)))
print("== 4/7 scored tactics == tactics tagged on the rules ==")
if layer_scored_tactics == rule_tactics:
ok("scored tactics match the rule set exactly (%s)"
% ", ".join(sorted(rule_tactics)))
else:
bad("tactic mismatch: layer=%s rules=%s"
% (sorted(layer_scored_tactics), sorted(rule_tactics)))
print("== 5/7 each scored technique is grounded in a rule file that exists ==")
for t in scored:
comment = t.get("comment", "")
refs = re.findall(r"sigma/[\w./-]+\.yml", comment)
grounded = False
for ref in refs:
if os.path.exists(os.path.join(DET, ref)):
grounded = True
else:
bad("%s comment cites a missing rule file: %s" % (t["techniqueID"], ref))
if "suricata" in comment.lower():
grounded = True
if grounded:
ok("%s grounded in an existing rule reference" % t["techniqueID"])
else:
bad("%s comment cites no existing rule file" % t["techniqueID"])
print("== 6/7 any score-less entry is a display parent of a scored sub-technique ==")
if not helpers:
ok("no display-only entries (nothing to check)")
for h in helpers:
hid = h.get("techniqueID", "")
children = [s for s in scored if s["techniqueID"].startswith(hid + ".")]
if children and h.get("showSubtechniques") is True:
ok("%s is a display parent of %s"
% (hid, ", ".join(c["techniqueID"] for c in children)))
else:
bad("score-less entry %s is not a valid display parent "
"(needs showSubtechniques:true and a scored child)" % hid)
print("== 7/7 scores stay within the gradient bounds ==")
grad = layer.get("gradient", {})
lo, hi = grad.get("minValue", 0), grad.get("maxValue", 100)
for t in scored:
s = t["score"]
if lo <= s <= hi:
ok("%s score %s within [%s, %s]" % (t["techniqueID"], s, lo, hi))
else:
bad("%s score %s outside gradient [%s, %s]" % (t["techniqueID"], s, lo, hi))
print()
print("reference: %d Sigma rule files scanned, %d scored techniques, %d display parents"
% (len(rule_files), len(scored), len(helpers)))
print("RESULT: %s" % ("PASS" if fail == 0 else "FAIL"))
sys.exit(fail)
+32
View File
@@ -0,0 +1,32 @@
#!/usr/bin/env bash
#
# Reproducible consistency test for the ARTEX ATT&CK coverage layer
# (../../attack/artex_navigator_layer.json). A coverage layer that drifts from the
# rules it claims to cover is worse than none, so this turns "these rules cover
# these ATT&CK techniques" from a claim into something a reviewer can re-run from
# source. It catches the realistic regression: a rule is added, removed, or
# retagged, but the Navigator layer is not updated to match.
#
# It proves (see check.py for the assertions) that the layer is a valid Navigator
# v4.x document and that its scored techniques and tactics are EXACTLY the attack.*
# tags on the Sigma rules — no rule technique missing from the layer, no layer
# technique absent from the rules — with every scored technique grounded in a rule
# file that exists.
#
# No host dependency beyond Docker: the check is pure Python standard library and
# runs in a container with the detections tree mounted read-only. Nothing is
# installed on the host and nothing is written to the repo.
#
# Usage: detections/tests/attack/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
DET_DIR="$REPO/detections"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
docker run --rm \
-v "$DET_DIR:/detections:ro" \
-v "$HERE:/src:ro" \
"$PYTHON_IMAGE" python3 /src/check.py
+177
View File
@@ -0,0 +1,177 @@
#!/usr/bin/env python3
#
# Harness self-consistency check for the detection test suites. check-harness-sync.sh
# launches this inside a Python container with the detection tree and the CI
# workflow mounted read-only under /repo. It proves the one thing the eight
# detection suites cannot: that run-all.sh, the CI workflow, and the suite
# directories on disk all name the same suites in the same order.
#
# run-all.sh, CONTRIBUTING, and this directory's README all promise that
# "run-all.sh runs the same suites as CI, in the same order." Nothing enforced
# that promise. A suite added to only one of the three places — a new CI step with
# no SUITES entry, or a new directory never wired into either — passes every
# per-suite test while quietly breaking the promise: run-all.sh and CI run
# different sets, so a green run-all.sh locally no longer implies a green CI. This
# check closes that gap the same way the indicator test closes the
# CI-paths / pre-commit-regex gap: it reads the three sources of truth and asserts
# they agree.
#
# A = the SUITES="..." list in run-all.sh (ordered)
# B = the per-suite `run: detections/tests/<x>/run.sh` steps (ordered)
# in .github/workflows/detections.yml
# C = the subdirectories of detections/tests/ that carry a (a set)
# run.sh
#
# It asserts A == B as ordered lists (so the documented "same order as CI" holds)
# and set(A) == C (so no directory is orphaned and no listed suite is missing on
# disk). It is deliberately not a detection suite: it is not in SUITES, not a
# `.../run.sh` CI step, and not a run.sh directory, so it never counts itself and
# the eight detection suites stay eight.
#
# It parses only the stable machine-readable lines (the SUITES assignment, the
# `run:` steps, the directory listing), never prose or the example-output blocks
# in the READMEs, so it cannot go brittle on documentation wording.
#
# Pure standard library (the slim image already ships python3); nothing is
# installed and nothing is written to the repo. Exits non-zero on any mismatch.
#
# Usage. check-harness-sync.sh runs this inside Docker with the repo mounted at
# /repo, which is why ROOT defaults to /repo below. To run it directly on the
# host instead, point ARTEX_REPO_ROOT at the repo root:
#
# ARTEX_REPO_ROOT="$(git rev-parse --show-toplevel)" \
# python3 detections/tests/check-harness-sync.py
import os
import re
import sys
# Default to /repo, the mount point check-harness-sync.sh uses inside Docker.
# Track whether the caller set the variable so a missing-path failure can tell a
# host-direct runner why ROOT is /repo (see main()).
ROOT = os.environ.get("ARTEX_REPO_ROOT", "/repo")
ROOT_FROM_ENV = "ARTEX_REPO_ROOT" in os.environ
TESTS_DIR = os.path.join(ROOT, "detections", "tests")
RUN_ALL = os.path.join(TESTS_DIR, "run-all.sh")
CI_WORKFLOW = os.path.join(ROOT, ".github", "workflows", "detections.yml")
def fail(msg):
print(f"FAIL: {msg}", file=sys.stderr)
sys.exit(1)
def suites_from_run_all(path):
"""Ordered suite names in the SUITES="..." assignment in run-all.sh."""
with open(path, encoding="utf-8") as f:
text = f.read()
m = re.search(r'^SUITES="([^"]*)"', text, re.MULTILINE)
if not m:
fail(f'could not find a SUITES="..." assignment in {path}')
names = m.group(1).split()
if not names:
fail(f"SUITES in {path} is empty")
return names
def suites_from_ci(path):
"""Ordered suite names in the per-suite `run:` steps of the CI workflow."""
names = []
with open(path, encoding="utf-8") as f:
for line in f:
m = re.search(r"run:\s*detections/tests/([^/]+)/run\.sh\s*$", line)
if m:
names.append(m.group(1))
if not names:
fail(f"found no `run: detections/tests/<suite>/run.sh` steps in {path}")
return names
def suites_from_dirs(path):
"""Suite directories under detections/tests/ that carry a run.sh."""
names = set()
for entry in sorted(os.listdir(path)):
d = os.path.join(path, entry)
if os.path.isdir(d) and os.path.isfile(os.path.join(d, "run.sh")):
names.add(entry)
if not names:
fail(f"found no suite directories with a run.sh under {path}")
return names
def main():
for p in (RUN_ALL, CI_WORKFLOW, TESTS_DIR):
if not os.path.exists(p):
hint = ""
if not ROOT_FROM_ENV:
hint = (
"\n ARTEX_REPO_ROOT is unset, so ROOT defaulted to /repo "
"(the path check-harness-sync.sh mounts the repo at inside Docker).\n"
" To run this script directly on the host, point it at the "
"repo root:\n"
' ARTEX_REPO_ROOT="$(git rev-parse --show-toplevel)" '
"python3 detections/tests/check-harness-sync.py\n"
" or use the Docker wrapper: "
"detections/tests/check-harness-sync.sh"
)
fail(f"missing expected path: {p}{hint}")
a = suites_from_run_all(RUN_ALL)
b = suites_from_ci(CI_WORKFLOW)
c = suites_from_dirs(TESTS_DIR)
errors = []
# A name repeated in an ordered source would make the comparisons below read
# misleadingly, so surface it on its own first.
for label, seq in (("run-all.sh SUITES", a), ("CI steps", b)):
if len(seq) != len(set(seq)):
dupes = sorted({x for x in seq if seq.count(x) > 1})
errors.append(f"{label} names a suite more than once: {dupes}")
if a != b:
errors.append(
"run-all.sh SUITES and the CI steps disagree (order matters: "
"run-all.sh promises the same order as CI):\n"
f" run-all.sh: {a}\n"
f" CI steps : {b}"
)
if set(a) != c:
detail = []
only_listed = sorted(set(a) - c)
only_on_disk = sorted(c - set(a))
if only_listed:
detail.append(
f" listed in run-all.sh but no run.sh directory: {only_listed}"
)
if only_on_disk:
detail.append(
f" run.sh directory present but not in run-all.sh: {only_on_disk}"
)
errors.append(
"run-all.sh SUITES and the suite directories disagree:\n"
+ "\n".join(detail)
)
if errors:
print("detection test harness is OUT OF SYNC:\n", file=sys.stderr)
for e in errors:
print(e + "\n", file=sys.stderr)
print(
"Wire the new suite into all three (SUITES in run-all.sh, a step in "
".github/workflows/detections.yml, and a run.sh directory) so a local "
"run-all.sh runs exactly what CI runs.",
file=sys.stderr,
)
sys.exit(1)
print(
f"harness sync OK: run-all.sh, CI, and {len(c)} suite directories "
"name the same suites in the same order:"
)
print(" " + " ".join(a))
if __name__ == "__main__":
main()
+33
View File
@@ -0,0 +1,33 @@
#!/usr/bin/env bash
#
# Harness self-consistency check for the detection test suites (see
# check-harness-sync.py for the assertions). It proves the one thing the eight
# detection suites cannot: that run-all.sh, the CI workflow, and the suite
# directories on disk all name the same suites in the same order, so a suite
# wired into only one of the three cannot silently break the "run-all.sh runs the
# same suites as CI" promise while every per-suite test stays green.
#
# It is a gate, not a suite: run-all.sh runs it before the suite loop and it does
# not appear in the per-suite summary, and it is not itself a SUITES entry, a
# `.../run.sh` CI step, or a run.sh directory — so the eight detection suites stay
# eight and this check never counts itself.
#
# No host dependency beyond Docker: the check is pure Python standard library and
# runs in a container with only the detection tree and the CI workflow mounted
# read-only (never work/ or anything else). Nothing is installed on the host and
# nothing is written to the repo.
#
# Usage: detections/tests/check-harness-sync.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../.." && pwd)"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
docker run --rm \
-e ARTEX_REPO_ROOT=/repo \
-v "$REPO/detections:/repo/detections:ro" \
-v "$REPO/.github:/repo/.github:ro" \
-v "$HERE:/src:ro" \
"$PYTHON_IMAGE" python3 /src/check-harness-sync.py
+378
View File
@@ -0,0 +1,378 @@
#!/usr/bin/env python3
#
# Source-of-truth consistency test for the ARTEX detection indicators. run.sh
# launches this inside a Python container with the detection rules and the
# upstream source packages they pin mounted read-only under /repo. It proves one
# property the other three detection tests do not: that each rule's pinned
# indicator is still the string ARTEX's own source actually emits.
#
# The Sigma test proves an indicator survives rule->query *compilation*; the
# ATT&CK test proves the layer matches the rules' tags; the Suricata test proves
# the network rule *fires*. None of them look back at the source the indicator
# claims to come from. So the realistic rot they miss is an upstream re-sync that
# bumps the prober User-Agent to "artex-enrich/2.0" or rewrites the guard marker:
# every rule still compiles, the layer still matches, the pcap test still fires on
# the synthesized capture — and the deployed rule silently stops matching real
# ARTEX traffic. This test turns detections/README's claim ("every indicator is
# grounded in a string verified in this repository's source, not inferred") and
# CONTRIBUTING's first contribution contract into a guard a reviewer can re-run.
#
# For every indicator it asserts, bidirectionally:
# - source drift: the value is still present in the upstream source file(s)
# that emit it (fails if an upstream re-sync changed the source but not the
# rule -> the rule is now stale);
# - rule drift: the value is still pinned in the rule(s) built on it (fails if a
# rule edit moved the indicator away from the source).
# The enrichment UA also carries a Suricata prefix check, because that rule
# matches the User-Agent by `startswith` and so pins a prefix of the full value.
#
# The destructive-command tokens are handled separately and honestly: they are
# generic hunting leads, not unique ARTEX fingerprints, so the test only asserts
# the correspondence the rule actually claims — each token appears both in ARTEX's
# guard deny-list (db/db.go) and in the hunting rule that mirrors it.
#
# It then validates the published, machine-readable indicator list
# (detections/indicators/artex_indicators.csv): every row's value must still be
# present in the source file(s) it cites and pinned in the rule(s) it cites, and
# every fingerprint this test grounds must appear in the list — so the artifact a
# defender imports cannot silently drift from the source it claims to come from.
#
# Finally it closes the loop on the two gates that fire this test: the CI workflow
# (.github/workflows/detections.yml push/pull_request paths) and the local
# pre-commit hook (.pre-commit-config.yaml files regex). Every upstream source file
# this test reads must be covered by both, or a change touching only a newly pinned
# source (as cmd/artex/main.go once was) would skip the test on one of them: on CI
# the drift sails through the merge gate green, on the hook it is never caught
# locally even though the hook's comment promises "the same source scope as CI".
# The check derives the required set from the indicators it already asserts, so
# pinning a new source without wiring it into *both* gates fails here until they
# stay in sync.
#
# Pure standard library (the slim image already ships python3); nothing is
# installed and nothing is written to the repo. Exits non-zero on any failure.
import csv
import io
import os
import re
import sys
ROOT = os.environ.get("ARTEX_REPO_ROOT", "/repo")
# --- exact ARTEX fingerprints ------------------------------------------------
# Each value is an operational string ARTEX emits; a rule is built on it. If an
# upstream re-sync changes the source string, the rule must change with it.
INDICATORS = [
{
"label": "enrichment prober User-Agent",
"value": "artex-enrich/1.0",
"sources": ["enrich/enrich.go"],
"rules": ["detections/sigma/artex_enrich_user_agent.yml"],
# Suricata matches the UA by `startswith`, so it pins a prefix of the
# full value rather than the whole string. (file, prefix)
"prefix_rules": [("detections/suricata/artex.rules", "artex-enrich/")],
},
{
"label": "self-update egress User-Agent",
"value": "artex-selfupdate",
"sources": ["selfupdate/github.go", "selfupdate/stage.go"],
"rules": ["detections/sigma/artex_selfupdate_egress.yml"],
},
{
"label": "platform-guard audit framing marker",
"value": "【ARTEX 平台管控·非目标防御】",
"sources": ["guard/guard.go"],
"rules": ["detections/sigma/artex_guard_audit_framing.yml"],
},
]
# --- generic destructive-command hunting leads -------------------------------
# NOT unique ARTEX fingerprints. These tokens are shared with ARTEX's own guard
# deny-list (db/db.go); the hunting rule mirrors that list. The test asserts only
# the correspondence the rule claims, so it catches an upstream re-sync that drops
# or renames a deny-list entry the rule says it mirrors.
DENYLIST = {
"source": "db/db.go",
"rule": "detections/sigma/destructive_command_hunting.yml",
"tokens": ["rm -rf", "mkfs", "DROP DATABASE", "FLUSHALL"],
}
fail = 0
def note(s):
print(" " + s)
def ok(s):
note("PASS " + s)
def bad(s):
global fail
note("FAIL " + s)
fail = 1
def read(rel):
"""Return the text of a repo-relative file, or None if it is missing."""
try:
with open(os.path.join(ROOT, rel), encoding="utf-8") as fh:
return fh.read()
except OSError:
return None
def contains(rel, needle):
text = read(rel)
if text is None:
return None # file missing -> distinct from "present but absent"
return needle in text
print("== 1/5 exact fingerprints are still emitted by the upstream source ==")
for ind in INDICATORS:
value, label = ind["value"], ind["label"]
present = [s for s in ind["sources"] if contains(s, value) is True]
missing_files = [s for s in ind["sources"] if contains(s, value) is None]
if present:
ok("%s: %r emitted by %s" % (label, value, ", ".join(present)))
elif missing_files:
bad("%s: source file(s) missing: %s (upstream moved the emitter?)"
% (label, ", ".join(missing_files)))
else:
bad("%s: %r NOT found in any source file %s "
"(upstream drift — update the rule to match)"
% (label, value, ind["sources"]))
print("== 2/5 each rule still pins the indicator it is built on ==")
for ind in INDICATORS:
value, label = ind["value"], ind["label"]
for rule in ind["rules"]:
hit = contains(rule, value)
if hit is True:
ok("%s pins %r" % (rule, value))
elif hit is None:
bad("rule file missing: %s" % rule)
else:
bad("%s no longer pins %r (rule drift from source)" % (rule, value))
for rfile, prefix in ind.get("prefix_rules", []):
if not value.startswith(prefix):
bad("%s: prefix %r is not a prefix of %r (internal inconsistency)"
% (rfile, prefix, value))
continue
hit = contains(rfile, prefix)
if hit is True:
ok("%s pins prefix %r of %r" % (rfile, prefix, value))
elif hit is None:
bad("rule file missing: %s" % rfile)
else:
bad("%s no longer pins prefix %r" % (rfile, prefix))
print("== 3/5 destructive hunting tokens match ARTEX's guard deny-list ==")
src, rule, tokens = DENYLIST["source"], DENYLIST["rule"], DENYLIST["tokens"]
for tok in tokens:
in_src = contains(src, tok)
in_rule = contains(rule, tok)
if in_src is None:
bad("deny-list source missing: %s" % src)
elif in_rule is None:
bad("hunting rule missing: %s" % rule)
elif in_src and in_rule:
ok("%r present in both %s and %s" % (tok, src, rule))
elif not in_src:
bad("%r pinned by the rule but absent from %s "
"(upstream dropped/renamed the deny-list entry)" % (tok, src))
else:
bad("%r in the deny-list but not pinned by %s" % (tok, rule))
print("== 4/5 the published indicator list matches source and rules ==")
CSV_REL = "detections/indicators/artex_indicators.csv"
EXPECTED_HEADER = ["id", "type", "value", "perspective", "source", "rule", "description"]
VALID_PERSPECTIVES = {"target", "forensic"}
csv_rows = []
csv_text = read(CSV_REL)
if csv_text is None:
bad("published indicator list missing: %s" % CSV_REL)
else:
rows = list(csv.reader(io.StringIO(csv_text)))
if not rows:
bad("%s is empty" % CSV_REL)
elif rows[0] != EXPECTED_HEADER:
bad("%s header is %r, expected %r" % (CSV_REL, rows[0], EXPECTED_HEADER))
else:
seen_ids = set()
for lineno, row in enumerate(rows[1:], start=2):
if len(row) != len(EXPECTED_HEADER):
bad("%s line %d: %d fields, expected %d"
% (CSV_REL, lineno, len(row), len(EXPECTED_HEADER)))
continue
rec = dict(zip(EXPECTED_HEADER, row))
csv_rows.append(rec)
rid, value = rec["id"], rec["value"]
if rid in seen_ids:
bad("%s: duplicate id %r" % (CSV_REL, rid))
seen_ids.add(rid)
if not value:
bad("%s: row %r has an empty value" % (CSV_REL, rid))
continue
if rec["perspective"] not in VALID_PERSPECTIVES:
bad("%s: row %r perspective %r not in %s"
% (CSV_REL, rid, rec["perspective"], sorted(VALID_PERSPECTIVES)))
src_files = [s for s in rec["source"].split(";") if s]
if not src_files:
bad("%s: row %r cites no source file" % (CSV_REL, rid))
for s in src_files:
hit = contains(s, value)
if hit is True:
ok("%s: %r grounded in %s" % (rid, value, s))
elif hit is None:
bad("%s: row %r source file missing: %s" % (CSV_REL, rid, s))
else:
bad("%s: row %r value %r not found in source %s (drift)"
% (CSV_REL, rid, value, s))
for r in [r for r in rec["rule"].split(";") if r]:
hit = contains(r, value)
if hit is True:
ok("%s: %r pinned in %s" % (rid, value, r))
elif hit is None:
bad("%s: row %r rule file missing: %s" % (CSV_REL, rid, r))
else:
bad("%s: row %r value %r not pinned in rule %s"
% (CSV_REL, rid, value, r))
published = {rec["value"] for rec in csv_rows}
for ind in INDICATORS:
if ind["value"] in published:
ok("tested fingerprint %r is published in the list" % ind["value"])
else:
bad("tested fingerprint %r is missing from %s" % (ind["value"], CSV_REL))
def paths_for_trigger(text, trigger):
"""Collect the quoted entries of `<trigger>: ... paths: [...]` in the detection
workflow. Returns the set of listed paths, or None if the trigger is absent.
A deliberately small parser for a known-shape file: it locates the trigger key
under `on:`, then the `paths:` list nested in it, and reads the `- "..."` items
until the indentation returns to the list's level."""
lines = text.splitlines()
t_indent = None
start = None
for idx, line in enumerate(lines):
if re.match(r"^\s{2,}%s:\s*$" % re.escape(trigger), line):
t_indent = len(line) - len(line.lstrip())
start = idx + 1
break
if start is None:
return None
items = set()
i = start
while i < len(lines):
line = lines[i]
if line.strip():
indent = len(line) - len(line.lstrip())
if indent <= t_indent:
break # left this trigger block
if re.match(r"^\s*paths:\s*$", line):
p_indent = indent
j = i + 1
while j < len(lines):
pl = lines[j]
if pl.strip():
pind = len(pl) - len(pl.lstrip())
if pind <= p_indent:
break
m = re.match(r"""^\s*-\s*['"]?([^'"\s]+)['"]?\s*$""", pl)
if m:
items.add(m.group(1))
j += 1
return items
i += 1
return items
print("== 5/5 CI and the pre-commit hook both fire this test on any pinned source ==")
# Two gates run this test only when a file they filter on changes: the CI workflow's
# paths filter and the pre-commit hook's files regex. Every upstream source this test
# reads must be covered by both, or a change touching only that source skips the test
# on the gate that misses it — on CI the drift above passes the merge gate green, on
# the hook it is never caught locally. The required set is derived from the indicators
# themselves, so pinning a new source without wiring it into both gates fails here.
# detections/** covers the rules, the CSV, and the tests, so only non-detections
# sources are required explicitly (plus a spot check that each gate still covers the
# detections/ tree at all).
WORKFLOW_REL = ".github/workflows/detections.yml"
PRECOMMIT_REL = ".pre-commit-config.yaml"
DETECTIONS_SAMPLE = "detections/sigma/artex_enrich_user_agent.yml"
needed_sources = set()
for ind in INDICATORS:
needed_sources.update(ind["sources"])
needed_sources.add(DENYLIST["source"])
for rec in csv_rows:
for s in rec["source"].split(";"):
if s:
needed_sources.add(s)
needed_sources = {s for s in needed_sources if not s.startswith("detections/")}
wf_text = read(WORKFLOW_REL)
if wf_text is None:
bad("CI workflow missing: %s" % WORKFLOW_REL)
else:
for trigger in ("push", "pull_request"):
listed = paths_for_trigger(wf_text, trigger)
if listed is None:
bad("%s has no %s: trigger" % (WORKFLOW_REL, trigger))
continue
if "detections/**" not in listed:
bad("%s %s paths is missing 'detections/**' "
"(rule/CSV/test changes would not trigger the detection tests)"
% (WORKFLOW_REL, trigger))
for s in sorted(needed_sources):
if s in listed:
ok("%s %s paths covers %s" % (WORKFLOW_REL, trigger, s))
else:
bad("%s %s paths is missing %s — a PR touching only that source "
"would skip this test and let source drift pass the merge gate"
% (WORKFLOW_REL, trigger, s))
# The local hook gates on a files regex, not a paths list. Its comment promises the
# "same source scope as CI", so the same required set must match that regex. This is
# the sibling drift the CI check above does not see: CI paths can carry a source the
# hook's regex omits (as cmd/artex/main.go once did), leaving the local gate a false
# promise even while the merge gate is sound.
pc_text = read(PRECOMMIT_REL)
if pc_text is None:
bad("pre-commit config missing: %s" % PRECOMMIT_REL)
else:
m = re.search(r"^\s*files:\s*(.+?)\s*$", pc_text, re.M)
if not m:
bad("%s has no files: pattern on the detections hook" % PRECOMMIT_REL)
else:
pattern_src = m.group(1).strip().strip("'\"")
try:
pat = re.compile(pattern_src)
except re.error as exc:
bad("%s files pattern does not compile: %s" % (PRECOMMIT_REL, exc))
pat = None
if pat is not None:
if pat.search(DETECTIONS_SAMPLE):
ok("%s files covers the detections/ tree" % PRECOMMIT_REL)
else:
bad("%s files does not cover detections/ "
"(rule/CSV/test changes would not fire the local hook)"
% PRECOMMIT_REL)
for s in sorted(needed_sources):
if pat.search(s):
ok("%s files covers %s" % (PRECOMMIT_REL, s))
else:
bad("%s files is missing %s — a commit touching only that source "
"would skip the local hook while CI still runs it (the hook's "
"'same source scope as CI' promise is false for this file)"
% (PRECOMMIT_REL, s))
print()
print("reference: %d exact fingerprints, %d deny-list tokens, %d published rows, "
"%d pinned sources checked against CI paths and the pre-commit files regex"
% (len(INDICATORS), len(tokens), len(csv_rows), len(needed_sources)))
print("RESULT: %s" % ("PASS" if fail == 0 else "FAIL"))
sys.exit(fail)
+36
View File
@@ -0,0 +1,36 @@
#!/usr/bin/env bash
#
# Source-of-truth consistency test for the ARTEX detection indicators
# (see check.py for the assertions). It proves the one thing the Sigma, Suricata,
# and ATT&CK tests do not: that each rule's pinned indicator is still the string
# ARTEX's own source actually emits. The realistic rot it catches is an upstream
# re-sync that bumps the prober User-Agent or rewrites the guard marker — every
# other test stays green while the deployed rule silently stops matching.
#
# No host dependency beyond Docker: the check is pure Python standard library and
# runs in a container with only the rule tree, the published indicator list, the
# source packages it pins, and the two configs that gate on them — the CI workflow
# and the pre-commit hook — mounted read-only (never work/ or anything else).
# Nothing is installed on the host and nothing is written to the repo.
#
# Usage: detections/tests/indicators/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
docker run --rm \
-e ARTEX_REPO_ROOT=/repo \
-v "$REPO/detections:/repo/detections:ro" \
-v "$REPO/enrich:/repo/enrich:ro" \
-v "$REPO/selfupdate:/repo/selfupdate:ro" \
-v "$REPO/guard:/repo/guard:ro" \
-v "$REPO/db:/repo/db:ro" \
-v "$REPO/cmd:/repo/cmd:ro" \
-v "$REPO/traffic:/repo/traffic:ro" \
-v "$REPO/.github:/repo/.github:ro" \
-v "$REPO/.pre-commit-config.yaml:/repo/.pre-commit-config.yaml:ro" \
-v "$HERE:/src:ro" \
"$PYTHON_IMAGE" python3 /src/check.py
+288
View File
@@ -0,0 +1,288 @@
#!/usr/bin/env python3
#
# Consistency test for the MISP-format export of the ARTEX detection indicators.
# run.sh launches this inside a Python container with pymisp installed and the
# detection tree plus the CI workflow mounted read-only under /repo. It proves
# two properties the other detection tests do not touch:
#
# 1. the published MISP event
# (detections/indicators/artex_indicators.misp.json) is a *valid MISP
# document* — pymisp parses it and accepts every attribute type/category,
# so a defender can import it into MISP (or export it on to STIX from
# there) without hand-fixing the format; and
# 2. that MISP event stays in sync with the source-of-truth CSV
# (detections/indicators/artex_indicators.csv) row for row — same values,
# the intended MISP type/category for each CSV indicator type, and a
# to_ids / disable_correlation flag that faithfully encodes the CSV's own
# honesty (a row with a detection rule is an actionable indicator; a
# host-forensic row without one is a triage hint, not a blocking IoC).
#
# The indicators source-of-truth test (../indicators/) already proves every CSV
# row is grounded in the upstream source and pinned in its rule; this test does
# not repeat that. It proves only that the MISP serialization a defender
# actually imports cannot silently drift away from that CSV — if a row is added,
# removed, retyped, or has its rule column changed, the MISP event must change
# with it or this test fails.
#
# Exits non-zero on any failed assertion.
import csv
import io
import json
import os
import re
import sys
ROOT = os.environ.get("ARTEX_REPO_ROOT", "/repo")
CSV_REL = "detections/indicators/artex_indicators.csv"
MISP_REL = "detections/indicators/artex_indicators.misp.json"
WORKFLOW_REL = ".github/workflows/detections.yml"
# The indicators README states the CSV `type` values "map onto the equivalent
# MISP/STIX attribute types". This is that mapping, made explicit and enforced:
# CSV indicator type -> (MISP attribute type, MISP attribute category).
TYPE_MAP = {
"http.user-agent": ("user-agent", "Network activity"),
"string": ("pattern-in-file", "Artifacts dropped"),
"port": ("port", "Network activity"),
"ip-dst|port": ("ip-dst|port", "Network activity"),
# a host artifact that fits no network/file slot (e.g. a DB schema object
# name); MISP's generic "other"/"Other" carries it as a triage lead.
"other": ("other", "Other"),
}
fail = 0
def note(s):
print(" " + s)
def ok(s):
note("PASS " + s)
def bad(s):
global fail
note("FAIL " + s)
fail = 1
def read(rel):
try:
with open(os.path.join(ROOT, rel), encoding="utf-8") as fh:
return fh.read()
except OSError:
return None
def misp_value_for(csv_type, csv_value):
"""The MISP value for a CSV row. MISP composite types join their parts with
`|`, so the CSV's `ip:port` becomes `ip|port`; every other type is verbatim."""
if csv_type == "ip-dst|port":
return csv_value.replace(":", "|", 1)
return csv_value
# --- load the source-of-truth CSV -------------------------------------------
EXPECTED_HEADER = ["id", "type", "value", "perspective", "source", "rule", "description"]
csv_rows = []
csv_text = read(CSV_REL)
if csv_text is None:
bad("source CSV missing: %s" % CSV_REL)
else:
rows = list(csv.reader(io.StringIO(csv_text)))
if not rows or rows[0] != EXPECTED_HEADER:
bad("%s header is %r, expected %r"
% (CSV_REL, rows[0] if rows else None, EXPECTED_HEADER))
else:
for row in rows[1:]:
if len(row) == len(EXPECTED_HEADER):
csv_rows.append(dict(zip(EXPECTED_HEADER, row)))
# --- load the MISP event (raw JSON) -----------------------------------------
misp_text = read(MISP_REL)
event = None
attrs = []
if misp_text is None:
bad("MISP event missing: %s" % MISP_REL)
else:
try:
doc = json.loads(misp_text)
except ValueError as exc:
bad("%s is not valid JSON: %s" % (MISP_REL, exc))
doc = None
if isinstance(doc, dict):
event = doc.get("Event")
if not isinstance(event, dict):
bad("%s has no top-level Event object" % MISP_REL)
else:
if not event.get("info"):
bad("%s Event has no info string" % MISP_REL)
if not event.get("uuid"):
bad("%s Event has no uuid" % MISP_REL)
attrs = event.get("Attribute") or []
if not isinstance(attrs, list) or not attrs:
bad("%s Event has no Attribute list" % MISP_REL)
attrs = []
print("== 1/5 the MISP event is a valid MISP document (pymisp parses it) ==")
# pymisp's object model rejects an unknown attribute type on load, so a parse
# here is a real check that every type we use is a genuine MISP type a MISP
# server would accept — not just a plausible-looking string.
if misp_text is None:
bad("cannot validate: MISP event missing")
else:
try:
from pymisp import MISPEvent
me = MISPEvent()
me.load_file(os.path.join(ROOT, MISP_REL))
ok("pymisp %s parsed the event (%d attributes, info=%r)"
% (__import__("pymisp").__version__, len(me.attributes), me.info))
if len(me.attributes) != len(attrs):
bad("pymisp parsed %d attributes but the JSON has %d"
% (len(me.attributes), len(attrs)))
except Exception as exc: # NewAttributeError, validation, import, ...
bad("pymisp rejected the MISP event: %s: %s"
% (type(exc).__name__, exc))
print("== 2/5 every published CSV row maps to one MISP attribute ==")
# value (transformed for composite types) -> list of matching MISP attributes
by_value = {}
for a in attrs:
by_value.setdefault(a.get("value"), []).append(a)
expected_misp_values = set()
for rec in csv_rows:
rid, ctype, cval = rec["id"], rec["type"], rec["value"]
if ctype not in TYPE_MAP:
bad("%s: CSV type %r has no MISP mapping (extend TYPE_MAP)" % (rid, ctype))
continue
want_type, want_cat = TYPE_MAP[ctype]
want_val = misp_value_for(ctype, cval)
expected_misp_values.add(want_val)
matches = by_value.get(want_val, [])
if not matches:
bad("%s: no MISP attribute with value %r (CSV row not exported)"
% (rid, want_val))
continue
if len(matches) > 1:
bad("%s: %d MISP attributes share value %r" % (rid, len(matches), want_val))
a = matches[0]
if a.get("type") == want_type:
ok("%s: %r is a %s" % (rid, want_val, want_type))
else:
bad("%s: value %r is type %r, expected %r"
% (rid, want_val, a.get("type"), want_type))
if a.get("category") != want_cat:
bad("%s: value %r category %r, expected %r"
% (rid, want_val, a.get("category"), want_cat))
# A row with a detection rule is an actionable indicator (to_ids on); a
# host-forensic row without one is a triage hint, not a blocking IoC
# (to_ids off, and correlation disabled so a common port / loopback does
# not pollute MISP correlations). This mirrors the CSV `rule` column.
want_ids = bool(rec["rule"].strip())
if bool(a.get("to_ids")) != want_ids:
bad("%s: to_ids=%r, expected %r (rule column=%r)"
% (rid, a.get("to_ids"), want_ids, rec["rule"]))
if bool(a.get("disable_correlation")) != (not want_ids):
bad("%s: disable_correlation=%r, expected %r"
% (rid, a.get("disable_correlation"), not want_ids))
if not (a.get("comment") or "").strip():
bad("%s: MISP attribute has an empty comment (grounding/caveat lost)" % rid)
print("== 3/5 no MISP attribute is unaccounted for (bijection) ==")
actual_values = [a.get("value") for a in attrs]
if len(actual_values) != len(set(actual_values)):
bad("the MISP event has duplicate attribute values")
extra = set(actual_values) - expected_misp_values
if extra:
bad("MISP attribute(s) with no CSV row: %s" % ", ".join(sorted(map(repr, extra))))
elif csv_rows and not fail:
ok("the %d MISP attributes are exactly the %d published CSV rows"
% (len(attrs), len(csv_rows)))
elif not extra:
ok("every MISP attribute corresponds to a CSV row")
print("== 4/5 the non-ASCII guard marker is preserved verbatim ==")
MARKER = "【ARTEX 平台管控·非目标防御】"
csv_has = any(r["value"] == MARKER for r in csv_rows)
misp_has = MARKER in actual_values
if csv_has and misp_has:
ok("guard audit marker exported byte-for-byte")
elif not csv_has:
bad("guard marker not found in the CSV (test assumption broke)")
else:
bad("guard marker in the CSV but not exported to the MISP event")
def paths_for_trigger(text, trigger):
"""The quoted entries of `<trigger>: ... paths: [...]` in the workflow, or
None if the trigger is absent. Small parser for a known-shape file."""
lines = text.splitlines()
t_indent = None
start = None
for idx, line in enumerate(lines):
if re.match(r"^\s{2,}%s:\s*$" % re.escape(trigger), line):
t_indent = len(line) - len(line.lstrip())
start = idx + 1
break
if start is None:
return None
items = set()
i = start
while i < len(lines):
line = lines[i]
if line.strip():
indent = len(line) - len(line.lstrip())
if indent <= t_indent:
break
if re.match(r"^\s*paths:\s*$", line):
p_indent = indent
j = i + 1
while j < len(lines):
pl = lines[j]
if pl.strip():
pind = len(pl) - len(pl.lstrip())
if pind <= p_indent:
break
m = re.match(r"""^\s*-\s*['"]?([^'"\s]+)['"]?\s*$""", pl)
if m:
items.add(m.group(1))
j += 1
return items
i += 1
return items
print("== 5/5 CI triggers this test when the published indicators change ==")
# The MISP event derives only from the CSV, and both live under detections/**,
# so detections/** in the paths filter is the required and sufficient wiring:
# a change to the CSV or the MISP event triggers the detection workflow, which
# runs this suite and re-checks the two stay in sync. (The upstream Go sources
# the indicators are grounded in are enforced by the indicators suite's own
# CI-paths check, not here.)
wf_text = read(WORKFLOW_REL)
if wf_text is None:
bad("CI workflow missing: %s" % WORKFLOW_REL)
else:
for trigger in ("push", "pull_request"):
listed = paths_for_trigger(wf_text, trigger)
if listed is None:
bad("%s has no %s: trigger" % (WORKFLOW_REL, trigger))
elif "detections/**" in listed:
ok("%s %s paths covers detections/** (CSV + MISP event)" % (WORKFLOW_REL, trigger))
else:
bad("%s %s paths is missing 'detections/**' — a change to the CSV or "
"the MISP event would skip this test" % (WORKFLOW_REL, trigger))
print()
print("reference: %d CSV rows, %d MISP attributes, %d type mappings"
% (len(csv_rows), len(attrs), len(TYPE_MAP)))
print("RESULT: %s" % ("PASS" if fail == 0 else "FAIL"))
sys.exit(fail)
+32
View File
@@ -0,0 +1,32 @@
#!/usr/bin/env bash
#
# Reproducible consistency test for the MISP-format export of the ARTEX
# indicators (see check.py for the assertions). It proves two things no other
# detection test does: that detections/indicators/artex_indicators.misp.json is
# a MISP document pymisp actually parses (every attribute type/category is a real
# MISP type a server would accept), and that it stays row-for-row in sync with
# the source-of-truth CSV it is generated from — same values, the intended MISP
# type/category per indicator, and a to_ids/disable_correlation flag that mirrors
# the CSV's own honesty (rule-backed = actionable; host-forensic = triage hint).
#
# No host dependency beyond Docker: pymisp is pinned and installed inside the
# container, and the detection tree and CI workflow are mounted read-only.
# Nothing is installed on the host and nothing is written to the repo tree.
#
# Usage: detections/tests/misp/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
# PYMISP_VERSION (default 2.5.34.4 — the pinned reference version)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
PYMISP_VERSION="${PYMISP_VERSION:-2.5.34.4}"
docker run --rm \
-e ARTEX_REPO_ROOT=/repo \
-e PYMISP_VERSION="$PYMISP_VERSION" \
-v "$REPO/detections:/repo/detections:ro" \
-v "$REPO/.github:/repo/.github:ro" \
-v "$HERE:/src:ro" \
"$PYTHON_IMAGE" sh -c 'pip install --quiet "pymisp==${PYMISP_VERSION}" && python3 /src/check.py'
+82
View File
@@ -0,0 +1,82 @@
#!/usr/bin/env bash
#
# Runs every detection test suite under this directory in one command — the
# local-developer and pre-commit counterpart to the per-suite CI steps in
# ../../.github/workflows/detections.yml. CONTRIBUTING.md and this directory's
# README.md promise that each suite "drops straight into CI or a pre-commit
# hook"; this is the single entry point that honours that promise for all of
# them at once, so a contributor does not have to invoke the eight run.sh scripts
# by hand (and reviewers do not have to improvise a loop).
#
# It runs the suites in the same order as CI, lets each suite's own output flow
# through, prints a one-line PASS/FAIL summary per suite at the end, and exits
# non-zero if any suite failed — so it is safe to drop into a CI step or a
# pre-commit hook. Every suite runs to completion even if an earlier one fails,
# so one invocation surfaces every regression rather than only the first.
#
# No host dependency beyond Docker: each suite runs its checks in a container and
# writes nothing to the repo tree (see the per-suite run.sh headers). The image
# and version overrides the child scripts honour (PYTHON_IMAGE, SIGMA_CLI_VERSION,
# SIGMAHQ_VALIDATORS_VERSION, SURICATA_IMAGE) are inherited from this process's
# environment, so exporting any of them here applies to every suite at once.
#
# Usage: detections/tests/run-all.sh
set -uo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
# Same order as the steps in .github/workflows/detections.yml.
SUITES="sigma sigma_match sigma_lint sigma_backends suricata attack indicators misp"
fail=0
harness_fail=0
triage_fail=0
results=""
# Before the suites, verify the harness itself is consistent: the SUITES list
# above, the per-suite steps in CI, and the suite directories on disk must all
# name the same suites in the same order. A suite wired into only one of the
# three (say a new CI step with no SUITES entry) passes every per-suite test yet
# silently breaks the "run-all.sh runs the same suites as CI" promise, which no
# other suite can see. This is a gate, not a suite: it runs first and stays out
# of the per-suite summary below, so that summary remains the detection suites.
printf '\n===== harness sync =====\n'
if ! "$HERE/check-harness-sync.sh"; then
harness_fail=1
fail=1
fi
# A second gate, not a suite: the host-triage tool's --self-test. The triage tool
# is a responder helper, not a detection rule, so it stays out of SUITES and the
# harness-sync registry (see triage-selftest.sh). It runs here and in the CI
# workflow so a local run-all.sh covers it too.
printf '\n===== triage self-test =====\n'
if ! "$HERE/triage-selftest.sh"; then
triage_fail=1
fail=1
fi
for suite in $SUITES; do
printf '\n===== %s =====\n' "$suite"
if "$HERE/$suite/run.sh"; then
results="${results} PASS ${suite}"$'\n'
else
rc=$?
results="${results} FAIL ${suite} (exit ${rc})"$'\n'
fail=1
fi
done
printf '\n===== detection suites summary =====\n'
printf '%s' "$results"
if [ "$harness_fail" -ne 0 ]; then
printf ' FAIL harness sync (run-all.sh / CI / directories out of sync: see above)\n'
fi
if [ "$triage_fail" -ne 0 ]; then
printf ' FAIL triage self-test (detections/triage/artex_host_triage.py --self-test: see above)\n'
fi
if [ "$fail" -ne 0 ]; then
printf 'RESULT: FAIL\n'
exit 1
fi
printf 'RESULT: PASS\n'
+87
View File
@@ -0,0 +1,87 @@
#!/bin/sh
#
# In-container half of the ARTEX Sigma rule test. run.sh launches this inside a
# Python container with the Sigma rule tree mounted read-only at /sigma. It
# installs a pinned sigma-cli (pySigma) plus the splunk backend, then asserts
# the properties the rule files and the defense guide claim:
#
# 1. structural + best-practice validation passes (sigma check == 0 errors)
# 2. the whole tree compiles to a backend query language (sigma convert -> splunk)
# 3. each atomic indicator string survives into the query (enrich UA, self-update UA, guard marker, CA file)
# 4. the correlation rules compile as correlations (event_count / value_count aggregations)
# 5. a correlation rule converted ALONE fails (it genuinely depends on its atomic base rule)
#
# POSIX sh (the slim image ships dash). Exits non-zero if any assertion fails.
set -eu
VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
pip install --quiet --disable-pip-version-check "sigma-cli==${VERSION}" >/dev/null 2>&1
sigma plugin install splunk >/dev/null 2>&1
fail=0
note() { printf ' %s\n' "$1"; }
pass() { note "PASS $1"; }
bad() { note "FAIL $1"; fail=1; }
echo "== 1/4 structural + best-practice validation (sigma check) =="
if check_out="$(sigma check /sigma 2>&1)" \
&& printf '%s' "$check_out" | grep -q 'Found 0 errors'; then
pass "sigma check: 0 errors, 0 condition errors, 0 issues"
else
bad "sigma check reported problems"
printf '%s\n' "$check_out" | sed 's/^/ /'
fi
echo "== 2/4 compile the whole tree to a backend (sigma convert -> splunk) =="
if tree_out="$(sigma convert -t splunk --without-pipeline /sigma 2>&1)"; then
pass "whole tree converts to splunk (exit 0)"
else
bad "whole-tree conversion failed"
printf '%s\n' "$tree_out" | sed 's/^/ /'
tree_out=""
fi
echo "== 3/4 each atomic indicator survives into the compiled query =="
# Grep the indicator VALUES, not backend field names or quoting, so the test is
# robust across splunk-backend releases. These strings come straight from the
# rule bodies, which are grounded in this repository's source. The last one is
# the recording-proxy CA filename, grounded in traffic/traffic.go.
for ind in 'artex-enrich/1.0' 'artex-selfupdate' '【ARTEX 平台管控·非目标防御】' 'mitmproxy-ca-cert.pem'; do
if printf '%s' "$tree_out" | grep -qF "$ind"; then
pass "indicator present: $ind"
else
bad "indicator missing from compiled query: $ind"
fi
done
echo "== 4/4 correlation rules compile as correlations, and depend on their base rules =="
# The event_count / value_count aggregation aliases prove the correlation rules
# were compiled as correlations (not dropped), using the whole tree so their
# base-rule references resolve.
if printf '%s' "$tree_out" | grep -q 'event_count' \
&& printf '%s' "$tree_out" | grep -q 'value_count'; then
pass "correlation aggregations present (event_count, value_count)"
else
bad "correlation aggregations missing from compiled query"
fi
# Specificity, mirrored from the Suricata test: converting one correlation rule
# ALONE must fail, because it references an atomic rule by id that is absent from
# a single-file input. A passing conversion here would mean the reference is
# decorative; this asserts it is load-bearing.
if sigma convert -t splunk --without-pipeline \
/sigma/correlation/artex_enrich_scan_velocity.yml >/dev/null 2>&1; then
bad "a correlation rule converted alone (its base-rule reference is not enforced)"
else
pass "correlation rule fails to convert alone — it requires its atomic base rule"
fi
echo
echo "reference: sigma-cli ${VERSION}, splunk backend (latest), pySigma"
if [ "$fail" -eq 0 ]; then
echo "RESULT: PASS"
else
echo "RESULT: FAIL"
fi
exit "$fail"
+42
View File
@@ -0,0 +1,42 @@
#!/usr/bin/env bash
#
# Reproducible regression test for the ARTEX Sigma rules (../../sigma/). It turns
# the "validated by sigma check and sigma convert" claim in the rule README into
# something a reviewer can re-run from source with one command, and it catches
# regressions: a malformed rule, a broken correlation reference, or an indicator
# string that silently dropped out of the compiled query.
#
# It proves five properties with no host dependency beyond Docker (sigma-cli
# runs in a container, nothing is installed on the host and nothing is written to
# the repo tree):
#
# 1. sigma check passes 0 errors / 0 condition errors / 0 issues
# 2. the whole tree compiles sigma convert -> splunk, exit 0
# 3. atomic indicators survive artex-enrich/1.0, artex-selfupdate, guard marker, mitmproxy-ca-cert.pem
# 4. correlations compile event_count / value_count aggregations present
# 5. correlations are load-bearing one correlation rule converted alone FAILS,
# because it references its atomic base rule by id
#
# Unlike a live event-matching harness (which needs a backend that normalizes the
# generic webserver/proxy/application fields — see ../README.md), this is the
# structural + compilation validation the Sigma README documents, made executable.
#
# Usage: detections/tests/sigma/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
# SIGMA_CLI_VERSION (default 3.1.0 — the pinned reference version)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
SIGMA_DIR="$REPO/detections/sigma"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
SIGMA_CLI_VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
# Everything runs inside the container: check.sh installs the pinned sigma-cli and
# the splunk backend, then asserts the five properties and exits non-zero on any
# failure. The rule tree and this directory are mounted read-only.
docker run --rm \
-v "$SIGMA_DIR:/sigma:ro" \
-v "$HERE:/src:ro" \
-e SIGMA_CLI_VERSION="$SIGMA_CLI_VERSION" \
"$PYTHON_IMAGE" sh /src/check.sh
+90
View File
@@ -0,0 +1,90 @@
#!/bin/sh
#
# In-container half of the ARTEX Sigma backend-portability test. run.sh launches
# this inside a Python container with the Sigma rule tree mounted read-only at
# /sigma. It installs a pinned sigma-cli (pySigma) plus four stable backends and
# proves that the rules convert beyond the single Splunk example the README used
# to show, and that the documented per-backend guidance is true for OUR rules.
#
# The base Sigma test (../sigma/) proves the rules are correct against Splunk.
# This test proves they are PORTABLE, and pins the two facts the README's
# "Validate and convert" section now documents:
#
# 1. Correlations are portable the WHOLE tree (atomic + correlation)
# beyond Splunk converts on splunk, Elasticsearch eql,
# and Grafana loki (exit 0), and the enrich
# indicator value survives into each query.
# 2. The atomic-only fallback works backends that do not support Sigma
# where correlations are not correlation conversion (Elasticsearch
# supported lucene, Microsoft kusto) still convert
# the five atomic rules (exit 0), with the
# enrich indicator surviving.
#
# Every assertion is POSITIVE (a capability that must keep working), so the test
# only fails on a genuine regression: a rule that stops converting, or a backend
# that drops support. It deliberately does not assert the negative "backend X
# cannot do correlations" — that would break when a backend improves. The honest
# limitation is documented in ../README.md, reproduced by this test's commands.
#
# POSIX sh (the slim image ships dash). Exits non-zero if any assertion fails.
set -eu
VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
pip install --quiet --disable-pip-version-check "sigma-cli==${VERSION}" >/dev/null 2>&1
# elasticsearch ships the lucene + eql targets; the others are one plugin each.
for plugin in splunk elasticsearch loki kusto; do
sigma plugin install "$plugin" >/dev/null 2>&1
done
ENRICH='artex-enrich/1.0'
ATOMICS='/sigma/artex_enrich_user_agent.yml /sigma/artex_selfupdate_egress.yml /sigma/artex_guard_audit_framing.yml /sigma/artex_recording_proxy_ca.yml /sigma/destructive_command_hunting.yml'
fail=0
note() { printf ' %s\n' "$1"; }
pass() { note "PASS $1"; }
bad() { note "FAIL $1"; fail=1; }
# Backends escape regex metacharacters differently (lucene: artex\-enrich\/1.0,
# loki: artex\-enrich/1\.0, splunk/eql/kusto: artex-enrich/1.0). Strip backslashes
# before matching so the indicator-survival check is robust across all of them
# without asserting any one backend's escaping syntax.
has_enrich() { printf '%s' "$1" | tr -d '\\' | grep -qF "$ENRICH"; }
echo "== 1/2 correlations are portable: the whole tree converts beyond Splunk =="
# Whole-tree conversion includes the four correlation rules, which reference
# their atomic base rules by id. If a backend compiles the whole tree at exit 0
# it supports Sigma correlation conversion for our rules.
for target in splunk eql loki; do
if out="$(sigma convert -t "$target" --without-pipeline /sigma 2>&1)" \
&& has_enrich "$out"; then
pass "whole tree (atomic + correlation) converts on '$target', enrich indicator survives"
else
bad "whole-tree conversion on '$target' failed or dropped the enrich indicator"
printf '%s' "$out" | grep -iE 'error|not supported' | head -2 | sed 's/^/ /'
fi
done
echo "== 2/2 atomic-only fallback: the five atomic rules convert where correlations are not supported =="
# Lucene and kusto (the Microsoft Sentinel / Defender backend) do not convert
# Sigma correlations at the pinned versions, so a defender deploys the five
# atomic rules and expresses the correlation logic natively. That fallback must
# work: all five atomic rules convert and the enrich indicator survives.
for target in lucene kusto; do
if out="$(sigma convert -t "$target" --without-pipeline $ATOMICS 2>&1)" \
&& has_enrich "$out"; then
pass "five atomic rules convert on '$target', enrich indicator survives"
else
bad "atomic-only conversion on '$target' failed or dropped the enrich indicator"
printf '%s' "$out" | grep -iE 'error|not supported' | head -2 | sed 's/^/ /'
fi
done
echo
echo "reference: sigma-cli ${VERSION}; backends splunk, elasticsearch (lucene/eql), loki, kusto (latest compatible), pySigma"
if [ "$fail" -eq 0 ]; then
echo "RESULT: PASS"
else
echo "RESULT: FAIL"
fi
exit "$fail"
+40
View File
@@ -0,0 +1,40 @@
#!/usr/bin/env bash
#
# Reproducible backend-portability test for the ARTEX Sigma rules (../../sigma/).
# The rule README claims the rules "convert to your own SIEM or EDR query
# language" and lists several supported targets. The base Sigma test (../sigma/)
# only exercises Splunk; this test turns the cross-backend claim into something a
# reviewer can re-run, and keeps the README's per-backend guidance honest.
#
# It proves two properties with no host dependency beyond Docker (sigma-cli and
# its backends run in a container, nothing is installed on the host and nothing
# is written to the repo tree):
#
# 1. correlations are portable the whole tree converts on splunk, the
# Elasticsearch eql target, and Grafana loki
# 2. the atomic-only fallback the five atomic rules convert on lucene and
# works kusto (Microsoft Sentinel / Defender), which
# do not support Sigma correlation conversion
#
# See ../README.md "Sigma backend portability" for the measured support matrix
# and the exact per-backend commands this test reproduces.
#
# Usage: detections/tests/sigma_backends/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
# SIGMA_CLI_VERSION (default 3.1.0 — the pinned reference version)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
SIGMA_DIR="$REPO/detections/sigma"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
SIGMA_CLI_VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
# Everything runs inside the container: check.sh installs the pinned sigma-cli and
# four backends, then asserts the two properties and exits non-zero on any
# failure. The rule tree and this directory are mounted read-only.
docker run --rm \
-v "$SIGMA_DIR:/sigma:ro" \
-v "$HERE:/src:ro" \
-e SIGMA_CLI_VERSION="$SIGMA_CLI_VERSION" \
"$PYTHON_IMAGE" sh /src/check.sh
+85
View File
@@ -0,0 +1,85 @@
#!/bin/sh
#
# In-container half of the ARTEX Sigma SigmaHQ-convention lint test. run.sh
# launches this inside a Python container with the Sigma rule tree mounted
# read-only at /sigma and this directory at /src. It installs a pinned sigma-cli
# plus the pinned SigmaHQ validator plugin, then asserts two properties:
#
# 1. baseline is clean sigma check with the documented validators.yml
# baseline reports 0 errors and 0 issues.
# 2. the full set is live running ALL SigmaHQ validators (no exclusions) still
# reports issues, and every issue type is one of the
# four documented, excluded categories — nothing else.
#
# Property 2 is the anti-vacuity guard. If the validator plugin failed to load,
# the "all" run would report zero issues and property 1 would pass vacuously;
# requiring the known exclusions to appear proves the full SigmaHQ set actually
# ran. It also fails the build the moment a rule picks up a NEW convention issue
# outside the documented baseline (e.g. a mis-cased title or an invalid field),
# because that issue type would not be in the allow-list below and property 1
# would stop being clean.
#
# POSIX sh (the slim image ships dash). Exits non-zero if any assertion fails.
set -eu
VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
SIGMAHQ_VALIDATORS_VERSION="${SIGMAHQ_VALIDATORS_VERSION:-0.21.0}"
pip install --quiet --disable-pip-version-check \
"sigma-cli==${VERSION}" "pySigma-validators-sigmahq==${SIGMAHQ_VALIDATORS_VERSION}" >/dev/null 2>&1
# The four issue types the documented baseline (validators.yml) intentionally
# excludes. Any issue outside this set must fail the build.
ALLOWED='SigmahqGithubLinkIssue SigmahqFilenamePrefixIssue SigmahqCorrelationFilenamePrefixIssue SigmahqLogsourceUnknownIssue'
printf 'validators:\n - all\n' > /tmp/all.yml
fail=0
note() { printf ' %s\n' "$1"; }
pass() { note "PASS $1"; }
bad() { note "FAIL $1"; fail=1; }
echo "== 1/2 documented SigmaHQ baseline is clean (validators.yml) =="
if base_out="$(sigma check --validation-config /src/validators.yml /sigma 2>&1)" \
&& printf '%s' "$base_out" | grep -q 'Found 0 errors, 0 condition errors and 0 issues'; then
pass "sigma check with the documented baseline: 0 errors, 0 issues"
else
bad "the documented baseline reported problems (a non-excluded convention issue, or an error)"
printf '%s\n' "$base_out" | sed 's/^/ /'
fi
echo "== 2/2 the full SigmaHQ validator set runs, and only the documented exclusions remain =="
all_out="$(sigma check --validation-config /tmp/all.yml /sigma 2>&1 || true)"
# Collect the distinct issue types the full set reports.
types="$(printf '%s' "$all_out" | grep -oE 'issue=Sigmahq[A-Za-z]+Issue' | sed 's/^issue=//' | sort -u)"
if [ -z "$types" ]; then
bad "the full validator set reported no SigmaHQ issues at all — the plugin did not load (vacuous)"
else
# Anti-vacuity: the two load-bearing exclusions must actually appear.
for must in SigmahqGithubLinkIssue SigmahqLogsourceUnknownIssue; do
if printf '%s\n' "$types" | grep -qx "$must"; then
pass "full set is live: $must present"
else
bad "expected $must from the full validator set but it was absent — plugin/version drift"
fi
done
# No issue type outside the documented allow-list may appear.
unexpected=0
for t in $types; do
case " $ALLOWED " in
*" $t "*) : ;;
*) bad "undocumented convention issue from the full set: $t"; unexpected=1 ;;
esac
done
[ "$unexpected" -eq 0 ] && pass "every reported issue is one of the four documented exclusions"
fi
echo
echo "reference: sigma-cli ${VERSION}, pySigma-validators-sigmahq ${SIGMAHQ_VALIDATORS_VERSION}"
if [ "$fail" -eq 0 ]; then
echo "RESULT: PASS"
else
echo "RESULT: FAIL"
fi
exit "$fail"
+41
View File
@@ -0,0 +1,41 @@
#!/usr/bin/env bash
#
# Reproducible SigmaHQ-convention lint for the ARTEX Sigma rules (../../sigma/).
# `sigma check` on its own runs only pySigma's core validators; this test runs
# the full SigmaHQ convention set (the pySigma-validators-sigmahq plugin) against
# the documented baseline in validators.yml, so the "passes sigma check cleanly"
# claim in the README and CONTRIBUTING covers SigmaHQ's conventions, not just the
# core checks.
#
# It proves two properties with no host dependency beyond Docker (everything runs
# in a container, nothing is installed on the host and nothing is written to the
# repo tree):
#
# 1. the documented baseline (validators.yml) reports 0 errors and 0 issues
# 2. the full validator set actually runs, and only the four documented
# exclusions remain — the anti-vacuity guard (see check.sh)
#
# The four exclusions and the rationale for each live in validators.yml.
#
# Usage: detections/tests/sigma_lint/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
# SIGMA_CLI_VERSION (default 3.1.0 — the pinned reference version)
# SIGMAHQ_VALIDATORS_VERSION (default 0.21.0 — the pinned validator plugin)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
SIGMA_DIR="$REPO/detections/sigma"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
SIGMA_CLI_VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
SIGMAHQ_VALIDATORS_VERSION="${SIGMAHQ_VALIDATORS_VERSION:-0.21.0}"
# Everything runs inside the container: check.sh installs the pinned sigma-cli and
# SigmaHQ validator plugin, then asserts the two properties and exits non-zero on
# any failure. The rule tree and this directory are mounted read-only.
docker run --rm \
-v "$SIGMA_DIR:/sigma:ro" \
-v "$HERE:/src:ro" \
-e SIGMA_CLI_VERSION="$SIGMA_CLI_VERSION" \
-e SIGMAHQ_VALIDATORS_VERSION="$SIGMAHQ_VALIDATORS_VERSION" \
"$PYTHON_IMAGE" sh /src/check.sh
@@ -0,0 +1,44 @@
# SigmaHQ validator baseline for the ARTEX detection rules.
#
# `sigma check` on its own runs only pySigma's core validators. This config turns
# on the full SigmaHQ convention set (the pySigma-validators-sigmahq plugin) and
# then disables four checks that encode SigmaHQ *monorepo* conventions which do
# not apply to this small, self-contained rule set. Every other SigmaHQ check is
# enforced, and detections/tests/sigma_lint/ fails the build if any enabled check
# reports an issue. Each exclusion below is a deliberate, documented decision, not
# a silenced defect.
#
# Run:
# pip install pySigma-validators-sigmahq
# sigma check --validation-config detections/tests/sigma_lint/validators.yml detections/sigma/
#
validators:
- all
# sigmahq_github_link wants every `references:` URL to be a commit permalink
# rather than a branch link. That check exists so rules citing external,
# third-party write-ups keep pointing at the exact revision they were written
# against. Our references point at *our own* living defense docs
# (docs/defense-ko.md, docs/defense-en.md) on `main`: we want them to track the
# current guide, not freeze to a snapshot that goes stale as the guide improves.
- -sigmahq_github_link
# sigmahq_filename_prefix and sigmahq_correlation_filename_prefix require
# logsource-prefixed filenames (web_*, proxy_*) and a correlation_* prefix, the
# filing scheme of SigmaHQ's single flat rules/ tree. This repository ships a
# small set under detections/sigma/ with descriptive artex_* names and a
# correlation/ subdirectory, referenced by the correlation rules' header
# comments, the reproduction tests, and the README index. Renaming to the
# monorepo prefixes would desynchronize those references for no gain on a
# standalone set.
- -sigmahq_filename_prefix
- -sigmahq_correlation_filename_prefix
# sigmahq_logsource_unknown flags `category: application` (the guard-marker
# forensic log search) and a product-less `category: process_creation` (the
# cross-platform destructive-command hunting lead) as outside the SigmaHQ
# taxonomy. Both logsources are intentionally generic: these indicators appear
# across heterogeneous application/audit and process-creation logs, and the
# README tells defenders to map them to their own pipeline. Pinning a single
# product would narrow the rules incorrectly.
- -sigmahq_logsource_unknown
+416
View File
@@ -0,0 +1,416 @@
#!/usr/bin/env python3
#
# Live event-matching test for the ARTEX Sigma rules (../../sigma/). check.sh
# installs a pinned pySigma inside a container and runs this script with the rule
# tree mounted read-only at /sigma and this directory at /src.
#
# WHAT THIS PROVES, AND WHY IT IS DIFFERENT FROM THE sigma/ SUITE
# --------------------------------------------------------------
# The sigma/ suite proves each rule is structurally valid and COMPILES to a
# backend query, and that its indicator strings survive into that query. It does
# NOT prove the rule actually fires on a matching event, or stays quiet on a
# benign one: a field renamed to something the log never carries, a wildcard that
# silently dropped, or an over-broad token would all still compile cleanly. The
# README's own principle is that "a detection you cannot run is only a claim," and
# the Suricata suite already backs its network rule with a real pcap replay
# (fires on the probe UA, silent on a benign browser). This suite closes the same
# gap for the host/log-layer Sigma rules on two levels:
# - ATOMIC rules (../../sigma/*.yml): for each rule a representative malicious
# event MATCHES and a benign event DOES NOT.
# - CORRELATION rules (../../sigma/correlation/*.yml): for each rule a positive
# timeline (threshold met, inside the window, within one group) FIRES and
# negative timelines (below threshold, threshold met but spread beyond the
# window, split across groups, or missing a leg) stay QUIET.
#
# HOW IT MATCHES (trust model)
# ----------------------------
# It does not hand-parse the YAML or re-implement Sigma's modifier logic. pySigma
# parses each rule and compiles its modifiers and condition into a tree:
# `|contains` becomes a wildcard-wrapped value, `|all` becomes an AND over values,
# `1 of selection_*` becomes an OR over the selection groups. This script only
# walks that compiled tree (AND / OR / NOT / field-equals / keyword) and tests
# each leaf against the event, so the authoritative parsing stays in pySigma. A
# leaf value or condition node this script does not explicitly support raises
# rather than passing silently (fail-closed), so a future rule using an
# unsupported construct surfaces loudly here instead of being waved through.
#
# For a correlation rule, pySigma likewise parses the aggregation spec — type
# (event_count / value_count / temporal), group-by fields, timespan, the
# threshold condition, and the resolved references to the atomic base rules. This
# script walks that parsed spec and applies it to a timeline, deciding which
# events feed each referenced rule with the very same atomic matcher above, so the
# Sigma logic again stays in pySigma; only the windowed aggregation is applied
# here. The correlation rules reference their atomics by id, so each is parsed in a
# collection that also holds every atomic rule (pySigma resolves the reference).
#
# SCOPE AND HONESTY (read before trusting a green run)
# ----------------------------------------------------
# - The CORRELATION window is the standard sliding-window interpretation: a
# window of `timespan` seconds anchored at each matching event, with inclusive
# bounds. Each timeline event carries an integer `ts` in relative seconds. A
# real SIEM's windowing (tumbling vs sliding, bound inclusivity, late arrival)
# may differ; this is a regression test for the rule's group-by / timespan /
# threshold logic — that it fires when they are satisfied and not when they are
# not — rather than a bit-exact model of any one backend's correlation engine.
# - Matching is CASE-INSENSITIVE. This mirrors the default of the splunk backend
# the sigma/ suite targets, and the destructive rule's own false-positive note
# assumes it (it warns that lowercase coreutils `truncate` shares the uppercase
# `TRUNCATE ` token and must be allow-listed). Your SIEM's case handling and
# field normalisation may differ; this is a regression test for the rules'
# field/value/condition logic, not a substitute for validating in your stack.
# - Keyword matching (the audit-framing rule) is modelled as a full-text
# substring search across all event field values, the common interpretation of
# an unbound Sigma keyword.
#
# Exits non-zero on any failure. Standard library only beyond pySigma.
import glob
import json
import os
import re
import sys
from sigma.collection import SigmaCollection
from sigma.conditions import (
ConditionAND,
ConditionFieldEqualsValueExpression,
ConditionNOT,
ConditionOR,
ConditionValueExpression,
)
from sigma.types import (
SigmaNull,
SigmaNumber,
SigmaRegularExpression,
SigmaString,
SpecialChars,
)
SIGMA_DIR = os.environ.get("SIGMA_DIR", "/sigma")
EVENTS_DIR = os.environ.get("EVENTS_DIR", "/src/events")
CORR_DIR = os.path.join(SIGMA_DIR, "correlation")
CORR_EVENTS_DIR = os.path.join(EVENTS_DIR, "correlation")
fail = 0
def note(msg):
print(f" {msg}")
def passed(msg):
note(f"PASS {msg}")
def bad(msg):
global fail
note(f"FAIL {msg}")
fail = 1
# --- matcher -------------------------------------------------------------------
def sigmastring_to_regex(value):
"""Compile a pySigma SigmaString (literal text plus wildcards) to an anchored,
case-insensitive regex. `|contains` already wrapped the value in multi
wildcards upstream, so a plain string compiles to an exact match and a
contains-value compiles to a substring match — exactly the Sigma semantics."""
parts = []
for part in value.s:
if part == SpecialChars.WILDCARD_MULTI:
parts.append(".*")
elif part == SpecialChars.WILDCARD_SINGLE:
parts.append(".")
elif isinstance(part, str):
parts.append(re.escape(part))
else:
raise ValueError(f"unsupported SigmaString part: {part!r}")
return re.compile("^" + "".join(parts) + "$", re.DOTALL | re.IGNORECASE)
def field_match(field, value, event):
if field not in event:
return False
observed = str(event[field])
if isinstance(value, SigmaString):
return sigmastring_to_regex(value).search(observed) is not None
if isinstance(value, SigmaNumber):
return observed == str(value.number)
if isinstance(value, SigmaNull):
return event.get(field) is None
if isinstance(value, SigmaRegularExpression):
return re.search(value.regexp, observed) is not None
raise ValueError(f"unsupported field value type: {type(value).__name__}")
def keyword_match(value, event):
"""Unbound keyword: full-text substring search across all field values."""
if not isinstance(value, SigmaString):
raise ValueError("unsupported keyword value type")
if any(not isinstance(p, str) for p in value.s):
raise ValueError("wildcard in keyword is not supported by this matcher")
token = "".join(value.s)
haystack = " ".join(str(v) for v in event.values())
return token.lower() in haystack.lower()
def evaluate(node, event):
if isinstance(node, ConditionAND):
return all(evaluate(a, event) for a in node.args)
if isinstance(node, ConditionOR):
return any(evaluate(a, event) for a in node.args)
if isinstance(node, ConditionNOT):
return not evaluate(node.args[0], event)
if isinstance(node, ConditionFieldEqualsValueExpression):
return field_match(node.field, node.value, event)
if isinstance(node, ConditionValueExpression):
return keyword_match(node.value, event)
raise ValueError(f"unsupported condition node: {type(node).__name__}")
def rule_matches(rule, event):
return any(evaluate(c.parsed, event) for c in rule.detection.parsed_condition)
# --- correlation evaluator -----------------------------------------------------
def group_key(event, fields):
if any(f not in event for f in fields):
return None
return tuple(event[f] for f in fields)
def correlation_fires(corr, timeline):
"""Apply a parsed SigmaCorrelationRule's aggregation to a timeline of events
(each carrying an integer `ts` in seconds). pySigma has parsed the rule into a
type, group-by fields, a timespan, a threshold condition, and resolved rule
references; this walks that parsed structure. Membership in a referenced rule
is decided by the same rule_matches the atomic suite uses, so the Sigma
detection logic stays in pySigma. The window is the standard sliding window:
`timespan` seconds anchored at each matching event, inclusive bounds."""
ctype = str(corr.type)
span = corr.timespan.seconds
group_by = corr.group_by or []
refs = [ref.rule for ref in corr.rules]
if ctype in ("event_count", "value_count"):
# A count correlation may reference several base rules; an event feeds the
# count if it matches ANY of them — the same union the temporal branch
# applies below. Looking at refs[0] alone would silently drop events
# matching the other referenced rules, a fail-open this suite's header
# forbids. With a single reference this reduces to the one-rule case, so
# the existing rules (each referencing one base rule) are unchanged.
matched = [e for e in timeline if any(rule_matches(r, e) for r in refs)]
groups = {}
for e in matched:
key = group_key(e, group_by)
if key is None:
continue
groups.setdefault(key, []).append(e)
threshold = corr.condition.count
fieldref = corr.condition.fieldref
for members in groups.values():
members = sorted(members, key=lambda e: e["ts"])
for anchor in members:
window = [
e for e in members if anchor["ts"] <= e["ts"] <= anchor["ts"] + span
]
if ctype == "event_count":
if len(window) >= threshold:
return True
else:
distinct = {e[fieldref] for e in window if fieldref in e}
if len(distinct) >= threshold:
return True
return False
if ctype == "temporal":
groups = {}
for e in timeline:
key = group_key(e, group_by)
if key is None:
continue
groups.setdefault(key, []).append(e)
for members in groups.values():
members = sorted(members, key=lambda e: e["ts"])
for anchor in members:
window = [
e for e in members if anchor["ts"] <= e["ts"] <= anchor["ts"] + span
]
if all(any(rule_matches(r, e) for e in window) for r in refs):
return True
return False
raise ValueError(f"unsupported correlation type: {ctype}")
def require_ts(events, stem, label):
for e in events:
if not isinstance(e.get("ts"), int):
raise ValueError(
f"{stem} ({label}): every timeline event needs an integer 'ts' "
f"(seconds); got {e!r}"
)
# --- loaders -------------------------------------------------------------------
def load_atomic_rules():
rules = {}
for path in sorted(glob.glob(os.path.join(SIGMA_DIR, "*.yml"))):
stem = os.path.splitext(os.path.basename(path))[0]
collection = SigmaCollection.from_yaml(open(path, encoding="utf-8").read())
for rule in collection.rules:
# Only plain atomic rules; correlation rules carry a `.type` and are
# handled separately below.
if type(rule).__name__ != "SigmaRule":
continue
rules[stem] = rule
return rules
def load_correlation_rules():
"""A correlation rule references its atomic base rules by id, so it must be
parsed in a collection that also contains those atomics. For each correlation
file, merge every atomic YAML with that one correlation YAML, parse the
collection (pySigma resolves the reference), and key the resulting
SigmaCorrelationRule by filename stem so it pairs with
events/correlation/<stem>.json."""
atomic_docs = [
open(p, encoding="utf-8").read()
for p in sorted(glob.glob(os.path.join(SIGMA_DIR, "*.yml")))
]
corrs = {}
for path in sorted(glob.glob(os.path.join(CORR_DIR, "*.yml"))):
stem = os.path.splitext(os.path.basename(path))[0]
merged = "\n---\n".join(atomic_docs + [open(path, encoding="utf-8").read()])
collection = SigmaCollection.from_yaml(merged)
found = [r for r in collection.rules if type(r).__name__ == "SigmaCorrelationRule"]
if len(found) != 1:
raise ValueError(f"{stem}: expected exactly 1 correlation rule, got {len(found)}")
corrs[stem] = found[0]
return corrs
def load_events(directory):
events = {}
for path in sorted(glob.glob(os.path.join(directory, "*.json"))):
stem = os.path.splitext(os.path.basename(path))[0]
events[stem] = json.load(open(path, encoding="utf-8"))
return events
def main():
rules = load_atomic_rules()
events = load_events(EVENTS_DIR)
print("== atomic 1/3 every atomic rule is paired with a sample-event file ==")
rule_stems = set(rules)
event_stems = set(events)
orphan_rules = sorted(rule_stems - event_stems)
orphan_events = sorted(event_stems - rule_stems)
if orphan_rules:
bad(f"atomic rules with no events/<name>.json: {orphan_rules}")
if orphan_events:
bad(f"event files with no matching atomic rule: {orphan_events}")
if not orphan_rules and not orphan_events:
passed(
f"rule/sample pairing: {len(rules)} atomic rules, "
f"{len(events)} event files, no orphans"
)
print("== atomic 2/3 each rule matches its malicious sample events (true positives) ==")
for stem in sorted(rule_stems & event_stems):
rule = rules[stem]
positives = events[stem].get("positive", [])
if not positives:
bad(f"{stem}: no positive sample events")
continue
missed = [e for e in positives if not rule_matches(rule, e)]
if missed:
bad(f"{stem}: {len(missed)}/{len(positives)} positive events did NOT match")
for e in missed:
note(f" unmatched: {json.dumps(e, ensure_ascii=False)}")
else:
passed(f"{stem}: {len(positives)}/{len(positives)} positive events matched")
print("== atomic 3/3 each rule rejects its benign sample events (true negatives) ==")
for stem in sorted(rule_stems & event_stems):
rule = rules[stem]
negatives = events[stem].get("negative", [])
if not negatives:
bad(f"{stem}: no negative sample events")
continue
fired = [e for e in negatives if rule_matches(rule, e)]
if fired:
bad(f"{stem}: {len(fired)}/{len(negatives)} benign events WRONGLY matched")
for e in fired:
note(f" wrongly matched: {json.dumps(e, ensure_ascii=False)}")
else:
passed(
f"{stem}: {len(negatives)}/{len(negatives)} benign events correctly "
"not matched"
)
corr_rules = load_correlation_rules()
corr_events = load_events(CORR_EVENTS_DIR)
print("== correlation 1/3 every correlation rule is paired with a timeline file ==")
corr_stems = set(corr_rules)
ce_stems = set(corr_events)
orphan_corr = sorted(corr_stems - ce_stems)
orphan_tl = sorted(ce_stems - corr_stems)
if orphan_corr:
bad(f"correlation rules with no events/correlation/<name>.json: {orphan_corr}")
if orphan_tl:
bad(f"timeline files with no matching correlation rule: {orphan_tl}")
if not orphan_corr and not orphan_tl:
passed(
f"rule/timeline pairing: {len(corr_rules)} correlation rules, "
f"{len(corr_events)} timeline files, no orphans"
)
print("== correlation 2/3 each rule fires on its positive timelines (true positives) ==")
for stem in sorted(corr_stems & ce_stems):
corr = corr_rules[stem]
positives = corr_events[stem].get("positive", [])
if not positives:
bad(f"{stem}: no positive timelines")
continue
for tl in positives:
require_ts(tl["events"], stem, tl["label"])
if correlation_fires(corr, tl["events"]):
passed(f"{stem}: fired — {tl['label']}")
else:
bad(f"{stem}: did NOT fire on a positive timeline — {tl['label']}")
print("== correlation 3/3 each rule stays quiet on its negative timelines (true negatives) ==")
for stem in sorted(corr_stems & ce_stems):
corr = corr_rules[stem]
negatives = corr_events[stem].get("negative", [])
if not negatives:
bad(f"{stem}: no negative timelines")
continue
for tl in negatives:
require_ts(tl["events"], stem, tl["label"])
if correlation_fires(corr, tl["events"]):
bad(f"{stem}: WRONGLY fired on a benign timeline — {tl['label']}")
else:
passed(f"{stem}: quiet — {tl['label']}")
print()
try:
import importlib.metadata as md
print(f"reference: pySigma {md.version('pysigma')}, atomic + correlation rules")
except Exception:
pass
print("RESULT: PASS" if fail == 0 else "RESULT: FAIL")
sys.exit(fail)
if __name__ == "__main__":
main()
+20
View File
@@ -0,0 +1,20 @@
#!/bin/sh
#
# In-container half of the ARTEX Sigma live event-matching test. run.sh launches
# this inside a Python container with the Sigma rule tree (atomic rules and the
# correlation/ subtree) mounted read-only at /sigma and this directory at /src. It
# installs a pinned pySigma, then hands off to check.py, which asserts that every
# atomic rule matches its malicious sample events and stays quiet on its benign
# ones, and that every correlation rule fires on its positive timeline and stays
# quiet on its negative ones (see check.py's header for the trust model and
# scope). pySigma does the parsing; check.py walks the compiled condition tree and
# aggregation spec and tests each sample event or timeline against it.
#
# POSIX sh (the slim image ships dash). Exits non-zero if any assertion fails.
set -eu
VERSION="${PYSIGMA_VERSION:-2.0.0}"
pip install --quiet --disable-pip-version-check "pysigma==${VERSION}" >/dev/null 2>&1
exec python3 /src/check.py
@@ -0,0 +1,10 @@
{
"note": "webserver access log. The rule matches cs-user-agent EXACTLY equal to 'artex-enrich/1.0' (enrich/enrich.go:233). A browser UA, and the same UA with a trailing suffix, must not match.",
"positive": [
{"cs-method": "GET", "cs-uri-stem": "/", "cs-user-agent": "artex-enrich/1.0", "c-ip": "203.0.113.7"}
],
"negative": [
{"cs-method": "GET", "cs-uri-stem": "/", "cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0", "c-ip": "203.0.113.8"},
{"cs-method": "GET", "cs-uri-stem": "/robots.txt", "cs-user-agent": "artex-enrich/1.0 (proxied)", "c-ip": "203.0.113.9"}
]
}
@@ -0,0 +1,9 @@
{
"note": "application log line. The rule is an unbound keyword matching the guard's audit-control framing marker (guard/guard.go). It should match wherever the marker appears in the message, and stay quiet on an ordinary log line.",
"positive": [
{"message": "2026-10-07T03:11:09Z guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"}
],
"negative": [
{"message": "2026-10-07T03:11:09Z auth: user login ok uid=42 ip=203.0.113.8"}
]
}
@@ -0,0 +1,10 @@
{
"note": "file-creation (file_event) telemetry. The rule needs TargetFilename to contain BOTH '_ca' AND 'mitmproxy-ca-cert.pem' (|all), which narrows it to ARTEX's '<dir>/_ca/mitmproxy-ca-cert.pem' layout (traffic/traffic.go). A standalone mitmproxy cert under .mitmproxy/ has the filename but not the '_ca' directory, so it must NOT match — that is the specificity the |all modifier buys.",
"positive": [
{"TargetFilename": "/home/ubuntu/.local/share/artex/data/_ca/mitmproxy-ca-cert.pem", "Image": "/opt/artex/artex"}
],
"negative": [
{"TargetFilename": "/home/ubuntu/.mitmproxy/mitmproxy-ca-cert.pem", "Image": "/usr/bin/mitmproxy"},
{"TargetFilename": "/etc/ssl/certs/ca-certificates.crt", "Image": "/usr/sbin/update-ca-certificates"}
]
}
@@ -0,0 +1,9 @@
{
"note": "forward-proxy egress log. The rule matches c-useragent EXACTLY equal to 'artex-selfupdate' (selfupdate/github.go), the UA ARTEX sets when it fetches its own release from GitHub. A generic client UA must not match.",
"positive": [
{"c-useragent": "artex-selfupdate", "cs-host": "github.com", "cs-uri-stem": "/Autumn-27/ARTEX/releases/latest"}
],
"negative": [
{"c-useragent": "curl/8.5.0", "cs-host": "github.com", "cs-uri-stem": "/"}
]
}
@@ -0,0 +1,864 @@
{
"note": "webserver access log timeline for the ARTEX Enrichment Fan-Out correlation (value_count of DISTINCT cs-host >= 20, grouped by c-ip, within a 10-minute window). 'ts' is relative seconds. Breadth — distinct hosts touched, not request volume — is the signal, so a high-volume/low-breadth burst must stay quiet.",
"positive": [
{
"label": "20 distinct hosts from one source within the 10-minute window",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 25,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 50,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 75,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.9",
"cs-host": "host04.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 125,
"c-ip": "10.0.0.9",
"cs-host": "host05.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 150,
"c-ip": "10.0.0.9",
"cs-host": "host06.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 175,
"c-ip": "10.0.0.9",
"cs-host": "host07.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "host08.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "host09.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 250,
"c-ip": "10.0.0.9",
"cs-host": "host10.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 275,
"c-ip": "10.0.0.9",
"cs-host": "host11.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 300,
"c-ip": "10.0.0.9",
"cs-host": "host12.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 325,
"c-ip": "10.0.0.9",
"cs-host": "host13.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 350,
"c-ip": "10.0.0.9",
"cs-host": "host14.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 375,
"c-ip": "10.0.0.9",
"cs-host": "host15.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 400,
"c-ip": "10.0.0.9",
"cs-host": "host16.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 425,
"c-ip": "10.0.0.9",
"cs-host": "host17.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 450,
"c-ip": "10.0.0.9",
"cs-host": "host18.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 475,
"c-ip": "10.0.0.9",
"cs-host": "host19.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
}
],
"negative": [
{
"label": "below the breadth threshold: only 19 distinct hosts",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 25,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 50,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 75,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.9",
"cs-host": "host04.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 125,
"c-ip": "10.0.0.9",
"cs-host": "host05.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 150,
"c-ip": "10.0.0.9",
"cs-host": "host06.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 175,
"c-ip": "10.0.0.9",
"cs-host": "host07.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "host08.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "host09.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 250,
"c-ip": "10.0.0.9",
"cs-host": "host10.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 275,
"c-ip": "10.0.0.9",
"cs-host": "host11.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 300,
"c-ip": "10.0.0.9",
"cs-host": "host12.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 325,
"c-ip": "10.0.0.9",
"cs-host": "host13.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 350,
"c-ip": "10.0.0.9",
"cs-host": "host14.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 375,
"c-ip": "10.0.0.9",
"cs-host": "host15.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 400,
"c-ip": "10.0.0.9",
"cs-host": "host16.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 425,
"c-ip": "10.0.0.9",
"cs-host": "host17.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 450,
"c-ip": "10.0.0.9",
"cs-host": "host18.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
},
{
"label": "20 distinct hosts but spread over ~13 minutes, so no single 10-minute window sees 20",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 40,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 80,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 120,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 160,
"c-ip": "10.0.0.9",
"cs-host": "host04.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "host05.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 240,
"c-ip": "10.0.0.9",
"cs-host": "host06.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 280,
"c-ip": "10.0.0.9",
"cs-host": "host07.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 320,
"c-ip": "10.0.0.9",
"cs-host": "host08.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 360,
"c-ip": "10.0.0.9",
"cs-host": "host09.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 400,
"c-ip": "10.0.0.9",
"cs-host": "host10.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 440,
"c-ip": "10.0.0.9",
"cs-host": "host11.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 480,
"c-ip": "10.0.0.9",
"cs-host": "host12.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 520,
"c-ip": "10.0.0.9",
"cs-host": "host13.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 560,
"c-ip": "10.0.0.9",
"cs-host": "host14.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 600,
"c-ip": "10.0.0.9",
"cs-host": "host15.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 640,
"c-ip": "10.0.0.9",
"cs-host": "host16.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 680,
"c-ip": "10.0.0.9",
"cs-host": "host17.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 720,
"c-ip": "10.0.0.9",
"cs-host": "host18.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 760,
"c-ip": "10.0.0.9",
"cs-host": "host19.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
},
{
"label": "high volume, low breadth: 25 requests from one source but only 4 distinct hosts",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 20,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 40,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 60,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 80,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 120,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 140,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 160,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 180,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 220,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 240,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 260,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 280,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 300,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 320,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 340,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 360,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 380,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 400,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 420,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 440,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 460,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 480,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
},
{
"label": "breadth split across two sources: 10 distinct hosts each, neither source reaches 20",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 25,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 50,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 75,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.9",
"cs-host": "host04.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 125,
"c-ip": "10.0.0.9",
"cs-host": "host05.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 150,
"c-ip": "10.0.0.9",
"cs-host": "host06.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 175,
"c-ip": "10.0.0.9",
"cs-host": "host07.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "host08.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "host09.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 0,
"c-ip": "10.0.0.10",
"cs-host": "host10.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 25,
"c-ip": "10.0.0.10",
"cs-host": "host11.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 50,
"c-ip": "10.0.0.10",
"cs-host": "host12.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 75,
"c-ip": "10.0.0.10",
"cs-host": "host13.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.10",
"cs-host": "host14.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 125,
"c-ip": "10.0.0.10",
"cs-host": "host15.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 150,
"c-ip": "10.0.0.10",
"cs-host": "host16.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 175,
"c-ip": "10.0.0.10",
"cs-host": "host17.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.10",
"cs-host": "host18.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.10",
"cs-host": "host19.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
}
]
}
@@ -0,0 +1,979 @@
{
"note": "webserver access log timeline for the ARTEX Enrichment Scan Velocity correlation (event_count >= 30 probes from one c-ip within a 5-minute window). 'ts' is relative seconds. Velocity (density in time), not total count, is the signal.",
"positive": [
{
"label": "30 enrichment probes from one source inside the 5-minute window",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 9,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 18,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 27,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 36,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 45,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 54,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 63,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 72,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 81,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 90,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 99,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 108,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 117,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 126,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 135,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 144,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 153,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 162,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 171,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 180,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 189,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 198,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 207,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 216,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 234,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 243,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 252,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 261,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
}
],
"negative": [
{
"label": "below the threshold: only 29 probes",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 9,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 18,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 27,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 36,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 45,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 54,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 63,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 72,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 81,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 90,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 99,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 108,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 117,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 126,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 135,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 144,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 153,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 162,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 171,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 180,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 189,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 198,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 207,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 216,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 234,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 243,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 252,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
},
{
"label": "30 probes total but spread over ~10 minutes, so no 5-minute window reaches 30",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 20,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 40,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 60,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 80,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 120,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 140,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 160,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 180,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 220,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 240,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 260,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 280,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 300,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 320,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 340,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 360,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 380,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 400,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 420,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 440,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 460,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 480,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 500,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 520,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 540,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 560,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 580,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
},
{
"label": "30 requests in the window but the User-Agent is a normal browser (base rule does not match)",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 9,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 18,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 27,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 36,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 45,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 54,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 63,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 72,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 81,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 90,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 99,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 108,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 117,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 126,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 135,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 144,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 153,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 162,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 171,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 180,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 189,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 198,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 207,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 216,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 234,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 243,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 252,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 261,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
}
]
}
]
}
@@ -0,0 +1,152 @@
{
"note": "application/audit log timeline for the ARTEX Guard-Block Burst correlation (event_count >= 5 guard-control markers on one host within a 10-minute window). 'ts' is relative seconds. A single marker can be a quoted string; a burst on one host indicates an actively engaged ARTEX run.",
"positive": [
{
"label": "5 guard-control markers on one host inside the 10-minute window",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 60,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 120,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 180,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 240,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
}
]
}
],
"negative": [
{
"label": "below the threshold: only 4 markers",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 60,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 120,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 180,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
}
]
},
{
"label": "5 markers but spread over ~13 minutes, so no 10-minute window holds 5",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 200,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 400,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 600,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 800,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
}
]
},
{
"label": "markers split across two hosts: 3 and 2, neither host reaches 5",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 60,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 120,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 0,
"host": "web02",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 60,
"host": "web02",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
}
]
},
{
"label": "ordinary log lines on one host, no guard marker (base rule does not match)",
"events": [
{
"ts": 0,
"host": "web01",
"message": "auth: user login ok uid=42 ip=203.0.113.8"
},
{
"ts": 60,
"host": "web01",
"message": "auth: user login ok uid=42 ip=203.0.113.8"
},
{
"ts": 120,
"host": "web01",
"message": "auth: user login ok uid=42 ip=203.0.113.8"
},
{
"ts": 180,
"host": "web01",
"message": "auth: user login ok uid=42 ip=203.0.113.8"
},
{
"ts": 240,
"host": "web01",
"message": "auth: user login ok uid=42 ip=203.0.113.8"
}
]
}
]
}
@@ -0,0 +1,81 @@
{
"note": "host timeline joining application/audit logs and process-creation logs for the ARTEX Guard Marker With Destructive Command temporal correlation (both referenced rules must fire on the SAME host within a 30-minute window). 'ts' is relative seconds.",
"positive": [
{
"label": "guard marker then a destructive command on the same host within 30 minutes",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 600,
"host": "web01",
"CommandLine": "rm -rf / --no-preserve-root",
"Image": "/usr/bin/rm"
}
]
}
],
"negative": [
{
"label": "only the guard marker, no destructive command",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 120,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
}
]
},
{
"label": "only a destructive command, no guard marker",
"events": [
{
"ts": 0,
"host": "web01",
"CommandLine": "rm -rf / --no-preserve-root",
"Image": "/usr/bin/rm"
}
]
},
{
"label": "both present but ~60 minutes apart, outside the 30-minute window",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 3600,
"host": "web01",
"CommandLine": "rm -rf / --no-preserve-root",
"Image": "/usr/bin/rm"
}
]
},
{
"label": "the two legs on different hosts",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 600,
"host": "db02",
"CommandLine": "rm -rf / --no-preserve-root",
"Image": "/usr/bin/rm"
}
]
}
]
}
@@ -0,0 +1,13 @@
{
"note": "process_creation telemetry (CommandLine). The rule hunts destructive shell/DB/availability commands via '1 of selection_*', so one representative command from each of the three selection groups must match. Benign commands — including a plain 'rm' without the recursive/force flags — must not. (The rule's own false-positive note documents that lowercase coreutils 'truncate' DOES share the TRUNCATE token and must be allow-listed, so it is intentionally not used here as a negative.)",
"positive": [
{"CommandLine": "rm -rf /var/www/html", "Image": "/usr/bin/rm"},
{"CommandLine": "mysql -u root -e 'DROP TABLE customers'", "Image": "/usr/bin/mysql"},
{"CommandLine": "iptables -F", "Image": "/usr/sbin/iptables"}
],
"negative": [
{"CommandLine": "ls -la /var/www/html", "Image": "/usr/bin/ls"},
{"CommandLine": "rm /tmp/scratch.txt", "Image": "/usr/bin/rm"},
{"CommandLine": "git status", "Image": "/usr/bin/git"}
]
}
+53
View File
@@ -0,0 +1,53 @@
#!/usr/bin/env bash
#
# Reproducible live event-matching test for the ARTEX Sigma rules — both the
# atomic rules (../../sigma/*.yml) and the correlation rules
# (../../sigma/correlation/*.yml). The sibling sigma/ suite proves those rules are
# valid and COMPILE to a backend query; this suite proves they actually FIRE on a
# matching event (or timeline) and stay quiet on a benign one — the same "a
# detection you cannot run is only a claim" guarantee the suricata/ suite already
# gives the network rule with a pcap replay.
#
# It proves six properties with no host dependency beyond Docker (pySigma runs in
# a container, nothing is installed on the host and nothing is written to the repo
# tree):
#
# atomic 1 rule/sample pairing every atomic rule has an events/<name>.json and
# every events file maps to a rule (no orphans)
# atomic 2 true positives each rule matches all of its malicious events
# atomic 3 true negatives each rule matches none of its benign events
# corr 1 rule/timeline pairing every correlation rule has an
# events/correlation/<name>.json (no orphans)
# corr 2 true positives each rule FIRES on its positive timeline
# (threshold met, inside the window, one group)
# corr 3 true negatives each rule stays QUIET on its negative timelines
# (below threshold, window exceeded, split group,
# or a missing leg)
#
# pySigma parses each rule — for an atomic rule its condition tree, for a
# correlation rule its aggregation spec (type, group-by, timespan, threshold, and
# the resolved references to the atomic base rules) — and check.py only walks that
# parsed structure, so the authoritative Sigma logic stays in pySigma (see
# check.py's header). The correlation window is the standard sliding-window model
# and matching is case-insensitive; see check.py for the full scope and honesty
# notes.
#
# Usage: detections/tests/sigma_match/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
# PYSIGMA_VERSION (default 2.0.0 — the pinned reference version)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
SIGMA_DIR="$REPO/detections/sigma"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
PYSIGMA_VERSION="${PYSIGMA_VERSION:-2.0.0}"
# Everything runs inside the container: check.sh installs the pinned pySigma and
# runs check.py, which asserts the three properties and exits non-zero on any
# failure. The rule tree and this directory are mounted read-only.
docker run --rm \
-v "$SIGMA_DIR:/sigma:ro" \
-v "$HERE:/src:ro" \
-e PYSIGMA_VERSION="$PYSIGMA_VERSION" \
"$PYTHON_IMAGE" sh /src/check.sh
+5
View File
@@ -0,0 +1,5 @@
# scratch captures created by run.sh (binary pcaps); never committed.
# Match both files and directories (.scratch.* , not .scratch.*/) so that any
# leftover .scratch.<name> is ignored even when an interrupted run (SIGKILL,
# power loss) skips the EXIT cleanup, and so `git check-ignore` reports it.
.scratch.*
+109
View File
@@ -0,0 +1,109 @@
#!/usr/bin/env python3
"""Deterministic pcap generator for the ARTEX Suricata rule tests.
Synthesizes N independent plaintext HTTP request/response flows from a single
source, each carrying a chosen User-Agent, so `suricata -r` can be run offline
to prove the rules in ../../suricata/artex.rules fire (or stay silent) exactly
as documented. Output is regenerated on every run and is never committed -- the
test ships as source, not as a binary capture.
Usage:
gen_pcap.py <out.pcap> <user-agent> [num_flows] [interval_seconds]
The capture is fully deterministic: fixed addresses, ports derived from the
flow index, a fixed base timestamp, and flows spaced `interval_seconds` apart.
Nothing here sends a packet or touches a network -- it only writes a file.
"""
import sys
from scapy.all import Ether, IP, TCP, Raw, wrpcap
# Fixed, private, non-routable endpoints. One source so Suricata's
# `detection_filter ... track by_src` on sid 1000002 counts per source.
SRC_MAC = "02:00:00:00:00:01"
DST_MAC = "02:00:00:00:00:02"
SRC_IP = "10.10.10.9"
DST_IP = "10.10.10.80"
DST_PORT = 80
BASE_EPOCH = 1_760_000_000.0 # fixed so timestamps never depend on wall clock
CLIENT_ISN = 1000
SERVER_ISN = 2000
def http_request(user_agent: str) -> bytes:
return (
"GET /products?category=all HTTP/1.1\r\n"
"Host: shop.example.test\r\n"
f"User-Agent: {user_agent}\r\n"
"Accept: */*\r\n"
"Connection: close\r\n"
"\r\n"
).encode()
HTTP_RESPONSE = (
"HTTP/1.1 200 OK\r\n"
"Content-Type: text/html\r\n"
"Content-Length: 13\r\n"
"Connection: close\r\n"
"\r\n"
"<html></html>"
).encode()
def flow(index: int, user_agent: str, t0: float):
"""One complete TCP+HTTP conversation; returns a list of timestamped packets."""
sport = 40000 + index
eth_c = Ether(src=SRC_MAC, dst=DST_MAC)
eth_s = Ether(src=DST_MAC, dst=SRC_MAC)
ip_c = IP(src=SRC_IP, dst=DST_IP)
ip_s = IP(src=DST_IP, dst=SRC_IP)
req = http_request(user_agent)
rlen = len(req)
slen = len(HTTP_RESPONSE)
pkts = []
def add(pkt, offset):
pkt.time = t0 + offset
pkts.append(pkt)
# Handshake
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="S", seq=CLIENT_ISN), 0.000)
add(eth_s / ip_s / TCP(sport=DST_PORT, dport=sport, flags="SA", seq=SERVER_ISN, ack=CLIENT_ISN + 1), 0.001)
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="A", seq=CLIENT_ISN + 1, ack=SERVER_ISN + 1), 0.002)
# Request
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="PA", seq=CLIENT_ISN + 1, ack=SERVER_ISN + 1) / Raw(req), 0.003)
add(eth_s / ip_s / TCP(sport=DST_PORT, dport=sport, flags="A", seq=SERVER_ISN + 1, ack=CLIENT_ISN + 1 + rlen), 0.004)
# Response
add(eth_s / ip_s / TCP(sport=DST_PORT, dport=sport, flags="PA", seq=SERVER_ISN + 1, ack=CLIENT_ISN + 1 + rlen) / Raw(HTTP_RESPONSE), 0.005)
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="A", seq=CLIENT_ISN + 1 + rlen, ack=SERVER_ISN + 1 + slen), 0.006)
# Teardown
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="FA", seq=CLIENT_ISN + 1 + rlen, ack=SERVER_ISN + 1 + slen), 0.007)
add(eth_s / ip_s / TCP(sport=DST_PORT, dport=sport, flags="A", seq=SERVER_ISN + 1 + slen, ack=CLIENT_ISN + 2 + rlen), 0.008)
add(eth_s / ip_s / TCP(sport=DST_PORT, dport=sport, flags="FA", seq=SERVER_ISN + 1 + slen, ack=CLIENT_ISN + 2 + rlen), 0.009)
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="A", seq=CLIENT_ISN + 2 + rlen, ack=SERVER_ISN + 2 + slen), 0.010)
return pkts
def main() -> int:
if len(sys.argv) < 3:
print(__doc__)
return 2
out = sys.argv[1]
user_agent = sys.argv[2]
num_flows = int(sys.argv[3]) if len(sys.argv) > 3 else 35
interval = float(sys.argv[4]) if len(sys.argv) > 4 else 1.0
packets = []
for i in range(num_flows):
packets.extend(flow(i, user_agent, BASE_EPOCH + i * interval))
wrpcap(out, packets)
print(f"wrote {len(packets)} packets across {num_flows} flows to {out} (UA={user_agent!r})")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+145
View File
@@ -0,0 +1,145 @@
#!/usr/bin/env bash
#
# Reproducible regression test for the ARTEX Suricata rules
# (../../suricata/artex.rules). It proves four properties with no committed
# binary capture and no host dependencies beyond Docker:
#
# 1. the whole rules file loads with zero errors (validity)
# `suricata -T --init-errors-fatal`; a rule that fails to parse or
# initialise is fatal even when no capture below exercises it
# 2. sid 1000001 fires exactly once per enrich probe (presence)
# 3. sid 1000002 fires once the 30-in-300s rate is hit (velocity)
# 4. sid 1000003 fires once per norma WebFetch request (presence)
# and the enrich sids stay silent on that capture (specificity)
# 5. an identical capture with a benign browser UA (specificity)
# produces zero alerts
#
# Everything runs in containers: `suricata -T` validates the ruleset, scapy
# synthesizes a deterministic pcap, then `suricata -r` reads it offline. The
# pcap is generated into a scratch dir that is removed on exit and is never
# committed.
#
# Usage: detections/tests/suricata/run.sh
# Env: SURICATA_IMAGE (default jasonish/suricata:latest)
# PYTHON_IMAGE (default python:3.12-slim)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
RULES_DIR="$REPO/detections/suricata"
SURICATA_IMAGE="${SURICATA_IMAGE:-jasonish/suricata:latest}"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
NUM_FLOWS=35
ENRICH_UA="artex-enrich/1.0"
# norma SDK WebFetch tool, hardcoded in github.com/Autumn-27/norma/tool/webfetch.go
# (literal "norma/0.4", verified in the go.sum-pinned v0.4.3 module source). Fewer flows
# than NUM_FLOWS because sid 1000003 is a single-hit presence rule with no rate component.
NORMA_UA="norma/0.4"
NORMA_FLOWS=8
BENIGN_UA="Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
# Scratch must live under the repo tree so Docker Desktop (macOS) can bind-mount
# it; /tmp and $TMPDIR are not shared by default. It is git-ignored and removed
# on exit.
SCRATCH="$(mktemp -d "$HERE/.scratch.XXXXXX")"
cleanup() { rm -rf "$SCRATCH"; }
trap cleanup EXIT
fail=0
note() { printf ' %s\n' "$1"; }
echo "== 1/5 validate the full ruleset loads (suricata -T) =="
# `suricata -T` loads the whole rules file in test mode and exits; --init-errors-fatal
# makes any rule that fails to parse or initialise a hard error. This catches a broken
# rule even when no capture below exercises it: plain `suricata -r` skips such a rule and
# still exits 0, so the firing checks would stay green while a signature silently fails to
# load. This is the Suricata analogue of the Sigma suite's `sigma check` validity assertion.
if docker run --rm -v "$RULES_DIR:/r:ro" "$SURICATA_IMAGE" \
suricata -T -S /r/artex.rules -l /tmp --init-errors-fatal >/dev/null 2>&1; then
note "PASS ruleset loads with zero parse/init errors (suricata -T)"
else
note "FAIL ruleset loads with zero parse/init errors (suricata -T)"
fail=1
fi
echo "== 2/5 synthesize deterministic captures (scapy) =="
docker run --rm -v "$SCRATCH:/out" -v "$HERE:/src:ro" "$PYTHON_IMAGE" sh -c "
pip install --quiet --disable-pip-version-check scapy >/dev/null 2>&1 &&
python /src/gen_pcap.py /out/enrich.pcap '$ENRICH_UA' $NUM_FLOWS &&
python /src/gen_pcap.py /out/norma.pcap '$NORMA_UA' $NORMA_FLOWS &&
python /src/gen_pcap.py /out/benign.pcap '$BENIGN_UA' $NUM_FLOWS
"
run_suricata() { # $1 = capture basename
local name="$1"
mkdir -p "$SCRATCH/$name-out"
# -k none: crafted packets carry no valid checksums; do not drop on them.
docker run --rm -v "$SCRATCH:/data" -v "$RULES_DIR:/r:ro" "$SURICATA_IMAGE" \
suricata -r "/data/$name.pcap" -S /r/artex.rules -k none -l "/data/$name-out" \
>/dev/null 2>&1
}
alerts() { # $1 = capture basename, $2 = sid (or "any")
python3 - "$SCRATCH/$1-out/eve.json" "$2" <<'PY'
import json, sys
path, sid = sys.argv[1], sys.argv[2]
n = 0
with open(path) as f:
for line in f:
line = line.strip()
if not line:
continue
try:
e = json.loads(line)
except ValueError:
continue
if e.get("event_type") != "alert":
continue
if sid == "any" or e.get("alert", {}).get("signature_id") == int(sid):
n += 1
print(n)
PY
}
expect() { # $1 label, $2 actual, $3 op (eq|ge), $4 expected
local label="$1" actual="$2" op="$3" expected="$4" ok
case "$op" in
eq) [ "$actual" -eq "$expected" ] && ok=1 || ok=0 ;;
ge) [ "$actual" -ge "$expected" ] && ok=1 || ok=0 ;;
esac
if [ "$ok" -eq 1 ]; then
note "PASS $label (got $actual, want $op $expected)"
else
note "FAIL $label (got $actual, want $op $expected)"
fail=1
fi
}
echo "== 3/5 run Suricata offline over the enrich capture =="
run_suricata enrich
e1="$(alerts enrich 1000001)"
e2="$(alerts enrich 1000002)"
expect "sid 1000001 presence: one alert per probe" "$e1" eq "$NUM_FLOWS"
expect "sid 1000002 velocity: fires past 30-in-300s" "$e2" ge 1
note "reference (Suricata 8.0.7): sid 1000002 = 5 (flows 31-35)"
echo "== 4/5 run Suricata offline over the norma WebFetch capture =="
run_suricata norma
n3="$(alerts norma 1000003)"
nenrich="$(( $(alerts norma 1000001) + $(alerts norma 1000002) ))"
expect "sid 1000003 presence: one alert per WebFetch request" "$n3" eq "$NORMA_FLOWS"
expect "enrich sids stay silent on norma traffic (specificity)" "$nenrich" eq 0
echo "== 5/5 run Suricata offline over the benign capture =="
run_suricata benign
b="$(alerts benign any)"
expect "benign browser UA produces no ARTEX alerts" "$b" eq 0
echo
if [ "$fail" -eq 0 ]; then
echo "RESULT: PASS"
else
echo "RESULT: FAIL"
fi
exit "$fail"
+31
View File
@@ -0,0 +1,31 @@
#!/usr/bin/env bash
#
# Non-suite gate: runs the host-triage tool's built-in --self-test in a container.
#
# The triage tool (detections/triage/artex_host_triage.py) is a responder helper,
# not a detection rule, so it is deliberately NOT one of the detections/tests/<x>/
# run.sh suites — that keeps the harness-sync registry exactly the eight rule
# suites (check-harness-sync.py counts only SUITES entries, `detections/tests/<x>/
# run.sh` CI steps, and directories carrying a run.sh; this file is none of them).
# It is a gate like check-harness-sync.sh: both run-all.sh and the detections CI
# workflow call it, so a broken triage check fails the same merge gate as the
# rule suites. A detection you cannot run is only a claim.
#
# The self-test builds its own synthetic host in a temporary directory, asserts
# every check fires on it and that a clean host produces zero findings, and exits
# non-zero on any failure. No host dependency beyond Docker: the tool is pure
# Python standard library and runs in a container with only detections/triage
# mounted read-only; nothing is installed on the host and nothing is written to
# the repo tree.
#
# Usage: detections/tests/triage-selftest.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../.." && pwd)"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
docker run --rm \
-v "$REPO/detections/triage:/triage:ro" \
"$PYTHON_IMAGE" python3 /triage/artex_host_triage.py --self-test