First Commit
ci / go (push) Waiting to run
ci / go-db (agent) (push) Waiting to run
ci / go-db (config) (push) Waiting to run
ci / go-db (db) (push) Waiting to run
ci / go-db (evidence) (push) Waiting to run
ci / go-db (llmrec) (push) Waiting to run
ci / go-db (server) (push) Waiting to run
detections / detections (push) Waiting to run
web / web (push) Waiting to run
docs / links (push) Canceled after 0s

This commit is contained in:
dela
2026-10-09 08:38:16 +08:00
commit 0335d572de
756 changed files with 201663 additions and 0 deletions
+260
View File
@@ -0,0 +1,260 @@
# ARTEX 탐지 규칙
한국어 · [English](README.md)
> 이 디렉터리는 [방어·탐지 가이드(../docs/defense-ko.md)](../docs/defense-ko.md) 4절 "탐지 규칙"의
> 의사 규칙을, 각자의 SIEM·EDR 질의 언어로 변환해 바로 배포할 수 있는 벤더 중립
> [Sigma](https://sigmahq.io) 형식으로 옮긴 것입니다. 여기 실린 모든 지표는 **추정이 아니라** 이
> 저장소 소스에서 실제로 확인한 문자열이나 행동에 근거합니다. 모든 규칙은 자신이 소유하거나 서면
> 허가를 받은 시스템을 지키는 **방어·탐지 목적에만** 사용하십시오.
## 원자(atomic) 규칙
- **`sigma/artex_enrich_user_agent.yml`**: ARTEX 자산 보강(`enrich/enrich.go`)이 보내는 인바운드
`artex-enrich/1.0` User-Agent 입니다. 대상 측에서 관측하는 보조 지표입니다. `level: high`.
- **`sigma/artex_selfupdate_egress.yml`**: 자가 업데이트 루틴(`selfupdate/github.go`)이 내보내는
아웃바운드 `artex-selfupdate` User-Agent 입니다. 호스트·포렌식 관점의 송신(egress) 지표입니다.
`level: medium`.
- **`sigma/artex_guard_audit_framing.yml`**: 도구 호출이 차단될 때 감사 로그에 기록되는 플랫폼
가드 통제 마커(`guard/guard.go`)입니다. 호스트·포렌식 지표입니다. `level: high`.
- **`sigma/destructive_command_hunting.yml`**: ARTEX 가드의 기본 차단 목록(`db/db.go` 시드)을
그대로 반영한 파괴적 셸·DB 명령입니다. ARTEX 고유 시그니처가 아니라 일반적인 헌팅 단서입니다.
`level: medium`.
- **`sigma/artex_recording_proxy_ca.yml`**: 기록용 프록시가 `_ca/mitmproxy-ca-cert.pem` 배치로
생성하는 MITM CA 인증서 파일(`traffic/traffic.go`)입니다. 호스트·포렌식 아티팩트이며, 파일명
자체는 단독 실행한 mitmproxy 와 공유되므로 헌팅 단서로 다룹니다. `level: medium`.
## 상관(correlation) 규칙: 행동 기반
정적 문자열은 바꿀 수 있지만 행동은 숨기기가 더 어렵습니다. [`sigma/correlation/`](sigma/correlation/)
의 Sigma **상관** 규칙은 방어 가이드(4.1~4.2절, 4.4절)의 행동 기반 계층을 담습니다. 각 상관 규칙은
위 원자 규칙을 `id` 로 참조하므로, 참조를 풀려면 단일 상관 파일이 아니라 `sigma/` 트리 전체를
변환해야 합니다(아래 참조).
- **`sigma/correlation/artex_enrich_scan_velocity.yml`**: 한 출처가 짧은 시간 창 안에서 쏟아내는
`artex-enrich/1.0` 프로브 묶음입니다(보강은 동시성 4 로 속도 제한 없이 돕니다). 단건 규칙이
놓치는 속도를 잡습니다. `event_count`, `level: high`.
- **`sigma/correlation/artex_enrich_fanout.yml`**: 한 출처가 보강 User-Agent 를 여러 **서로 다른**
호스트로 실어 나르는 경우입니다. 자산 목록 전체로 기계 속도로 퍼지는 양상으로, 양(volume)만이
아니라 폭(breadth)이 단서입니다. `value_count`, `level: high`.
- **`sigma/correlation/artex_guard_block_burst.yml`**: 한 호스트에서 플랫폼 가드 통제 마커가 반복해
찍히는 경우입니다. 단지 마커를 인용한 문서가 아니라, 돌고 있는 ARTEX 실행이 자기 가드를 건드리고
있다는 신호입니다. `event_count`, `level: high`.
- **`sigma/correlation/artex_guard_marker_then_destructive.yml`**: 한 호스트에서 시간 창 안에 가드
마커와 파괴적 명령이 함께 나타나는 경우입니다(방어 가이드 §4.2, 다단계). ARTEX 고유 마커를, 그
자체로는 일반적인 파괴 명령 신호와 결합해 특이도를 높입니다. `temporal`, `level: high`.
임계값과 시간 창은 보수적인 기본값입니다. 각자의 기준선(baseline)에 맞게 조정하십시오. §4.2 의 순수
웹 다단계 사례(열거 → 프로브 → 인증)는 그 패턴이 단일 ARTEX 고유 User-Agent 로 환원되지 않으므로,
여전히 환경별 기본 규칙이 따로 필요합니다. 그 출발점으로 쓸 수 있는 일반 행동 기반 Sigma 베이스
템플릿을 [방어 가이드 §4.2](../docs/defense-ko.md#42-siem-상관-규칙)에 두었습니다. ARTEX 소스로 근거를
고정할 수 없어 여기 테스트되는 규칙 트리에는 넣지 않았습니다.
## 네트워크 규칙 (Suricata)
Sigma 는 호스트와 로그 텔레메트리를 다룹니다. 네트워크 선에서 관측되는 ARTEX 고유 User-Agent 는 두
가지이고, 둘 다 [`suricata/`](suricata/)에 [Suricata](https://suricata.io) 규칙으로 들어 있습니다. 보강
프로버의 `artex-enrich/1.0`(`enrich/enrich.go`)에는 존재 시그니처 하나와 고속 열거 변형 하나(sid
1000001·1000002)가, norma SDK 의 WebFetch 도구가 공격 단계에 보내는 `norma/0.4`(`github.com/Autumn-27/norma/tool/webfetch.go`)에는
존재 시그니처 하나(sid 1000003)가 대응합니다. 그 밖의 worker 도구(Bash 로 실행하는 `curl`·`nmap` 등)는
자체 User-Agent 를 쓰므로 ARTEX 고유 지문이 없어, 네트워크 계층은 의도적으로 이 두 UA 로만 좁게
잡았습니다. 범위와 TLS 유의점, `suricata -T` 와 참조 pcap 으로 검증하는 방법은
[`suricata/README.ko.md`](suricata/README.ko.md)를 참조하십시오.
## ATT&CK 커버리지
이 규칙들이 태그하는 기법은 [MITRE ATT&CK](https://attack.mitre.org/) Navigator 레이어
[`attack/artex_navigator_layer.json`](attack/)에 모았습니다. 여섯 전술(정찰, 명령·제어, 실행, 임팩트,
자격 증명 접근, 수집)에 걸친 여덟 기법으로, 각 기법은 규칙의 `attack.*` 태그에 근거하고 탐지 강도(ARTEX 고유 시그니처인지,
일반 헌팅 단서인지)로 점수를 매겼습니다. [ATT&CK Navigator](https://mitre-attack.github.io/attack-navigator/)
에서 열면 어떤 ARTEX 행동을 어떤 규칙이 덮는지 볼 수 있습니다. 점수 산정과 기법↔규칙 대응, 그리고
정직한 범위(커버리지는 완전성이 아닙니다)는 [`attack/README.ko.md`](attack/README.ko.md)를 참조하십시오.
[일관성 테스트](tests/attack/run.sh)가 레이어와 규칙 집합이 서로 어긋나지 않게 지킵니다.
## 침해지표 목록 (기계가 읽는)
탐지 로직이 아니라 원자 지표 자체를 원하는 방어자를 위해,
[`indicators/artex_indicators.csv`](indicators/)는 ARTEX 가 내보내는 고유 지문을 CSV 한 파일에
모았습니다. 위협 인텔리전스 플랫폼이나 SIEM 조회 테이블, 호스트 분류(triage) 체크리스트에 바로
넣을 수 있도록 보강·자가 업데이트 User-Agent, 가드 감사 마커, 서버·프록시 기본 엔드포인트, 기록
프록시 CA 인증서, 그리고 PostgreSQL 탐색 그래프 스키마 지문을 담고, 각 행에는 근거가 된 소스
파일과 (있다면) 그 위에 세운 규칙을 함께 적었습니다. 같은 지표를 바로
가져올 수 있는 [MISP](https://www.misp-project.org/) 이벤트
([`indicators/artex_indicators.misp.json`](indicators/))로도 제공하므로, MISP 를 쓰거나 거기서
STIX 로 내보내는 방어자는 CSV 열을 손으로 매핑할 필요가 없습니다. 규칙에 근거한 지문은 `to_ids`
로 표시했고, 호스트 포렌식용 포트와 스키마 지문은 표시하지 않았습니다. 일반 헌팅 단서(파괴 명령)와
norma SDK 가 공유하는 `norma/0.4` WebFetch User-Agent(Suricata sid 1000003 이 잡는 네트워크 서명이지
ARTEX 고유 문자열이 아닙니다)는 오탐을 피하려 가져오기용 목록에서 의도적으로 뺐습니다. 열 구성, MISP 타입 매핑, 정직한 유의점, 그리고 CSV 와
MISP 이벤트가 어긋나지 않게 지키는 일관성 테스트는 [`indicators/README.ko.md`](indicators/README.ko.md)를
참조하십시오.
## 호스트 분류(triage)
위 규칙은 SIEM·네트워크 센서·위협 인텔리전스 플랫폼을 쓰는 방어자를 위한 것입니다. 그와 다른 대응자, 곧
SIEM 없이 의심 호스트 한 대의 셸 앞에 선 사람을 위해 [`triage/artex_host_triage.py`](triage/)를 둡니다. 로컬
상태만으로 "여기서 ARTEX 가 돌았는가"를 답하는 읽기 전용 스크립트입니다. 같은 지문을 점검하고, 여기에 더해
**CSV 가 의도적으로 Sigma 규칙 없이 둔 세 가지 호스트·DB 지표**(서버 리슨 포트, 기록 프록시 엔드포인트,
PostgreSQL 탐색 스키마)까지 점검합니다. 이 세 가지는 로그나 네트워크로 관측되지 않아 호스트에서 직접 확인할
수밖에 없습니다. 또한 기록기가 자식 프로세스에 주입하는 환경변수 흔적, 곧 실행 중인 프로세스가 프록시 변수와
mitmproxy CA 신뢰 변수를 함께 지니는지(`agent/worker.go`)를 `/proc` 나 `--proc-from` 덤프에서 확인합니다.
각 발견은 해당 침해지표 행과 같은 한계를 지닌 분류 단서입니다. 자세한 내용은
[`triage/README.ko.md`](triage/README.ko.md)에 있고, 내장된 `--self-test` 가 아래 머지 게이트로 돌아갑니다.
## 테스트
규칙에는 Docker 만 있으면 돌릴 수 있는 재현 테스트가 [`tests/`](tests/)에 함께 들어 있습니다.
- **Suricata** ([`tests/suricata/run.sh`](tests/suricata/run.sh)): 먼저 규칙 파일 전체가
`suricata -T --init-errors-fatal` 로 적재되는지 검증하고(어떤 캡처도 건드리지 않는 규칙이라도
파싱 실패는 잡힙니다), scapy 로 결정적 캡처를 합성한 뒤 `suricata -r` 로 그 위를 돌려, 존재 규칙이
프로브마다 한 번씩 발화하고 속도 규칙이 임계를 넘으면 걸리며 양성(benign) User-Agent 캡처에서는
경보가 0 인지 단언합니다. 이진 캡처는 커밋하지 않고 매 실행마다 다시 생성합니다.
- **Sigma** ([`tests/sigma/run.sh`](tests/sigma/run.sh)): 아래 "검증과 변환"의 `sigma check` 와
`sigma convert` 검증을 실행 가능한 테스트로 돌립니다. 오류 0 과, 트리 전체가 백엔드 질의로
컴파일됨과, 각 원자 지표 문자열이 그 질의까지 살아남음과, 상관 규칙이 홀로는 변환에 실패함을
단언합니다. 마지막 단언은 상관 규칙이 참조하는 원자 규칙에 실제로 의존함을 증명합니다.
- **Sigma 실시간 이벤트 매칭** ([`tests/sigma_match/run.sh`](tests/sigma_match/run.sh)): 위 Sigma 테스트가
규칙의 유효성과 컴파일을 증명한다면, 이 테스트는 원자 규칙과 상관 규칙이 실제로 발화하는지를 증명합니다.
각 원자 규칙마다 대표적인 악성 샘플 이벤트가 규칙을 발화시키고 정상 샘플 이벤트는 발화시키지 않음을
단언합니다(예: `.mitmproxy/` 아래 단독 CA 파일은 `_ca/` 디렉터리까지 함께 요구하는 기록용 프록시 규칙을
발화시키지 않습니다). 각 상관 규칙에 대해서는 임계를 시간 창 안에서 한 그룹이 채우는 양성 타임라인에
발화하고, 임계 미달·창 초과·그룹 분할·레그 누락 타임라인에는 침묵함을 단언합니다. 파싱은 전부 pySigma 에
맡기고 테스트는 컴파일된 조건 트리와 집계 명세만 걸으며, 상관 규칙이 어느 이벤트를 먹는지는 원자 규칙과
같은 매처로 판정합니다. "돌려 볼 수 없는 탐지 규칙은 주장일 뿐"이라는 원칙을 Suricata 처럼 Sigma 쪽에도
적용합니다.
- **ATT&CK 레이어** ([`tests/attack/run.sh`](tests/attack/run.sh)): ATT&CK 커버리지 레이어가
규칙과 일관되게 유지되는지 확인합니다. 점수를 매긴 기법·전술이 정확히 규칙 집합의 `attack.*`
태그여야 하고, 각 기법은 실재하는 규칙 파일을 지목해야 합니다. 레이어를 갱신하지 않고 규칙을
추가하면(또는 그 반대면) 테스트가 실패합니다.
- **지표 근거(source-of-truth)** ([`tests/indicators/run.sh`](tests/indicators/run.sh)): 각 규칙이
고정한 지표가 여전히 상류 소스가 내보내는 바로 그 문자열인지 확인합니다. `enrich/enrich.go` 의
`artex-enrich/1.0`, `selfupdate/` 의 `artex-selfupdate`, `guard/guard.go` 의 가드 마커, `db/db.go`
의 파괴 토큰이 규칙에도 여전히 고정돼 있는지 봅니다. 다른 세 테스트가 놓치는 드리프트, 즉 모든
규칙이 컴파일되고 발화하는 와중에 상류 재동기화가 User-Agent 나 마커를 바꿔 버리는 경우를
잡습니다. 같은 테스트가 기계가 읽는 [`indicators/artex_indicators.csv`](indicators/artex_indicators.csv)
를 다시 읽어, 발행된 모든 행이 여전히 소스와 규칙에 근거함을 단언하므로 방어자가 가져오는 산출물도
낡지 않습니다. 끝으로, 읽는 모든 상류 소스가 CI 워크플로의 `push`·`pull_request` 경로 필터에
들어 있음을 단언해, 새로 고정한 소스 하나만 건드린 PR 이 테스트를 건너뛰어 그 드리프트가 머지
게이트를 통과하지 못하게 합니다. 이로써 "추정이 아니라 이 저장소 소스에서 확인한 문자열에
근거한다"(위)는 약속이 말이 아니라 가드가 됩니다.
- **MISP 내보내기 일관성** ([`tests/misp/run.sh`](tests/misp/run.sh)): MISP 이벤트
([`indicators/artex_indicators.misp.json`](indicators/artex_indicators.misp.json))가 유효한 MISP
문서임을 증명합니다. [pymisp](https://github.com/MISP/PyMISP) 로 적재되는데, pymisp 의 객체 모델은
실재하지 않는 속성 타입을 거부하므로 이 산출물은 MISP 처럼 보이기만 하는 것이 아니라 실제로
가져와집니다. 또한 위 CSV 와 행 단위로 동기화됨을 단언합니다. 같은 값, 지표별로 의도한 MISP
타입·카테고리, 그리고 CSV 의 정직함을 그대로 반영하는 `to_ids`·`disable_correlation` 플래그가
일치해야 합니다(규칙에 근거 = 조치 가능이므로 `to_ids` on, 호스트 포렌식용 포트 = 분류 힌트이므로
`to_ids` off 에 상관 비활성화). 이 이벤트는 CSV 와 나란히 손으로 관리됩니다. CSV 에 없는 설명
주석·UUID·태그를 지니므로, CSV 행을 더하거나 빼거나 타입을 바꿀 때 같은 커밋에서 MISP 이벤트도
고쳐야 하고, 둘이 일치할 때까지 이 테스트가 실패합니다.
- **Sigma 백엔드 이식성** ([`tests/sigma_backends/run.sh`](tests/sigma_backends/run.sh)): 규칙이
Splunk 예시 하나를 넘어 변환됨을 증명합니다. 트리 전체(원자 + 상관)가 Splunk, Elasticsearch `eql`
타깃, Grafana Loki 로 컴파일되고, 다섯 원자 규칙은 Sigma 상관을 지원하지 않는 백엔드(Elasticsearch
`lucene`, Microsoft `kusto` 백엔드)에서도 여전히 컴파일됩니다. 아래 "검증과 변환"의 백엔드별 지원
표를 다시 돌릴 수 있는 점검으로 뒷받침합니다.
- **SigmaHQ 관례 린트** ([`tests/sigma_lint/run.sh`](tests/sigma_lint/run.sh)): SigmaHQ 검증기
전체(`pySigma-validators-sigmahq` 플러그인으로, 평범한 `sigma check` 는 적재하지 않습니다)를
[`tests/sigma_lint/validators.yml`](tests/sigma_lint/validators.yml)에 문서화한 기준선에 맞춰
돌리고 이슈 0 을 단언합니다. 또한 검증기 전체가 실제로 돌았고 의도적으로 제외한, 문서화된 네
검사만 남아 있음을 확인하므로, 규칙이 새 관례 이슈(잘못 대소문자를 쓴 제목, 분류 체계를 벗어난
필드)를 하나라도 들이면 빌드가 실패합니다.
여덟 규칙 테스트와 별개로, 규칙이 아닌 두 게이트가 같은 CI 워크플로와 [`tests/run-all.sh`](tests/run-all.sh)
에서 함께 돕니다. 하나는 하네스 동기 검사(run-all.sh·CI·스위트 디렉터리가 같은 스위트를 같은 순서로 부르는지
확인)이고, 다른 하나는 [호스트 분류 도구](triage/)의 `--self-test`([`tests/triage-selftest.sh`](tests/triage-selftest.sh))
로, 합성 호스트를 만들어 모든 분류 점검이 발화하는지와 깨끗한 호스트에서는 발견이 0 건인지 단언합니다.
각 스크립트는 단언이 하나라도 실패하면 0 이 아닌 코드로 종료합니다. [`tests/README.ko.md`](tests/README.ko.md)
를 참조하십시오.
## 이 규칙들을 정직하게 읽는 법
- **정적 지표는 바꿀 수 있습니다.** 운영자가 User-Agent 를 다른 값으로 설정할 수 있으므로,
`artex-enrich/1.0` 이나 `artex-selfupdate` 가 **없다고 해서 안전하다는 뜻은 아닙니다.** 오래가는
신호는 *행동* 입니다. 한 출처가 정찰 → 열거 → 프로브 → 인증·주입 시도로 이어지며 응답에 적응하고
쉼 없이 도는 양상입니다. 그 계층은 방어 가이드(1절·2절·4.1~4.2절)에 설명했고, 위
`sigma/correlation/` 규칙이 배포 가능한 상관(속도, 팬아웃, 가드 차단 묶음, 그리고 가드
마커+파괴명령 다단계)으로 담았으며, 순수 웹 다단계 사례는 여전히 환경별 기본 규칙이 필요합니다.
- **파괴 명령 규칙은 일반 헌팅입니다.** ARTEX 가드의 차단 목록을 반영하지만, 같은 명령은 정당한
관리자도 실행합니다. 적중은 단서로 다루고, 환경에 맞게 허용 목록을 두며, 그것만으로 ARTEX 라고
단정하지 마십시오.
- **포트 지표는 Sigma 가 아니라 호스트 포렌식용입니다.** ARTEX 서버 기본 `:8787` 과 기록 프록시
`127.0.0.1:8788`(`cmd/artex/main.go`)은 의심되는 호스트에서 `ss`·`netstat` 로 확인하는 편이
낫습니다. 그래서 시끄러운 네트워크 규칙으로 싣지 않고, 방어 가이드에 문서화하고 분류용으로
[지표 CSV](indicators/)에 올렸습니다. [호스트 분류 스크립트](triage/)는 바로 이런 호스트 로컬 점검(포트,
기록 프록시 아티팩트, 로그 마커, PostgreSQL 스키마)을 셸 접근은 있으나 SIEM 이 없는 대응자를 위해
대신 돌려 줍니다.
## 검증과 변환
이 규칙들은 [sigma-cli](https://github.com/SigmaHQ/sigma-cli)(pySigma)로 검증했습니다. 재현하려면:
```sh
python3 -m venv .venv && . .venv/bin/activate
pip install sigma-cli
# 구조 + 모범 사례 검증 (기대: 오류 0, 이슈 0)
sigma check detections/sigma/
# 이 독립 규칙 집합을 위한 문서화된 기준선으로 SigmaHQ 관례 전체 검사 (기대: 이슈 0).
# 위의 평범한 `sigma check` 는 이 검증기들을 적재하지 않습니다.
pip install pySigma-validators-sigmahq
sigma check --validation-config detections/tests/sigma_lint/validators.yml detections/sigma/
# 대상 질의 언어로 컴파일, 예: Splunk
sigma plugin install splunk
sigma convert -t splunk --without-pipeline detections/sigma/artex_enrich_user_agent.yml
# 상관 규칙이 id 로 참조하는 원자 규칙을 풀 수 있도록 트리 전체를 변환
sigma convert -t splunk --without-pipeline detections/sigma/
```
기준선은 SigmaHQ 관례를 모두 강제하되, SigmaHQ 모노레포의 파일 정리 체계와 분류 체계를 담은 네
검사만은 이 독립 규칙 집합에 해당하지 않으므로 제외합니다. 각 제외와 그 근거는
[`tests/sigma_lint/validators.yml`](tests/sigma_lint/validators.yml)에 문서화했고 위 린트 테스트가
강제합니다.
### Sigma 백엔드 이식성
`sigma/correlation/` 규칙은 원자 기본 규칙을 `id` 로 참조하므로, Sigma 상관 변환을 지원하는
백엔드에서만 변환됩니다. 그 지원은 백엔드마다 다르므로 `-t` 선택이 중요합니다. 아래 표는 고정한
기준(`sigma-cli` 3.1.0, 호환되는 최신 백엔드)에 맞춰 측정했고
[`tests/sigma_backends/run.sh`](tests/sigma_backends/run.sh)가 재현합니다.
- **트리 전체(원자 + 상관) 변환:** Splunk(`-t splunk`), Elasticsearch EQL(`-t eql`),
Grafana Loki(`-t loki`). `detections/sigma/` 를 바로 변환하면 상관 질의까지 함께 얻습니다.
- **원자 규칙만(상관 아직 미지원):** Elasticsearch Lucene(`-t lucene`),
OpenSearch(`-t opensearch_lucene`), 그리고 Sentinel·Defender XDR 를 겨냥하는 Microsoft `kusto`
백엔드(`-t kusto`). 이들에서는 다섯 원자 규칙을 변환하고 상관 시간 창은 제품 안에서 네이티브로
표현합니다(예: Sentinel 예약 분석의 `summarize ... by bin(TimeGenerated, 30m)`). 디렉터리 전체를
넘기면 "Backend does not support correlation rules" 로 변환이 멈춥니다.
```sh
# 원자 규칙만, 예: Microsoft Sentinel / Defender (kusto 백엔드)
sigma plugin install kusto
sigma convert -t kusto --without-pipeline \
detections/sigma/artex_enrich_user_agent.yml \
detections/sigma/artex_selfupdate_egress.yml \
detections/sigma/artex_guard_audit_framing.yml \
detections/sigma/artex_recording_proxy_ca.yml \
detections/sigma/destructive_command_hunting.yml
```
고정한 버전에서 알려진 경계: Elasticsearch ES|QL 타깃(`-t esql`)은 가드 마커 규칙을 거부하므로
(`String value expressions are not supported`) 거기서는 나머지 세 원자 규칙을 변환하십시오. 그리고
IBM QRadar 플러그인(`ibm-qradar-aql`)은 고정한 pySigma 와 호환되지 않아 `--force-install` 이
필요하므로 테스트에서 다루지 않습니다. 환경에 설치된 백엔드는 `sigma list targets` 로 확인하십시오.
위 예시는 `--without-pipeline` 을 써서 규칙 본문의 일반 필드명(`cs-user-agent`·`cs-host`·
`CommandLine`)을 그대로 내보냅니다. 제품 스키마에 맞추려면 그 플래그를 빼고 `-p` 로 처리
파이프라인을 적용하십시오(`sigma list pipelines` 참조). 다만 제품 파이프라인은 필드명을 매핑하되
규칙의 일반 `logsource` 가 지정하지 않는 대상 테이블을 추가로 요구할 수 있습니다. 예를 들어
`-p sentinel_asim` 은 데이터에 맞는 `query_table` 을 설정하기 전까지 "Unable to determine table name"
으로 멈추므로, 배포 전에 필드와 목적지 테이블을 환경에 맞게 매핑하십시오.
## 기여
탐지와 하드닝 기여를 환영합니다. 새 규칙은 모든 지표를 관측 가능한 사실에 근거해 두고, 한계를
`description` 에 밝히며, SigmaHQ 검증기 기준선을 깨끗이 통과하고
(`sigma check --validation-config tests/sigma_lint/validators.yml`), 공격 안내로 읽히는 내용을
담지 않아야 합니다. [`../CONTRIBUTING.md`](../CONTRIBUTING.md)를 참조하십시오.
+267
View File
@@ -0,0 +1,267 @@
# ARTEX detection rules
English · [한국어](README.ko.md)
> 한국어: 이 디렉터리는 [방어·탐지 가이드(docs/defense-ko.md)](../docs/defense-ko.md)의 4절 "탐지 규칙"을
> 실제로 배포 가능한 [Sigma](https://sigmahq.io) 규칙으로 옮긴 것입니다. 모든 규칙은 자신이 소유하거나 서면
> 허가를 받은 시스템을 지키는 **방어·탐지 목적에만** 사용하십시오. 한국어 전체 문서는
> **[detections/README.ko.md](README.ko.md)** 를 보십시오.
Deployable [Sigma](https://sigmahq.io) rules that formalize the pseudo-rules in the defense guide
([Korean](../docs/defense-ko.md) · [English](../docs/defense-en.md), section 4) into a vendor-neutral
format you can convert to your own SIEM or EDR query language. Every indicator here is grounded in a
string or behaviour verified in this repository's source, not inferred.
## Atomic rules
- **`sigma/artex_enrich_user_agent.yml`** — inbound `artex-enrich/1.0` User-Agent from ARTEX asset
enrichment (`enrich/enrich.go`). Target-side, supporting indicator. `level: high`.
- **`sigma/artex_selfupdate_egress.yml`** — outbound `artex-selfupdate` User-Agent from the self-update
routine (`selfupdate/github.go`). Host/forensic egress indicator. `level: medium`.
- **`sigma/artex_guard_audit_framing.yml`** — the platform-guard control marker written to the audit log
on a blocked tool call (`guard/guard.go`). Host/forensic indicator. `level: high`.
- **`sigma/destructive_command_hunting.yml`** — destructive shell/DB commands mirroring the ARTEX guard's
built-in deny list (`db/db.go` seed). Generic hunting lead, not an ARTEX signature. `level: medium`.
- **`sigma/artex_recording_proxy_ca.yml`** — creation of the recording proxy's MITM CA file under the
`_ca/mitmproxy-ca-cert.pem` layout (`traffic/traffic.go`). Host/forensic artifact; the bare filename is
shared with standalone mitmproxy, so it is a hunting lead. `level: medium`.
## Correlation rules (behaviour)
Static strings can be changed; behaviour is harder to hide. These Sigma **correlation** rules in
[`sigma/correlation/`](sigma/correlation/) encode the behaviour-based layer of the defense guide
(sections 4.1–4.2 and 4.4). Each references an atomic rule above by its `id`, so convert the whole
`sigma/` tree — not a single correlation file — to resolve the reference (see below).
- **`sigma/correlation/artex_enrich_scan_velocity.yml`** — a burst of `artex-enrich/1.0` probes from one
source in a short window (enrichment runs at concurrency 4 with no rate limit). The velocity the
single-request rule misses. `event_count`, `level: high`.
- **`sigma/correlation/artex_enrich_fanout.yml`** — one source carrying the enrichment User-Agent to many
*distinct* hosts: machine-speed fan-out across an asset list, where breadth (not just volume) is the
tell. `value_count`, `level: high`.
- **`sigma/correlation/artex_guard_block_burst.yml`** — repeated platform-guard control markers on one
host, i.e. an active ARTEX run tripping its guard rather than a document that merely quotes the marker.
`event_count`, `level: high`.
- **`sigma/correlation/artex_guard_marker_then_destructive.yml`** — the guard marker and a destructive
command co-occurring on one host within a window (defense guide §4.2, multi-stage). Combining an
ARTEX-specific marker with the otherwise-generic destructive-command signal raises specificity.
`temporal`, `level: high`.
Thresholds and windows are conservative defaults — tune them to your baseline. The pure web multi-stage
case in §4.2 (enumerate → probe → authenticate) still needs base rules specific to your environment,
because that pattern does not reduce to a single ARTEX-unique User-Agent. A generic behavioral Sigma base
template to start from is provided in [defense guide §4.2](../docs/defense-en.md#42-siem-correlation-rules);
it is kept out of this tested rule tree because it cannot be grounded in ARTEX source.
## Network rules (Suricata)
Sigma covers host and log telemetry. The two ARTEX User-Agents observable on the wire both ship as
[Suricata](https://suricata.io) rules in [`suricata/`](suricata/): the enrichment prober's `artex-enrich/1.0`
(`enrich/enrich.go`) with a presence signature plus a high-rate enumeration variant (sid 1000001–1000002),
and the norma SDK WebFetch tool's attack-phase `norma/0.4` (`github.com/Autumn-27/norma/tool/webfetch.go`) with a presence signature
(sid 1000003). Other worker tools (Bash-run `curl`, `nmap`) use their own User-Agents and carry no
ARTEX-unique fingerprint, so the network layer is intentionally narrow to these two UAs; see
[`suricata/README.md`](suricata/README.md) for the scope, the TLS caveat, and how to validate with
`suricata -T` and a reference pcap.
## ATT&CK coverage
The techniques these rules tag are collected into a [MITRE ATT&CK](https://attack.mitre.org/) Navigator
layer in [`attack/artex_navigator_layer.json`](attack/) — eight techniques across six tactics
(Reconnaissance, Command and Control, Execution, Impact, Credential Access, Collection), each grounded in a rule's `attack.*` tags and
scored by detection strength (ARTEX-specific signature vs. generic hunting lead). Open it in the
[ATT&CK Navigator](https://mitre-attack.github.io/attack-navigator/) to see which ARTEX behaviour each
rule covers; see [`attack/README.md`](attack/README.md) for the scoring, the technique-to-rule map, and
the honest scope (coverage is not completeness). A [consistency test](tests/attack/run.sh) keeps the layer
from drifting away from the rule set.
## Indicator list (machine-readable)
For defenders who want the atomic indicators rather than the detection logic,
[`indicators/artex_indicators.csv`](indicators/) collects the unique fingerprints ARTEX emits into one
CSV to drop into a threat-intelligence platform, a SIEM lookup, or a host-triage checklist — the enrichment
and self-update User-Agents, the guard audit marker, the server/proxy default endpoints, the recording-proxy
CA certificate, and the PostgreSQL exploration-graph schema fingerprint — each row recording the source file
it is grounded in and the rule (if any) built on it. The same indicators ship as a
ready-to-import [MISP](https://www.misp-project.org/) event
([`indicators/artex_indicators.misp.json`](indicators/)), so a defender running MISP (or exporting on to
STIX from it) does not have to map the CSV columns by hand — the rule-backed fingerprints are flagged
`to_ids`, the host-forensic ports and schema fingerprint are not. Generic hunting leads (the
destructive commands) and the norma SDK's shared `norma/0.4` WebFetch User-Agent — a wire signature carried
by Suricata sid 1000003, not an ARTEX-unique string — are deliberately kept out of the import-ready list to
avoid false positives; see
[`indicators/README.md`](indicators/README.md) for the columns, the MISP type mapping, the honest caveats,
and the consistency test that keeps both the CSV and the MISP event from drifting.
## Host triage
The rules above serve defenders with a SIEM, a network sensor, or a threat-intel platform. For the other
responder — the one at a single suspected host's shell, with no SIEM — [`triage/artex_host_triage.py`](triage/)
is a read-only script that answers "did ARTEX run here?" from local state. It operationalizes the same
fingerprints, **plus the three host/DB indicators the CSV deliberately carries without a Sigma rule**
(the server listen port, the recording-proxy endpoint, and the PostgreSQL exploration schema),
which are not log- or network-observable and can only be checked on the box. It also flags the recorder's
subprocess env-injection — a running process carrying a proxy var together with a mitmproxy CA-trust var
(`agent/worker.go`), read from `/proc` or a `--proc-from` dump. Every finding is a triage lead carrying the
same caveat as its indicator row. See [`triage/README.md`](triage/README.md); a built-in `--self-test`
runs as a merge-gate (below).
## Tests
The rules ship with reproducible tests in [`tests/`](tests/), each needing only Docker:
- **Suricata** ([`tests/suricata/run.sh`](tests/suricata/run.sh)) first validates that the whole rules file
loads under `suricata -T --init-errors-fatal` (a rule that fails to parse is caught even if no capture
exercises it), then synthesizes a deterministic capture with scapy, runs `suricata -r` over it, and asserts
that the presence rule fires once per probe, the velocity rule trips past its rate threshold, and a
benign-User-Agent capture produces zero alerts. No binary capture is committed — the test regenerates it on
every run.
- **Sigma** ([`tests/sigma/run.sh`](tests/sigma/run.sh)) runs the `sigma check` and `sigma convert` validation
below as an executable test: it asserts 0 errors, that the whole tree compiles to a backend query, that each
atomic indicator string survives into that query, and that a correlation rule fails to convert on its own —
proving it genuinely depends on the atomic rule it references.
- **Sigma live event-matching** ([`tests/sigma_match/run.sh`](tests/sigma_match/run.sh)) extends the Sigma suite
above from validity/compilation to actual firing, for both the atomic and the correlation rules: for every
atomic rule it asserts a representative malicious sample event matches and a benign one does not (for example,
a standalone CA file under `.mitmproxy/` does not trip the recording-proxy rule, whose `|all` also requires the
`_ca/` directory); for every correlation rule it asserts a positive timeline fires (threshold met inside the
window within one group) and negative ones stay quiet (below threshold, window exceeded, split group, or a
missing leg). pySigma does all parsing; the test only walks the compiled condition tree and aggregation spec,
deciding which events feed a referenced rule with the same atomic matcher. It brings "a detection you cannot
run is only a claim" to the Sigma side the way Suricata has it.
- **ATT&CK layer** ([`tests/attack/run.sh`](tests/attack/run.sh)) checks that the ATT&CK coverage layer stays
consistent with the rules: its scored techniques and tactics must be exactly the `attack.*` tags on the
rule set, and each technique must name a rule file that exists. Adding a rule without updating the layer
(or vice versa) fails the test.
- **Indicator source-of-truth** ([`tests/indicators/run.sh`](tests/indicators/run.sh)) checks that each rule's
pinned indicator is still the string the upstream source emits — `artex-enrich/1.0` in `enrich/enrich.go`,
`artex-selfupdate` in `selfupdate/`, the guard marker in `guard/guard.go`, the destructive tokens in
`db/db.go` — and is still pinned in the rule. It catches the drift the other three miss: an upstream re-sync
that changes a User-Agent or marker while every rule still compiles and fires. The same test re-reads the
machine-readable [`indicators/artex_indicators.csv`](indicators/artex_indicators.csv) and asserts every
published row is still grounded in its source and rule, so the artifact a defender imports cannot drift
either. Finally it asserts that every upstream source it reads is listed in the CI workflow's `push` and
`pull_request` paths filter, so a PR touching only a newly pinned source cannot skip the test and let that
drift pass the merge gate. This makes "grounded in a string verified in this repository's source, not
inferred" (above) a guard, not a promise.
- **MISP export consistency** ([`tests/misp/run.sh`](tests/misp/run.sh)) proves the MISP event
([`indicators/artex_indicators.misp.json`](indicators/artex_indicators.misp.json)) is a valid MISP document —
it loads under [pymisp](https://github.com/MISP/PyMISP), whose object model rejects any attribute type that
is not a real MISP type, so the artifact really imports rather than merely looking like MISP — and that it
stays row-for-row in sync with the CSV above: same values, the intended MISP type/category per indicator,
and a `to_ids`/`disable_correlation` flag that mirrors the CSV's honesty (rule-backed = actionable, so
`to_ids` on; host-forensic port = triage hint, so `to_ids` off and correlation disabled). The event is
hand-maintained alongside the CSV — it carries curated comments, UUIDs, and tags the CSV does not — so
when you add, remove, or retype a CSV row you update the MISP event in the same commit, and this test
fails until the two agree.
- **Sigma backend portability** ([`tests/sigma_backends/run.sh`](tests/sigma_backends/run.sh)) proves the rules
convert beyond the single Splunk example: the whole tree (atomic + correlation) compiles on Splunk, the
Elasticsearch `eql` target, and Grafana Loki, and the five atomic rules still compile on backends that do not
support Sigma correlations (Elasticsearch `lucene`, the Microsoft `kusto` backend). It backs the per-backend
support matrix in [Validate and convert](#validate-and-convert) below with a re-runnable check.
- **SigmaHQ convention lint** ([`tests/sigma_lint/run.sh`](tests/sigma_lint/run.sh)) runs the full SigmaHQ
validator set (the `pySigma-validators-sigmahq` plugin, which plain `sigma check` does not load) against the
documented baseline in [`tests/sigma_lint/validators.yml`](tests/sigma_lint/validators.yml) and asserts 0
issues. It also checks that the full set actually ran and that only four deliberately excluded, documented
checks remain, so a rule that picks up a new convention issue (a mis-cased title, an off-taxonomy field)
fails the build.
Alongside the eight rule suites, two non-suite gates run in the same CI workflow and in
[`tests/run-all.sh`](tests/run-all.sh): a harness-sync check (that `run-all.sh`, CI, and the suite
directories name the same suites in the same order) and the [host-triage tool](triage/)'s `--self-test`
([`tests/triage-selftest.sh`](tests/triage-selftest.sh)), which builds a synthetic host and asserts every
triage check fires on it while a clean host produces zero findings.
Each script exits non-zero on any failed assertion. See [`tests/README.md`](tests/README.md).
## How to read these honestly
- **Static indicators can be changed.** An operator can set a different User-Agent, so the absence of
`artex-enrich/1.0` or `artex-selfupdate` does **not** mean safety. The durable signal is *behaviour* —
a single source chaining recon → enumeration → probing → auth/injection attempts, adapting to responses,
running without pause. That layer is described in the defense guide (sections 1, 2, and 4.1–4.2); the
`sigma/correlation/` rules above ship it as deployable correlations (velocity, fan-out, guard-block
burst, and a guard-marker-with-destructive-command multi-stage), and the pure web multi-stage case
still needs base rules specific to your environment.
- **The destructive-command rule is generic hunting.** It mirrors ARTEX's guard deny list, but the same
commands are run by legitimate administrators. Treat a hit as a lead, allow-list your environment, and
do not attribute it to ARTEX on its own.
- **Port indicators are host-forensic, not Sigma.** The ARTEX server default `:8787` and the recording
proxy `127.0.0.1:8788` (`cmd/artex/main.go`) are best checked on a suspected host with `ss`/`netstat`,
so they are documented in the defense guide and listed in the [indicator CSV](indicators/) for triage,
rather than shipped as a noisy network rule. The [host-triage script](triage/) runs exactly those
host-local checks (ports, recording-proxy artifacts, log markers, and the PostgreSQL schema) for a
responder who has shell access but no SIEM.
## Validate and convert
These rules are validated with [sigma-cli](https://github.com/SigmaHQ/sigma-cli) (pySigma). To reproduce:
```sh
python3 -m venv .venv && . .venv/bin/activate
pip install sigma-cli
# structural + best-practice validation (expect: 0 errors, 0 issues)
sigma check detections/sigma/
# full SigmaHQ convention set with the documented baseline for this standalone
# rule set (expect: 0 issues). Plain `sigma check` above does not load these.
pip install pySigma-validators-sigmahq
sigma check --validation-config detections/tests/sigma_lint/validators.yml detections/sigma/
# compile to a target query language, e.g. Splunk
sigma plugin install splunk
sigma convert -t splunk --without-pipeline detections/sigma/artex_enrich_user_agent.yml
# convert the whole tree so the correlation rules can resolve the atomic rules they reference by id
sigma convert -t splunk --without-pipeline detections/sigma/
```
The baseline enforces every SigmaHQ convention except four checks that encode SigmaHQ's monorepo filing scheme
and taxonomy, which do not apply to a standalone rule set; each exclusion and its rationale is documented in
[`tests/sigma_lint/validators.yml`](tests/sigma_lint/validators.yml) and enforced by the lint test above.
### Sigma backend portability
The `sigma/correlation/` rules reference their atomic base rules by `id`, so they only convert on backends
that support Sigma correlation conversion. That support varies by backend, so `-t` choice matters. The matrix
below is measured against the pinned reference (`sigma-cli` 3.1.0, latest compatible backends) and reproduced
by [`tests/sigma_backends/run.sh`](tests/sigma_backends/run.sh):
- **Converts the whole tree (atomic + correlation):** Splunk (`-t splunk`), Elasticsearch EQL (`-t eql`),
Grafana Loki (`-t loki`). Convert `detections/sigma/` directly and you get the correlation queries too.
- **Atomic rules only (correlations not yet supported):** Elasticsearch Lucene (`-t lucene`), OpenSearch
(`-t opensearch_lucene`), and the Microsoft `kusto` backend that targets Sentinel and Defender XDR
(`-t kusto`). On these, convert the five atomic rules and express the correlation window natively in the
product (e.g. a Sentinel scheduled-analytics `summarize ... by bin(TimeGenerated, 30m)`). Pass the whole
directory and the conversion stops with "Backend does not support correlation rules."
```sh
# atomic rules only, e.g. for Microsoft Sentinel / Defender (kusto backend)
sigma plugin install kusto
sigma convert -t kusto --without-pipeline \
detections/sigma/artex_enrich_user_agent.yml \
detections/sigma/artex_selfupdate_egress.yml \
detections/sigma/artex_guard_audit_framing.yml \
detections/sigma/artex_recording_proxy_ca.yml \
detections/sigma/destructive_command_hunting.yml
```
Known edges at the pinned versions: the Elasticsearch ES|QL target (`-t esql`) rejects the guard-marker rule
(`String value expressions are not supported`), so convert the other three atomic rules there; and the IBM
QRadar plugin (`ibm-qradar-aql`) is not compatible with the pinned pySigma and needs `--force-install`, so it
is not covered by the test. Run `sigma list targets` for the backends installed in your environment.
The examples above use `--without-pipeline`, which emits the generic field names from the rule bodies
(`cs-user-agent`, `cs-host`, `CommandLine`). To match your product's schema, drop that flag and apply a
processing pipeline with `-p` (see `sigma list pipelines`). Note that a product pipeline maps field names but
may also need a target table the rules' generic `logsource` does not specify — e.g. `-p sentinel_asim` stops
with "Unable to determine table name" until you set `query_table` for your data, so map the fields and the
destination table to your environment before deploying.
## Contributing
Detection and hardening contributions are welcome. New rules should keep every indicator grounded in an
observable fact, state limitations in the `description`, pass the SigmaHQ validator baseline cleanly
(`sigma check --validation-config tests/sigma_lint/validators.yml`), and avoid any content that reads as
attack guidance. See [`../CONTRIBUTING.en.md`](../CONTRIBUTING.en.md).
+79
View File
@@ -0,0 +1,79 @@
# ARTEX ATT&CK 커버리지
한국어 · [English](README.md)
이 저장소의 탐지 규칙이 태그하는 [MITRE ATT&CK](https://attack.mitre.org/)(Enterprise) 기법을
Navigator 레이어로 정리한 것입니다. [Sigma 규칙](../sigma/)의 `attack.*` 태그에서 손으로 만들었고,
모든 기법은 지표가 이 저장소 소스에서 확인한 문자열이나 행동인 규칙에 근거합니다. 추정으로 넣은
항목은 없으며, [일관성 테스트](../tests/attack/run.sh)가 레이어와 규칙이 서로 어긋나지 않게 지킵니다.
- **`artex_navigator_layer.json`**: ATT&CK Navigator v4.5 형식의 레이어입니다.
## 점수의 의미
여기서 커버리지는 "이 저장소가 이 기법을 태그하는 탐지를 제공한다"는 뜻이지, "이 기법이 완전히
덮인다"는 뜻이 아닙니다. 점수는 탐지 강도를 일부러 정직하게 매겼습니다.
- **100: ARTEX 고유 시그니처 또는 행동.** ARTEX 에만 있는 정적 지표(`artex-enrich/1.0`·
`artex-selfupdate` User-Agent, 가드 감사 마커)이거나, 그 위에 세운 행동 규칙(보강 속도·팬아웃,
가드 차단 묶음)입니다.
- **50–65: 일반 헌팅 단서.** ARTEX 가드의 차단 목록을 반영한 파괴적 명령 헌팅입니다. 같은 명령은
정당한 관리자도 실행하므로 양성(benign) 활동에서도 발화합니다. 적중은 단서로 다루고 단정의
근거로 삼지 마십시오. 65 는 상관 규칙이 그 명령을 ARTEX 가드 마커와 결합해 특이도를 높인
경우를 가리킵니다.
## 다루는 기법
여섯 전술에 걸친 여덟 기법입니다. 각 기법은 그것을 태그하는 규칙에 대응합니다.
- **정찰(Reconnaissance): T1595 (Active Scanning), T1592 (Gather Victim Host Information).**
[`sigma/artex_enrich_user_agent.yml`](../sigma/artex_enrich_user_agent.yml),
[`sigma/correlation/artex_enrich_scan_velocity.yml`](../sigma/correlation/artex_enrich_scan_velocity.yml),
[`sigma/correlation/artex_enrich_fanout.yml`](../sigma/correlation/artex_enrich_fanout.yml), 그리고
[Suricata 규칙](../suricata/artex.rules)(sid 1000001 / 1000002)입니다.
- **명령·제어(Command and Control): T1105 (Ingress Tool Transfer).**
[`sigma/artex_selfupdate_egress.yml`](../sigma/artex_selfupdate_egress.yml)입니다.
- **실행(Execution): T1059 (Command and Scripting Interpreter).**
[`sigma/artex_guard_audit_framing.yml`](../sigma/artex_guard_audit_framing.yml),
[`sigma/correlation/artex_guard_block_burst.yml`](../sigma/correlation/artex_guard_block_burst.yml),
[`sigma/correlation/artex_guard_marker_then_destructive.yml`](../sigma/correlation/artex_guard_marker_then_destructive.yml)입니다.
- **임팩트(Impact): T1485 (Data Destruction), T1561.002 (Disk Wipe: Disk Structure Wipe), T1489 (Service Stop).**
[`sigma/destructive_command_hunting.yml`](../sigma/destructive_command_hunting.yml)이며, T1485 는
[`sigma/correlation/artex_guard_marker_then_destructive.yml`](../sigma/correlation/artex_guard_marker_then_destructive.yml)로도 보강됩니다.
- **자격 증명 접근·수집(Credential Access / Collection): T1557 (Adversary-in-the-Middle).**
[`sigma/artex_recording_proxy_ca.yml`](../sigma/artex_recording_proxy_ca.yml)이며, 워커 도구의 트래픽을
복호화·기록하려고 ARTEX 내장 트래픽 기록기(`traffic/traffic.go`)가 설치하는 MITM 루트 CA 아티팩트를
겨냥한 호스트·포렌식 헌팅 단서입니다.
## 사용법
1. [ATT&CK Navigator](https://mitre-attack.github.io/attack-navigator/)를 엽니다.
2. **Open Existing Layer → Upload from local** 을 골라 `artex_navigator_layer.json` 을 선택합니다
(또는 이 저장소의 raw 파일 URL 을 가리킵니다).
3. 점수를 매긴 기법이 탐지 강도에 따라 색으로 구분되어 나타나고, 각 기법에는 근거가 된 규칙 파일과
방어 가이드 절을 적은 주석이 붙어 있습니다.
## 범위와 정직함
- **커버리지는 완전성이 아닙니다.** 여기서 점수를 받은 기법은 규칙이 그것을 태그한다는 뜻이지, 그
기법의 모든 변형을 탐지한다는 뜻이 아닙니다. 네트워크 선에서 ARTEX 고유 User-Agent 로 잡히는 신호는
정찰 단계의 보강 프로버(`artex-enrich/1.0`)와 공격 단계의 norma SDK WebFetch(`norma/0.4`) 둘뿐이고,
그 밖의 공격 트래픽은 도구 기본 지문을 따릅니다. 오래가는 탐지는 행동 기반입니다
(방어 가이드 [한국어](../../docs/defense-ko.md) · [English](../../docs/defense-en.md) 1~2절·4.1~4.2절 참조). 순수 웹 다단계 사례는 여전히
환경별 기본 규칙이 필요합니다.
- **정적 지표는 바꿀 수 있습니다.** 운영자가 User-Agent 를 다른 값으로 설정할 수 있으므로, 태그된
지표가 없다고 해서 안전하다는 뜻은 아닙니다. 규칙 파일에도 같은 유의점을 달아 두었습니다.
## 검증과 기여
[일관성 테스트](../tests/attack/run.sh)를 돌리십시오. Docker 만 있으면 되며, 레이어가 점수를 매긴
기법·전술이 정확히 규칙의 `attack.*` 태그와 같은지, 그리고 모든 기법이 실재하는 규칙 파일에
근거하는지 단언합니다.
```sh
detections/tests/attack/run.sh
```
규칙을 추가하거나 다시 태그하면 이 레이어도 맞춰 갱신하십시오. 규칙의 기법이 레이어에 없거나
레이어의 기법이 규칙에 없으면 테스트가 실패합니다. [`../README.ko.md`](../README.ko.md)와
[`../../CONTRIBUTING.md`](../../CONTRIBUTING.md)를 참조하십시오.
+90
View File
@@ -0,0 +1,90 @@
# ARTEX ATT&CK coverage
English · [한국어](README.ko.md)
> 한국어: 이 디렉터리는 [`../`](../)의 ARTEX 탐지 규칙(Sigma·Suricata)이 다루는 공격 기법을
> [MITRE ATT&CK](https://attack.mitre.org/) 전술·기법으로 정리한 **커버리지 레이어**입니다.
> 각 기법은 저장소 소스에 근거가 있는 규칙의 `attack.*` 태그에서만 가져왔고, 추정으로 넣은 항목은
> 없습니다. [ATT&CK Navigator](https://mitre-attack.github.io/attack-navigator/)에 그대로 올려
> 어떤 ARTEX 행위에 어떤 규칙이 걸리는지 한눈에 볼 수 있습니다. 이 레이어는 자신이 소유하거나 서면
> 허가를 받은 시스템을 지키는 **방어·탐지 목적에만** 쓰십시오. 한국어 전체 문서는
> **[README.ko.md](README.ko.md)** 를 보십시오.
A [MITRE ATT&CK](https://attack.mitre.org/) Navigator layer that maps the detection rules in this
repository to the ATT&CK (Enterprise) techniques they tag. It is built by hand from the `attack.*` tags on
the [Sigma rules](../sigma/) — every technique is grounded in a rule whose indicator is a string or
behaviour verified in this repository's source, and the [consistency test](../tests/attack/run.sh)
keeps the layer and the rules from drifting apart.
- **`artex_navigator_layer.json`** — the layer, in ATT&CK Navigator v4.5 format.
## What the score means
Coverage here means "this repository ships a detection that tags this technique", not "this technique is
fully covered". The score is deliberately honest about detection strength:
- **100 — ARTEX-specific signature or behaviour.** A static indicator unique to ARTEX (the
`artex-enrich/1.0` / `artex-selfupdate` User-Agents, the guard audit marker) or a behaviour rule built
on one (enrichment velocity / fan-out, guard-block burst).
- **50–65 — generic hunting lead.** Destructive-command hunting mirrored from the ARTEX guard deny list.
The same commands are run by legitimate administrators, so these fire on benign activity too; treat a
hit as a lead, not an attribution. 65 marks the case where a correlation rule raises specificity by
pairing the command with the ARTEX guard marker.
## Techniques covered
Eight techniques across six tactics. Each maps to the rule(s) that tag it:
- **Reconnaissance — T1595 (Active Scanning), T1592 (Gather Victim Host Information).**
[`sigma/artex_enrich_user_agent.yml`](../sigma/artex_enrich_user_agent.yml),
[`sigma/correlation/artex_enrich_scan_velocity.yml`](../sigma/correlation/artex_enrich_scan_velocity.yml),
[`sigma/correlation/artex_enrich_fanout.yml`](../sigma/correlation/artex_enrich_fanout.yml), and the
[Suricata rules](../suricata/artex.rules) (sid 1000001 / 1000002).
- **Command and Control — T1105 (Ingress Tool Transfer).**
[`sigma/artex_selfupdate_egress.yml`](../sigma/artex_selfupdate_egress.yml).
- **Execution — T1059 (Command and Scripting Interpreter).**
[`sigma/artex_guard_audit_framing.yml`](../sigma/artex_guard_audit_framing.yml),
[`sigma/correlation/artex_guard_block_burst.yml`](../sigma/correlation/artex_guard_block_burst.yml),
[`sigma/correlation/artex_guard_marker_then_destructive.yml`](../sigma/correlation/artex_guard_marker_then_destructive.yml).
- **Impact — T1485 (Data Destruction), T1561.002 (Disk Wipe: Disk Structure Wipe), T1489 (Service Stop).**
[`sigma/destructive_command_hunting.yml`](../sigma/destructive_command_hunting.yml), with T1485 also
reinforced by
[`sigma/correlation/artex_guard_marker_then_destructive.yml`](../sigma/correlation/artex_guard_marker_then_destructive.yml).
- **Credential Access / Collection — T1557 (Adversary-in-the-Middle).**
[`sigma/artex_recording_proxy_ca.yml`](../sigma/artex_recording_proxy_ca.yml) — the MITM root-CA artifact
ARTEX's embedded traffic recorder installs (`traffic/traffic.go`) to decrypt and log the worker tools'
traffic. A host/forensic hunting lead.
## How to use it
1. Open the [ATT&CK Navigator](https://mitre-attack.github.io/attack-navigator/).
2. Choose **Open Existing Layer → Upload from local**, and select `artex_navigator_layer.json` (or point
it at the raw file URL from this repository).
3. The scored techniques appear colour-graded by detection strength, each with a comment naming the rule
file(s) and the defense-guide section behind it.
## Scope and honesty
- **Coverage is not completeness.** A technique scored here means a rule tags it, not that every variant
of the technique is detected. Only two ARTEX-unique User-Agents are visible on the wire — the enrichment
prober (`artex-enrich/1.0`) in the reconnaissance phase and the norma SDK WebFetch tool (`norma/0.4`) in
the attack phase — while the rest of the attack traffic follows tool-default fingerprints; the durable
detection is behavioural
(see the defense guide, [Korean](../../docs/defense-ko.md) · [English](../../docs/defense-en.md), sections 1–2 and 4.1–4.2). The pure web multi-stage
case still needs base rules specific to your environment.
- **Static indicators can be changed.** An operator can set a different User-Agent, so the absence of a
tagged indicator does not imply safety. This is the same caveat the rule files carry.
## Validate and contribute
Run the [consistency test](../tests/attack/run.sh) — it needs only Docker and asserts that the layer's
scored techniques and tactics are exactly the `attack.*` tags on the rules, with every technique grounded
in a rule file that exists:
```sh
detections/tests/attack/run.sh
```
When you add or retag a rule, update this layer to match — the test fails if a rule technique is missing
from the layer or a layer technique is absent from the rules. See [`../README.md`](../README.md) and
[`../../CONTRIBUTING.en.md`](../../CONTRIBUTING.en.md).
@@ -0,0 +1,149 @@
{
"name": "ARTEX detection coverage",
"versions": {
"attack": "16",
"navigator": "5.1.0",
"layer": "4.5"
},
"domain": "enterprise-attack",
"description": "MITRE ATT&CK (Enterprise) coverage of the ARTEX detection rules in this repository (detections/sigma, detections/suricata). Every technique below is drawn from the attack.* tags of a rule whose indicator is grounded in this repository's source; nothing is inferred. Score reflects detection strength: 100 = ARTEX-specific signature or behaviour, 50-65 = generic hunting lead that also catches legitimate administration. Maintained by hand from those tags and checked for rule<->layer consistency by detections/tests/attack/run.sh.",
"filters": {
"platforms": [
"PRE",
"Windows",
"Linux",
"macOS",
"Network",
"Containers"
]
},
"sorting": 0,
"layout": {
"layout": "side",
"aggregateFunction": "average",
"showID": true,
"showName": true,
"showAggregateScores": false,
"countUnscored": false,
"expandedSubtechniques": "annotated"
},
"hideDisabled": false,
"techniques": [
{
"techniqueID": "T1595",
"tactic": "reconnaissance",
"score": 100,
"comment": "ARTEX asset-enrichment probe (User-Agent artex-enrich/1.0). sigma/artex_enrich_user_agent.yml; behaviour via sigma/correlation/artex_enrich_scan_velocity.yml and artex_enrich_fanout.yml; network via suricata sid 1000001/1000002. Defense guide section 2 (target view), 4.1, 4.4.",
"enabled": true,
"showSubtechniques": false
},
{
"techniqueID": "T1592",
"tactic": "reconnaissance",
"score": 100,
"comment": "ARTEX auto-enrichment gathers victim host info (DNS/HTTP, reads <title>) under User-Agent artex-enrich/1.0. sigma/artex_enrich_user_agent.yml; behaviour via sigma/correlation/artex_enrich_scan_velocity.yml and artex_enrich_fanout.yml; network via suricata sid 1000001/1000002. Defense guide section 2 (target view), 4.4.",
"enabled": true,
"showSubtechniques": false
},
{
"techniqueID": "T1105",
"tactic": "command-and-control",
"score": 100,
"comment": "ARTEX self-update egress (User-Agent artex-selfupdate) fetching a newer binary from a code-hosting host. sigma/artex_selfupdate_egress.yml. Operator/forensic, not target-side. Defense guide section 2 (operator view), 4.4.",
"enabled": true,
"showSubtechniques": false
},
{
"techniqueID": "T1059",
"tactic": "execution",
"score": 100,
"comment": "ARTEX platform-guard control marker written to the audit log on a blocked tool call. sigma/artex_guard_audit_framing.yml; behaviour via sigma/correlation/artex_guard_block_burst.yml and artex_guard_marker_then_destructive.yml. Forensic/host-side. Defense guide section 2 (operator view), 4.2.",
"enabled": true,
"showSubtechniques": false
},
{
"techniqueID": "T1485",
"tactic": "impact",
"score": 65,
"comment": "Destructive data commands (rm -rf, DROP DATABASE, FLUSHALL, ...) mirrored from the ARTEX guard deny list. Generic hunting via sigma/destructive_command_hunting.yml; specificity raised when co-occurring with the guard marker in sigma/correlation/artex_guard_marker_then_destructive.yml. Expect legitimate-admin false positives. Defense guide section 2 (operator view), 4.2.",
"enabled": true,
"showSubtechniques": false
},
{
"techniqueID": "T1561",
"tactic": "impact",
"comment": "Parent shown only to surface the scored subtechnique below.",
"enabled": true,
"showSubtechniques": true
},
{
"techniqueID": "T1561.002",
"tactic": "impact",
"score": 50,
"comment": "Disk-structure wipe commands (mkfs, dd of=/dev/, shred) from the ARTEX guard deny list. Generic hunting lead, not an ARTEX-specific signature. sigma/destructive_command_hunting.yml. Defense guide section 2 (operator view).",
"enabled": true,
"showSubtechniques": false
},
{
"techniqueID": "T1489",
"tactic": "impact",
"score": 50,
"comment": "Service/availability stop commands (kill -9 -1, killall -9, iptables -F, nft flush ruleset) from the ARTEX guard deny list. Generic hunting lead. sigma/destructive_command_hunting.yml. Defense guide section 2 (operator view).",
"enabled": true,
"showSubtechniques": false
},
{
"techniqueID": "T1557",
"tactic": "credential-access",
"score": 50,
"comment": "ARTEX embedded recording proxy is an adversary-in-the-middle: it installs a MITM root CA (_ca/mitmproxy-ca-cert.pem) to decrypt and log its worker tools' HTTP(S) traffic, including any credentials in transit. The CA file on a host evidences the recorder having run. sigma/artex_recording_proxy_ca.yml. Forensic/host-side; the bare CA filename is shared with standalone mitmproxy, so it is a hunting lead. Defense guide section 2 (operator view).",
"enabled": true,
"showSubtechniques": false
},
{
"techniqueID": "T1557",
"tactic": "collection",
"score": 50,
"comment": "ARTEX embedded recording proxy is an adversary-in-the-middle: it installs a MITM root CA (_ca/mitmproxy-ca-cert.pem) to decrypt and log (collect) all of its worker tools' HTTP(S) traffic. The CA file on a host evidences the recorder having run. sigma/artex_recording_proxy_ca.yml. Forensic/host-side; the bare CA filename is shared with standalone mitmproxy, so it is a hunting lead. Defense guide section 2 (operator view).",
"enabled": true,
"showSubtechniques": false
}
],
"gradient": {
"colors": [
"#f0f0f0",
"#ffe766",
"#1a9850"
],
"minValue": 0,
"maxValue": 100
},
"legendItems": [
{
"label": "ARTEX-specific signature or behaviour (high)",
"color": "#1a9850"
},
{
"label": "Generic hunting lead, also catches legit admin (medium)",
"color": "#ffe766"
}
],
"metadata": [
{
"name": "repository",
"value": "https://github.com/jiwoochris/artex-ko"
},
{
"name": "rules",
"value": "detections/sigma (9), detections/suricata (2)"
},
{
"name": "consistency-test",
"value": "detections/tests/attack/run.sh"
}
],
"showTacticRowBackground": true,
"tacticRowBackground": "#205b8f",
"selectTechniquesAcrossTactics": true,
"selectSubtechniquesWithParent": false
}
+105
View File
@@ -0,0 +1,105 @@
# ARTEX 침해지표 (기계가 읽는)
한국어 · [English](README.md)
ARTEX 가 스스로 내보내는 고유 지문을 한 파일로 모은, 기계가 읽는 목록입니다. 탐지 로직이 아니라
원자 지표 자체를 원하는 방어자를 위한 것입니다: [`artex_indicators.csv`](artex_indicators.csv)를
위협 인텔리전스 플랫폼이나 SIEM 조회 테이블, 호스트 분류(triage) 체크리스트에 바로 넣으십시오.
모든 값은 이 저장소 소스에서 확인한 문자열입니다. [탐지 규칙](../README.ko.md)과 방어 가이드
([한국어](../../docs/defense-ko.md) · [English](../../docs/defense-en.md) 2절)가 근거로 삼는 바로
그 문자열입니다. 각 행은 그 값이 어디서 오는지와 (있다면) 그 위에 세운 규칙을 적습니다.
## 열 구성
- **`id`**: 지표의 안정적인 슬러그입니다.
- **`type`**: 지표의 종류입니다: `http.user-agent`, `string`(로그·파일에서 찾을 리터럴), `port`,
`ip-dst|port`, `other`(위 범주에 들지 않는 호스트 아티팩트, 예: 데이터베이스 스키마 객체 이름).
이들은 대응하는 MISP/STIX 속성 타입에 매핑됩니다.
- **`value`**: 정확한 지표입니다. 비 ASCII 가드 마커를 포함해 원문 그대로 보존합니다.
- **`perspective`**: `target`(ARTEX 가 탐침하는 시스템을 *향하는* 트래픽에서 관측) 또는
`forensic`(ARTEX 가 실행됐거나 경유한 호스트 *위에서* 관측)입니다. 방어 가이드는 이 둘을 일부러
구분합니다. 섞으면 틀린 결론이 나옵니다.
- **`source`**: 그 값을 내보내는, 저장소 기준 상대 경로 소스 파일입니다(`;` 로 구분). 이것이
근거입니다: 상류 재동기화가 내보내는 쪽을 바꾸면 여기 지표도 함께 바뀌어야 합니다.
- **`rule`**: 그 정확한 값 위에 세운 탐지 규칙입니다(`;` 로 구분). (시끄러운) 규칙으로 내보내지
않고 직접 분류하는 호스트 포렌식 지표는 비어 있습니다.
- **`description`**: 한 줄 설명이며, 해당하는 경우 정직한 유의점을 함께 적습니다.
## MISP 이벤트 내보내기
같은 지표를 바로 가져올 수 있는 [MISP](https://www.misp-project.org/) 이벤트
[`artex_indicators.misp.json`](artex_indicators.misp.json)로도 제공합니다. MISP 인스턴스를 운영하는
(또는 MISP 형식을 적재하는 위협 인텔리전스 플랫폼을 쓰는) 방어자는 CSV 열을 손으로 매핑하지 않고
지문을 바로 가져올 수 있습니다. STIX 2.1 은 MISP 자체 변환기로 한 번 내보내면 되므로, 저장소가
손실 있는 두 번째 형식을 따로 만들지 않습니다.
- **타입 매핑.** 각 CSV `type` 은 대응하는 MISP 속성 타입이 됩니다: `http.user-agent` →
`user-agent`, 가드 마커 `string` → `pattern-in-file`(카테고리 *Artifacts dropped*), `port` →
`port`, `ip-dst|port` → `ip-dst|port`(합성 값은 MISP 의 `ip|port` 형식을 쓰므로 `127.0.0.1:8788`
은 `127.0.0.1|8788` 로 저장됩니다), 탐색 그래프 스키마 지문 `other` → `other`(카테고리 *Other*).
- **`to_ids` 는 `rule` 열을 정직하게 따릅니다.** 탐지 규칙이 세워진 행은 조치 가능한 지표이므로
`to_ids: true` 로 표시합니다. 규칙이 없는 호스트 포렌식 행(기본 수신 포트와 루프백 프록시
엔드포인트, 그리고 탐색 그래프 스키마 지문)은 차단용 IoC 가 아니라 분류 힌트이므로
`to_ids: false` 에 `disable_correlation: true` 로 둡니다(흔한 포트나 `127.0.0.1`, 또는 범용 테이블
이름이 MISP 상관을 오염시키면 안 됩니다). 이는 CSV 의 `rule` 열과 아래 유의점이 이미 담고
있는 것과 같은 구분입니다.
- **가져오기.** [pymisp](https://github.com/MISP/PyMISP)로
`MISPEvent().load_file("artex_indicators.misp.json")`, 또는 *Add event → Populate from … → MISP
format* UI, 또는 REST API 로 가져옵니다. 이벤트는 미발행 상태이고 `tlp:clear` 태그가 붙어
있습니다. 가져올 때 인스턴스에 맞는 배포 범위와 발행 상태를 설정하십시오.
## 이 목록을 정직하게 읽는 법
- **이것들은 바뀔 수 있는 지문이지 안전의 증거가 아닙니다.** 운영자가 User-Agent 를 다른 값으로
설정하거나 기본 포트를 바꿀 수 있으므로, 여기 있는 어떤 값이 *없다고* 해서 ARTEX 가 없다는 뜻은
**아닙니다**. 오래가는 신호는 행동입니다. 상관 규칙과 방어 가이드 1·2·4.1~4.2절을 참조하십시오.
- **일반 헌팅 단서는 의도적으로 뺐습니다.** 파괴적 셸·DB 명령(`rm -rf`, `DROP DATABASE`, …)은 ARTEX
지문이 *아닙니다*. 정당한 관리자도 실행합니다. 가져오기용 지표가 아니라 헌팅 단서이므로, 이
목록이 아니라 [`destructive_command_hunting.yml`](../sigma/destructive_command_hunting.yml)과 방어
가이드에 둡니다. 이것들을 차단용 지표로 가져오면 오탐이 생깁니다.
- **norma 의 WebFetch User-Agent 는 네트워크 서명이지 가져오기용 원자 지표가 아닙니다.** 워커의
페이지 가져오기 도구는 공격 단계에서 `norma/0.4` 를 보내고 Suricata 규칙 sid 1000003 이 `norma/`
접두사에 발화하지만, 그 문자열은 norma SDK 에 하드코딩된 자체 User-Agent(`github.com/Autumn-27/norma/tool/webfetch.go`)여서
norma 위에 세운 모든 도구가 똑같이 내보내는 값이지 ARTEX 고유 지문이 아닙니다. 이 값을 차단용
지표로 이 목록에 넣으면 모든 norma SDK 트래픽에 경보가 울리는데, 이는 파괴적 명령을 뺀 것과 같은
오탐 함정입니다. 그래서 norma UA 는 이 목록에서 의도적으로 빼고 네트워크 규칙으로만 싣습니다
([`../suricata/README.ko.md`](../suricata/README.ko.md), sid 1000003). 또한 이 값은 이 저장소의
소스가 아니라 고정된 의존성(`github.com/Autumn-27/norma`)에 근거를 두므로, 아래의 근거 테스트가
ARTEX 자신이 내보내는 문자열을 다시 읽듯이 이 값을 다시 읽을 수는 없습니다.
- **호스트 포렌식 포트는 차단이 아니라 분류용입니다.** `:8787` 과 `127.0.0.1:8788` 은 ARTEX 를
돌리고 있을 수 있는 호스트를 가리킵니다. `ss`·`netstat` 로 확인하고, 맹목적으로 방화벽을 걸지
마십시오.
- **탐색 그래프 스키마 지문은 네트워크·파일 IoC 가 아니라 DB 조사용입니다.** `exploration_nodes`
테이블은 ARTEX 가 PostgreSQL 에 두는 탐색 그래프의 핵심 테이블입니다. 단독 적중으로 단정하지
말고, 형제 테이블(`exploration_edges`·`exploration_anchors`·`assets`·`companies`·`activity`)과
`agent_prompts` 시드가 같은 데이터베이스에 함께 있는지로 확인하십시오. 운영자가 테이블을 바꾸거나
지울 수 있으므로 부재가 안전을 뜻하지는 않습니다.
- **`rule` 이 비어 있는 호스트·DB 행에는 실행기가 있습니다.** Sigma 규칙으로 싣지 않고 직접 분류하는 세
지표, 곧 리슨 포트·기록 프록시 엔드포인트·이 스키마 지문은 [호스트 분류 스크립트](../triage/)가 의심
호스트에서 모두 점검합니다. 그래서 셸 접근은 있으나 SIEM 이 없는 대응자가 `ss`·`netstat`·`psql`
을 손으로 돌리지 않아도 됩니다.
## 검증
이 목록은 [지표 근거(source-of-truth) 테스트](../tests/indicators/run.sh)가 덮습니다. 이 CSV 를 다시
읽어 모든 행에 대해, 그 값이 인용한 소스 파일에 여전히 있고 인용한 규칙에 고정돼 있는지, 그리고
테스트가 근거로 삼는 모든 지표가 목록에 나타나는지 단언합니다. 소스에서 어긋난 행이나 목록에서 빠진
알려진 지문이 있으면 테스트가 실패합니다. 다음으로 돌리십시오.
```sh
detections/tests/indicators/run.sh
```
MISP 이벤트는 자체 [MISP 내보내기 일관성 테스트](../tests/misp/run.sh)가 덮습니다. 이벤트를 pymisp
로 적재해(모든 속성 타입이 서버가 받아들이는 실재 MISP 타입이 되도록) 이 CSV 와 행 단위로
동기화됨을 단언합니다. 곧 같은 값, 의도한 타입·카테고리, 그리고 `rule` 열에 맞춘 `to_ids` 플래그입니다.
이벤트는 CSV 와 나란히 손으로 관리합니다. CSV 가 지니지 않는, 속성별로 정리한 주석·안정적인
UUID·이벤트 수준 태그도 함께 지니므로, 손실 있는 기본값으로 이것들을 덮어쓸 생성기가 없습니다.
CSV 행을 더하거나 빼거나 타입을 바꿀 때는 같은 커밋에서
[`artex_indicators.misp.json`](artex_indicators.misp.json)도 맞춰 고치십시오(새 속성에는 새 `uuid` 와
근거가 되는 `comment` 를 주십시오). 둘이 일치할 때까지 이 테스트가 실패하므로, 갱신을 조용히 잊을
수 없습니다. 다음으로 돌리십시오.
```sh
detections/tests/misp/run.sh
```
+116
View File
@@ -0,0 +1,116 @@
# ARTEX indicators (machine-readable)
English · [한국어](README.ko.md)
> 한국어: [`artex_indicators.csv`](artex_indicators.csv) 는 ARTEX 가 실제로 내보내는 고유 지문(침해지표,
> IoC)을 한 파일로 모은 것입니다. 위협 인텔리전스 플랫폼·SIEM 조회 테이블·호스트 분류 작업에 바로
> 넣을 수 있게 기계가 읽는 CSV 로 둡니다. 모든 값은 이 저장소 소스에서 확인한 문자열이며, 각 행의
> 출처 파일과 탐지 규칙을 함께 적습니다. 배경 설명은 [방어·탐지 가이드(docs/defense-ko.md)](../../docs/defense-ko.md)
> 2절 "방어자가 관측할 수 있는 지문"에 있습니다. 자신이 소유하거나 서면 허가를 받은 시스템을 지키는
> **방어·탐지 목적에만** 사용하십시오. 한국어 전체 문서는 **[README.ko.md](README.ko.md)** 를
> 보십시오.
A single, machine-readable list of the unique fingerprints ARTEX itself emits, for defenders who want the
atomic indicators rather than the detection logic: drop [`artex_indicators.csv`](artex_indicators.csv)
into a threat-intelligence platform, a SIEM lookup table, or a host-triage checklist. Every value is a
string verified in this repository's source — the same grounding the
[detection rules](../README.md) and the defense guide
([Korean](../../docs/defense-ko.md) · [English](../../docs/defense-en.md), section 2) rely on — and each
row records where it comes from and which rule (if any) is built on it.
## Columns
- **`id`** — a stable slug for the indicator.
- **`type`** — the kind of indicator: `http.user-agent`, `string` (a literal to hunt for in logs/files),
`port`, `ip-dst|port`, or `other` (a host artifact that fits none of the above, e.g. a database schema
object name). These map onto the equivalent MISP/STIX attribute types.
- **`value`** — the exact indicator. Preserved verbatim, including the non-ASCII guard marker.
- **`perspective`** — `target` (observable in traffic *toward* a system ARTEX probes) or `forensic`
(observable *on* a host where ARTEX ran or was relayed through). The defense guide keeps these apart on
purpose; mixing them produces false conclusions.
- **`source`** — the repository-relative source file(s) that emit the value, `;`-separated. This is the
grounding: if an upstream re-sync changes the emitter, the indicator here must change with it.
- **`rule`** — the detection rule(s) built on the exact value, `;`-separated, or empty for host-forensic
indicators that are triaged directly rather than shipped as a (noisy) rule.
- **`description`** — a one-line note, including the honest caveat where one applies.
## MISP event export
The same indicators ship as a ready-to-import [MISP](https://www.misp-project.org/) event,
[`artex_indicators.misp.json`](artex_indicators.misp.json), so a defender running a MISP instance (or a
threat-intelligence platform that ingests the MISP format) can import the fingerprints directly instead of
mapping the CSV columns by hand. STIX 2.1 is then one export away using MISP's own converter, so the
repository does not hand-roll a second, lossy format.
- **Type mapping.** Each CSV `type` becomes the equivalent MISP attribute type: `http.user-agent` →
`user-agent`, the guard marker `string` → `pattern-in-file` (category *Artifacts dropped*), `port` →
`port`, `ip-dst|port` → `ip-dst|port` (the composite value uses MISP's `ip|port` form, so
`127.0.0.1:8788` is stored as `127.0.0.1|8788`), and the exploration-graph schema fingerprint `other` →
`other` (category *Other*).
- **`to_ids` follows the `rule` column, honestly.** A row that a detection rule is built on is an actionable
indicator and is flagged `to_ids: true`. A host-forensic row with no rule — the default listen port, the
loopback proxy endpoint, and the exploration-graph schema fingerprint — is a triage hint, not a blocking
IoC, so it is `to_ids: false` with `disable_correlation: true` (a common port, `127.0.0.1`, or a generic
table name should not pollute MISP correlations). This is the same distinction the CSV `rule` column and
the caveats below already carry.
- **Import.** `MISPEvent().load_file("artex_indicators.misp.json")` with
[pymisp](https://github.com/MISP/PyMISP), the *Add event → Populate from … → MISP format* UI, or the REST
API. The event is unpublished and tagged `tlp:clear`; set the distribution and publish state your instance
needs on import.
## How to read this honestly
- **These are changeable fingerprints, not proof of safety.** An operator can set a different User-Agent
or change a default port, so the *absence* of any value here does **not** mean ARTEX is absent. The
durable signal is behaviour — see the correlation rules and sections 1, 2, and 4.1–4.2 of the defense
guide.
- **Generic hunting leads are deliberately excluded.** Destructive shell/DB commands (`rm -rf`, `DROP
DATABASE`, …) are *not* ARTEX fingerprints — legitimate administrators run them too. They are a hunting
lead, not an import-ready indicator, so they live in
[`destructive_command_hunting.yml`](../sigma/destructive_command_hunting.yml) and the defense guide, not
in this list. Importing them as blocking indicators would cause false positives.
- **The norma WebFetch User-Agent is a wire signature, not an atomic indicator.** The worker's page-fetch
tool sends `norma/0.4` during the attack phase, and Suricata sid 1000003 fires on the `norma/` prefix, but
that string is the norma SDK's own hardcoded User-Agent (`github.com/Autumn-27/norma/tool/webfetch.go`), shared by every tool built on
norma rather than an ARTEX-unique fingerprint. Importing it here as a blocking indicator would alert on all
norma-SDK traffic — the same false-positive trap the destructive commands sit in — so it is deliberately
kept out of this list and shipped only as the network rule
([`../suricata/README.md`](../suricata/README.md), sid 1000003). It is also grounded in a pinned dependency
(`github.com/Autumn-27/norma`), not this repository's own source, so the source-of-truth test below cannot
re-read it the way it re-reads ARTEX's own emitters.
- **Host-forensic ports are for triage, not blocking.** `:8787` and `127.0.0.1:8788` describe a host that
may be running ARTEX; check them with `ss`/`netstat`, do not firewall them blindly.
- **The exploration-graph schema fingerprint is for DB inspection, not a network/file IoC.** The
`exploration_nodes` table is the core of the exploration graph ARTEX keeps in PostgreSQL. Do not conclude
from a single hit; confirm that the sibling tables (`exploration_edges`, `exploration_anchors`, `assets`,
`companies`, `activity`) and the `agent_prompts` seed sit in the same database. An operator can rename or
drop tables, so absence does not mean safety.
- **The host/DB rows with no `rule` have a runner.** The three indicators triaged directly rather than
shipped as a Sigma rule — the listen ports, the recording-proxy endpoint, and this schema fingerprint —
are all checked by the [host-triage script](../triage/) on a suspected host, so a responder with
shell access but no SIEM does not have to run `ss`/`netstat`/`psql` by hand.
## Verification
The list is covered by the [indicator source-of-truth test](../tests/indicators/run.sh): it re-reads this
CSV and asserts, for every row, that the value is still present in the cited source file(s) and pinned in
the cited rule(s), and that every indicator the test grounds appears in the list. A row that drifts from
the source, or a known fingerprint dropped from the list, fails the test. Run it with:
```sh
detections/tests/indicators/run.sh
```
The MISP event is covered by its own [MISP export consistency test](../tests/misp/run.sh): it loads the
event under pymisp (so every attribute type is a real MISP type a server accepts) and asserts it stays
row-for-row in sync with this CSV — same values, the intended type/category, and the `to_ids` flag matching
the `rule` column. The event is maintained by hand alongside the CSV — it also carries curated per-attribute
comments, stable UUIDs, and event-level tags that the CSV does not hold, so there is no generator that would
overwrite them with lossy defaults. When you add, remove, or retype a CSV row, edit
[`artex_indicators.misp.json`](artex_indicators.misp.json) to match in the same commit (give a new attribute a
fresh `uuid` and a grounding `comment`); this test fails until the two agree, so the update cannot be silently
forgotten. Run it with:
```sh
detections/tests/misp/run.sh
```
@@ -0,0 +1,8 @@
id,type,value,perspective,source,rule,description
enrich-user-agent,http.user-agent,artex-enrich/1.0,target,enrich/enrich.go,detections/sigma/artex_enrich_user_agent.yml,"HTTP User-Agent of ARTEX asset-enrichment probes (short single GET, no redirects, reads only the title). Suricata also matches it by the prefix artex-enrich/. An operator can change it, so absence is not safety."
selfupdate-user-agent,http.user-agent,artex-selfupdate,forensic,selfupdate/github.go;selfupdate/stage.go,detections/sigma/artex_selfupdate_egress.yml,"HTTP User-Agent of the self-update egress call to the release host; seen in outbound logs from a host running ARTEX."
guard-audit-marker,string,【ARTEX 平台管控·非目标防御】,forensic,guard/guard.go,detections/sigma/artex_guard_audit_framing.yml,"Control-framing prefix written to the audit log on a blocked tool call; its presence in audit records supports an ARTEX-execution finding."
server-listen-port,port,8787,forensic,cmd/artex/main.go,,"Default ARTEX server HTTP listen port (flag --addr). An internal host serving its admin UI here warrants triage; best checked on the host with ss or netstat, not as a network rule."
recording-proxy-endpoint,ip-dst|port,127.0.0.1:8788,forensic,cmd/artex/main.go,,"Default loopback traffic-recording MITM proxy endpoint (flag --proxy). Check with ss or netstat on a suspected host."
recording-proxy-ca,string,mitmproxy-ca-cert.pem,forensic,traffic/traffic.go,detections/sigma/artex_recording_proxy_ca.yml,"MITM CA certificate file the ARTEX recording proxy writes on first start (traffic/traffic.go, under <dir>/_ca/); injected into spawned worker tools via SSL_CERT_FILE/CURL_CA_BUNDLE/REQUESTS_CA_BUNDLE/NODE_EXTRA_CA_CERTS with a loopback HTTP_PROXY. Its presence on a host evidences the recorder having run. The bare filename is shared with standalone go-mitmproxy/mitmproxy, so treat it as a host-triage lead, not a unique fingerprint."
postgres-exploration-schema,other,exploration_nodes,forensic,db/schema.sql,,"ARTEX exploration-graph table in its PostgreSQL store (db/schema.sql). With exploration_edges/exploration_anchors/assets/companies/activity and an agent_prompts seed it forms the ARTEX dual-graph schema; their presence together is a strong host-forensic tell. A triage lead checked by inspecting the database, not a network or file IoC, so to_ids is off."
1 id type value perspective source rule description
2 enrich-user-agent http.user-agent artex-enrich/1.0 target enrich/enrich.go detections/sigma/artex_enrich_user_agent.yml HTTP User-Agent of ARTEX asset-enrichment probes (short single GET, no redirects, reads only the title). Suricata also matches it by the prefix artex-enrich/. An operator can change it, so absence is not safety.
3 selfupdate-user-agent http.user-agent artex-selfupdate forensic selfupdate/github.go;selfupdate/stage.go detections/sigma/artex_selfupdate_egress.yml HTTP User-Agent of the self-update egress call to the release host; seen in outbound logs from a host running ARTEX.
4 guard-audit-marker string 【ARTEX 平台管控·非目标防御】 forensic guard/guard.go detections/sigma/artex_guard_audit_framing.yml Control-framing prefix written to the audit log on a blocked tool call; its presence in audit records supports an ARTEX-execution finding.
5 server-listen-port port 8787 forensic cmd/artex/main.go Default ARTEX server HTTP listen port (flag --addr). An internal host serving its admin UI here warrants triage; best checked on the host with ss or netstat, not as a network rule.
6 recording-proxy-endpoint ip-dst|port 127.0.0.1:8788 forensic cmd/artex/main.go Default loopback traffic-recording MITM proxy endpoint (flag --proxy). Check with ss or netstat on a suspected host.
7 recording-proxy-ca string mitmproxy-ca-cert.pem forensic traffic/traffic.go detections/sigma/artex_recording_proxy_ca.yml MITM CA certificate file the ARTEX recording proxy writes on first start (traffic/traffic.go, under <dir>/_ca/); injected into spawned worker tools via SSL_CERT_FILE/CURL_CA_BUNDLE/REQUESTS_CA_BUNDLE/NODE_EXTRA_CA_CERTS with a loopback HTTP_PROXY. Its presence on a host evidences the recorder having run. The bare filename is shared with standalone go-mitmproxy/mitmproxy, so treat it as a host-triage lead, not a unique fingerprint.
8 postgres-exploration-schema other exploration_nodes forensic db/schema.sql ARTEX exploration-graph table in its PostgreSQL store (db/schema.sql). With exploration_edges/exploration_anchors/assets/companies/activity and an agent_prompts seed it forms the ARTEX dual-graph schema; their presence together is a strong host-forensic tell. A triage lead checked by inspecting the database, not a network or file IoC, so to_ids is off.
@@ -0,0 +1,84 @@
{
"Event": {
"uuid": "1a5aa723-0e87-4cbe-97f4-c84f93e4efeb",
"info": "ARTEX (autonomous AI pentest framework) — defensive host/network fingerprints",
"date": "2026-10-07",
"threat_level_id": "4",
"analysis": "2",
"distribution": "3",
"published": false,
"Orgc": {
"name": "artex-ko",
"uuid": "fff4e6e6-a076-433d-9a12-873ff60de4e6"
},
"Tag": [
{ "name": "tlp:clear" },
{ "name": "type:OSINT" }
],
"Attribute": [
{
"uuid": "cb9d8f4e-cef3-44a4-8499-457620861b74",
"type": "user-agent",
"category": "Network activity",
"to_ids": true,
"disable_correlation": false,
"value": "artex-enrich/1.0",
"comment": "ARTEX asset-enrichment prober User-Agent (enrich/enrich.go). Target-side. An operator can change it, so absence is not safety. Sigma: artex_enrich_user_agent.yml."
},
{
"uuid": "dbe975a6-0a12-43a7-a8e8-4bd0ba65daea",
"type": "user-agent",
"category": "Network activity",
"to_ids": true,
"disable_correlation": false,
"value": "artex-selfupdate",
"comment": "ARTEX self-update egress User-Agent to the release host (selfupdate/github.go, selfupdate/stage.go). Seen in outbound logs from an ARTEX host. Sigma: artex_selfupdate_egress.yml."
},
{
"uuid": "053cbbbd-c345-49f0-b6e3-1b6f6075a218",
"type": "pattern-in-file",
"category": "Artifacts dropped",
"to_ids": true,
"disable_correlation": false,
"value": "【ARTEX 平台管控·非目标防御】",
"comment": "Control-framing prefix ARTEX writes to its audit log on a blocked tool call (guard/guard.go). Its presence in audit records supports an ARTEX-execution finding. Sigma: artex_guard_audit_framing.yml."
},
{
"uuid": "d5fe7761-c297-4e35-acfe-a3a4a5f01f7b",
"type": "port",
"category": "Network activity",
"to_ids": false,
"disable_correlation": true,
"value": "8787",
"comment": "Default ARTEX server HTTP listen port (cmd/artex/main.go --addr). Host-triage hint, not a blocking indicator; check with ss/netstat on a suspected host."
},
{
"uuid": "b575c98a-629d-42d4-bac0-3280f6c6c6b8",
"type": "ip-dst|port",
"category": "Network activity",
"to_ids": false,
"disable_correlation": true,
"value": "127.0.0.1|8788",
"comment": "Default loopback traffic-recording MITM proxy endpoint (cmd/artex/main.go --proxy). Loopback — a host-triage hint, not a network block; check with ss/netstat."
},
{
"uuid": "7628f0dc-d5f8-45ea-8aec-64843667658e",
"type": "pattern-in-file",
"category": "Artifacts dropped",
"to_ids": true,
"disable_correlation": false,
"value": "mitmproxy-ca-cert.pem",
"comment": "MITM CA certificate file the ARTEX recording proxy writes on first start (traffic/traffic.go, under <dir>/_ca/); injected into spawned tools via SSL_CERT_FILE/CURL_CA_BUNDLE/REQUESTS_CA_BUNDLE/NODE_EXTRA_CA_CERTS with a loopback HTTP_PROXY. Evidences the recorder having run; the bare filename is shared with standalone mitmproxy, so it is a host-triage lead. Sigma: artex_recording_proxy_ca.yml."
},
{
"uuid": "134f14d2-0af0-4a4b-89bb-805ab5f2b1a7",
"type": "other",
"category": "Other",
"to_ids": false,
"disable_correlation": true,
"value": "exploration_nodes",
"comment": "ARTEX exploration-graph table in its PostgreSQL store (db/schema.sql); with exploration_edges/exploration_anchors/assets/companies/activity and an agent_prompts seed it forms the ARTEX dual-graph schema. Host-triage lead checked by inspecting the database, not a blocking IoC."
}
]
}
}
+124
View File
@@ -0,0 +1,124 @@
# ARTEX 탐지 규칙 (Sigma / 호스트·로그·SIEM)
한국어 · [English](README.md)
이 디렉터리는 ARTEX 탐지 묶음에서 호스트·로그·SIEM 계층을 맡습니다. 여기 실린
[Sigma](https://sigmahq.io) 규칙은 방어 가이드([한국어](../../docs/defense-ko.md) ·
[English](../../docs/defense-en.md)) 4절의 의사 규칙을 벤더 중립 형식으로 정식화한 것이고, 각자의
SIEM·EDR 질의 언어로 변환해 씁니다. 모든 지표는 추정이 아니라 이 저장소 소스에서 실제로 확인한
문자열이나 행동에 근거합니다. 네트워크 계층은 [`../suricata/`](../suricata/)에 있고, ATT&CK 레이어와
지표 CSV·MISP 내보내기, 호스트 분류 스크립트를 포함한 전체 탐지 묶음은
[`../README.ko.md`](../README.ko.md)가 색인합니다. 모든 규칙은 자신이 소유하거나 서면 허가를 받은
시스템을 지키는 **방어·탐지 목적에만** 사용하십시오.
## 원자(atomic) 규칙
규칙 하나가 관측 가능한 사실 하나에 대응합니다. 개별로 변환해도 되고, 트리 전체의 일부로 변환해도
됩니다.
- **[`artex_enrich_user_agent.yml`](artex_enrich_user_agent.yml)**: *ARTEX Asset Enrichment Probe
User-Agent*. 자산 보강(`enrich/enrich.go`)이 보내는 인바운드 `artex-enrich/1.0` User-Agent 입니다.
대상 측에서 관측하는 보조 지표입니다. `level: high`.
- **[`artex_selfupdate_egress.yml`](artex_selfupdate_egress.yml)**: *ARTEX Self-Update Egress
User-Agent*. 자가 업데이트 루틴(`selfupdate/github.go`)이 내보내는 아웃바운드 `artex-selfupdate`
User-Agent 입니다. 호스트·포렌식 egress 지표입니다. `level: medium`.
- **[`artex_guard_audit_framing.yml`](artex_guard_audit_framing.yml)**: *ARTEX Platform Guard
Audit-Log Framing*. 도구 호출이 차단될 때 감사 로그에 기록되는 플랫폼 가드 통제 마커입니다
(`guard/guard.go`). 호스트·포렌식 지표입니다. `level: high`.
- **[`artex_recording_proxy_ca.yml`](artex_recording_proxy_ca.yml)**: *ARTEX Recording-Proxy MITM CA
Certificate Artifact*. 기록 프록시가 `_ca/mitmproxy-ca-cert.pem` 배치로 생성하는 MITM CA 파일입니다
(`traffic/traffic.go`). 호스트·포렌식 산출물이며, 파일명 자체는 단독 실행 mitmproxy 와도 공유되므로
사냥 단서(hunting lead)로 취급합니다. `level: medium`.
- **[`destructive_command_hunting.yml`](destructive_command_hunting.yml)**: *Destructive Command
Execution (ARTEX Guard-List Hunting)*. ARTEX 가드의 내장 거부 목록(`db/db.go` 시드)을 반영한 파괴적
셸·DB 명령입니다. ARTEX 고유 시그니처가 **아니라** 일반 사냥 단서입니다. `level: medium`.
## 상관(correlation) 규칙 (행동 기반) · [`correlation/`](correlation/)
정적 문자열은 바꿀 수 있지만 행동은 숨기기가 더 어렵습니다. 이 Sigma **상관** 규칙들은 방어 가이드
4.1~4.2절과 4.4절의 행동 기반 계층을 정식화합니다. 각 규칙은 위의 원자 규칙 하나를 `id` 로 참조하므로,
상관 파일 하나가 아니라 **`sigma/` 트리 전체를 변환해야** 참조가 풀립니다([Sigma 테스트](../tests/sigma/)가
바로 이 의존 관계를 단언합니다).
- **[`correlation/artex_enrich_scan_velocity.yml`](correlation/artex_enrich_scan_velocity.yml)**:
*Enrichment Scan Velocity*. 한 출처가 짧은 창 안에 `artex-enrich/1.0` 프로브를 몰아치는 경우입니다
(보강은 동시성 4로, 속도 제한 없이 돕니다). 단건 규칙이 놓치는 속도를 잡습니다. `event_count`,
`level: high`.
- **[`correlation/artex_enrich_fanout.yml`](correlation/artex_enrich_fanout.yml)**: *Enrichment
Fan-Out*. 한 출처가 보강 User-Agent 를 서로 다른 여러 호스트로 퍼뜨리는 경우입니다. 요청량이 아니라
접촉한 서로 다른 호스트 수가 신호이며, 자산 목록을 기계 속도로 훑는 폭을 잡습니다. `value_count`,
`level: high`.
- **[`correlation/artex_guard_block_burst.yml`](correlation/artex_guard_block_burst.yml)**:
*Guard-Block Burst*. 한 호스트에서 플랫폼 가드 통제 마커가 반복되는 경우입니다. 마커를 인용만 한
문서가 아니라, 실제로 가동 중인 ARTEX 실행이 자기 가드를 건드리는 상황을 가리킵니다. `event_count`,
`level: high`.
- **[`correlation/artex_guard_marker_then_destructive.yml`](correlation/artex_guard_marker_then_destructive.yml)**:
*Guard Marker With Destructive Command*. 가드 마커와 파괴적 명령이 한 호스트에서 한 창 안에 함께
나타나는 경우입니다(방어 가이드 4.2절, 다단계). ARTEX 고유 마커를 원래 일반적인 파괴적 명령 신호와
결합하므로 특이도가 올라갑니다. `temporal`, `level: high`.
임계값과 창은 보수적인 기본값이므로, 자신의 기준선에 맞게 조정하십시오. 순수 웹 다단계 경우(열거 →
프로빙 → 인증)는 여전히 환경별 기본 규칙이 필요합니다. 그 패턴은 ARTEX 고유 User-Agent 하나로
환원되지 않기 때문입니다. 시작점으로 쓸 일반 행동 기반 기본 템플릿은 [방어 가이드 4.2절](../../docs/defense-ko.md)에
있으며, ARTEX 소스에 근거를 둘 수 없어 이 검증된 트리에서는 의도적으로 뺐습니다.
## 범위와 정직함: 배포 전에 읽으십시오
- **정적 지표는 바꿀 수 있습니다.** 운영자가 User-Agent 를 바꾸거나 CA 파일을 지울 수 있으므로, 원자
지표가 없다고 해서 안전하다는 뜻은 **아닙니다**. 오래가는 신호는 `correlation/` 규칙이 기준으로 삼는
행동입니다. 한 출처가 정찰에서 열거, 프로빙, 인증·주입 시도로 이어 가며, 응답에 적응하고, 쉬지 않고
도는 흐름이 그것입니다.
- **파괴적 명령 규칙은 일반 사냥입니다.** ARTEX 가드 거부 목록을 반영하지만, 같은 명령을 정당한
관리자도 실행합니다. 적중은 단서로 다루고, 자신의 환경을 허용 목록으로 걸러 내며, 그것만으로 ARTEX
라고 단정하지 마십시오.
- **포트와 스키마는 네트워크가 아니라 호스트 포렌식입니다.** 서버 기본 포트 `:8787` 과 기록 프록시
`127.0.0.1:8788`(`cmd/artex/main.go`), 그리고 PostgreSQL 탐색 그래프 스키마는 의심 호스트에서 직접
확인하는 편이 낫습니다. 그래서 시끄러운 규칙 대신 [지표 CSV](../indicators/)와
[호스트 분류 스크립트](../triage/)로 제공합니다.
- **`logsource` 와 필드명은 일반값입니다.** 규칙은 일반 `category`·`product` 로그 소스와 필드명
(`cs-user-agent`, `CommandLine`, `TargetFilename`)을 씁니다. 변환 시 파이프라인(`-p`)으로 자신의
제품 스키마에 매핑하십시오. 아래 백엔드 설명을 참조하십시오.
## 검증과 변환
[sigma-cli](https://github.com/SigmaHQ/sigma-cli)(pySigma)로 검증했습니다. 저장소 루트에서 실행합니다.
```sh
python3 -m venv .venv && . .venv/bin/activate
pip install sigma-cli
# 구조 + 모범 사례 검증 (기대: 0 errors, 0 issues)
sigma check detections/sigma/
# 이 규칙 세트의 문서화된 기준선으로 SigmaHQ 관례 전체를 검사 (기대: 0 issues)
pip install pySigma-validators-sigmahq
sigma check --validation-config detections/tests/sigma_lint/validators.yml detections/sigma/
# 상관 규칙이 참조하는 원자 규칙을 풀 수 있도록 트리 전체를 변환
sigma plugin install splunk
sigma convert -t splunk --without-pipeline detections/sigma/
```
백엔드마다 상관 규칙 지원이 다르므로 `-t` 선택이 중요합니다. Splunk, Elasticsearch EQL, Grafana Loki 는
트리 전체를 변환하고, Elasticsearch Lucene, OpenSearch, 마이크로소프트 `kusto` 백엔드는 원자 규칙
다섯 개만 변환합니다(창은 제품에서 네이티브로 표현합니다). 백엔드별 실측 표와 `--without-pipeline` ·
`-p` 필드 매핑 설명은 [`../README.ko.md`](../README.ko.md)에 있고,
[`../tests/sigma_backends/`](../tests/sigma_backends/)가 재현합니다.
## 테스트
[`../tests/`](../tests/) 아래 재현 가능한 네 스위트가 이 규칙들을 다루며, 각각 Docker 만 있으면 됩니다.
[`sigma/`](../tests/sigma/)는 검증과 트리 전체 컴파일, 그리고 상관 규칙이 단독으로는 변환에 실패함을
단언하고, [`sigma_match/`](../tests/sigma_match/)는 규칙이 악성 샘플에 실제로 발화하고 양성 샘플에는
침묵하는지 확인하며, [`sigma_backends/`](../tests/sigma_backends/)는 백엔드 다섯 종의 이식성을,
[`sigma_lint/`](../tests/sigma_lint/)는 SigmaHQ 검증기 기준선 전체(0 issues)를 확인합니다.
[`../tests/README.ko.md`](../tests/README.ko.md)를 참조하십시오.
## 기여
탐지 기여를 환영합니다. 새 규칙은 모든 지표를 관측 가능한 사실에 근거해 두고, 한계를 `description` 에
밝히며, SigmaHQ 검증기 기준선을 깨끗이 통과하고
(`sigma check --validation-config ../tests/sigma_lint/validators.yml .`), 공격 안내로 읽히는 내용을
담지 않아야 합니다. [`../../CONTRIBUTING.md`](../../CONTRIBUTING.md)와
[`../suricata/`](../suricata/)의 네트워크 계층, 그리고 [`../README.ko.md`](../README.ko.md)를
참조하십시오.
+127
View File
@@ -0,0 +1,127 @@
# ARTEX detection rules (Sigma / host · log · SIEM)
English · [한국어](README.ko.md)
> 한국어: 이 디렉터리는 [방어·탐지 가이드(docs/defense-ko.md)](../../docs/defense-ko.md) 4절
> "탐지 규칙"의 의사 규칙을 실제로 배포 가능한 [Sigma](https://sigmahq.io) 규칙으로 옮긴 것입니다.
> 네트워크 계층은 [`../suricata/`](../suricata/)가 담당합니다. 모든 규칙은 자신이 소유하거나 서면
> 허가를 받은 시스템을 지키는 **방어·탐지 목적에만** 사용하십시오. 한국어 전체 문서는
> **[README.ko.md](README.ko.md)** 를, 전체 탐지 묶음 개요는 **[../README.ko.md](../README.ko.md)** 를 보십시오.
The host, log, and SIEM layer of the ARTEX detection set. These [Sigma](https://sigmahq.io) rules
formalize the pseudo-rules in the defense guide ([Korean](../../docs/defense-ko.md) ·
[English](../../docs/defense-en.md), section 4) into a vendor-neutral format you convert to your own
SIEM or EDR query language. Every indicator is grounded in a string or behaviour verified in this
repository's source, not inferred. The network layer lives under [`../suricata/`](../suricata/); the
full detection set — the ATT&CK coverage layer, the indicator CSV / MISP export, and the host-triage
script — is indexed in [`../README.md`](../README.md).
## Atomic rules
One rule, one observable fact. Convert them individually or as part of the whole tree.
- **[`artex_enrich_user_agent.yml`](artex_enrich_user_agent.yml)** — *ARTEX Asset Enrichment Probe
User-Agent*. Inbound `artex-enrich/1.0` User-Agent from asset enrichment (`enrich/enrich.go`).
Target-side, supporting indicator. `level: high`.
- **[`artex_selfupdate_egress.yml`](artex_selfupdate_egress.yml)** — *ARTEX Self-Update Egress
User-Agent*. Outbound `artex-selfupdate` User-Agent from the self-update routine
(`selfupdate/github.go`). Host/forensic egress indicator. `level: medium`.
- **[`artex_guard_audit_framing.yml`](artex_guard_audit_framing.yml)** — *ARTEX Platform Guard
Audit-Log Framing*. The platform-guard control marker written to the audit log on a blocked tool
call (`guard/guard.go`). Host/forensic indicator. `level: high`.
- **[`artex_recording_proxy_ca.yml`](artex_recording_proxy_ca.yml)** — *ARTEX Recording-Proxy MITM CA
Certificate Artifact*. Creation of the recording proxy's MITM CA file under the
`_ca/mitmproxy-ca-cert.pem` layout (`traffic/traffic.go`). Host/forensic artifact; the bare filename
is shared with standalone mitmproxy, so it is a hunting lead. `level: medium`.
- **[`destructive_command_hunting.yml`](destructive_command_hunting.yml)** — *Destructive Command
Execution (ARTEX Guard-List Hunting)*. Destructive shell/DB commands mirroring the ARTEX guard's
built-in deny list (`db/db.go` seed). Generic hunting lead, **not** an ARTEX signature. `level: medium`.
## Correlation rules (behaviour) — [`correlation/`](correlation/)
Static strings can be changed; behaviour is harder to hide. These Sigma **correlation** rules encode the
behaviour-based layer of the defense guide (sections 4.1–4.2 and 4.4). Each references an atomic rule
above by its `id`, so **convert the whole `sigma/` tree, not a single correlation file**, or the
reference will not resolve (the [Sigma test](../tests/sigma/) asserts exactly this dependency).
- **[`correlation/artex_enrich_scan_velocity.yml`](correlation/artex_enrich_scan_velocity.yml)** —
*Enrichment Scan Velocity*. A burst of `artex-enrich/1.0` probes from one source in a short window
(enrichment runs at concurrency 4 with no rate limit) — the velocity the single-request rule misses.
`event_count`, `level: high`.
- **[`correlation/artex_enrich_fanout.yml`](correlation/artex_enrich_fanout.yml)** — *Enrichment
Fan-Out*. One source carrying the enrichment User-Agent to many *distinct* hosts: machine-speed breadth
across an asset list, where the distinct-host count, not request volume, is the tell. `value_count`,
`level: high`.
- **[`correlation/artex_guard_block_burst.yml`](correlation/artex_guard_block_burst.yml)** —
*Guard-Block Burst*. Repeated platform-guard control markers on one host — an actively engaged ARTEX
run tripping its own guard, not a document that merely quotes the marker. `event_count`, `level: high`.
- **[`correlation/artex_guard_marker_then_destructive.yml`](correlation/artex_guard_marker_then_destructive.yml)**
— *Guard Marker With Destructive Command*. The guard marker and a destructive command co-occurring on
one host within a window (defense guide §4.2, multi-stage): combining an ARTEX-specific marker with the
otherwise-generic destructive-command signal raises specificity. `temporal`, `level: high`.
Thresholds and windows are conservative defaults — tune them to your baseline. The pure web multi-stage
case (enumerate → probe → authenticate) still needs base rules specific to your environment, because that
pattern does not reduce to a single ARTEX-unique User-Agent; a generic behavioural base template to start
from is in the [defense guide §4.2](../../docs/defense-en.md), kept out of this tested tree because it
cannot be grounded in ARTEX source.
## Scope and honesty — read before deploying
- **Static indicators can be changed.** An operator can set a different User-Agent or clean up the CA
file, so the absence of an atomic indicator does **not** mean safety. The durable signal is the
behaviour the `correlation/` rules key on — one source chaining recon → enumeration → probing →
auth/injection attempts, adapting to responses, running without pause.
- **The destructive-command rule is generic hunting.** It mirrors ARTEX's guard deny list, but the same
commands are run by legitimate administrators. Treat a hit as a lead, allow-list your environment, and
do not attribute it to ARTEX on its own.
- **Ports and schema are host-forensic, not Sigma.** The server default `:8787` and recording proxy
`127.0.0.1:8788` (`cmd/artex/main.go`), and the PostgreSQL exploration-graph schema, are best checked on
a suspected host, so they ship in the [indicator CSV](../indicators/) and the
[host-triage script](../triage/) rather than as noisy rules.
- **`logsource` and field names are generic.** The rules use generic `category`/`product` log sources and
field names (`cs-user-agent`, `CommandLine`, `TargetFilename`). Map them to your product's schema with a
pipeline (`-p`) at convert time; see the backend notes below.
## Validate and convert
Validated with [sigma-cli](https://github.com/SigmaHQ/sigma-cli) (pySigma). From the repository root:
```sh
python3 -m venv .venv && . .venv/bin/activate
pip install sigma-cli
# structural + best-practice validation (expect: 0 errors, 0 issues)
sigma check detections/sigma/
# full SigmaHQ convention set with this rule set's documented baseline (expect: 0 issues)
pip install pySigma-validators-sigmahq
sigma check --validation-config detections/tests/sigma_lint/validators.yml detections/sigma/
# compile the WHOLE tree so the correlation rules resolve the atomic rules they reference by id
sigma plugin install splunk
sigma convert -t splunk --without-pipeline detections/sigma/
```
Backends vary in correlation support, so the `-t` choice matters: Splunk, Elasticsearch EQL, and Grafana
Loki convert the whole tree, while Elasticsearch Lucene, OpenSearch, and the Microsoft `kusto` backend
convert the five atomic rules only (express the window natively in the product). The measured per-backend
matrix and the `--without-pipeline` / `-p` field-mapping notes are in [`../README.md`](../README.md), and
they are reproduced by [`../tests/sigma_backends/`](../tests/sigma_backends/).
## Tests
Four reproducible suites under [`../tests/`](../tests/) cover these rules, each needing only Docker:
[`sigma/`](../tests/sigma/) (validation, whole-tree compilation, and that a correlation rule fails to
convert alone), [`sigma_match/`](../tests/sigma_match/) (the rules actually fire on malicious samples and
stay quiet on benign ones), [`sigma_backends/`](../tests/sigma_backends/) (portability across five
backends), and [`sigma_lint/`](../tests/sigma_lint/) (the full SigmaHQ validator baseline, 0 issues). See
[`../tests/README.md`](../tests/README.md).
## Contributing
Detection contributions are welcome. New rules should keep every indicator grounded in an observable fact,
state limitations in the `description`, pass the SigmaHQ validator baseline cleanly
(`sigma check --validation-config ../tests/sigma_lint/validators.yml .`), and avoid any content that reads
as attack guidance. See [`../../CONTRIBUTING.en.md`](../../CONTRIBUTING.en.md) and the network layer in
[`../suricata/`](../suricata/) / [`../README.md`](../README.md).
@@ -0,0 +1,31 @@
title: ARTEX Asset Enrichment Probe User-Agent
id: 34adfa15-1696-4322-afc0-f69988e9cc1e
status: experimental
description: |
Detects inbound HTTP requests whose User-Agent is "artex-enrich/1.0", set by the ARTEX
autonomous penetration-testing framework when it auto-enriches assets (DNS/HTTP checks)
and reads a target's <title>. This probe is generated by ARTEX itself, independent of the
LLM: it does not follow redirects, disables keep-alive, and reads only the beginning of the
response. Default concurrency is 4, so several assets may be probed at once. An operator can
change this User-Agent, so its ABSENCE does not imply safety. Treat it as a supporting
indicator and combine it with the behaviour-based detection in the defense guide, section 4.
references:
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-ko.md
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-en.md
- https://github.com/jiwoochris/artex-ko
author: artex-ko defense guide
date: 2026-10-05
tags:
- attack.reconnaissance
- attack.t1595
- attack.t1592
logsource:
category: webserver
detection:
selection:
cs-user-agent: 'artex-enrich/1.0'
condition: selection
falsepositives:
- Unlikely; this User-Agent string is specific to the ARTEX enrichment client, but an
operator who changed it will not be caught here.
level: high
@@ -0,0 +1,27 @@
title: ARTEX Platform Guard Audit-Log Framing
id: 3add497e-36cb-47c4-8993-df8f98585ef7
status: experimental
description: |
Detects the control-framing string the ARTEX platform guard writes to its audit log when it
blocks a tool call. Blocked calls are recorded with a message beginning with the literal
marker shown below (ARTEX platform control, non-target defence). Finding this marker in a
host's application or audit logs strongly supports that ARTEX ran on that host. This is a
host and forensic indicator, not a target-side signal.
references:
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-ko.md
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-en.md
- https://github.com/jiwoochris/artex-ko
author: artex-ko defense guide
date: 2026-10-05
tags:
- attack.execution
- attack.t1059
logsource:
category: application
detection:
keywords:
- '【ARTEX 平台管控·非目标防御】'
condition: keywords
falsepositives:
- Logs that quote this defense guide or the ARTEX source code for documentation purposes.
level: high
@@ -0,0 +1,37 @@
title: ARTEX Recording-Proxy MITM CA Certificate Artifact
id: 3bac40a5-a780-4d1f-a7e8-5d0daa29ef47
status: experimental
description: |
Detects creation of the man-in-the-middle certificate-authority file the ARTEX recording proxy writes
when it starts. ARTEX embeds a go-mitmproxy traffic recorder that decrypts and logs every HTTP(S)
exchange its worker tools make; on first start the recorder generates a CA under its data directory
(traffic/traffic.go writes "<dir>/_ca/mitmproxy-ca-cert.pem") and injects it into spawned tools through
SSL_CERT_FILE / CURL_CA_BUNDLE / REQUESTS_CA_BUNDLE / NODE_EXTRA_CA_CERTS together with an
HTTP(S)_PROXY pointing at the loopback recorder (agent/worker.go). The file appearing on a host is a
forensic artifact of that recording proxy having run: an adversary-in-the-middle traffic recorder
(ATT&CK T1557) whose trust anchor is an installed root certificate. The "_ca/mitmproxy-ca-cert.pem"
layout narrows it to ARTEX's data directory; a bare mitmproxy-ca-cert.pem is shared with standalone
go-mitmproxy / mitmproxy, so treat a hit as a host-triage lead to correlate with the loopback proxy
endpoint and server port (see the indicators list), not a standalone alert.
references:
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-ko.md
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-en.md
- https://github.com/jiwoochris/artex-ko
author: artex-ko defense guide
date: 2026-10-07
tags:
- attack.credential-access
- attack.collection
- attack.t1557
logsource:
category: file_event
detection:
selection:
TargetFilename|contains|all:
- '_ca'
- 'mitmproxy-ca-cert.pem'
condition: selection
falsepositives:
- Standalone go-mitmproxy or mitmproxy deployments that write the same CA filename.
- Developers intentionally running a recording or debugging proxy on the host.
level: medium
@@ -0,0 +1,27 @@
title: ARTEX Self-Update Egress User-Agent
id: e96380a3-2a34-4237-be2b-088ad9cc947d
status: experimental
description: |
Detects outbound (egress) HTTP requests whose User-Agent is "artex-selfupdate", used by the
ARTEX self-update routine when it queries code-repository hosts (for example GitHub releases)
for a newer binary. Seeing this User-Agent leave an internal host toward a code-hosting
service suggests an ARTEX binary is installed on that host. This is primarily an operator and
forensic indicator on a (possibly compromised relay) host, not a target-side signal.
references:
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-ko.md
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-en.md
- https://github.com/jiwoochris/artex-ko
author: artex-ko defense guide
date: 2026-10-05
tags:
- attack.command-and-control
- attack.t1105
logsource:
category: proxy
detection:
selection:
c-useragent: 'artex-selfupdate'
condition: selection
falsepositives:
- Unlikely; this User-Agent string is specific to the ARTEX self-update client.
level: medium
@@ -0,0 +1,37 @@
# references base rule: ../artex_enrich_user_agent.yml (ARTEX Asset Enrichment Probe User-Agent)
title: ARTEX Enrichment Fan-Out (One Source, Many Distinct Hosts)
id: 6fba3b7c-1dd9-4e18-bf55-d93fc1d2e2e0
status: experimental
description: |
Correlates a single client carrying the ARTEX enrichment User-Agent to many DISTINCT
destination hosts within a short window (distinct count of cs-host). Autonomous enrichment
fans out across an asset list at machine speed, so breadth — the number of different hosts
touched, not just request volume — is what separates it from a person browsing a few pages.
Evaluate this where your telemetry spans multiple hosts (CDN, WAF, reverse proxy, or shared
hosting) or at an egress point that sees outbound enrichment. As with the base rule, the
User-Agent can be changed; the durable signal is the fan-out behaviour, so pair this with the
defense guide section 4 and tune the distinct-host threshold and window to your environment.
references:
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-ko.md
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-en.md
- https://github.com/jiwoochris/artex-ko
author: artex-ko defense guide
date: 2026-10-05
tags:
- attack.reconnaissance
- attack.t1595
- attack.t1592
correlation:
type: value_count
rules:
- 34adfa15-1696-4322-afc0-f69988e9cc1e
group-by:
- c-ip
timespan: 10m
condition:
gte: 20
field: cs-host
falsepositives:
- Shared egress (NAT/proxy) where many users appear as one source; a legitimate scanner or
uptime monitor that fronts many hosts. Allow-list known sources and raise the threshold.
level: high
@@ -0,0 +1,35 @@
# references base rule: ../artex_enrich_user_agent.yml (ARTEX Asset Enrichment Probe User-Agent)
title: ARTEX Enrichment Scan Velocity (Burst From One Source)
id: 1d07c5f1-e0ff-4a6a-8017-6008fa790889
status: experimental
description: |
Correlates a burst of ARTEX asset-enrichment probes from a single client within a short
window. The base rule matches the "artex-enrich/1.0" User-Agent; this correlation adds the
behaviour the single-request rule misses — velocity. ARTEX enriches assets with a default
concurrency of 4 and ships no built-in rate limit (enrich/enrich.go), so an active run emits
many enrichment requests in quick succession rather than one. An operator can change the
User-Agent, so a quiet result is inconclusive, not proof of safety; see the behaviour-based
detection in the defense guide, section 4. Tune the count and window to your own baseline.
references:
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-ko.md
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-en.md
- https://github.com/jiwoochris/artex-ko
author: artex-ko defense guide
date: 2026-10-05
tags:
- attack.reconnaissance
- attack.t1595
- attack.t1592
correlation:
type: event_count
rules:
- 34adfa15-1696-4322-afc0-f69988e9cc1e
group-by:
- c-ip
timespan: 5m
condition:
gte: 30
falsepositives:
- Legitimate asset-management or monitoring tools that set this User-Agent and poll many
assets; allow-list their source addresses.
level: high
@@ -0,0 +1,34 @@
# references base rule: ../artex_guard_audit_framing.yml (ARTEX Platform Guard Audit-Log Framing)
title: ARTEX Guard-Block Burst (Active Engaged Session)
id: 065c11b5-4dcf-48ae-84fe-e8301346d5b4
status: experimental
description: |
Correlates repeated ARTEX platform-guard control markers in one host's application or audit
log within a short window. The base rule matches a single marker, which can also appear in a
document or source file that merely quotes it. A burst of the same marker on one host instead
indicates an ARTEX run that is actively and repeatedly tripping its guard — which both
confirms execution on that host and removes the quote-a-marker false positive of the single
event. This remains a host and forensic indicator, not a target-side signal. Tune the count
and window to your environment.
references:
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-ko.md
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-en.md
- https://github.com/jiwoochris/artex-ko
author: artex-ko defense guide
date: 2026-10-05
tags:
- attack.execution
- attack.t1059
correlation:
type: event_count
rules:
- 3add497e-36cb-47c4-8993-df8f98585ef7
group-by:
- host
timespan: 10m
condition:
gte: 5
falsepositives:
- A log pipeline that repeatedly ingests or re-processes a document quoting this marker;
scope the rule to runtime application/audit logs, not documentation stores.
level: high
@@ -0,0 +1,40 @@
# references base rules:
# ../artex_guard_audit_framing.yml (ARTEX Platform Guard Audit-Log Framing)
# ../destructive_command_hunting.yml (Destructive Command Execution — ARTEX Guard-List Hunting)
title: ARTEX Guard Marker With Destructive Command on One Host
id: 1a03bfb8-7893-4730-9422-dd50c2dd91e8
status: experimental
description: |
Temporal correlation for the multi-stage behaviour in the defense guide, section 4.2: on a
single host, the ARTEX platform-guard control marker (confirming ARTEX executed there) AND a
destructive command from the guard deny-list hunting rule occur within the same window.
Pairing an ARTEX-specific marker with the otherwise-generic destructive-command signal raises
specificity. A destructive command on its own is routine administration and a known source of
false positives; one that co-occurs with an active ARTEX run on the same host is worth
investigating. This assumes a normalized host field is present across both the application or
audit logs and the process-creation logs; map it in your pipeline. It does not by itself prove
ARTEX issued the command — the guard blocks these patterns, so a hit means the destructive
command ran outside or despite the guard near an ARTEX run. Treat it as a high-priority lead.
references:
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-ko.md
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-en.md
- https://github.com/jiwoochris/artex-ko
author: artex-ko defense guide
date: 2026-10-05
tags:
- attack.execution
- attack.t1059
- attack.impact
- attack.t1485
correlation:
type: temporal
rules:
- 3add497e-36cb-47c4-8993-df8f98585ef7
- f510564f-2958-4dc8-a188-3300a2f6f5a7
group-by:
- host
timespan: 30m
falsepositives:
- Maintenance windows where an operator legitimately runs ARTEX and, separately, performs
destructive administration on the same host within the window. Confirm with change tickets.
level: high
@@ -0,0 +1,64 @@
title: Destructive Command Execution (ARTEX Guard-List Hunting)
id: f510564f-2958-4dc8-a188-3300a2f6f5a7
status: experimental
description: |
Hunts for destructive shell and database commands on a host. The pattern set mirrors the
built-in deny rules the ARTEX guard ships with (db/db.go seed): because the guard blocks
these, they are the inverse image of the destructive actions an autonomous agent could
attempt if the guard were disabled or bypassed. This is GENERIC destructive-command hunting
informed by that list, not an ARTEX-specific signature, and matches are expected from
legitimate administration. Tune and allow-list for your environment and treat a hit as a
hunting lead, not a standalone alert. High-noise availability commands the guard also blocks
(bare shutdown/reboot) are intentionally omitted here; hunt those separately. The database
patterns likewise track data-destroying objects (DROP DATABASE/TABLE/SCHEMA) rather than the
guard's wider DROP set (INDEX/VIEW/USER/ROLE/TABLESPACE), which alter structure or access
rather than destroy data and are noisy in routine migrations.
references:
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-ko.md
- https://github.com/jiwoochris/artex-ko/blob/main/docs/defense-en.md
- https://github.com/jiwoochris/artex-ko
author: artex-ko defense guide
date: 2026-10-05
tags:
- attack.impact
- attack.t1485
- attack.t1561.002
- attack.t1489
logsource:
category: process_creation
detection:
selection_filesystem:
CommandLine|contains:
- 'rm -rf'
- 'rm -fr'
- 'rm --recursive'
- '--no-preserve-root'
- 'mkfs'
- 'dd of=/dev/'
- 'shred '
- 'wipe /dev/'
selection_database:
CommandLine|contains:
- 'DROP DATABASE'
- 'DROP TABLE'
- 'DROP SCHEMA'
- 'TRUNCATE '
- '.dropDatabase('
- '.dropCollection('
- 'FLUSHALL'
- 'FLUSHDB'
selection_availability:
CommandLine|contains:
- 'curl -X DELETE'
- 'curl --request DELETE'
- 'wget --method=DELETE'
- 'iptables -F'
- 'nft flush ruleset'
- 'kill -9 -1'
- 'killall -9'
condition: 1 of selection_*
falsepositives:
- Routine system administration, maintenance scripts, and container teardown.
- CI/CD pipelines that drop and recreate test databases or caches.
- The GNU coreutils `truncate` command (e.g. log rotation `truncate -s 0 file`) shares the TRUNCATE token; allow-list it, since it is followed by a flag rather than a table name.
level: medium
+84
View File
@@ -0,0 +1,84 @@
# ARTEX 탐지 규칙 (Suricata / 네트워크)
한국어 · [English](README.md)
[Sigma 규칙](../sigma/)의 네트워크 계층 짝입니다. 이 [Suricata](https://suricata.io) 시그니처는
네트워크 선에서 관측되는 두 가지 ARTEX 산출물을 다루며, 모든 지표는 추정이 아니라 이 저장소
소스에서 확인한 문자열이나 행동에 근거합니다. 호스트·로그·SIEM 계층은 [`../sigma/`](../sigma/)에
있고, 전체 그림은 방어 가이드([한국어](../../docs/defense-ko.md) · [English](../../docs/defense-en.md))가
설명합니다.
## 규칙: [`artex.rules`](artex.rules)
- **sid 1000001**: `ARTEX enrichment prober User-Agent`. User-Agent 가 `artex-enrich/` 로 시작하는
인바운드 HTTP `GET` 입니다(자산 보강 프로버 `enrich/enrich.go:233`). 단건 요청 존재 지표입니다.
`classtype: attempted-recon`.
- **sid 1000002**: `ARTEX enrichment prober high-rate enumeration`. 같은 User-Agent 가
`detection_filter` 임계인 **출처당 300초에 30요청** 을 넘는 경우입니다. 단건 규칙이 놓치는, 기계
속도로 쏟아내는 빈도입니다. Sigma 상관 규칙 `artex_enrich_scan_velocity` 를 반영합니다.
`classtype: attempted-recon`.
- **sid 1000003**: `ARTEX worker WebFetch User-Agent`. User-Agent 가 `norma/` 로 시작하는 인바운드
HTTP 요청입니다(norma SDK 의 WebFetch 도구 `github.com/Autumn-27/norma/tool/webfetch.go`). 이 UA 는 norma 전 버전
(v0.1.0–v0.4.3, 검증 완료)에 걸쳐 하드코딩되어 있으며, 기록 프록시가 요청 헤더를 수정하지
않으므로(`traffic/traffic.go`) 대상 호스트 와이어에 그대로 도달합니다. 보강 프로버와 달리
**공격 단계**(능동적 취약점 프로빙) 에서 발화합니다. `classtype: attempted-recon`.
## 범위와 정직함: 배포 전에 읽으십시오
- **두 가지 ARTEX User-Agent 가 네트워크에서 관측됩니다.** 보강 프로버는 정찰 단계에서
`artex-enrich/1.0`(`enrich/enrich.go:233`)을, norma SDK 의 WebFetch 도구는 공격 단계에서
`norma/0.4`(`github.com/Autumn-27/norma/tool/webfetch.go`)를 보냅니다. 기록 프록시(`traffic/traffic.go`)는 요청 헤더를
변경하지 않으므로 두 UA 모두 대상 와이어에 도달합니다. 그 외 worker 도구(Bash 하위 프로세스인
`curl`, `nmap` 등)는 자체 User-Agent 를 사용하므로, 일반 스캐너 시그니처와
[`../sigma/`](../sigma/)의 행동 기반 SIEM 규칙으로 탐지하십시오.
- **User-Agent 는 평문에서만 보입니다.** 트래픽이 평문 HTTP 이거나 TLS 를 종단하는 프록시·WAF 에서
검사될 때 나타납니다. 종단 간 TLS 는 이것을 암호화하므로, 실제로 HTTP 요청 버퍼를 볼 수 있는
자리에 배포하십시오.
- **정적 User-Agent 는 운영자가 바꿀 수 있으므로**, 없다고 해서 안전하다는 뜻은 **아닙니다**.
오래가는 신호는 행동, 곧 속도와 폭입니다. sid 1000002(그리고 Sigma 상관 계층)가 속도를 기준으로
삼는 이유, 그리고 순수 웹 다단계 탐지가 환경별 기본 규칙을 필요로 하는 이유가 여기에 있습니다.
- **의도적으로 뺀 것.** 자가 업데이트 User-Agent `artex-selfupdate` 는 GitHub 로 HTTPS 를 타고 가므로
네트워크에서 관측되지 않습니다(TLS SNI 만으로는 경보를 걸기에 너무 흔합니다). 감사 통제 마커는
대상을 향하는 트래픽이 아니라 운영자 측 로그 산출물이므로,
[`../sigma/artex_guard_audit_framing.yml`](../sigma/artex_guard_audit_framing.yml)로 탐지하십시오.
서버 포트 `:8787` 과 기록 프록시 `127.0.0.1:8788`(`cmd/artex/main.go`)은 네트워크 시그니처가
아니라 호스트 포렌식용(`ss`·`netstat`)입니다.
## 검증과 테스트
Suricata 8 로 검증했습니다. 적재 테스트는 트래픽이 필요 없고 항상 돌아갑니다.
```sh
# 문법 + 엔진 적재 테스트 (기대: "Configuration provided was successfully loaded")
docker run --rm -v "$PWD/detections/suricata":/r -w /r jasonish/suricata:latest \
suricata -T -S artex.rules -l /tmp --init-errors-fatal
```
`--init-errors-fatal` 은 파싱은 되지만 초기화에 실패하는 규칙도 하드 에러로 만들어, 조용히 버려진
시그니처가 있으면 적재 테스트가 통과하지 못하게 합니다.
규칙이 실제로 발화하는지 확인하려면, 재현 가능한 회귀 테스트가 [`../tests/suricata/`](../tests/suricata/)에
있습니다. 먼저 같은 적재 점검을 돌리고, scapy 로 결정적 캡처를 합성한 뒤 그 위에서 `suricata -r` 를
돌려 경보 수를 단언합니다. Docker 만 있으면 됩니다.
```sh
detections/tests/suricata/run.sh
```
sid 1000001 이 프로브마다 정확히 한 번씩 발화하고(35플로 캡처에서 35회), sid 1000002 가 300초에 30
임계를 넘으며(Suricata 8.0.7 에서 **5**회 경보, 31~35번째 플로), 같은 캡처를 양성(benign) 브라우저
User-Agent 로 돌리면 경보가 **0** 임을 단언합니다. 시그니처가 특이함을 확인하는 것입니다.
[`../tests/README.ko.md`](../tests/README.ko.md)를 참조하십시오. 대신 자신의 트래픽으로 확인하려면, 로컬
서버에 대해 루프백 `curl -A 'artex-enrich/1.0'` 을 캡처해 경보를 읽으십시오.
```sh
suricata -r enrich.pcap -S artex.rules -l out && \
grep -c '"signature_id":1000001' out/eve.json # 존재: 프로브마다 한 번
```
## 기여
탐지 기여를 환영합니다. 새 규칙은 모든 지표를 관측 가능한 사실에 근거해 두고, 한계를 주석에
밝히며, `suricata -T` 를 깨끗이 통과하고, 공격 안내로 읽히는 내용을 담지 않아야 합니다.
[`../../CONTRIBUTING.md`](../../CONTRIBUTING.md)와 [`../sigma/`](../sigma/)의 Sigma 계층 /
[`../README.ko.md`](../README.ko.md)를 참조하십시오.
+90
View File
@@ -0,0 +1,90 @@
# ARTEX detection rules (Suricata / network)
English · [한국어](README.ko.md)
> 한국어: 이 디렉터리는 [방어·탐지 가이드(docs/defense-ko.md)](../../docs/defense-ko.md) 2·4절의
> 네트워크 관측 지문을 실제로 배포 가능한 [Suricata](https://suricata.io) 규칙으로 옮긴 것입니다.
> 로그·호스트 계층은 [`../sigma/`](../sigma/)(Sigma)가 담당합니다. 모든 규칙은 자신이 소유하거나
> 서면 허가를 받은 시스템을 지키는 **방어·탐지 목적에만** 사용하십시오. 한국어 전체 문서는
> **[README.ko.md](README.ko.md)** 를 보십시오.
The network-layer companion to the [Sigma rules](../sigma/). These [Suricata](https://suricata.io)
signatures cover the two ARTEX artifacts that are observable on the wire, and every indicator is grounded in a
string or behaviour verified in this repository's source, not inferred. The host, log, and SIEM layers live
under [`../sigma/`](../sigma/); the defense guide ([Korean](../../docs/defense-ko.md) ·
[English](../../docs/defense-en.md)) explains the full picture.
## Rules — [`artex.rules`](artex.rules)
- **sid 1000001** — `ARTEX enrichment prober User-Agent`. An inbound HTTP `GET` whose User-Agent starts with
`artex-enrich/` — the asset-enrichment prober (`enrich/enrich.go:233`). The single-request presence
indicator. `classtype: attempted-recon`.
- **sid 1000002** — `ARTEX enrichment prober high-rate enumeration`. The same User-Agent crossing a
`detection_filter` rate of **30 requests in 300 s per source** — the machine-speed velocity a single-hit
rule misses. Mirrors the Sigma correlation `artex_enrich_scan_velocity`. `classtype: attempted-recon`.
- **sid 1000003** — `ARTEX worker WebFetch User-Agent`. An inbound HTTP request whose User-Agent starts with
`norma/` — the norma SDK's WebFetch tool (`github.com/Autumn-27/norma/tool/webfetch.go`). This UA is hardcoded across all norma
versions (v0.1.0–v0.4.3, verified) and reaches the target through the recording proxy, which does not
modify request headers (`traffic/traffic.go`). Unlike the enrich prober, this fires during the **attack
phase** (active vulnerability probing). `classtype: attempted-recon`.
## Scope and honesty — read before deploying
- **Two ARTEX User-Agents are network-observable.** The enrich prober sends `artex-enrich/1.0`
(`enrich/enrich.go:233`) during reconnaissance; the norma SDK's WebFetch tool sends `norma/0.4`
(`github.com/Autumn-27/norma/tool/webfetch.go`) during the attack phase. The recording proxy (`traffic/traffic.go`) does not
modify request headers, so both UAs reach the target on the wire. Other worker tools (Bash subprocesses
like `curl`, `nmap`) use their own User-Agents — detect those with generic scanner signatures and the
behavioural SIEM rules under [`../sigma/`](../sigma/).
- **The User-Agent is only visible in plaintext.** It appears where traffic is plaintext HTTP or inspected at
a TLS-terminating proxy / WAF. End-to-end TLS encrypts it, so deploy these where you actually see the HTTP
request buffer.
- **A static User-Agent can be changed** by the operator, so its absence does **not** mean safety. The durable
signal is behaviour — rate and breadth — which is why sid 1000002 (and the Sigma correlation layer) key on
velocity, and why pure web multi-stage detection needs base rules specific to your environment.
- **Deliberately omitted.** The self-update User-Agent `artex-selfupdate` travels over HTTPS to GitHub and is
not network-observable (TLS SNI alone is too common to alert on). The audit-control marker is an
operator-side log artifact, not target-facing traffic — detect it with
[`../sigma/artex_guard_audit_framing.yml`](../sigma/artex_guard_audit_framing.yml). The server port `:8787`
and recording proxy `127.0.0.1:8788` (`cmd/artex/main.go`) are host-forensic (`ss`/`netstat`), not a
network signature.
## Validate and test
Validated with Suricata 8. The load test needs no traffic and always runs:
```sh
# syntax + engine load test (expect: "Configuration provided was successfully loaded")
docker run --rm -v "$PWD/detections/suricata":/r -w /r jasonish/suricata:latest \
suricata -T -S artex.rules -l /tmp --init-errors-fatal
```
`--init-errors-fatal` makes a rule that parses but fails to initialise a hard error too, so the load test
cannot pass with a silently dropped signature.
To confirm the rules actually fire, a reproducible regression test lives in
[`../tests/suricata/`](../tests/suricata/). It runs this same load check first, then synthesizes a
deterministic capture with scapy, runs `suricata -r` over it, and asserts the alert counts — needing only
Docker:
```sh
detections/tests/suricata/run.sh
```
It asserts that sid 1000001 fires exactly once per probe (35 over a 35-flow capture), that sid 1000002
trips past the 30-in-300 s rate (**5** alerts on Suricata 8.0.7, flows 31–35), and that the same capture
with a benign browser User-Agent produces **0** alerts — confirming the signatures are specific. See
[`../tests/README.md`](../tests/README.md). To check against your own traffic instead, capture a loopback
`curl -A 'artex-enrich/1.0'` against a local server and read the alerts:
```sh
suricata -r enrich.pcap -S artex.rules -l out && \
grep -c '"signature_id":1000001' out/eve.json # presence: one per probe
```
## Contributing
Detection contributions are welcome. New rules should keep every indicator grounded in an observable fact,
state limitations in a comment, pass `suricata -T` cleanly, and avoid any content that reads as attack
guidance. See [`../../CONTRIBUTING.en.md`](../../CONTRIBUTING.en.md) and the Sigma layer in
[`../sigma/`](../sigma/) / [`../README.md`](../README.md).
+25
View File
@@ -0,0 +1,25 @@
# ARTEX network detection - Suricata rules
# Repo: https://github.com/jiwoochris/artex-ko
# Guide: ../../docs/defense-en.md (English) / ../../docs/defense-ko.md (Korean)
# Index & how-to-test: detections/suricata/README.md
#
# SCOPE AND HONESTY - read before deploying:
# - These rules target two ARTEX artifacts observable on the wire:
# (1) the enrichment prober's HTTP User-Agent "artex-enrich/1.0" (enrich/enrich.go:233),
# (2) the norma SDK's WebFetch User-Agent "norma/0.4" (github.com/Autumn-27/norma/tool/webfetch.go).
# The recording proxy (traffic/traffic.go) does not modify request headers, so both
# UAs reach the target on the wire. Other worker tools (curl, nmap, etc.) use their own
# default User-Agents — detect those with generic scanner signatures and ../sigma/.
# - The enrich User-Agent is only visible where traffic is plaintext HTTP or inspected
# at a TLS-terminating proxy / WAF. End-to-end TLS encrypts it.
# - A static User-Agent can be changed by the operator; its absence does NOT imply safety.
# - The self-update User-Agent "artex-selfupdate" travels over HTTPS to GitHub and is not
# network-observable (TLS SNI alone is too common to alert on) - intentionally omitted.
# - The audit-control marker is an operator-side log artifact, not target-facing traffic -
# detect it with ../sigma/artex_guard_audit_framing.yml instead.
alert http any any -> any any (msg:"ARTEX enrichment prober User-Agent (artex-enrich)"; flow:established,to_server; http.method; content:"GET"; http.user_agent; content:"artex-enrich/"; startswith; fast_pattern; classtype:attempted-recon; reference:url,github.com/jiwoochris/artex-ko/tree/main/detections/suricata; metadata:created_at 2026_10_05; sid:1000001; rev:1;)
alert http any any -> any any (msg:"ARTEX enrichment prober high-rate enumeration (artex-enrich)"; flow:established,to_server; http.user_agent; content:"artex-enrich/"; startswith; fast_pattern; detection_filter:track by_src, count 30, seconds 300; classtype:attempted-recon; reference:url,github.com/jiwoochris/artex-ko/tree/main/detections/suricata; metadata:created_at 2026_10_05; sid:1000002; rev:1;)
alert http any any -> any any (msg:"ARTEX worker WebFetch User-Agent (norma)"; flow:established,to_server; http.user_agent; content:"norma/"; startswith; fast_pattern; classtype:attempted-recon; reference:url,github.com/jiwoochris/artex-ko/tree/main/detections/suricata; metadata:created_at 2026_10_07; sid:1000003; rev:1;)
+455
View File
@@ -0,0 +1,455 @@
# ARTEX 탐지 규칙 테스트
한국어 · [English](README.md)
[`../`](../) 아래의 탐지 규칙이 실제로 발화하는지, 그리고 그에 못지않게 중요한, 양성(benign)
트래픽에는 침묵하는지를 재현 가능하게 증명하는 회귀 테스트입니다. 돌려 볼 수 없는 탐지 규칙은
주장에 지나지 않습니다. 이 테스트들은 규칙 파일과 방어 가이드에 적힌 주장을 검토자가 소스에서
다시 돌려 볼 수 있는 것으로 바꿉니다.
이진 패킷 캡처는 저장소에 넣지 않습니다. 캡처는 **매 실행마다 결정론적으로 생성**했다가 끝난 뒤
지우므로, 테스트는 불투명한 고정 파일이 아니라 읽을 수 있는 소스로 배포되며 저장소를 불리지
않습니다.
## 모든 스위트를 한 번에 실행: [`run-all.sh`](run-all.sh)
[`run-all.sh`](run-all.sh) 는 아래 여덟 스위트를 CI 와 같은 순서로 한 명령에 전부 돌리므로, 여덟 개
`run.sh` 스크립트를 손으로 하나씩 호출하지 않아도 됩니다. 앞 스위트가 실패해도 각 스위트는 끝까지
돌고, 스크립트는 마지막에 스위트마다 PASS/FAIL 한 줄 요약을 출력하며, 하나라도 실패하면 0 이 아닌
코드로 종료합니다.
스위트를 돌리기 전에 하네스 자기 점검([`check-harness-sync.sh`](check-harness-sync.sh))을 먼저 실행합니다.
이 점검은 위의 스위트 목록, [CI](../../.github/workflows/detections.yml) 의 스위트별 스텝, 디스크의 스위트
디렉터리 이 셋이 서로 다른 스위트나 다른 순서를 가리키면 실행을 실패로 끝냅니다. 이것은 여덟 스위트가
스스로 보지 못하는 유일한 공백입니다. 세 곳 중 한 곳에만 배선된 스위트(예: `run-all.sh` 항목 없이 CI
스텝만 추가하거나, 어느 쪽에도 넣지 않은 디렉터리)는 스위트별 테스트를 모두 통과하면서도, 로컬에서
초록이던 `run-all.sh` 가 더는 초록 CI 를 뜻하지 않게 만듭니다. 이 점검은 아홉째 스위트가 아니라 게이트라서
아래 요약에는 나타나지 않으므로, 탐지 스위트는 여덟 그대로입니다.
```sh
detections/tests/run-all.sh
```
예상 출력(축약):
```
===== detection suites summary =====
PASS sigma
PASS sigma_match
PASS sigma_lint
PASS sigma_backends
PASS suricata
PASS attack
PASS indicators
PASS misp
RESULT: PASS
```
실패가 하나라도 있으면 0 이 아닌 코드로 종료하므로 pre-commit 훅에 그대로 넣을 수 있습니다. 바로 쓸
수 있는 예시가 저장소 최상위 [`.pre-commit-config.yaml`](../../.pre-commit-config.yaml) 에 있습니다.
`pip install pre-commit && pre-commit install` 로 설치하면, 탐지 규칙이나 그 규칙이 고정한 상류 소스
파일을 건드리는 커밋에서 러너가 발화합니다. CI 와 같은 범위입니다. 개별 스위트가 인식하는 이미지·버전
재정의(`PYTHON_IMAGE`, `SIGMA_CLI_VERSION`, `SIGMAHQ_VALIDATORS_VERSION`, `SURICATA_IMAGE`)는 러너가
그대로 물려받으므로, 그중 어느 것을 export 해도 모든 스위트에 한꺼번에 적용됩니다.
## Suricata: [`suricata/`](suricata/)
[`suricata/run.sh`](suricata/run.sh) 는 [`../suricata/artex.rules`](../suricata/artex.rules) 의
네트워크 규칙을 종단으로 돌려 다섯 가지 속성을 단언합니다:
- **유효성**: 규칙 파일 전체가 `suricata -T --init-errors-fatal` 로 적재되므로, 아래 어떤 캡처도
건드리지 않는 규칙이라도 파싱·초기화에 실패하면 잡아냅니다. 그냥 `suricata -r` 는 그런 규칙을
건너뛰고도 0 으로 종료하므로, 이 적재 검사는 Sigma 스위트의 `sigma check` 유효성 단언에 해당하는
Suricata 쪽 장치입니다.
- **존재성(보강 프로버)**: sid `1000001` 이 보강 프로브마다 정확히 한 번 발화합니다.
- **속도**: 소스당 300 초에 30 요청이라는 `detection_filter` 임계를 넘으면 sid `1000002` 가
발화합니다.
- **존재성(WebFetch)**: sid `1000003` 이 norma WebFetch 요청마다 정확히 한 번 발화하고, 같은 캡처에서
보강 프로버 sid 는 침묵합니다. 두 네트워크 시그니처가 각자 발화할 뿐 아니라 서로 특이적임을
확인합니다.
- **특이성**: 다른 것은 같고 User-Agent 만 양성(benign) 브라우저로 바꾼 캡처는 ARTEX 경보를
**하나도** 내지 않습니다.
[`suricata/gen_pcap.py`](suricata/gen_pcap.py) 는 [scapy](https://scapy.net) 로 캡처를 만듭니다.
고정된 한 소스에서 나오는 N 개의 독립적인 평문 HTTP 요청/응답 흐름을, 각각 지정한 User-Agent 를
실어, 고정된 기준 타임스탬프에서 1 초 간격으로 배치합니다. 파일을 쓰기만 할 뿐, 패킷을 보내거나
네트워크를 건드리지 않습니다.
### 실행
Docker 만 있으면 됩니다. scapy 와 Suricata 모두 컨테이너에서 돕니다.
```sh
detections/tests/suricata/run.sh
```
예상 출력(축약):
```
PASS ruleset loads with zero parse/init errors (suricata -T)
PASS sid 1000001 presence: one alert per probe (got 35, want eq 35)
PASS sid 1000002 velocity: fires past 30-in-300s (got 5, want ge 1)
PASS sid 1000003 presence: one alert per WebFetch request (got 8, want eq 8)
PASS enrich sids stay silent on norma traffic (specificity) (got 0, want eq 0)
PASS benign browser UA produces no ARTEX alerts (got 0, want eq 0)
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료하므로 CI 나 pre-commit 훅에 그대로 넣을 수
있습니다. 내부에 미러를 두었다면 `SURICATA_IMAGE` / `PYTHON_IMAGE` 로 이미지를 재정의하십시오.
### 속도 경보 수를 정확한 값이 아니라 하한으로 단언하는 이유
`run.sh` 는 존재성 경보 수(`1000001 == 35`·`1000003 == 8`)와 양성 경보 수(`== 0`)를 정확히 단언합니다.
이들은 엔진 버전과 무관하기 때문입니다. 일치하는 요청마다 경보 하나, 다른 User-Agent 에는 불일치입니다. 속도
규칙의 경보 수는 특정 Suricata 릴리스가 경계에서 `detection_filter` 임계를 어떻게 처리하느냐에 달려
있으므로, 테스트는 `>= 1` 로 단언하고 기준값은 따로 기록합니다. **Suricata 8.0.7** 에서는 기준
실행이 sid `1000002` 에 경보 **5** 개를 냅니다(300 초에 30 임계를 넘긴 뒤의 31–35 번째 흐름).
## Sigma: [`sigma/`](sigma/)
[`sigma/run.sh`](sigma/run.sh) 는 [`../sigma/`](../sigma/) 아래의 Sigma 규칙을 구조적으로, 그리고
[sigma-cli](https://github.com/SigmaHQ/sigma-cli)(pySigma) 로 컴파일해 검증하며, 다섯 가지 속성을
단언합니다:
- **유효성**: `sigma check` 가 트리 전체에서 오류 0, 조건 오류 0, 이슈 0 을 보고합니다.
- **컴파일**: `sigma convert -t splunk` 가 트리 전체를 오류 없이 백엔드 질의 언어로 변환합니다.
- **지표 보존**: 각 원자 지표 문자열(`artex-enrich/1.0`, `artex-selfupdate`, 가드 마커, 그리고 기록용
프록시 CA 파일명 `mitmproxy-ca-cert.pem`)이 컴파일된 질의에 그대로 남아 있으므로, 규칙이 자신이
기반한 문자열을 조용히 잃을 수 없습니다.
- **상관 규칙 컴파일**: [`../sigma/correlation/`](../sigma/correlation/) 의 행동 규칙이 버려지지
않고 `event_count` / `value_count` 집계를 내보냅니다.
- **상관 규칙이 실제로 작동함**: 상관 규칙 하나만 *단독으로* 변환하면 실패합니다. 그 규칙이 원자
기반 규칙을 `id` 로 참조하기 때문이며, 이 참조는 장식이 아니라 강제됩니다. 이는 위 Suricata
특이성 단언에 해당하는 Sigma 쪽 장치입니다.
이는 [`../README.ko.md`](../README.ko.md) 에 설명한 구조 + 컴파일 검증을, 실행 가능하고 단언하는 형태로 만든
것입니다. 아래의 짝 스위트 [`sigma_match/`](sigma_match/) 가 원자 규칙과 상관 규칙 양쪽에 *매칭* 절반을
더합니다. 대표적인 악성 이벤트(또는 타임라인)가 각 규칙을 발화시키고 정상 이벤트는 발화시키지 않음을
확인하므로, 이제 Sigma 규칙도 Suricata 규칙처럼 재현 가능한 검증 테스트와 재현 가능한 매칭 테스트를 함께
갖습니다. (엉성하게 손으로 짠 매처가 규칙을
깎아내릴 수 있다는 기존 우려는, 파싱을 전부 pySigma 에 위임해 해소했습니다. 신뢰 모델은 다음 절에서 설명합니다.)
### 실행
Docker 만 있으면 됩니다. sigma-cli 와 splunk 백엔드가 컨테이너에서 돌고 저장소에는 아무것도 쓰지
않습니다.
```sh
detections/tests/sigma/run.sh
```
예상 출력(축약):
```
PASS sigma check: 0 errors, 0 condition errors, 0 issues
PASS whole tree converts to splunk (exit 0)
PASS indicator present: artex-enrich/1.0
PASS correlation rule fails to convert alone — it requires its atomic base rule
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료하므로 CI 나 pre-commit 훅에 그대로 넣을 수
있습니다. sigma-cli 는 기준 버전(`3.1.0`)으로 고정돼 있습니다. 내부에 미러를 두었다면
`SIGMA_CLI_VERSION` 으로 버전을, `PYTHON_IMAGE` 로 이미지를 재정의하십시오.
## Sigma 실시간 이벤트 매칭: [`sigma_match/`](sigma_match/)
[`sigma_match/run.sh`](sigma_match/run.sh) 는 [`../sigma/`](../sigma/) 아래의 Sigma 규칙이, 원자 규칙과
[`../sigma/correlation/`](../sigma/correlation/) 의 상관 규칙을 모두 포함해, 매칭되는 이벤트에 실제로 *발화*하고
정상 이벤트에는 침묵함을 증명합니다. Suricata 스위트가 네트워크 규칙에 주는 "돌려 볼 수 없는 탐지 규칙은
주장일 뿐"이라는 보증을, 호스트·로그 계층 규칙으로 확장한 것입니다. 원자 규칙 셋과 상관 규칙 셋, 모두 여섯
속성을 단언합니다:
- **규칙·샘플 짝짓기**: 모든 원자 규칙에는 [`events/<이름>.json`](sigma_match/events/) 샘플 파일이 있고,
모든 샘플 파일은 규칙으로 되짚어집니다. 샘플 없이 추가한 규칙은 검증을 못 받고 넘어가는 대신 여기서
실패합니다.
- **참 양성(true positive)**: 각 규칙이 자신의 악성 샘플 이벤트를 전부 매칭합니다.
- **참 음성(true negative)**: 각 규칙이 자신의 정상 샘플 이벤트를 하나도 매칭하지 않습니다. 예를 들어
`.mitmproxy/` 아래의 단독 `mitmproxy-ca-cert.pem` 은 기록용 프록시 규칙을 발화시키지 **않습니다**. 그
규칙의 `|all` 수식자가 ARTEX 가 쓰는 `_ca/` 디렉터리까지 함께 요구하기 때문이며, 이 판별을 증명하는 것이
바로 매칭 테스트입니다.
- **상관 규칙·타임라인 짝짓기**: 모든 상관 규칙에는 [`events/correlation/<이름>.json`](sigma_match/events/correlation/)
타임라인 파일이 있고, 모든 타임라인은 규칙으로 되짚어집니다. 타임라인의 각 이벤트는 상대 초를 담은 `ts`
필드를 지닙니다.
- **상관 규칙의 참 양성**: 임계를 시간 창 안에서 한 그룹이 채우는 양성 타임라인에 각 규칙이 발화합니다.
예를 들어 한 출처(`c-ip`)에서 10분 안에 서로 다른 20개 호스트로 퍼지는 요청이 수집 팬아웃 규칙을
발화시킵니다.
- **상관 규칙의 참 음성**: 임계 미달, 임계는 채웠지만 시간 창을 벗어난 경우, 그룹이 갈린 경우,
시간 상관에서 한쪽 레그가 빠진 경우에는 침묵합니다. 특히 요청량은 많아도 폭(서로 다른 호스트 수)이 작은
버스트는 팬아웃 규칙을 발화시키지 **않습니다**. 폭이 신호이지 양이 신호가 아니며, 이 판별을 증명하는 것이
바로 매칭 테스트입니다.
신뢰 모델은 이렇습니다. 손으로 짠 코드가 아니라 pySigma 가 각 규칙을 파싱합니다. 원자 규칙은 수식자와
조건을 트리로 컴파일하고(`|contains` → 와일드카드 값, `|all` → AND, `1 of selection_*` → OR), 상관 규칙은
집계 명세(유형·group-by·시간 창·임계 조건·참조하는 원자 규칙)로 컴파일합니다. [`check.py`](sigma_match/check.py)
는 그 트리와 명세를 따라 걷을 뿐이고, 상관 규칙이 어느 이벤트를 먹는지는 원자 규칙과 똑같은 매처로
판정하므로 권위 있는 Sigma 로직은 pySigma 안에 남습니다. 명시적으로 지원하지 않는 구문을 만나면 조용히
통과시키지 않고 예외를 던집니다(fail-closed). 범위와 한계는 스크립트 머리말에 밝혀 둡니다. 상관 규칙의
시간 창은 표준 슬라이딩 윈도(매칭 이벤트마다 `timespan` 길이의 창을 잡는) 해석이며, 실제 SIEM 의 윈도
방식은 다를 수 있습니다. 매칭은 **대소문자를 무시**하고(`sigma/` 스위트가 겨냥하는 splunk 백엔드의
기본값이며, 파괴 명령 규칙의 오탐 주석 자체가 이를 전제합니다), 키워드 매칭은 전문 부분 문자열 검색입니다.
이것은 규칙의 필드·값·조건·집계 로직에 대한 회귀 테스트이지, 필드 정규화가 다를 수 있는 각자의 SIEM 에서
검증하는 일을 대신하지는 않습니다.
### 실행
Docker 만 있으면 됩니다. pySigma 가 컨테이너에서 돌고 저장소에는 아무것도 쓰지 않습니다.
```sh
detections/tests/sigma_match/run.sh
```
예상 출력(축약):
```
PASS rule/sample pairing: 5 atomic rules, 5 event files, no orphans
PASS artex_enrich_user_agent: 1/1 positive events matched
PASS artex_recording_proxy_ca: 2/2 benign events correctly not matched
PASS rule/timeline pairing: 4 correlation rules, 4 timeline files, no orphans
PASS artex_enrich_fanout: fired — 20 distinct hosts from one source within the 10-minute window
PASS artex_enrich_fanout: quiet — high volume, low breadth: 25 requests from one source but only 4 distinct hosts
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료하므로 CI 나 pre-commit 훅에 그대로 넣을 수
있습니다. pySigma 는 기준 버전(`2.0.0`)으로 고정돼 있습니다. 내부에 미러를 두었다면 `PYSIGMA_VERSION`
으로 버전을, `PYTHON_IMAGE` 로 이미지를 재정의하십시오.
## Sigma 백엔드 이식성: [`sigma_backends/`](sigma_backends/)
[`sigma_backends/run.sh`](sigma_backends/run.sh) 는 규칙이 Sigma 테스트가 돌려 보는 단일 Splunk
예시를 넘어서도 변환됨을 증명하고, [`../README.ko.md`](../README.ko.md) 의 백엔드별 지원 표를 정직하게
유지합니다. Sigma 상관 규칙 변환은 백엔드에 따라 다르므로, README 는 어떤 `-t` 대상이 트리 전체를
받고 어떤 대상이 원자 규칙만 받는지 방어자에게 알려 줍니다. 다시 돌려 봐야만 믿을 수 있는
주장입니다. 두 가지 속성을 단언하는데, 둘 다 긍정형이라 실제 회귀가 있을 때만 실패합니다:
- **상관 규칙의 이식성**: 트리 전체(원자 + 상관)가 Splunk, Elasticsearch `eql` 대상, Grafana
`loki` 에서 변환되며, 보강 지표가 각 질의에 그대로 살아남습니다. 상관 규칙이 Splunk 전용이 아님을
보여 줍니다.
- **원자 전용 폴백 동작**: 다섯 개의 원자 규칙은 `lucene` 과 Microsoft `kusto` 백엔드에서도
변환됩니다. 이 백엔드들은 고정된 버전에서 Sigma 상관 규칙 변환을 지원하지 않으므로, 해당 백엔드를
쓰는 방어자는 원자 규칙을 배포하고 상관 윈도우는 그 백엔드 고유 기능으로 표현할 수 있습니다.
"백엔드 X 는 상관 규칙을 처리하지 못한다"는 부정형은 일부러 단언하지 않습니다. 그렇게 하면 백엔드가
*개선되는* 것이 빨간 빌드가 되기 때문입니다. 정직한 한계는 README 에 적어 두었고, 이 테스트의 명령이
그것을 재현합니다. [`sigma_backends/check.sh`](sigma_backends/check.sh) 는 컨테이너 안쪽 절반입니다.
고정된 sigma-cli 와 네 백엔드를 설치하고, 읽기 전용으로 마운트한 규칙 트리를 읽습니다.
### 실행
Docker 만 있으면 됩니다. sigma-cli 와 백엔드들이 컨테이너에서 돌고 저장소에는 아무것도 쓰지
않습니다.
```sh
detections/tests/sigma_backends/run.sh
```
예상 출력(축약):
```
PASS whole tree (atomic + correlation) converts on 'eql', enrich indicator survives
PASS five atomic rules convert on 'kusto', enrich indicator survives
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료합니다. sigma-cli 는 고정돼 있고(`3.1.0`,
`SIGMA_CLI_VERSION` 으로 재정의), 백엔드 플러그인은 호환되는 최신 버전으로 설치됩니다. 그래서 이
스위트는 상류 백엔드 릴리스에 가장 민감합니다. 지원을 떨어뜨린 플러그인은 빌드를 빨갛게 만들고, 이는
고정 버전과 README 표를 함께 갱신하라는 신호입니다.
## SigmaHQ 관례 린트: [`sigma_lint/`](sigma_lint/)
[`sigma_lint/run.sh`](sigma_lint/run.sh) 는 README 와 `CONTRIBUTING.md` 의 "`sigma check` 를 깨끗이
통과한다"는 약속이 pySigma 의 핵심 검사뿐 아니라 SigmaHQ 의 관례까지 포함하도록 만듭니다. 그냥
`sigma check` 는 `pySigma-validators-sigmahq` 플러그인을 적재하지 않으므로, 제목 대소문자·필드명 분류
체계·로그소스 분류 체계·참조 링크 관례가 검사되지 않고 지나갑니다. 이 스위트는 그 플러그인을 설치하고,
[`sigma_lint/validators.yml`](sigma_lint/validators.yml) 에 문서화한 기준선에 맞춰 전체 검사 집합을
돌립니다. 두 가지 속성을 단언합니다:
- **문서화한 기준선이 깨끗함**: `validators.yml` 과 함께 `sigma check` 를 돌리면 오류 0, 이슈 0 을
보고합니다.
- **전체 집합이 살아 있고, 문서화한 제외만 남음**: 제외 없이 모든 SigmaHQ 검증기를 돌려도 이슈가
보고되며, 그 각각은 `validators.yml` 이 일부러 끄는 네 검사 중 하나입니다(그 외에는 없음). 이것은
공허 방지 가드입니다. 플러그인이 적재에 실패했다면 전체 실행이 아무것도 보고하지 않아 첫 번째 속성이
잘못된 이유로 통과할 것이므로, 알려진 제외 항목이 반드시 나타나도록 요구합니다.
네 제외 항목은 SigmaHQ 의 모노레포 파일 정리 방식(로그소스 접두어가 붙은 파일명과 `correlation_`
파일명)과 분류 체계(일반 `application` 로그소스, 제품명 없는 `process_creation`), 그리고 브랜치 대
영구링크(permalink) 참조 관례를 담습니다. 어느 것도 자기 저장소의 살아 있는 문서를 참조하는 작고
독립적인 규칙 집합에는 맞지 않습니다. 각 제외 항목은 그 근거를 `validators.yml` 안에 함께 적어
두었습니다. *나머지* 모든 SigmaHQ 검사는 강제되므로, 새 관례 이슈를 들인 규칙(대소문자가 틀린 제목,
분류 체계를 벗어난 필드명)은 빌드를 빨갛게 만듭니다. `pySigma-validators-sigmahq` 는 고정돼 있고
(`0.21.0`, `SIGMAHQ_VALIDATORS_VERSION` 으로 재정의), 버전을 올리면 새 관례가 드러날 수 있는데, 이는
규칙이나 문서화한 기준선을 갱신하라는 신호입니다.
### 실행
Docker 만 있으면 됩니다. sigma-cli 와 검증기 플러그인이 컨테이너에서 돌고 저장소에는 아무것도 쓰지
않습니다.
```sh
detections/tests/sigma_lint/run.sh
```
예상 출력(축약):
```
PASS sigma check with the documented baseline: 0 errors, 0 issues
PASS every reported issue is one of the four documented exclusions
RESULT: PASS
```
## ATT&CK 레이어: [`attack/`](attack/)
[`attack/run.sh`](attack/run.sh) 는 [`../attack/artex_navigator_layer.json`](../attack/artex_navigator_layer.json)
의 [ATT&CK 커버리지 레이어](../attack/)가 커버한다고 주장하는 규칙과 어긋나지 않는지 확인합니다. 규칙
집합에서 어긋난 커버리지 레이어는 없느니만 못하므로, 이 테스트는 "이 규칙들이 이 ATT&CK 기법들을
커버한다"를 검토자가 소스에서 다시 돌려 볼 수 있는 것으로 바꿉니다. 단언하는 것:
- **유효한 레이어**: 파일이 JSON 으로 파싱되고 필수 ATT&CK Navigator v4.x 필드를 지니며, 모든
항목에 올바른 형식의 기법 ID 와 유효한 ATT&CK 전술이 있습니다.
- **양방향 일치**: 점수가 매겨진 기법이 Sigma 규칙의 `attack.*` 기법 태그와 *정확히* 일치합니다.
레이어에 빠진 규칙 기법도, 규칙에 없는 레이어 기법도 없습니다. 전술도 같은 방식으로 일치합니다.
- **근거 있음**: 점수가 매겨진 모든 기법의 주석이 실재하는 규칙 파일을 가리키므로, 레이어가 이름이
바뀌거나 삭제된 규칙을 인용할 수 없습니다.
이것은 발화 테스트가 아니라 일관성 검사입니다. 탐지 백엔드가 필요 없고 Python 표준 라이브러리만 있으면
되므로, Sigma·Suricata 테스트와 달리 버전에 의존하는 경보 수가 없습니다. [`attack/check.py`](attack/check.py)
는 컨테이너 안쪽 절반입니다. 읽기 전용으로 마운트한 탐지 트리를 읽고 아무것도 쓰지 않습니다.
### 실행
Docker 만 있으면 됩니다. 검사가 Python 컨테이너에서 돌고 저장소에는 아무것도 쓰지 않습니다.
```sh
detections/tests/attack/run.sh
```
예상 출력(축약):
```
PASS scored techniques match the rule set exactly (8: T1059, T1105, T1485, T1489, T1557, T1561.002, T1592, T1595)
PASS scored tactics match the rule set exactly (collection, command-and-control, credential-access, execution, impact, reconnaissance)
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료하므로 CI 나 pre-commit 훅에 그대로 넣을 수
있습니다. 내부에 미러를 두었다면 `PYTHON_IMAGE` 로 이미지를 재정의하십시오.
## 지표 근거(source-of-truth): [`indicators/`](indicators/)
[`indicators/run.sh`](indicators/run.sh) 는 위 세 테스트가 하지 못하는 한 가지를 증명합니다. 각 규칙이
고정한 지표가 여전히 ARTEX 자신의 소스가 실제로 내보내는 문자열인지입니다. Sigma 테스트는 지표가
규칙→질의 *컴파일*을 거쳐 살아남음을 증명하고, ATT&CK 테스트는 레이어가 규칙 태그와 일치함을
증명하며, Suricata 테스트는 네트워크 규칙이 생성한 캡처에서 *발화*함을 증명합니다. 어느 것도 지표가
유래했다고 주장하는 소스 파일을 되짚어 보지는 않습니다. 이들이 모두 놓치는 부패는, 프로버 User-Agent
를 `artex-enrich/2.0` 으로 올리거나 가드 마커를 다시 쓰는 상류 재동기화입니다. 그래도 규칙은 모두
컴파일되고, 레이어는 여전히 일치하고, pcap 테스트도 여전히 발화합니다. 그런데 배포된 규칙은 실제
ARTEX 트래픽에 조용히 매칭을 멈춥니다. 각 지표에 대해 양방향으로 단언합니다:
- **소스가 여전히 내보냄**: 값이 그것을 만들어 내는 상류 소스 파일에 존재합니다(`enrich/enrich.go`
의 `artex-enrich/1.0`, `selfupdate/` 의 `artex-selfupdate`, `guard/guard.go` 의 가드 마커). 값이
없다는 것은 규칙이 아직 따라잡지 못한 상류 변경을 뜻합니다.
- **규칙이 여전히 고정함**: 값이 그것을 기반으로 세운 규칙에 존재하므로, 규칙 편집이 지표를 소스에서
조용히 떼어 놓을 수 없습니다. Suricata 규칙은 `startswith` 접두어로 확인하는데, 이는 그 규칙이
실제로 와이어를 매칭하는 방식과 같습니다.
- **차단 목록 대응**: 파괴적 명령 토큰(`rm -rf`, `mkfs`, `DROP DATABASE`, `FLUSHALL`)이 ARTEX
가드의 차단 목록(`db/db.go`)과 그것을 반영한 헌팅 규칙 양쪽에 나타납니다. 이것들은 고유 지문이
아니라 일반 헌팅 단서이므로, 테스트는 규칙이 실제로 주장하는 대응 관계만 단언합니다.
- **공개 목록이 근거를 유지함**: 방어자가 가져다 쓰는 산출물인 기계가 읽는 지표 목록
[`detections/indicators/artex_indicators.csv`](../indicators/artex_indicators.csv) 을 행 단위로 다시
읽습니다. 모든 값은 인용한 소스 파일에 여전히 존재하고 인용한 규칙에 고정돼 있어야 하며, 테스트가
근거를 확인한 모든 지문은 이 목록에 나타나야 합니다. 그래서 공개된 CSV 는 자신이 유래했다고 주장하는
소스에서 어느 방향으로도 조용히 어긋날 수 없습니다.
- **두 관문이 고정된 각 소스에서 발화함**: 테스트가 읽는 모든 상류 소스는 그것을 돌리는 두 관문에
포함됩니다. CI 워크플로의 `push`·`pull_request` paths 필터([`.github/workflows/detections.yml`](../../.github/workflows/detections.yml))와
로컬 pre-commit 훅의 `files` 정규식([`.pre-commit-config.yaml`](../../.pre-commit-config.yaml))입니다.
필요한 집합은 지표 자체에서 파생되므로, 새 소스를 고정하면서(예전의 `cmd/artex/main.go` 포트가
그랬듯) *두* 관문에 모두 배선하지 않으면 여기서 실패합니다. 그러지 않으면 그 소스만 건드린 변경이
그 소스를 빠뜨린 관문에서 테스트를 건너뜁니다. CI 에서는 머지 게이트를 초록으로 통과하고, 훅에서는
"CI 와 같은 소스 범위"라고 약속해 놓고도 로컬에서 끝내 잡히지 않습니다.
이는 [`../README.ko.md`](../README.ko.md) 의 약속("여기 모든 지표는 추정이 아니라 이 저장소 소스에서 확인한
문자열에 근거한다")과 CONTRIBUTING 의 첫 번째 기여 계약을, 검토자가 다시 돌려 볼 수 있는 가드로
바꿉니다. ATT&CK 테스트처럼 탐지 백엔드가 필요 없고 Python 표준 라이브러리만 있으면 됩니다.
[`indicators/check.py`](indicators/check.py) 는 규칙 트리, 공개 지표 목록, 고정된 소스 패키지, 그리고
그것을 발화시키는 두 관문(CI 워크플로와 pre-commit 설정)을 읽기 전용으로 마운트해 읽고, 아무것도 쓰지
않습니다.
### 실행
Docker 만 있으면 됩니다. 검사가 Python 컨테이너에서 돌고 저장소에는 아무것도 쓰지 않습니다.
```sh
detections/tests/indicators/run.sh
```
예상 출력(축약):
```
PASS enrichment prober User-Agent: 'artex-enrich/1.0' emitted by enrich/enrich.go
PASS detections/sigma/artex_enrich_user_agent.yml pins 'artex-enrich/1.0'
PASS 'FLUSHALL' present in both db/db.go and detections/sigma/destructive_command_hunting.yml
PASS enrich-user-agent: 'artex-enrich/1.0' grounded in enrich/enrich.go
PASS tested fingerprint 'artex-enrich/1.0' is published in the list
PASS .github/workflows/detections.yml push paths covers cmd/artex/main.go
PASS .pre-commit-config.yaml files covers cmd/artex/main.go
RESULT: PASS
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료하므로 CI 나 pre-commit 훅에 그대로 넣을 수
있습니다. 내부에 미러를 두었다면 `PYTHON_IMAGE` 로 이미지를 재정의하십시오.
## MISP 내보내기 일관성: [`misp/`](misp/)
[`misp/run.sh`](misp/run.sh) 는 지표의 두 번째 공개 형태, 바로 가져올 수 있는 MISP 이벤트
[`detections/indicators/artex_indicators.misp.json`](../indicators/artex_indicators.misp.json) 를
다룹니다. 위 지표 테스트가 CSV 를 소스에 근거하게 유지한다면, 이 테스트는 방어자가 실제로 위협
인텔리전스 플랫폼에 적재하는 산출물인 MISP 이벤트가 그 CSV 에서 어긋나지 않게 유지합니다. 단언하는 것:
- **정말로 MISP 임**: 이벤트가 [pymisp](https://github.com/MISP/PyMISP) 로 적재되는데, 그 객체
모델은 `type` 이 진짜 MISP 타입이 아닌 속성을 거부합니다. 그럴듯해 보여도 유효하지 않은 타입은
여기서 실패하므로, "유효한 MISP"는 그냥 주장하는 것이 아니라 MISP 서버가 쓰는 라이브러리로
증명됩니다.
- **CSV 와 행 단위 동기화**: 모든 CSV 행이 의도한 타입·카테고리를 가진 MISP 속성 정확히 하나로
대응되고(`http.user-agent` → `user-agent`, 가드 마커 `string` → `pattern-in-file`, `port` →
`port`, `ip-dst|port` → 합성 `ip|port` 값을 가진 `ip-dst|port`, 탐색 스키마 `other` → `other`), CSV 행 없이 남는 MISP 속성이
하나도 없습니다. 이벤트는 CSV 와 함께 손으로 유지하므로, `artex_indicators.misp.json` 을 같은
커밋에서 맞춰 갱신하지 않은 채 CSV 행을 추가·삭제·타입 변경하면 실패합니다.
- **`to_ids` 가 `rule` 열을 반영함**: 규칙이 뒷받침하는 지표는 `to_ids: true` 이고, 규칙이 없는
호스트 포렌식 행은 `disable_correlation: true` 와 함께 `to_ids: false` 입니다. CSV 가 함의하는
것과 다르게 플래그를 뒤집으면 실패하므로, MISP 이벤트는 어떤 지문이 실행 가능한지를 조용히
부풀리거나 줄여 주장할 수 없습니다.
- **가드 마커가 바이트 단위로 보존되고** `detections/**` 가 CI paths 필터에 있어, CSV 나 이벤트를
바꾸면 이 스위트가 발화합니다.
위의 순수 표준 라이브러리 테스트들과 달리, 이 스위트는 컨테이너 안에 고정된 `pymisp` 를
설치합니다(호스트에는 아무것도 설치하지 않음). [`misp/check.py`](misp/check.py) 는 CSV, MISP 이벤트,
CI 워크플로를 읽기 전용으로 마운트해 읽고, 아무것도 쓰지 않습니다.
### 실행
Docker 만 있으면 됩니다. pymisp 가 컨테이너에 설치되고 저장소에는 아무것도 쓰지 않습니다.
```sh
detections/tests/misp/run.sh
```
단언이 하나라도 실패하면 스크립트가 0 이 아닌 코드로 종료합니다. 내부에 미러를 두었다면
`PYTHON_IMAGE` 로 이미지를, `PYMISP_VERSION` 으로 고정된 라이브러리를 재정의하십시오.
## 기여
새 탐지 규칙은 그것이 발화함을 보여 주는 테스트가 있을 때 더 강합니다. 테스트는 자기 입력을 결정론적으로
생성하고, 엔진 버전과 무관한 속성은 정확히 단언하며(그보다 무른 속성은 기준값을 기록한 하한으로), 공격
안내로 읽힐 수 있는 내용은 피해야 합니다. [`../../CONTRIBUTING.md`](../../CONTRIBUTING.md) 와
[`../README.ko.md`](../README.ko.md) 의 규칙 색인을 보십시오.
여덟 스위트는 모두 `detections/` 를 건드리는 모든 push 나 pull request 에서 CI 로 돕니다
([`../../.github/workflows/detections.yml`](../../.github/workflows/detections.yml) 참조). 그리고 지표
테스트는 그것이 고정한 상류 소스 파일(`enrich/`, `selfupdate/`, `guard/`, `db/`, `cmd/artex/main.go`)이
바뀔 때도 돕니다. 그래서 지표를 떨어뜨리거나, ATT&CK 레이어에서 어긋나거나, 문서화한 백엔드에서
변환이 멈추거나, SigmaHQ 관례를 깨거나, 소스와 동기화가 어긋나거나, MISP 이벤트가 CSV 에서 어긋나게
두거나, 워크플로가 아직 감시하지 않는 새 소스를 고정하는 규칙 변경은 머지되기 전에 빌드를 빨갛게
만듭니다.
+456
View File
@@ -0,0 +1,456 @@
# ARTEX detection rule tests
English · [한국어](README.ko.md)
> 한국어: 이 디렉터리는 [`../`](../)의 탐지 규칙이 실제로 발화하는지를 재현 가능하게 증명하는
> 회귀 테스트입니다. 바이너리 캡처를 저장소에 넣지 않고, 패킷 캡처를 매번 결정론적으로 생성한 뒤
> [Suricata](https://suricata.io)로 직접 돌려 경보 수를 확인합니다. 모든 테스트는 자신이 소유하거나
> 서면 허가를 받은 시스템을 지키는 **방어·탐지 목적에만** 쓰십시오. 한국어 전체 문서는
> **[README.ko.md](README.ko.md)** 를 보십시오.
Reproducible regression tests that prove the rules under [`../`](../) actually fire — and, just as
important, stay silent on benign traffic. A detection rule you cannot run is a claim; these tests turn the
claims in the rule files and the defense guide into something a reviewer can re-run from source.
No binary packet capture is committed. The capture is **synthesized deterministically on every run** and
removed afterwards, so the test ships as readable source, not as an opaque fixture, and never bloats the
repository.
## Run every suite at once — [`run-all.sh`](run-all.sh)
[`run-all.sh`](run-all.sh) runs all eight suites below in one command, in the same order as CI, so you do not
have to invoke the eight `run.sh` scripts by hand. Each suite runs to completion even if an earlier one fails,
the script prints a one-line PASS/FAIL summary per suite at the end, and it exits non-zero if any suite failed.
Before the suites, it runs a harness self-check ([`check-harness-sync.sh`](check-harness-sync.sh)) that fails
the run if this suite list, the per-suite steps in [CI](../../.github/workflows/detections.yml), and the suite
directories on disk ever name different suites or a different order. That is the one gap the eight suites
cannot see on their own: a suite wired into only one of the three (a new CI step with no `run-all.sh` entry, or
a directory never added to either) would otherwise pass every per-suite test while a green local `run-all.sh`
quietly stopped meaning a green CI. The check is a gate, not a ninth suite: it stays out of the summary
below, so the eight detection suites stay eight.
```sh
detections/tests/run-all.sh
```
Expected output (abridged):
```
===== detection suites summary =====
PASS sigma
PASS sigma_match
PASS sigma_lint
PASS sigma_backends
PASS suricata
PASS attack
PASS indicators
PASS misp
RESULT: PASS
```
Because it exits non-zero on any failure, it drops straight into a pre-commit hook. A ready-to-use example
lives in [`.pre-commit-config.yaml`](../../.pre-commit-config.yaml) at the repository root: install it with
`pip install pre-commit && pre-commit install`, and the runner then fires on commits that touch the detection
rules or the upstream source files they pin — the same scope as CI. The image and version overrides the
individual suites honour (`PYTHON_IMAGE`, `SIGMA_CLI_VERSION`, `SIGMAHQ_VALIDATORS_VERSION`, `SURICATA_IMAGE`)
are inherited by the runner, so exporting any of them applies to every suite at once.
## Suricata — [`suricata/`](suricata/)
[`suricata/run.sh`](suricata/run.sh) exercises the network rules in
[`../suricata/artex.rules`](../suricata/artex.rules) end to end and asserts five properties:
- **Valid** — the whole rules file loads under `suricata -T --init-errors-fatal`, so a rule that fails to
parse or initialise is caught even when no capture below exercises it. Plain `suricata -r` skips such a
rule and still exits 0, so this load check is the Suricata analogue of the Sigma suite's `sigma check`
validity assertion.
- **Presence (enrich)** — sid `1000001` fires exactly once per enrichment probe.
- **Velocity** — sid `1000002` fires once the `detection_filter` rate of 30 requests in 300 s per source is
crossed.
- **Presence (WebFetch)** — sid `1000003` fires exactly once per norma WebFetch request, and the enrich sids
stay silent on that same capture — so the two network signatures are mutually specific, not just each
present.
- **Specificity** — an identical capture whose only change is a benign browser User-Agent produces **zero**
ARTEX alerts.
[`suricata/gen_pcap.py`](suricata/gen_pcap.py) builds the capture with [scapy](https://scapy.net): N
independent plaintext HTTP request/response flows from one fixed source, each carrying a chosen
User-Agent, at a fixed base timestamp spaced one second apart. It only writes a file — it never sends a
packet or touches a network.
### Run it
Needs only Docker; scapy and Suricata both run in containers.
```sh
detections/tests/suricata/run.sh
```
Expected output (abridged):
```
PASS ruleset loads with zero parse/init errors (suricata -T)
PASS sid 1000001 presence: one alert per probe (got 35, want eq 35)
PASS sid 1000002 velocity: fires past 30-in-300s (got 5, want ge 1)
PASS sid 1000003 presence: one alert per WebFetch request (got 8, want eq 8)
PASS enrich sids stay silent on norma traffic (specificity) (got 0, want eq 0)
PASS benign browser UA produces no ARTEX alerts (got 0, want eq 0)
RESULT: PASS
```
The script exits non-zero if any assertion fails, so it drops straight into CI or a pre-commit hook.
Override the images with `SURICATA_IMAGE` / `PYTHON_IMAGE` if you mirror them internally.
### Why the velocity count is a floor, not an exact match
`run.sh` asserts the presence counts (`1000001 == 35`, `1000003 == 8`) and the benign count (`== 0`) exactly,
because those are engine-version-independent: one alert per matching request, and no match on a different
User-Agent. The
velocity rule's count depends on how a given Suricata release resolves the `detection_filter` threshold at
the boundary, so the test asserts `>= 1` and records the reference value separately. On **Suricata 8.0.7**
the reference run produces **5** alerts on sid `1000002` (flows 31–35, after the 30-in-300 s threshold is
crossed).
## Sigma — [`sigma/`](sigma/)
[`sigma/run.sh`](sigma/run.sh) validates the Sigma rules under [`../sigma/`](../sigma/) structurally and by
compilation with [sigma-cli](https://github.com/SigmaHQ/sigma-cli) (pySigma), and asserts five properties:
- **Valid** — `sigma check` reports 0 errors, 0 condition errors, and 0 issues over the whole tree.
- **Compiles** — `sigma convert -t splunk` turns the whole tree into a backend query language without error.
- **Indicators survive** — each atomic indicator string (`artex-enrich/1.0`, `artex-selfupdate`, the guard
marker, and the recording-proxy CA filename `mitmproxy-ca-cert.pem`) is still present in the compiled query,
so a rule cannot silently lose the string it is built on.
- **Correlations compile** — the behaviour rules in [`../sigma/correlation/`](../sigma/correlation/) emit their
`event_count` / `value_count` aggregations rather than being dropped.
- **Correlations are load-bearing** — converting one correlation rule *alone* fails, because it references its
atomic base rule by `id`; the reference is enforced, not decorative. This is the Sigma analogue of the
Suricata specificity assertion above.
This is the structural + compilation validation documented in [`../README.md`](../README.md), made executable
and assertive. The companion [`sigma_match/`](sigma_match/) suite below adds the *matching* half for both the
atomic and the correlation rules — a representative malicious event (or timeline) fires each rule and a benign
one does not — so the Sigma rules now get both a reproducible validation test and a reproducible matching test,
the way the Suricata rule does. (The
earlier concern that a weak hand-written matcher would undercut the rules is addressed by delegating all parsing
to pySigma; see the trust model in the next section.)
### Run it
Needs only Docker; sigma-cli and the splunk backend run in a container and nothing is written to the repo.
```sh
detections/tests/sigma/run.sh
```
Expected output (abridged):
```
PASS sigma check: 0 errors, 0 condition errors, 0 issues
PASS whole tree converts to splunk (exit 0)
PASS indicator present: artex-enrich/1.0
PASS correlation rule fails to convert alone — it requires its atomic base rule
RESULT: PASS
```
The script exits non-zero if any assertion fails, so it drops straight into CI or a pre-commit hook. sigma-cli
is pinned to a reference version (`3.1.0`); override it with `SIGMA_CLI_VERSION`, or the image with
`PYTHON_IMAGE`, if you mirror them internally.
## Sigma live event-matching — [`sigma_match/`](sigma_match/)
[`sigma_match/run.sh`](sigma_match/run.sh) proves the Sigma rules under [`../sigma/`](../sigma/) — both the
atomic rules and the correlation rules under [`../sigma/correlation/`](../sigma/correlation/) — actually *fire*
on a matching event (or timeline) and stay quiet on a benign one — the "a detection you cannot run is only a
claim" guarantee the Suricata suite gives the network rule, extended here to the host/log-layer rules. It
asserts six properties, three for the atomic rules and three for the correlations:
- **Rule/sample pairing** — every atomic rule has an [`events/<name>.json`](sigma_match/events/) sample file and
every sample file maps back to a rule, so a rule added without samples fails here rather than going untested.
- **True positives** — each rule matches every one of its malicious sample events.
- **True negatives** — each rule matches none of its benign sample events. For example, a standalone
`mitmproxy-ca-cert.pem` under `.mitmproxy/` does **not** trip the recording-proxy rule, because its `|all`
modifier also requires the `_ca/` directory ARTEX writes — the matching test is what proves that discrimination.
- **Correlation rule/timeline pairing** — every correlation rule has an
[`events/correlation/<name>.json`](sigma_match/events/correlation/) timeline file and every timeline maps back
to a rule. Each timeline event carries a `ts` field in relative seconds.
- **Correlation true positives** — each rule fires on a positive timeline where the threshold is met inside the
window within one group. For example, requests from one source (`c-ip`) fanning out to 20 distinct hosts
within 10 minutes trip the enrichment fan-out rule.
- **Correlation true negatives** — each rule stays quiet when the threshold is not met, when it is met but the
events are spread beyond the window, when they are split across groups, or when a temporal rule is missing a
leg. In particular a high-volume, low-breadth burst does **not** trip the fan-out rule: breadth, not volume, is
the signal, and the matching test is what proves that discrimination.
The trust model is that pySigma — not hand-written code — parses each rule: an atomic rule into a condition tree
(`|contains` → a wildcard value, `|all` → an AND, `1 of selection_*` → an OR), and a correlation rule into its
aggregation spec (type, group-by, timespan, threshold, and the resolved references to the atomic base rules).
[`check.py`](sigma_match/check.py) only walks that tree and spec, deciding which events feed a referenced rule
with the very same atomic matcher, so the authoritative Sigma logic stays in pySigma; it raises rather than
passing on any construct it does not explicitly support (fail-closed). Scope and limits are stated in the script
header: the correlation window is the standard sliding-window interpretation (a `timespan`-second window
anchored at each matching event) and a real SIEM's windowing may differ; matching is **case-insensitive** (the
splunk-backend default the `sigma/` suite targets, which the destructive rule's own false-positive note
assumes); and keyword matching is a full-text substring search. It is a regression test for the rules'
field/value/condition/aggregation logic, not a substitute for validating in your own SIEM, whose field
normalization may differ.
### Run it
Needs only Docker; pySigma runs in a container and nothing is written to the repo.
```sh
detections/tests/sigma_match/run.sh
```
Expected output (abridged):
```
PASS rule/sample pairing: 5 atomic rules, 5 event files, no orphans
PASS artex_enrich_user_agent: 1/1 positive events matched
PASS artex_recording_proxy_ca: 2/2 benign events correctly not matched
PASS rule/timeline pairing: 4 correlation rules, 4 timeline files, no orphans
PASS artex_enrich_fanout: fired — 20 distinct hosts from one source within the 10-minute window
PASS artex_enrich_fanout: quiet — high volume, low breadth: 25 requests from one source but only 4 distinct hosts
RESULT: PASS
```
The script exits non-zero if any assertion fails, so it drops straight into CI or a pre-commit hook. pySigma is
pinned to a reference version (`2.0.0`); override it with `PYSIGMA_VERSION`, or the image with `PYTHON_IMAGE`,
if you mirror them internally.
## Sigma backend portability — [`sigma_backends/`](sigma_backends/)
[`sigma_backends/run.sh`](sigma_backends/run.sh) proves the rules convert beyond the single Splunk example the
Sigma test exercises, and keeps the per-backend support matrix in [`../README.md`](../README.md) honest. Sigma
correlation conversion is backend-dependent, so the README tells a defender which `-t` targets take the whole
tree and which take only the atomic rules — a claim that is only trustworthy if it is re-run. It asserts two
properties, both positive so the test fails only on a real regression:
- **Correlations are portable** — the whole tree (atomic + correlation) converts on Splunk, the Elasticsearch
`eql` target, and Grafana `loki`, with the enrich indicator surviving into each query. This shows the
correlation rules are not Splunk-only.
- **Atomic-only fallback works** — the five atomic rules still convert on `lucene` and the Microsoft `kusto`
backend, which do not support Sigma correlation conversion at the pinned versions, so a defender on those
backends can deploy the atomic rules and express the correlation window natively.
It deliberately does not assert the negative "backend X cannot do correlations": that would turn a backend
*improving* into a red build. The honest limitation lives in the README, reproduced by this test's commands.
[`sigma_backends/check.sh`](sigma_backends/check.sh) is the in-container half; it installs the pinned sigma-cli
plus four backends and reads the rule tree mounted read-only.
### Run it
Needs only Docker; sigma-cli and the backends run in a container and nothing is written to the repo.
```sh
detections/tests/sigma_backends/run.sh
```
Expected output (abridged):
```
PASS whole tree (atomic + correlation) converts on 'eql', enrich indicator survives
PASS five atomic rules convert on 'kusto', enrich indicator survives
RESULT: PASS
```
The script exits non-zero if any assertion fails. sigma-cli is pinned (`3.1.0`, override with
`SIGMA_CLI_VERSION`); the backend plugins install at their latest compatible version, so this suite is the one
most sensitive to an upstream backend release — a plugin that drops support turns the build red, which is the
signal to update the pin and the README matrix together.
## SigmaHQ convention lint — [`sigma_lint/`](sigma_lint/)
[`sigma_lint/run.sh`](sigma_lint/run.sh) makes the "passes `sigma check` cleanly" promise in the README and
`CONTRIBUTING.md` cover SigmaHQ's conventions, not just pySigma's core checks. Plain `sigma check` does not
load the `pySigma-validators-sigmahq` plugin, so title casing, field-name taxonomy, logsource taxonomy, and
reference-link conventions go unchecked. This suite installs that plugin and runs the full set against the
documented baseline in [`sigma_lint/validators.yml`](sigma_lint/validators.yml). It asserts two properties:
- **The documented baseline is clean** — `sigma check` with `validators.yml` reports 0 errors and 0 issues.
- **The full set is live, and only the documented exclusions remain** — running every SigmaHQ validator with
no exclusions still reports issues, and each one is among the four checks `validators.yml` deliberately
disables (nothing else). This is the anti-vacuity guard: if the plugin failed to load, the full run would
report nothing and the first property would pass for the wrong reason, so the known exclusions are required
to appear.
The four exclusions encode SigmaHQ's monorepo filing scheme (logsource-prefixed and `correlation_` filenames)
and taxonomy (a generic `application` logsource and a product-less `process_creation`), plus the
branch-vs-permalink reference convention — none of which fit a small standalone rule set that references its
own living docs. Each exclusion carries its rationale inline in `validators.yml`. Because every *other*
SigmaHQ check is enforced, a rule that picks up a new convention issue — a mis-cased title, an off-taxonomy
field name — turns the build red. `pySigma-validators-sigmahq` is pinned (`0.21.0`, override with
`SIGMAHQ_VALIDATORS_VERSION`); bumping it may surface new conventions, which is the signal to update the rules
or the documented baseline.
### Run it
Needs only Docker; sigma-cli and the validator plugin run in a container and nothing is written to the repo.
```sh
detections/tests/sigma_lint/run.sh
```
Expected output (abridged):
```
PASS sigma check with the documented baseline: 0 errors, 0 issues
PASS every reported issue is one of the four documented exclusions
RESULT: PASS
```
## ATT&CK layer — [`attack/`](attack/)
[`attack/run.sh`](attack/run.sh) checks that the [ATT&CK coverage layer](../attack/) in
[`../attack/artex_navigator_layer.json`](../attack/artex_navigator_layer.json) stays consistent with the
rules it claims to cover. A coverage layer that drifts from its rule set is worse than none, so this turns
"these rules cover these ATT&CK techniques" into something a reviewer can re-run from source. It asserts:
- **Valid layer** — the file parses as JSON and carries the required ATT&CK Navigator v4.x fields, with a
well-formed technique ID and a valid ATT&CK tactic on every entry.
- **Bidirectional match** — the scored techniques are *exactly* the `attack.*` technique tags on the Sigma
rules: no rule technique missing from the layer, no layer technique absent from the rules. The tactics
match the same way.
- **Grounded** — every scored technique's comment names a rule file that exists, so the layer cannot cite a
rule that was renamed or removed.
This is a consistency check, not a firing test: it needs no detection backend, only the Python standard
library, so unlike the Sigma and Suricata tests it carries no version-dependent counts. [`attack/check.py`](attack/check.py)
is the in-container half; it reads the detections tree mounted read-only and writes nothing.
### Run it
Needs only Docker; the check runs in a Python container and nothing is written to the repo.
```sh
detections/tests/attack/run.sh
```
Expected output (abridged):
```
PASS scored techniques match the rule set exactly (8: T1059, T1105, T1485, T1489, T1557, T1561.002, T1592, T1595)
PASS scored tactics match the rule set exactly (collection, command-and-control, credential-access, execution, impact, reconnaissance)
RESULT: PASS
```
The script exits non-zero if any assertion fails, so it drops straight into CI or a pre-commit hook.
Override the image with `PYTHON_IMAGE` if you mirror it internally.
## Indicator source-of-truth — [`indicators/`](indicators/)
[`indicators/run.sh`](indicators/run.sh) proves the one thing the three tests above do not: that each rule's
pinned indicator is still the string ARTEX's own source actually emits. The Sigma test proves an indicator
survives rule→query *compilation*; the ATT&CK test proves the layer matches the rules' tags; the Suricata
test proves the network rule *fires* on a synthesized capture. None of them look back at the source file the
indicator claims to come from. The rot they all miss is an upstream re-sync that bumps the prober User-Agent
to `artex-enrich/2.0` or rewrites the guard marker: every rule still compiles, the layer still matches, the
pcap test still fires — and the deployed rule silently stops matching real ARTEX traffic. It asserts, for
each indicator, bidirectionally:
- **Source still emits it** — the value is present in the upstream source file(s) that produce it
(`artex-enrich/1.0` in `enrich/enrich.go`, `artex-selfupdate` in `selfupdate/`, the guard marker in
`guard/guard.go`). A missing value means an upstream change the rule has not caught up with.
- **Rule still pins it** — the value is present in the rule built on it, so a rule edit cannot quietly move
the indicator away from its source. The Suricata rule is checked by its `startswith` prefix, matching how
it actually matches the wire.
- **Deny-list correspondence** — the destructive-command tokens (`rm -rf`, `mkfs`, `DROP DATABASE`,
`FLUSHALL`) appear both in ARTEX's guard deny-list (`db/db.go`) and in the hunting rule that mirrors it.
These are generic hunting leads, not unique fingerprints, so the test asserts only the correspondence the
rule actually claims.
- **Published list stays grounded** — the machine-readable indicator list
[`detections/indicators/artex_indicators.csv`](../indicators/artex_indicators.csv), the artifact a
defender imports, is re-read row by row: every value must still be present in the source file(s) it cites
and pinned in the rule(s) it cites, and every fingerprint the test grounds must appear in the list. So the
published CSV cannot silently drift from the source it claims to come from, in either direction.
- **Both gates fire on each pinned source** — every upstream source the test reads is covered by the two
gates that run it: the CI workflow's `push` and `pull_request` paths filter
([`.github/workflows/detections.yml`](../../.github/workflows/detections.yml)) and the local pre-commit
hook's `files` regex ([`.pre-commit-config.yaml`](../../.pre-commit-config.yaml)). The required set is
derived from the indicators themselves, so pinning a new source (as the `cmd/artex/main.go` ports once were)
without wiring it into *both* gates fails here — otherwise a change touching only that source skips the test
on whichever gate omits it: on CI it passes the merge gate green, on the hook it is never caught locally
even though the hook promises "the same source scope as CI".
This turns [`../README.md`](../README.md)'s promise — "every indicator here is grounded in a string verified
in this repository's source, not inferred" — and CONTRIBUTING's first contribution contract into a guard a
reviewer can re-run. Like the ATT&CK test it needs no detection backend, only the Python standard library;
[`indicators/check.py`](indicators/check.py) reads the rule tree, the published indicator list, the pinned
source packages, and the two gates that fire it (the CI workflow and the pre-commit config) mounted read-only
and writes nothing.
### Run it
Needs only Docker; the check runs in a Python container and nothing is written to the repo.
```sh
detections/tests/indicators/run.sh
```
Expected output (abridged):
```
PASS enrichment prober User-Agent: 'artex-enrich/1.0' emitted by enrich/enrich.go
PASS detections/sigma/artex_enrich_user_agent.yml pins 'artex-enrich/1.0'
PASS 'FLUSHALL' present in both db/db.go and detections/sigma/destructive_command_hunting.yml
PASS enrich-user-agent: 'artex-enrich/1.0' grounded in enrich/enrich.go
PASS tested fingerprint 'artex-enrich/1.0' is published in the list
PASS .github/workflows/detections.yml push paths covers cmd/artex/main.go
PASS .pre-commit-config.yaml files covers cmd/artex/main.go
RESULT: PASS
```
The script exits non-zero if any assertion fails, so it drops straight into CI or a pre-commit hook.
Override the image with `PYTHON_IMAGE` if you mirror it internally.
## MISP export consistency — [`misp/`](misp/)
[`misp/run.sh`](misp/run.sh) covers the second published form of the indicators — the ready-to-import MISP
event [`detections/indicators/artex_indicators.misp.json`](../indicators/artex_indicators.misp.json). The
indicator test above keeps the CSV grounded in the source; this test keeps the MISP event, the artifact a
defender actually loads into a threat-intelligence platform, from drifting away from that CSV. It asserts:
- **It is really MISP** — the event loads under [pymisp](https://github.com/MISP/PyMISP), whose object model
rejects any attribute whose `type` is not a genuine MISP type. A plausible-looking but invalid type fails
here, so "valid MISP" is proven by the library a MISP server uses, not asserted.
- **Row-for-row sync with the CSV** — every CSV row maps to exactly one MISP attribute with the intended type
and category (`http.user-agent` → `user-agent`, the guard marker `string` → `pattern-in-file`, `port` →
`port`, `ip-dst|port` → `ip-dst|port` with the composite `ip|port` value, and the exploration-schema `other` → `other`), and no MISP attribute is left
without a CSV row. The event is hand-maintained alongside the CSV, so adding, removing, or retyping a CSV
row without updating `artex_indicators.misp.json` to match in the same commit fails.
- **`to_ids` mirrors the `rule` column** — a rule-backed indicator is `to_ids: true`; a host-forensic row
with no rule is `to_ids: false` with `disable_correlation: true`. Flipping a flag away from what the CSV
implies fails, so the MISP event cannot quietly over- or under-claim which fingerprints are actionable.
- **The guard marker survives byte-for-byte** and `detections/**` is in the CI paths filter, so a change to
the CSV or the event triggers this suite.
Unlike the pure-standard-library tests above, this suite installs a pinned `pymisp` inside its container
(nothing is installed on the host); [`misp/check.py`](misp/check.py) reads the CSV, the MISP event, and the
CI workflow mounted read-only and writes nothing.
### Run it
Needs only Docker; pymisp is installed in the container and nothing is written to the repo.
```sh
detections/tests/misp/run.sh
```
The script exits non-zero if any assertion fails. Override the image with `PYTHON_IMAGE` and the pinned
library with `PYMISP_VERSION` if you mirror them internally.
## Contributing
A new detection rule is stronger with a test that shows it firing. Tests should synthesize their own input
deterministically, assert engine-version-independent properties exactly (and softer ones as floors with a
recorded reference), and avoid any content that reads as attack guidance. See
[`../../CONTRIBUTING.en.md`](../../CONTRIBUTING.en.md) and the rule indexes in [`../README.md`](../README.md).
All eight suites run in CI (see [`../../.github/workflows/detections.yml`](../../.github/workflows/detections.yml))
on every push or pull request that touches `detections/` — and the indicator test also runs when the upstream
source files it pins (`enrich/`, `selfupdate/`, `guard/`, `db/`, `cmd/artex/main.go`) change — so a rule
change that drops an indicator, drifts from the ATT&CK layer, stops converting on a documented backend,
breaks a SigmaHQ convention, falls out of sync with the source, lets the MISP event drift from the CSV, or
pins a new source the workflow does not yet watch turns the build red before it can merge.
+191
View File
@@ -0,0 +1,191 @@
#!/usr/bin/env python3
#
# In-container half of the ARTEX ATT&CK coverage-layer test. run.sh launches this
# inside a Python container with the detections tree mounted read-only at
# /detections. It proves that the ATT&CK Navigator layer in
# detections/attack/artex_navigator_layer.json stays consistent with the rules it
# claims to cover, so the layer cannot silently drift from the Sigma rule set:
#
# 1. the layer is valid JSON with the required Navigator v4.x fields
# 2. every technique entry has a well-formed ID and a valid ATT&CK tactic
# 3. the scored techniques are EXACTLY the attack.* techniques tagged on the
# rules (bidirectional: no rule technique missing from the layer, no layer
# technique absent from the rules)
# 4. the scored tactics are exactly the attack.* tactics tagged on the rules
# 5. every scored technique's comment grounds it in a rule file that exists
# 6. any score-less entry is a display-only parent of a scored sub-technique
# 7. scores stay within the gradient bounds
#
# Pure standard library (the slim image already ships python3); nothing is
# installed and nothing is written to the repo. Exits non-zero on any failure.
import glob
import json
import os
import re
import sys
DET = "/detections"
LAYER = os.path.join(DET, "attack", "artex_navigator_layer.json")
SIGMA = os.path.join(DET, "sigma")
# ATT&CK Enterprise tactic shortnames (the Navigator "tactic" field uses these).
VALID_TACTICS = {
"reconnaissance", "resource-development", "initial-access", "execution",
"persistence", "privilege-escalation", "defense-evasion", "credential-access",
"discovery", "lateral-movement", "collection", "command-and-control",
"exfiltration", "impact",
}
TECHNIQUE_RE = re.compile(r"^T\d{4}(\.\d{3})?$")
TAG_RE = re.compile(r"attack\.(t\d{4}(?:\.\d{3})?)", re.IGNORECASE)
TACTIC_TAG_RE = re.compile(r"attack\.([a-z][a-z-]+)")
fail = 0
def note(s):
print(" " + s)
def ok(s):
note("PASS " + s)
def bad(s):
global fail
note("FAIL " + s)
fail = 1
def rule_tags():
"""Techniques and tactics tagged across every Sigma rule file."""
techs, tactics = set(), set()
files = sorted(glob.glob(os.path.join(SIGMA, "**", "*.yml"), recursive=True))
for path in files:
with open(path, encoding="utf-8") as fh:
for line in fh:
m = TAG_RE.search(line)
if m:
techs.add(m.group(1).upper())
continue
t = TACTIC_TAG_RE.search(line)
if t and t.group(1) in VALID_TACTICS:
tactics.add(t.group(1))
return techs, tactics, files
print("== 1/7 layer parses as JSON with the required Navigator fields ==")
try:
with open(LAYER, encoding="utf-8") as fh:
layer = json.load(fh)
ok("artex_navigator_layer.json is valid JSON")
except Exception as exc: # noqa: BLE001
print(" FAIL cannot parse layer: %s" % exc)
print("RESULT: FAIL")
sys.exit(1)
for key in ("name", "versions", "domain", "techniques", "gradient"):
if key in layer:
ok("top-level key present: %s" % key)
else:
bad("top-level key missing: %s" % key)
for vkey in ("attack", "navigator", "layer"):
if vkey in layer.get("versions", {}):
ok("versions.%s present (%s)" % (vkey, layer["versions"][vkey]))
else:
bad("versions.%s missing" % vkey)
if layer.get("domain") == "enterprise-attack":
ok("domain is enterprise-attack")
else:
bad("domain is not enterprise-attack: %r" % layer.get("domain"))
techniques = layer.get("techniques", [])
scored = [t for t in techniques if "score" in t]
helpers = [t for t in techniques if "score" not in t]
print("== 2/7 every technique entry has a valid ID and tactic ==")
for t in techniques:
tid = t.get("techniqueID", "")
if TECHNIQUE_RE.match(tid):
ok("well-formed techniqueID: %s" % tid)
else:
bad("malformed techniqueID: %r" % tid)
tac = t.get("tactic", "")
if tac in VALID_TACTICS:
ok("valid tactic for %s: %s" % (tid, tac))
else:
bad("invalid tactic for %s: %r" % (tid, tac))
rule_techs, rule_tactics, rule_files = rule_tags()
layer_scored_ids = {t["techniqueID"] for t in scored}
layer_scored_tactics = {t["tactic"] for t in scored}
print("== 3/7 scored techniques == techniques tagged on the rules (bidirectional) ==")
if not rule_techs:
bad("found no attack.* technique tags in %s" % SIGMA)
missing_in_layer = rule_techs - layer_scored_ids
extra_in_layer = layer_scored_ids - rule_techs
if not missing_in_layer and not extra_in_layer:
ok("scored techniques match the rule set exactly (%d: %s)"
% (len(rule_techs), ", ".join(sorted(rule_techs))))
else:
if missing_in_layer:
bad("rule techniques missing from the layer: %s"
% ", ".join(sorted(missing_in_layer)))
if extra_in_layer:
bad("layer techniques not tagged on any rule: %s"
% ", ".join(sorted(extra_in_layer)))
print("== 4/7 scored tactics == tactics tagged on the rules ==")
if layer_scored_tactics == rule_tactics:
ok("scored tactics match the rule set exactly (%s)"
% ", ".join(sorted(rule_tactics)))
else:
bad("tactic mismatch: layer=%s rules=%s"
% (sorted(layer_scored_tactics), sorted(rule_tactics)))
print("== 5/7 each scored technique is grounded in a rule file that exists ==")
for t in scored:
comment = t.get("comment", "")
refs = re.findall(r"sigma/[\w./-]+\.yml", comment)
grounded = False
for ref in refs:
if os.path.exists(os.path.join(DET, ref)):
grounded = True
else:
bad("%s comment cites a missing rule file: %s" % (t["techniqueID"], ref))
if "suricata" in comment.lower():
grounded = True
if grounded:
ok("%s grounded in an existing rule reference" % t["techniqueID"])
else:
bad("%s comment cites no existing rule file" % t["techniqueID"])
print("== 6/7 any score-less entry is a display parent of a scored sub-technique ==")
if not helpers:
ok("no display-only entries (nothing to check)")
for h in helpers:
hid = h.get("techniqueID", "")
children = [s for s in scored if s["techniqueID"].startswith(hid + ".")]
if children and h.get("showSubtechniques") is True:
ok("%s is a display parent of %s"
% (hid, ", ".join(c["techniqueID"] for c in children)))
else:
bad("score-less entry %s is not a valid display parent "
"(needs showSubtechniques:true and a scored child)" % hid)
print("== 7/7 scores stay within the gradient bounds ==")
grad = layer.get("gradient", {})
lo, hi = grad.get("minValue", 0), grad.get("maxValue", 100)
for t in scored:
s = t["score"]
if lo <= s <= hi:
ok("%s score %s within [%s, %s]" % (t["techniqueID"], s, lo, hi))
else:
bad("%s score %s outside gradient [%s, %s]" % (t["techniqueID"], s, lo, hi))
print()
print("reference: %d Sigma rule files scanned, %d scored techniques, %d display parents"
% (len(rule_files), len(scored), len(helpers)))
print("RESULT: %s" % ("PASS" if fail == 0 else "FAIL"))
sys.exit(fail)
+32
View File
@@ -0,0 +1,32 @@
#!/usr/bin/env bash
#
# Reproducible consistency test for the ARTEX ATT&CK coverage layer
# (../../attack/artex_navigator_layer.json). A coverage layer that drifts from the
# rules it claims to cover is worse than none, so this turns "these rules cover
# these ATT&CK techniques" from a claim into something a reviewer can re-run from
# source. It catches the realistic regression: a rule is added, removed, or
# retagged, but the Navigator layer is not updated to match.
#
# It proves (see check.py for the assertions) that the layer is a valid Navigator
# v4.x document and that its scored techniques and tactics are EXACTLY the attack.*
# tags on the Sigma rules — no rule technique missing from the layer, no layer
# technique absent from the rules — with every scored technique grounded in a rule
# file that exists.
#
# No host dependency beyond Docker: the check is pure Python standard library and
# runs in a container with the detections tree mounted read-only. Nothing is
# installed on the host and nothing is written to the repo.
#
# Usage: detections/tests/attack/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
DET_DIR="$REPO/detections"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
docker run --rm \
-v "$DET_DIR:/detections:ro" \
-v "$HERE:/src:ro" \
"$PYTHON_IMAGE" python3 /src/check.py
+177
View File
@@ -0,0 +1,177 @@
#!/usr/bin/env python3
#
# Harness self-consistency check for the detection test suites. check-harness-sync.sh
# launches this inside a Python container with the detection tree and the CI
# workflow mounted read-only under /repo. It proves the one thing the eight
# detection suites cannot: that run-all.sh, the CI workflow, and the suite
# directories on disk all name the same suites in the same order.
#
# run-all.sh, CONTRIBUTING, and this directory's README all promise that
# "run-all.sh runs the same suites as CI, in the same order." Nothing enforced
# that promise. A suite added to only one of the three places — a new CI step with
# no SUITES entry, or a new directory never wired into either — passes every
# per-suite test while quietly breaking the promise: run-all.sh and CI run
# different sets, so a green run-all.sh locally no longer implies a green CI. This
# check closes that gap the same way the indicator test closes the
# CI-paths / pre-commit-regex gap: it reads the three sources of truth and asserts
# they agree.
#
# A = the SUITES="..." list in run-all.sh (ordered)
# B = the per-suite `run: detections/tests/<x>/run.sh` steps (ordered)
# in .github/workflows/detections.yml
# C = the subdirectories of detections/tests/ that carry a (a set)
# run.sh
#
# It asserts A == B as ordered lists (so the documented "same order as CI" holds)
# and set(A) == C (so no directory is orphaned and no listed suite is missing on
# disk). It is deliberately not a detection suite: it is not in SUITES, not a
# `.../run.sh` CI step, and not a run.sh directory, so it never counts itself and
# the eight detection suites stay eight.
#
# It parses only the stable machine-readable lines (the SUITES assignment, the
# `run:` steps, the directory listing), never prose or the example-output blocks
# in the READMEs, so it cannot go brittle on documentation wording.
#
# Pure standard library (the slim image already ships python3); nothing is
# installed and nothing is written to the repo. Exits non-zero on any mismatch.
#
# Usage. check-harness-sync.sh runs this inside Docker with the repo mounted at
# /repo, which is why ROOT defaults to /repo below. To run it directly on the
# host instead, point ARTEX_REPO_ROOT at the repo root:
#
# ARTEX_REPO_ROOT="$(git rev-parse --show-toplevel)" \
# python3 detections/tests/check-harness-sync.py
import os
import re
import sys
# Default to /repo, the mount point check-harness-sync.sh uses inside Docker.
# Track whether the caller set the variable so a missing-path failure can tell a
# host-direct runner why ROOT is /repo (see main()).
ROOT = os.environ.get("ARTEX_REPO_ROOT", "/repo")
ROOT_FROM_ENV = "ARTEX_REPO_ROOT" in os.environ
TESTS_DIR = os.path.join(ROOT, "detections", "tests")
RUN_ALL = os.path.join(TESTS_DIR, "run-all.sh")
CI_WORKFLOW = os.path.join(ROOT, ".github", "workflows", "detections.yml")
def fail(msg):
print(f"FAIL: {msg}", file=sys.stderr)
sys.exit(1)
def suites_from_run_all(path):
"""Ordered suite names in the SUITES="..." assignment in run-all.sh."""
with open(path, encoding="utf-8") as f:
text = f.read()
m = re.search(r'^SUITES="([^"]*)"', text, re.MULTILINE)
if not m:
fail(f'could not find a SUITES="..." assignment in {path}')
names = m.group(1).split()
if not names:
fail(f"SUITES in {path} is empty")
return names
def suites_from_ci(path):
"""Ordered suite names in the per-suite `run:` steps of the CI workflow."""
names = []
with open(path, encoding="utf-8") as f:
for line in f:
m = re.search(r"run:\s*detections/tests/([^/]+)/run\.sh\s*$", line)
if m:
names.append(m.group(1))
if not names:
fail(f"found no `run: detections/tests/<suite>/run.sh` steps in {path}")
return names
def suites_from_dirs(path):
"""Suite directories under detections/tests/ that carry a run.sh."""
names = set()
for entry in sorted(os.listdir(path)):
d = os.path.join(path, entry)
if os.path.isdir(d) and os.path.isfile(os.path.join(d, "run.sh")):
names.add(entry)
if not names:
fail(f"found no suite directories with a run.sh under {path}")
return names
def main():
for p in (RUN_ALL, CI_WORKFLOW, TESTS_DIR):
if not os.path.exists(p):
hint = ""
if not ROOT_FROM_ENV:
hint = (
"\n ARTEX_REPO_ROOT is unset, so ROOT defaulted to /repo "
"(the path check-harness-sync.sh mounts the repo at inside Docker).\n"
" To run this script directly on the host, point it at the "
"repo root:\n"
' ARTEX_REPO_ROOT="$(git rev-parse --show-toplevel)" '
"python3 detections/tests/check-harness-sync.py\n"
" or use the Docker wrapper: "
"detections/tests/check-harness-sync.sh"
)
fail(f"missing expected path: {p}{hint}")
a = suites_from_run_all(RUN_ALL)
b = suites_from_ci(CI_WORKFLOW)
c = suites_from_dirs(TESTS_DIR)
errors = []
# A name repeated in an ordered source would make the comparisons below read
# misleadingly, so surface it on its own first.
for label, seq in (("run-all.sh SUITES", a), ("CI steps", b)):
if len(seq) != len(set(seq)):
dupes = sorted({x for x in seq if seq.count(x) > 1})
errors.append(f"{label} names a suite more than once: {dupes}")
if a != b:
errors.append(
"run-all.sh SUITES and the CI steps disagree (order matters: "
"run-all.sh promises the same order as CI):\n"
f" run-all.sh: {a}\n"
f" CI steps : {b}"
)
if set(a) != c:
detail = []
only_listed = sorted(set(a) - c)
only_on_disk = sorted(c - set(a))
if only_listed:
detail.append(
f" listed in run-all.sh but no run.sh directory: {only_listed}"
)
if only_on_disk:
detail.append(
f" run.sh directory present but not in run-all.sh: {only_on_disk}"
)
errors.append(
"run-all.sh SUITES and the suite directories disagree:\n"
+ "\n".join(detail)
)
if errors:
print("detection test harness is OUT OF SYNC:\n", file=sys.stderr)
for e in errors:
print(e + "\n", file=sys.stderr)
print(
"Wire the new suite into all three (SUITES in run-all.sh, a step in "
".github/workflows/detections.yml, and a run.sh directory) so a local "
"run-all.sh runs exactly what CI runs.",
file=sys.stderr,
)
sys.exit(1)
print(
f"harness sync OK: run-all.sh, CI, and {len(c)} suite directories "
"name the same suites in the same order:"
)
print(" " + " ".join(a))
if __name__ == "__main__":
main()
+33
View File
@@ -0,0 +1,33 @@
#!/usr/bin/env bash
#
# Harness self-consistency check for the detection test suites (see
# check-harness-sync.py for the assertions). It proves the one thing the eight
# detection suites cannot: that run-all.sh, the CI workflow, and the suite
# directories on disk all name the same suites in the same order, so a suite
# wired into only one of the three cannot silently break the "run-all.sh runs the
# same suites as CI" promise while every per-suite test stays green.
#
# It is a gate, not a suite: run-all.sh runs it before the suite loop and it does
# not appear in the per-suite summary, and it is not itself a SUITES entry, a
# `.../run.sh` CI step, or a run.sh directory — so the eight detection suites stay
# eight and this check never counts itself.
#
# No host dependency beyond Docker: the check is pure Python standard library and
# runs in a container with only the detection tree and the CI workflow mounted
# read-only (never work/ or anything else). Nothing is installed on the host and
# nothing is written to the repo.
#
# Usage: detections/tests/check-harness-sync.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../.." && pwd)"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
docker run --rm \
-e ARTEX_REPO_ROOT=/repo \
-v "$REPO/detections:/repo/detections:ro" \
-v "$REPO/.github:/repo/.github:ro" \
-v "$HERE:/src:ro" \
"$PYTHON_IMAGE" python3 /src/check-harness-sync.py
+378
View File
@@ -0,0 +1,378 @@
#!/usr/bin/env python3
#
# Source-of-truth consistency test for the ARTEX detection indicators. run.sh
# launches this inside a Python container with the detection rules and the
# upstream source packages they pin mounted read-only under /repo. It proves one
# property the other three detection tests do not: that each rule's pinned
# indicator is still the string ARTEX's own source actually emits.
#
# The Sigma test proves an indicator survives rule->query *compilation*; the
# ATT&CK test proves the layer matches the rules' tags; the Suricata test proves
# the network rule *fires*. None of them look back at the source the indicator
# claims to come from. So the realistic rot they miss is an upstream re-sync that
# bumps the prober User-Agent to "artex-enrich/2.0" or rewrites the guard marker:
# every rule still compiles, the layer still matches, the pcap test still fires on
# the synthesized capture — and the deployed rule silently stops matching real
# ARTEX traffic. This test turns detections/README's claim ("every indicator is
# grounded in a string verified in this repository's source, not inferred") and
# CONTRIBUTING's first contribution contract into a guard a reviewer can re-run.
#
# For every indicator it asserts, bidirectionally:
# - source drift: the value is still present in the upstream source file(s)
# that emit it (fails if an upstream re-sync changed the source but not the
# rule -> the rule is now stale);
# - rule drift: the value is still pinned in the rule(s) built on it (fails if a
# rule edit moved the indicator away from the source).
# The enrichment UA also carries a Suricata prefix check, because that rule
# matches the User-Agent by `startswith` and so pins a prefix of the full value.
#
# The destructive-command tokens are handled separately and honestly: they are
# generic hunting leads, not unique ARTEX fingerprints, so the test only asserts
# the correspondence the rule actually claims — each token appears both in ARTEX's
# guard deny-list (db/db.go) and in the hunting rule that mirrors it.
#
# It then validates the published, machine-readable indicator list
# (detections/indicators/artex_indicators.csv): every row's value must still be
# present in the source file(s) it cites and pinned in the rule(s) it cites, and
# every fingerprint this test grounds must appear in the list — so the artifact a
# defender imports cannot silently drift from the source it claims to come from.
#
# Finally it closes the loop on the two gates that fire this test: the CI workflow
# (.github/workflows/detections.yml push/pull_request paths) and the local
# pre-commit hook (.pre-commit-config.yaml files regex). Every upstream source file
# this test reads must be covered by both, or a change touching only a newly pinned
# source (as cmd/artex/main.go once was) would skip the test on one of them: on CI
# the drift sails through the merge gate green, on the hook it is never caught
# locally even though the hook's comment promises "the same source scope as CI".
# The check derives the required set from the indicators it already asserts, so
# pinning a new source without wiring it into *both* gates fails here until they
# stay in sync.
#
# Pure standard library (the slim image already ships python3); nothing is
# installed and nothing is written to the repo. Exits non-zero on any failure.
import csv
import io
import os
import re
import sys
ROOT = os.environ.get("ARTEX_REPO_ROOT", "/repo")
# --- exact ARTEX fingerprints ------------------------------------------------
# Each value is an operational string ARTEX emits; a rule is built on it. If an
# upstream re-sync changes the source string, the rule must change with it.
INDICATORS = [
{
"label": "enrichment prober User-Agent",
"value": "artex-enrich/1.0",
"sources": ["enrich/enrich.go"],
"rules": ["detections/sigma/artex_enrich_user_agent.yml"],
# Suricata matches the UA by `startswith`, so it pins a prefix of the
# full value rather than the whole string. (file, prefix)
"prefix_rules": [("detections/suricata/artex.rules", "artex-enrich/")],
},
{
"label": "self-update egress User-Agent",
"value": "artex-selfupdate",
"sources": ["selfupdate/github.go", "selfupdate/stage.go"],
"rules": ["detections/sigma/artex_selfupdate_egress.yml"],
},
{
"label": "platform-guard audit framing marker",
"value": "【ARTEX 平台管控·非目标防御】",
"sources": ["guard/guard.go"],
"rules": ["detections/sigma/artex_guard_audit_framing.yml"],
},
]
# --- generic destructive-command hunting leads -------------------------------
# NOT unique ARTEX fingerprints. These tokens are shared with ARTEX's own guard
# deny-list (db/db.go); the hunting rule mirrors that list. The test asserts only
# the correspondence the rule claims, so it catches an upstream re-sync that drops
# or renames a deny-list entry the rule says it mirrors.
DENYLIST = {
"source": "db/db.go",
"rule": "detections/sigma/destructive_command_hunting.yml",
"tokens": ["rm -rf", "mkfs", "DROP DATABASE", "FLUSHALL"],
}
fail = 0
def note(s):
print(" " + s)
def ok(s):
note("PASS " + s)
def bad(s):
global fail
note("FAIL " + s)
fail = 1
def read(rel):
"""Return the text of a repo-relative file, or None if it is missing."""
try:
with open(os.path.join(ROOT, rel), encoding="utf-8") as fh:
return fh.read()
except OSError:
return None
def contains(rel, needle):
text = read(rel)
if text is None:
return None # file missing -> distinct from "present but absent"
return needle in text
print("== 1/5 exact fingerprints are still emitted by the upstream source ==")
for ind in INDICATORS:
value, label = ind["value"], ind["label"]
present = [s for s in ind["sources"] if contains(s, value) is True]
missing_files = [s for s in ind["sources"] if contains(s, value) is None]
if present:
ok("%s: %r emitted by %s" % (label, value, ", ".join(present)))
elif missing_files:
bad("%s: source file(s) missing: %s (upstream moved the emitter?)"
% (label, ", ".join(missing_files)))
else:
bad("%s: %r NOT found in any source file %s "
"(upstream drift — update the rule to match)"
% (label, value, ind["sources"]))
print("== 2/5 each rule still pins the indicator it is built on ==")
for ind in INDICATORS:
value, label = ind["value"], ind["label"]
for rule in ind["rules"]:
hit = contains(rule, value)
if hit is True:
ok("%s pins %r" % (rule, value))
elif hit is None:
bad("rule file missing: %s" % rule)
else:
bad("%s no longer pins %r (rule drift from source)" % (rule, value))
for rfile, prefix in ind.get("prefix_rules", []):
if not value.startswith(prefix):
bad("%s: prefix %r is not a prefix of %r (internal inconsistency)"
% (rfile, prefix, value))
continue
hit = contains(rfile, prefix)
if hit is True:
ok("%s pins prefix %r of %r" % (rfile, prefix, value))
elif hit is None:
bad("rule file missing: %s" % rfile)
else:
bad("%s no longer pins prefix %r" % (rfile, prefix))
print("== 3/5 destructive hunting tokens match ARTEX's guard deny-list ==")
src, rule, tokens = DENYLIST["source"], DENYLIST["rule"], DENYLIST["tokens"]
for tok in tokens:
in_src = contains(src, tok)
in_rule = contains(rule, tok)
if in_src is None:
bad("deny-list source missing: %s" % src)
elif in_rule is None:
bad("hunting rule missing: %s" % rule)
elif in_src and in_rule:
ok("%r present in both %s and %s" % (tok, src, rule))
elif not in_src:
bad("%r pinned by the rule but absent from %s "
"(upstream dropped/renamed the deny-list entry)" % (tok, src))
else:
bad("%r in the deny-list but not pinned by %s" % (tok, rule))
print("== 4/5 the published indicator list matches source and rules ==")
CSV_REL = "detections/indicators/artex_indicators.csv"
EXPECTED_HEADER = ["id", "type", "value", "perspective", "source", "rule", "description"]
VALID_PERSPECTIVES = {"target", "forensic"}
csv_rows = []
csv_text = read(CSV_REL)
if csv_text is None:
bad("published indicator list missing: %s" % CSV_REL)
else:
rows = list(csv.reader(io.StringIO(csv_text)))
if not rows:
bad("%s is empty" % CSV_REL)
elif rows[0] != EXPECTED_HEADER:
bad("%s header is %r, expected %r" % (CSV_REL, rows[0], EXPECTED_HEADER))
else:
seen_ids = set()
for lineno, row in enumerate(rows[1:], start=2):
if len(row) != len(EXPECTED_HEADER):
bad("%s line %d: %d fields, expected %d"
% (CSV_REL, lineno, len(row), len(EXPECTED_HEADER)))
continue
rec = dict(zip(EXPECTED_HEADER, row))
csv_rows.append(rec)
rid, value = rec["id"], rec["value"]
if rid in seen_ids:
bad("%s: duplicate id %r" % (CSV_REL, rid))
seen_ids.add(rid)
if not value:
bad("%s: row %r has an empty value" % (CSV_REL, rid))
continue
if rec["perspective"] not in VALID_PERSPECTIVES:
bad("%s: row %r perspective %r not in %s"
% (CSV_REL, rid, rec["perspective"], sorted(VALID_PERSPECTIVES)))
src_files = [s for s in rec["source"].split(";") if s]
if not src_files:
bad("%s: row %r cites no source file" % (CSV_REL, rid))
for s in src_files:
hit = contains(s, value)
if hit is True:
ok("%s: %r grounded in %s" % (rid, value, s))
elif hit is None:
bad("%s: row %r source file missing: %s" % (CSV_REL, rid, s))
else:
bad("%s: row %r value %r not found in source %s (drift)"
% (CSV_REL, rid, value, s))
for r in [r for r in rec["rule"].split(";") if r]:
hit = contains(r, value)
if hit is True:
ok("%s: %r pinned in %s" % (rid, value, r))
elif hit is None:
bad("%s: row %r rule file missing: %s" % (CSV_REL, rid, r))
else:
bad("%s: row %r value %r not pinned in rule %s"
% (CSV_REL, rid, value, r))
published = {rec["value"] for rec in csv_rows}
for ind in INDICATORS:
if ind["value"] in published:
ok("tested fingerprint %r is published in the list" % ind["value"])
else:
bad("tested fingerprint %r is missing from %s" % (ind["value"], CSV_REL))
def paths_for_trigger(text, trigger):
"""Collect the quoted entries of `<trigger>: ... paths: [...]` in the detection
workflow. Returns the set of listed paths, or None if the trigger is absent.
A deliberately small parser for a known-shape file: it locates the trigger key
under `on:`, then the `paths:` list nested in it, and reads the `- "..."` items
until the indentation returns to the list's level."""
lines = text.splitlines()
t_indent = None
start = None
for idx, line in enumerate(lines):
if re.match(r"^\s{2,}%s:\s*$" % re.escape(trigger), line):
t_indent = len(line) - len(line.lstrip())
start = idx + 1
break
if start is None:
return None
items = set()
i = start
while i < len(lines):
line = lines[i]
if line.strip():
indent = len(line) - len(line.lstrip())
if indent <= t_indent:
break # left this trigger block
if re.match(r"^\s*paths:\s*$", line):
p_indent = indent
j = i + 1
while j < len(lines):
pl = lines[j]
if pl.strip():
pind = len(pl) - len(pl.lstrip())
if pind <= p_indent:
break
m = re.match(r"""^\s*-\s*['"]?([^'"\s]+)['"]?\s*$""", pl)
if m:
items.add(m.group(1))
j += 1
return items
i += 1
return items
print("== 5/5 CI and the pre-commit hook both fire this test on any pinned source ==")
# Two gates run this test only when a file they filter on changes: the CI workflow's
# paths filter and the pre-commit hook's files regex. Every upstream source this test
# reads must be covered by both, or a change touching only that source skips the test
# on the gate that misses it — on CI the drift above passes the merge gate green, on
# the hook it is never caught locally. The required set is derived from the indicators
# themselves, so pinning a new source without wiring it into both gates fails here.
# detections/** covers the rules, the CSV, and the tests, so only non-detections
# sources are required explicitly (plus a spot check that each gate still covers the
# detections/ tree at all).
WORKFLOW_REL = ".github/workflows/detections.yml"
PRECOMMIT_REL = ".pre-commit-config.yaml"
DETECTIONS_SAMPLE = "detections/sigma/artex_enrich_user_agent.yml"
needed_sources = set()
for ind in INDICATORS:
needed_sources.update(ind["sources"])
needed_sources.add(DENYLIST["source"])
for rec in csv_rows:
for s in rec["source"].split(";"):
if s:
needed_sources.add(s)
needed_sources = {s for s in needed_sources if not s.startswith("detections/")}
wf_text = read(WORKFLOW_REL)
if wf_text is None:
bad("CI workflow missing: %s" % WORKFLOW_REL)
else:
for trigger in ("push", "pull_request"):
listed = paths_for_trigger(wf_text, trigger)
if listed is None:
bad("%s has no %s: trigger" % (WORKFLOW_REL, trigger))
continue
if "detections/**" not in listed:
bad("%s %s paths is missing 'detections/**' "
"(rule/CSV/test changes would not trigger the detection tests)"
% (WORKFLOW_REL, trigger))
for s in sorted(needed_sources):
if s in listed:
ok("%s %s paths covers %s" % (WORKFLOW_REL, trigger, s))
else:
bad("%s %s paths is missing %s — a PR touching only that source "
"would skip this test and let source drift pass the merge gate"
% (WORKFLOW_REL, trigger, s))
# The local hook gates on a files regex, not a paths list. Its comment promises the
# "same source scope as CI", so the same required set must match that regex. This is
# the sibling drift the CI check above does not see: CI paths can carry a source the
# hook's regex omits (as cmd/artex/main.go once did), leaving the local gate a false
# promise even while the merge gate is sound.
pc_text = read(PRECOMMIT_REL)
if pc_text is None:
bad("pre-commit config missing: %s" % PRECOMMIT_REL)
else:
m = re.search(r"^\s*files:\s*(.+?)\s*$", pc_text, re.M)
if not m:
bad("%s has no files: pattern on the detections hook" % PRECOMMIT_REL)
else:
pattern_src = m.group(1).strip().strip("'\"")
try:
pat = re.compile(pattern_src)
except re.error as exc:
bad("%s files pattern does not compile: %s" % (PRECOMMIT_REL, exc))
pat = None
if pat is not None:
if pat.search(DETECTIONS_SAMPLE):
ok("%s files covers the detections/ tree" % PRECOMMIT_REL)
else:
bad("%s files does not cover detections/ "
"(rule/CSV/test changes would not fire the local hook)"
% PRECOMMIT_REL)
for s in sorted(needed_sources):
if pat.search(s):
ok("%s files covers %s" % (PRECOMMIT_REL, s))
else:
bad("%s files is missing %s — a commit touching only that source "
"would skip the local hook while CI still runs it (the hook's "
"'same source scope as CI' promise is false for this file)"
% (PRECOMMIT_REL, s))
print()
print("reference: %d exact fingerprints, %d deny-list tokens, %d published rows, "
"%d pinned sources checked against CI paths and the pre-commit files regex"
% (len(INDICATORS), len(tokens), len(csv_rows), len(needed_sources)))
print("RESULT: %s" % ("PASS" if fail == 0 else "FAIL"))
sys.exit(fail)
+36
View File
@@ -0,0 +1,36 @@
#!/usr/bin/env bash
#
# Source-of-truth consistency test for the ARTEX detection indicators
# (see check.py for the assertions). It proves the one thing the Sigma, Suricata,
# and ATT&CK tests do not: that each rule's pinned indicator is still the string
# ARTEX's own source actually emits. The realistic rot it catches is an upstream
# re-sync that bumps the prober User-Agent or rewrites the guard marker — every
# other test stays green while the deployed rule silently stops matching.
#
# No host dependency beyond Docker: the check is pure Python standard library and
# runs in a container with only the rule tree, the published indicator list, the
# source packages it pins, and the two configs that gate on them — the CI workflow
# and the pre-commit hook — mounted read-only (never work/ or anything else).
# Nothing is installed on the host and nothing is written to the repo.
#
# Usage: detections/tests/indicators/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
docker run --rm \
-e ARTEX_REPO_ROOT=/repo \
-v "$REPO/detections:/repo/detections:ro" \
-v "$REPO/enrich:/repo/enrich:ro" \
-v "$REPO/selfupdate:/repo/selfupdate:ro" \
-v "$REPO/guard:/repo/guard:ro" \
-v "$REPO/db:/repo/db:ro" \
-v "$REPO/cmd:/repo/cmd:ro" \
-v "$REPO/traffic:/repo/traffic:ro" \
-v "$REPO/.github:/repo/.github:ro" \
-v "$REPO/.pre-commit-config.yaml:/repo/.pre-commit-config.yaml:ro" \
-v "$HERE:/src:ro" \
"$PYTHON_IMAGE" python3 /src/check.py
+288
View File
@@ -0,0 +1,288 @@
#!/usr/bin/env python3
#
# Consistency test for the MISP-format export of the ARTEX detection indicators.
# run.sh launches this inside a Python container with pymisp installed and the
# detection tree plus the CI workflow mounted read-only under /repo. It proves
# two properties the other detection tests do not touch:
#
# 1. the published MISP event
# (detections/indicators/artex_indicators.misp.json) is a *valid MISP
# document* — pymisp parses it and accepts every attribute type/category,
# so a defender can import it into MISP (or export it on to STIX from
# there) without hand-fixing the format; and
# 2. that MISP event stays in sync with the source-of-truth CSV
# (detections/indicators/artex_indicators.csv) row for row — same values,
# the intended MISP type/category for each CSV indicator type, and a
# to_ids / disable_correlation flag that faithfully encodes the CSV's own
# honesty (a row with a detection rule is an actionable indicator; a
# host-forensic row without one is a triage hint, not a blocking IoC).
#
# The indicators source-of-truth test (../indicators/) already proves every CSV
# row is grounded in the upstream source and pinned in its rule; this test does
# not repeat that. It proves only that the MISP serialization a defender
# actually imports cannot silently drift away from that CSV — if a row is added,
# removed, retyped, or has its rule column changed, the MISP event must change
# with it or this test fails.
#
# Exits non-zero on any failed assertion.
import csv
import io
import json
import os
import re
import sys
ROOT = os.environ.get("ARTEX_REPO_ROOT", "/repo")
CSV_REL = "detections/indicators/artex_indicators.csv"
MISP_REL = "detections/indicators/artex_indicators.misp.json"
WORKFLOW_REL = ".github/workflows/detections.yml"
# The indicators README states the CSV `type` values "map onto the equivalent
# MISP/STIX attribute types". This is that mapping, made explicit and enforced:
# CSV indicator type -> (MISP attribute type, MISP attribute category).
TYPE_MAP = {
"http.user-agent": ("user-agent", "Network activity"),
"string": ("pattern-in-file", "Artifacts dropped"),
"port": ("port", "Network activity"),
"ip-dst|port": ("ip-dst|port", "Network activity"),
# a host artifact that fits no network/file slot (e.g. a DB schema object
# name); MISP's generic "other"/"Other" carries it as a triage lead.
"other": ("other", "Other"),
}
fail = 0
def note(s):
print(" " + s)
def ok(s):
note("PASS " + s)
def bad(s):
global fail
note("FAIL " + s)
fail = 1
def read(rel):
try:
with open(os.path.join(ROOT, rel), encoding="utf-8") as fh:
return fh.read()
except OSError:
return None
def misp_value_for(csv_type, csv_value):
"""The MISP value for a CSV row. MISP composite types join their parts with
`|`, so the CSV's `ip:port` becomes `ip|port`; every other type is verbatim."""
if csv_type == "ip-dst|port":
return csv_value.replace(":", "|", 1)
return csv_value
# --- load the source-of-truth CSV -------------------------------------------
EXPECTED_HEADER = ["id", "type", "value", "perspective", "source", "rule", "description"]
csv_rows = []
csv_text = read(CSV_REL)
if csv_text is None:
bad("source CSV missing: %s" % CSV_REL)
else:
rows = list(csv.reader(io.StringIO(csv_text)))
if not rows or rows[0] != EXPECTED_HEADER:
bad("%s header is %r, expected %r"
% (CSV_REL, rows[0] if rows else None, EXPECTED_HEADER))
else:
for row in rows[1:]:
if len(row) == len(EXPECTED_HEADER):
csv_rows.append(dict(zip(EXPECTED_HEADER, row)))
# --- load the MISP event (raw JSON) -----------------------------------------
misp_text = read(MISP_REL)
event = None
attrs = []
if misp_text is None:
bad("MISP event missing: %s" % MISP_REL)
else:
try:
doc = json.loads(misp_text)
except ValueError as exc:
bad("%s is not valid JSON: %s" % (MISP_REL, exc))
doc = None
if isinstance(doc, dict):
event = doc.get("Event")
if not isinstance(event, dict):
bad("%s has no top-level Event object" % MISP_REL)
else:
if not event.get("info"):
bad("%s Event has no info string" % MISP_REL)
if not event.get("uuid"):
bad("%s Event has no uuid" % MISP_REL)
attrs = event.get("Attribute") or []
if not isinstance(attrs, list) or not attrs:
bad("%s Event has no Attribute list" % MISP_REL)
attrs = []
print("== 1/5 the MISP event is a valid MISP document (pymisp parses it) ==")
# pymisp's object model rejects an unknown attribute type on load, so a parse
# here is a real check that every type we use is a genuine MISP type a MISP
# server would accept — not just a plausible-looking string.
if misp_text is None:
bad("cannot validate: MISP event missing")
else:
try:
from pymisp import MISPEvent
me = MISPEvent()
me.load_file(os.path.join(ROOT, MISP_REL))
ok("pymisp %s parsed the event (%d attributes, info=%r)"
% (__import__("pymisp").__version__, len(me.attributes), me.info))
if len(me.attributes) != len(attrs):
bad("pymisp parsed %d attributes but the JSON has %d"
% (len(me.attributes), len(attrs)))
except Exception as exc: # NewAttributeError, validation, import, ...
bad("pymisp rejected the MISP event: %s: %s"
% (type(exc).__name__, exc))
print("== 2/5 every published CSV row maps to one MISP attribute ==")
# value (transformed for composite types) -> list of matching MISP attributes
by_value = {}
for a in attrs:
by_value.setdefault(a.get("value"), []).append(a)
expected_misp_values = set()
for rec in csv_rows:
rid, ctype, cval = rec["id"], rec["type"], rec["value"]
if ctype not in TYPE_MAP:
bad("%s: CSV type %r has no MISP mapping (extend TYPE_MAP)" % (rid, ctype))
continue
want_type, want_cat = TYPE_MAP[ctype]
want_val = misp_value_for(ctype, cval)
expected_misp_values.add(want_val)
matches = by_value.get(want_val, [])
if not matches:
bad("%s: no MISP attribute with value %r (CSV row not exported)"
% (rid, want_val))
continue
if len(matches) > 1:
bad("%s: %d MISP attributes share value %r" % (rid, len(matches), want_val))
a = matches[0]
if a.get("type") == want_type:
ok("%s: %r is a %s" % (rid, want_val, want_type))
else:
bad("%s: value %r is type %r, expected %r"
% (rid, want_val, a.get("type"), want_type))
if a.get("category") != want_cat:
bad("%s: value %r category %r, expected %r"
% (rid, want_val, a.get("category"), want_cat))
# A row with a detection rule is an actionable indicator (to_ids on); a
# host-forensic row without one is a triage hint, not a blocking IoC
# (to_ids off, and correlation disabled so a common port / loopback does
# not pollute MISP correlations). This mirrors the CSV `rule` column.
want_ids = bool(rec["rule"].strip())
if bool(a.get("to_ids")) != want_ids:
bad("%s: to_ids=%r, expected %r (rule column=%r)"
% (rid, a.get("to_ids"), want_ids, rec["rule"]))
if bool(a.get("disable_correlation")) != (not want_ids):
bad("%s: disable_correlation=%r, expected %r"
% (rid, a.get("disable_correlation"), not want_ids))
if not (a.get("comment") or "").strip():
bad("%s: MISP attribute has an empty comment (grounding/caveat lost)" % rid)
print("== 3/5 no MISP attribute is unaccounted for (bijection) ==")
actual_values = [a.get("value") for a in attrs]
if len(actual_values) != len(set(actual_values)):
bad("the MISP event has duplicate attribute values")
extra = set(actual_values) - expected_misp_values
if extra:
bad("MISP attribute(s) with no CSV row: %s" % ", ".join(sorted(map(repr, extra))))
elif csv_rows and not fail:
ok("the %d MISP attributes are exactly the %d published CSV rows"
% (len(attrs), len(csv_rows)))
elif not extra:
ok("every MISP attribute corresponds to a CSV row")
print("== 4/5 the non-ASCII guard marker is preserved verbatim ==")
MARKER = "【ARTEX 平台管控·非目标防御】"
csv_has = any(r["value"] == MARKER for r in csv_rows)
misp_has = MARKER in actual_values
if csv_has and misp_has:
ok("guard audit marker exported byte-for-byte")
elif not csv_has:
bad("guard marker not found in the CSV (test assumption broke)")
else:
bad("guard marker in the CSV but not exported to the MISP event")
def paths_for_trigger(text, trigger):
"""The quoted entries of `<trigger>: ... paths: [...]` in the workflow, or
None if the trigger is absent. Small parser for a known-shape file."""
lines = text.splitlines()
t_indent = None
start = None
for idx, line in enumerate(lines):
if re.match(r"^\s{2,}%s:\s*$" % re.escape(trigger), line):
t_indent = len(line) - len(line.lstrip())
start = idx + 1
break
if start is None:
return None
items = set()
i = start
while i < len(lines):
line = lines[i]
if line.strip():
indent = len(line) - len(line.lstrip())
if indent <= t_indent:
break
if re.match(r"^\s*paths:\s*$", line):
p_indent = indent
j = i + 1
while j < len(lines):
pl = lines[j]
if pl.strip():
pind = len(pl) - len(pl.lstrip())
if pind <= p_indent:
break
m = re.match(r"""^\s*-\s*['"]?([^'"\s]+)['"]?\s*$""", pl)
if m:
items.add(m.group(1))
j += 1
return items
i += 1
return items
print("== 5/5 CI triggers this test when the published indicators change ==")
# The MISP event derives only from the CSV, and both live under detections/**,
# so detections/** in the paths filter is the required and sufficient wiring:
# a change to the CSV or the MISP event triggers the detection workflow, which
# runs this suite and re-checks the two stay in sync. (The upstream Go sources
# the indicators are grounded in are enforced by the indicators suite's own
# CI-paths check, not here.)
wf_text = read(WORKFLOW_REL)
if wf_text is None:
bad("CI workflow missing: %s" % WORKFLOW_REL)
else:
for trigger in ("push", "pull_request"):
listed = paths_for_trigger(wf_text, trigger)
if listed is None:
bad("%s has no %s: trigger" % (WORKFLOW_REL, trigger))
elif "detections/**" in listed:
ok("%s %s paths covers detections/** (CSV + MISP event)" % (WORKFLOW_REL, trigger))
else:
bad("%s %s paths is missing 'detections/**' — a change to the CSV or "
"the MISP event would skip this test" % (WORKFLOW_REL, trigger))
print()
print("reference: %d CSV rows, %d MISP attributes, %d type mappings"
% (len(csv_rows), len(attrs), len(TYPE_MAP)))
print("RESULT: %s" % ("PASS" if fail == 0 else "FAIL"))
sys.exit(fail)
+32
View File
@@ -0,0 +1,32 @@
#!/usr/bin/env bash
#
# Reproducible consistency test for the MISP-format export of the ARTEX
# indicators (see check.py for the assertions). It proves two things no other
# detection test does: that detections/indicators/artex_indicators.misp.json is
# a MISP document pymisp actually parses (every attribute type/category is a real
# MISP type a server would accept), and that it stays row-for-row in sync with
# the source-of-truth CSV it is generated from — same values, the intended MISP
# type/category per indicator, and a to_ids/disable_correlation flag that mirrors
# the CSV's own honesty (rule-backed = actionable; host-forensic = triage hint).
#
# No host dependency beyond Docker: pymisp is pinned and installed inside the
# container, and the detection tree and CI workflow are mounted read-only.
# Nothing is installed on the host and nothing is written to the repo tree.
#
# Usage: detections/tests/misp/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
# PYMISP_VERSION (default 2.5.34.4 — the pinned reference version)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
PYMISP_VERSION="${PYMISP_VERSION:-2.5.34.4}"
docker run --rm \
-e ARTEX_REPO_ROOT=/repo \
-e PYMISP_VERSION="$PYMISP_VERSION" \
-v "$REPO/detections:/repo/detections:ro" \
-v "$REPO/.github:/repo/.github:ro" \
-v "$HERE:/src:ro" \
"$PYTHON_IMAGE" sh -c 'pip install --quiet "pymisp==${PYMISP_VERSION}" && python3 /src/check.py'
+82
View File
@@ -0,0 +1,82 @@
#!/usr/bin/env bash
#
# Runs every detection test suite under this directory in one command — the
# local-developer and pre-commit counterpart to the per-suite CI steps in
# ../../.github/workflows/detections.yml. CONTRIBUTING.md and this directory's
# README.md promise that each suite "drops straight into CI or a pre-commit
# hook"; this is the single entry point that honours that promise for all of
# them at once, so a contributor does not have to invoke the eight run.sh scripts
# by hand (and reviewers do not have to improvise a loop).
#
# It runs the suites in the same order as CI, lets each suite's own output flow
# through, prints a one-line PASS/FAIL summary per suite at the end, and exits
# non-zero if any suite failed — so it is safe to drop into a CI step or a
# pre-commit hook. Every suite runs to completion even if an earlier one fails,
# so one invocation surfaces every regression rather than only the first.
#
# No host dependency beyond Docker: each suite runs its checks in a container and
# writes nothing to the repo tree (see the per-suite run.sh headers). The image
# and version overrides the child scripts honour (PYTHON_IMAGE, SIGMA_CLI_VERSION,
# SIGMAHQ_VALIDATORS_VERSION, SURICATA_IMAGE) are inherited from this process's
# environment, so exporting any of them here applies to every suite at once.
#
# Usage: detections/tests/run-all.sh
set -uo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
# Same order as the steps in .github/workflows/detections.yml.
SUITES="sigma sigma_match sigma_lint sigma_backends suricata attack indicators misp"
fail=0
harness_fail=0
triage_fail=0
results=""
# Before the suites, verify the harness itself is consistent: the SUITES list
# above, the per-suite steps in CI, and the suite directories on disk must all
# name the same suites in the same order. A suite wired into only one of the
# three (say a new CI step with no SUITES entry) passes every per-suite test yet
# silently breaks the "run-all.sh runs the same suites as CI" promise, which no
# other suite can see. This is a gate, not a suite: it runs first and stays out
# of the per-suite summary below, so that summary remains the detection suites.
printf '\n===== harness sync =====\n'
if ! "$HERE/check-harness-sync.sh"; then
harness_fail=1
fail=1
fi
# A second gate, not a suite: the host-triage tool's --self-test. The triage tool
# is a responder helper, not a detection rule, so it stays out of SUITES and the
# harness-sync registry (see triage-selftest.sh). It runs here and in the CI
# workflow so a local run-all.sh covers it too.
printf '\n===== triage self-test =====\n'
if ! "$HERE/triage-selftest.sh"; then
triage_fail=1
fail=1
fi
for suite in $SUITES; do
printf '\n===== %s =====\n' "$suite"
if "$HERE/$suite/run.sh"; then
results="${results} PASS ${suite}"$'\n'
else
rc=$?
results="${results} FAIL ${suite} (exit ${rc})"$'\n'
fail=1
fi
done
printf '\n===== detection suites summary =====\n'
printf '%s' "$results"
if [ "$harness_fail" -ne 0 ]; then
printf ' FAIL harness sync (run-all.sh / CI / directories out of sync: see above)\n'
fi
if [ "$triage_fail" -ne 0 ]; then
printf ' FAIL triage self-test (detections/triage/artex_host_triage.py --self-test: see above)\n'
fi
if [ "$fail" -ne 0 ]; then
printf 'RESULT: FAIL\n'
exit 1
fi
printf 'RESULT: PASS\n'
+87
View File
@@ -0,0 +1,87 @@
#!/bin/sh
#
# In-container half of the ARTEX Sigma rule test. run.sh launches this inside a
# Python container with the Sigma rule tree mounted read-only at /sigma. It
# installs a pinned sigma-cli (pySigma) plus the splunk backend, then asserts
# the properties the rule files and the defense guide claim:
#
# 1. structural + best-practice validation passes (sigma check == 0 errors)
# 2. the whole tree compiles to a backend query language (sigma convert -> splunk)
# 3. each atomic indicator string survives into the query (enrich UA, self-update UA, guard marker, CA file)
# 4. the correlation rules compile as correlations (event_count / value_count aggregations)
# 5. a correlation rule converted ALONE fails (it genuinely depends on its atomic base rule)
#
# POSIX sh (the slim image ships dash). Exits non-zero if any assertion fails.
set -eu
VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
pip install --quiet --disable-pip-version-check "sigma-cli==${VERSION}" >/dev/null 2>&1
sigma plugin install splunk >/dev/null 2>&1
fail=0
note() { printf ' %s\n' "$1"; }
pass() { note "PASS $1"; }
bad() { note "FAIL $1"; fail=1; }
echo "== 1/4 structural + best-practice validation (sigma check) =="
if check_out="$(sigma check /sigma 2>&1)" \
&& printf '%s' "$check_out" | grep -q 'Found 0 errors'; then
pass "sigma check: 0 errors, 0 condition errors, 0 issues"
else
bad "sigma check reported problems"
printf '%s\n' "$check_out" | sed 's/^/ /'
fi
echo "== 2/4 compile the whole tree to a backend (sigma convert -> splunk) =="
if tree_out="$(sigma convert -t splunk --without-pipeline /sigma 2>&1)"; then
pass "whole tree converts to splunk (exit 0)"
else
bad "whole-tree conversion failed"
printf '%s\n' "$tree_out" | sed 's/^/ /'
tree_out=""
fi
echo "== 3/4 each atomic indicator survives into the compiled query =="
# Grep the indicator VALUES, not backend field names or quoting, so the test is
# robust across splunk-backend releases. These strings come straight from the
# rule bodies, which are grounded in this repository's source. The last one is
# the recording-proxy CA filename, grounded in traffic/traffic.go.
for ind in 'artex-enrich/1.0' 'artex-selfupdate' '【ARTEX 平台管控·非目标防御】' 'mitmproxy-ca-cert.pem'; do
if printf '%s' "$tree_out" | grep -qF "$ind"; then
pass "indicator present: $ind"
else
bad "indicator missing from compiled query: $ind"
fi
done
echo "== 4/4 correlation rules compile as correlations, and depend on their base rules =="
# The event_count / value_count aggregation aliases prove the correlation rules
# were compiled as correlations (not dropped), using the whole tree so their
# base-rule references resolve.
if printf '%s' "$tree_out" | grep -q 'event_count' \
&& printf '%s' "$tree_out" | grep -q 'value_count'; then
pass "correlation aggregations present (event_count, value_count)"
else
bad "correlation aggregations missing from compiled query"
fi
# Specificity, mirrored from the Suricata test: converting one correlation rule
# ALONE must fail, because it references an atomic rule by id that is absent from
# a single-file input. A passing conversion here would mean the reference is
# decorative; this asserts it is load-bearing.
if sigma convert -t splunk --without-pipeline \
/sigma/correlation/artex_enrich_scan_velocity.yml >/dev/null 2>&1; then
bad "a correlation rule converted alone (its base-rule reference is not enforced)"
else
pass "correlation rule fails to convert alone — it requires its atomic base rule"
fi
echo
echo "reference: sigma-cli ${VERSION}, splunk backend (latest), pySigma"
if [ "$fail" -eq 0 ]; then
echo "RESULT: PASS"
else
echo "RESULT: FAIL"
fi
exit "$fail"
+42
View File
@@ -0,0 +1,42 @@
#!/usr/bin/env bash
#
# Reproducible regression test for the ARTEX Sigma rules (../../sigma/). It turns
# the "validated by sigma check and sigma convert" claim in the rule README into
# something a reviewer can re-run from source with one command, and it catches
# regressions: a malformed rule, a broken correlation reference, or an indicator
# string that silently dropped out of the compiled query.
#
# It proves five properties with no host dependency beyond Docker (sigma-cli
# runs in a container, nothing is installed on the host and nothing is written to
# the repo tree):
#
# 1. sigma check passes 0 errors / 0 condition errors / 0 issues
# 2. the whole tree compiles sigma convert -> splunk, exit 0
# 3. atomic indicators survive artex-enrich/1.0, artex-selfupdate, guard marker, mitmproxy-ca-cert.pem
# 4. correlations compile event_count / value_count aggregations present
# 5. correlations are load-bearing one correlation rule converted alone FAILS,
# because it references its atomic base rule by id
#
# Unlike a live event-matching harness (which needs a backend that normalizes the
# generic webserver/proxy/application fields — see ../README.md), this is the
# structural + compilation validation the Sigma README documents, made executable.
#
# Usage: detections/tests/sigma/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
# SIGMA_CLI_VERSION (default 3.1.0 — the pinned reference version)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
SIGMA_DIR="$REPO/detections/sigma"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
SIGMA_CLI_VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
# Everything runs inside the container: check.sh installs the pinned sigma-cli and
# the splunk backend, then asserts the five properties and exits non-zero on any
# failure. The rule tree and this directory are mounted read-only.
docker run --rm \
-v "$SIGMA_DIR:/sigma:ro" \
-v "$HERE:/src:ro" \
-e SIGMA_CLI_VERSION="$SIGMA_CLI_VERSION" \
"$PYTHON_IMAGE" sh /src/check.sh
+90
View File
@@ -0,0 +1,90 @@
#!/bin/sh
#
# In-container half of the ARTEX Sigma backend-portability test. run.sh launches
# this inside a Python container with the Sigma rule tree mounted read-only at
# /sigma. It installs a pinned sigma-cli (pySigma) plus four stable backends and
# proves that the rules convert beyond the single Splunk example the README used
# to show, and that the documented per-backend guidance is true for OUR rules.
#
# The base Sigma test (../sigma/) proves the rules are correct against Splunk.
# This test proves they are PORTABLE, and pins the two facts the README's
# "Validate and convert" section now documents:
#
# 1. Correlations are portable the WHOLE tree (atomic + correlation)
# beyond Splunk converts on splunk, Elasticsearch eql,
# and Grafana loki (exit 0), and the enrich
# indicator value survives into each query.
# 2. The atomic-only fallback works backends that do not support Sigma
# where correlations are not correlation conversion (Elasticsearch
# supported lucene, Microsoft kusto) still convert
# the five atomic rules (exit 0), with the
# enrich indicator surviving.
#
# Every assertion is POSITIVE (a capability that must keep working), so the test
# only fails on a genuine regression: a rule that stops converting, or a backend
# that drops support. It deliberately does not assert the negative "backend X
# cannot do correlations" — that would break when a backend improves. The honest
# limitation is documented in ../README.md, reproduced by this test's commands.
#
# POSIX sh (the slim image ships dash). Exits non-zero if any assertion fails.
set -eu
VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
pip install --quiet --disable-pip-version-check "sigma-cli==${VERSION}" >/dev/null 2>&1
# elasticsearch ships the lucene + eql targets; the others are one plugin each.
for plugin in splunk elasticsearch loki kusto; do
sigma plugin install "$plugin" >/dev/null 2>&1
done
ENRICH='artex-enrich/1.0'
ATOMICS='/sigma/artex_enrich_user_agent.yml /sigma/artex_selfupdate_egress.yml /sigma/artex_guard_audit_framing.yml /sigma/artex_recording_proxy_ca.yml /sigma/destructive_command_hunting.yml'
fail=0
note() { printf ' %s\n' "$1"; }
pass() { note "PASS $1"; }
bad() { note "FAIL $1"; fail=1; }
# Backends escape regex metacharacters differently (lucene: artex\-enrich\/1.0,
# loki: artex\-enrich/1\.0, splunk/eql/kusto: artex-enrich/1.0). Strip backslashes
# before matching so the indicator-survival check is robust across all of them
# without asserting any one backend's escaping syntax.
has_enrich() { printf '%s' "$1" | tr -d '\\' | grep -qF "$ENRICH"; }
echo "== 1/2 correlations are portable: the whole tree converts beyond Splunk =="
# Whole-tree conversion includes the four correlation rules, which reference
# their atomic base rules by id. If a backend compiles the whole tree at exit 0
# it supports Sigma correlation conversion for our rules.
for target in splunk eql loki; do
if out="$(sigma convert -t "$target" --without-pipeline /sigma 2>&1)" \
&& has_enrich "$out"; then
pass "whole tree (atomic + correlation) converts on '$target', enrich indicator survives"
else
bad "whole-tree conversion on '$target' failed or dropped the enrich indicator"
printf '%s' "$out" | grep -iE 'error|not supported' | head -2 | sed 's/^/ /'
fi
done
echo "== 2/2 atomic-only fallback: the five atomic rules convert where correlations are not supported =="
# Lucene and kusto (the Microsoft Sentinel / Defender backend) do not convert
# Sigma correlations at the pinned versions, so a defender deploys the five
# atomic rules and expresses the correlation logic natively. That fallback must
# work: all five atomic rules convert and the enrich indicator survives.
for target in lucene kusto; do
if out="$(sigma convert -t "$target" --without-pipeline $ATOMICS 2>&1)" \
&& has_enrich "$out"; then
pass "five atomic rules convert on '$target', enrich indicator survives"
else
bad "atomic-only conversion on '$target' failed or dropped the enrich indicator"
printf '%s' "$out" | grep -iE 'error|not supported' | head -2 | sed 's/^/ /'
fi
done
echo
echo "reference: sigma-cli ${VERSION}; backends splunk, elasticsearch (lucene/eql), loki, kusto (latest compatible), pySigma"
if [ "$fail" -eq 0 ]; then
echo "RESULT: PASS"
else
echo "RESULT: FAIL"
fi
exit "$fail"
+40
View File
@@ -0,0 +1,40 @@
#!/usr/bin/env bash
#
# Reproducible backend-portability test for the ARTEX Sigma rules (../../sigma/).
# The rule README claims the rules "convert to your own SIEM or EDR query
# language" and lists several supported targets. The base Sigma test (../sigma/)
# only exercises Splunk; this test turns the cross-backend claim into something a
# reviewer can re-run, and keeps the README's per-backend guidance honest.
#
# It proves two properties with no host dependency beyond Docker (sigma-cli and
# its backends run in a container, nothing is installed on the host and nothing
# is written to the repo tree):
#
# 1. correlations are portable the whole tree converts on splunk, the
# Elasticsearch eql target, and Grafana loki
# 2. the atomic-only fallback the five atomic rules convert on lucene and
# works kusto (Microsoft Sentinel / Defender), which
# do not support Sigma correlation conversion
#
# See ../README.md "Sigma backend portability" for the measured support matrix
# and the exact per-backend commands this test reproduces.
#
# Usage: detections/tests/sigma_backends/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
# SIGMA_CLI_VERSION (default 3.1.0 — the pinned reference version)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
SIGMA_DIR="$REPO/detections/sigma"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
SIGMA_CLI_VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
# Everything runs inside the container: check.sh installs the pinned sigma-cli and
# four backends, then asserts the two properties and exits non-zero on any
# failure. The rule tree and this directory are mounted read-only.
docker run --rm \
-v "$SIGMA_DIR:/sigma:ro" \
-v "$HERE:/src:ro" \
-e SIGMA_CLI_VERSION="$SIGMA_CLI_VERSION" \
"$PYTHON_IMAGE" sh /src/check.sh
+85
View File
@@ -0,0 +1,85 @@
#!/bin/sh
#
# In-container half of the ARTEX Sigma SigmaHQ-convention lint test. run.sh
# launches this inside a Python container with the Sigma rule tree mounted
# read-only at /sigma and this directory at /src. It installs a pinned sigma-cli
# plus the pinned SigmaHQ validator plugin, then asserts two properties:
#
# 1. baseline is clean sigma check with the documented validators.yml
# baseline reports 0 errors and 0 issues.
# 2. the full set is live running ALL SigmaHQ validators (no exclusions) still
# reports issues, and every issue type is one of the
# four documented, excluded categories — nothing else.
#
# Property 2 is the anti-vacuity guard. If the validator plugin failed to load,
# the "all" run would report zero issues and property 1 would pass vacuously;
# requiring the known exclusions to appear proves the full SigmaHQ set actually
# ran. It also fails the build the moment a rule picks up a NEW convention issue
# outside the documented baseline (e.g. a mis-cased title or an invalid field),
# because that issue type would not be in the allow-list below and property 1
# would stop being clean.
#
# POSIX sh (the slim image ships dash). Exits non-zero if any assertion fails.
set -eu
VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
SIGMAHQ_VALIDATORS_VERSION="${SIGMAHQ_VALIDATORS_VERSION:-0.21.0}"
pip install --quiet --disable-pip-version-check \
"sigma-cli==${VERSION}" "pySigma-validators-sigmahq==${SIGMAHQ_VALIDATORS_VERSION}" >/dev/null 2>&1
# The four issue types the documented baseline (validators.yml) intentionally
# excludes. Any issue outside this set must fail the build.
ALLOWED='SigmahqGithubLinkIssue SigmahqFilenamePrefixIssue SigmahqCorrelationFilenamePrefixIssue SigmahqLogsourceUnknownIssue'
printf 'validators:\n - all\n' > /tmp/all.yml
fail=0
note() { printf ' %s\n' "$1"; }
pass() { note "PASS $1"; }
bad() { note "FAIL $1"; fail=1; }
echo "== 1/2 documented SigmaHQ baseline is clean (validators.yml) =="
if base_out="$(sigma check --validation-config /src/validators.yml /sigma 2>&1)" \
&& printf '%s' "$base_out" | grep -q 'Found 0 errors, 0 condition errors and 0 issues'; then
pass "sigma check with the documented baseline: 0 errors, 0 issues"
else
bad "the documented baseline reported problems (a non-excluded convention issue, or an error)"
printf '%s\n' "$base_out" | sed 's/^/ /'
fi
echo "== 2/2 the full SigmaHQ validator set runs, and only the documented exclusions remain =="
all_out="$(sigma check --validation-config /tmp/all.yml /sigma 2>&1 || true)"
# Collect the distinct issue types the full set reports.
types="$(printf '%s' "$all_out" | grep -oE 'issue=Sigmahq[A-Za-z]+Issue' | sed 's/^issue=//' | sort -u)"
if [ -z "$types" ]; then
bad "the full validator set reported no SigmaHQ issues at all — the plugin did not load (vacuous)"
else
# Anti-vacuity: the two load-bearing exclusions must actually appear.
for must in SigmahqGithubLinkIssue SigmahqLogsourceUnknownIssue; do
if printf '%s\n' "$types" | grep -qx "$must"; then
pass "full set is live: $must present"
else
bad "expected $must from the full validator set but it was absent — plugin/version drift"
fi
done
# No issue type outside the documented allow-list may appear.
unexpected=0
for t in $types; do
case " $ALLOWED " in
*" $t "*) : ;;
*) bad "undocumented convention issue from the full set: $t"; unexpected=1 ;;
esac
done
[ "$unexpected" -eq 0 ] && pass "every reported issue is one of the four documented exclusions"
fi
echo
echo "reference: sigma-cli ${VERSION}, pySigma-validators-sigmahq ${SIGMAHQ_VALIDATORS_VERSION}"
if [ "$fail" -eq 0 ]; then
echo "RESULT: PASS"
else
echo "RESULT: FAIL"
fi
exit "$fail"
+41
View File
@@ -0,0 +1,41 @@
#!/usr/bin/env bash
#
# Reproducible SigmaHQ-convention lint for the ARTEX Sigma rules (../../sigma/).
# `sigma check` on its own runs only pySigma's core validators; this test runs
# the full SigmaHQ convention set (the pySigma-validators-sigmahq plugin) against
# the documented baseline in validators.yml, so the "passes sigma check cleanly"
# claim in the README and CONTRIBUTING covers SigmaHQ's conventions, not just the
# core checks.
#
# It proves two properties with no host dependency beyond Docker (everything runs
# in a container, nothing is installed on the host and nothing is written to the
# repo tree):
#
# 1. the documented baseline (validators.yml) reports 0 errors and 0 issues
# 2. the full validator set actually runs, and only the four documented
# exclusions remain — the anti-vacuity guard (see check.sh)
#
# The four exclusions and the rationale for each live in validators.yml.
#
# Usage: detections/tests/sigma_lint/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
# SIGMA_CLI_VERSION (default 3.1.0 — the pinned reference version)
# SIGMAHQ_VALIDATORS_VERSION (default 0.21.0 — the pinned validator plugin)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
SIGMA_DIR="$REPO/detections/sigma"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
SIGMA_CLI_VERSION="${SIGMA_CLI_VERSION:-3.1.0}"
SIGMAHQ_VALIDATORS_VERSION="${SIGMAHQ_VALIDATORS_VERSION:-0.21.0}"
# Everything runs inside the container: check.sh installs the pinned sigma-cli and
# SigmaHQ validator plugin, then asserts the two properties and exits non-zero on
# any failure. The rule tree and this directory are mounted read-only.
docker run --rm \
-v "$SIGMA_DIR:/sigma:ro" \
-v "$HERE:/src:ro" \
-e SIGMA_CLI_VERSION="$SIGMA_CLI_VERSION" \
-e SIGMAHQ_VALIDATORS_VERSION="$SIGMAHQ_VALIDATORS_VERSION" \
"$PYTHON_IMAGE" sh /src/check.sh
@@ -0,0 +1,44 @@
# SigmaHQ validator baseline for the ARTEX detection rules.
#
# `sigma check` on its own runs only pySigma's core validators. This config turns
# on the full SigmaHQ convention set (the pySigma-validators-sigmahq plugin) and
# then disables four checks that encode SigmaHQ *monorepo* conventions which do
# not apply to this small, self-contained rule set. Every other SigmaHQ check is
# enforced, and detections/tests/sigma_lint/ fails the build if any enabled check
# reports an issue. Each exclusion below is a deliberate, documented decision, not
# a silenced defect.
#
# Run:
# pip install pySigma-validators-sigmahq
# sigma check --validation-config detections/tests/sigma_lint/validators.yml detections/sigma/
#
validators:
- all
# sigmahq_github_link wants every `references:` URL to be a commit permalink
# rather than a branch link. That check exists so rules citing external,
# third-party write-ups keep pointing at the exact revision they were written
# against. Our references point at *our own* living defense docs
# (docs/defense-ko.md, docs/defense-en.md) on `main`: we want them to track the
# current guide, not freeze to a snapshot that goes stale as the guide improves.
- -sigmahq_github_link
# sigmahq_filename_prefix and sigmahq_correlation_filename_prefix require
# logsource-prefixed filenames (web_*, proxy_*) and a correlation_* prefix, the
# filing scheme of SigmaHQ's single flat rules/ tree. This repository ships a
# small set under detections/sigma/ with descriptive artex_* names and a
# correlation/ subdirectory, referenced by the correlation rules' header
# comments, the reproduction tests, and the README index. Renaming to the
# monorepo prefixes would desynchronize those references for no gain on a
# standalone set.
- -sigmahq_filename_prefix
- -sigmahq_correlation_filename_prefix
# sigmahq_logsource_unknown flags `category: application` (the guard-marker
# forensic log search) and a product-less `category: process_creation` (the
# cross-platform destructive-command hunting lead) as outside the SigmaHQ
# taxonomy. Both logsources are intentionally generic: these indicators appear
# across heterogeneous application/audit and process-creation logs, and the
# README tells defenders to map them to their own pipeline. Pinning a single
# product would narrow the rules incorrectly.
- -sigmahq_logsource_unknown
+416
View File
@@ -0,0 +1,416 @@
#!/usr/bin/env python3
#
# Live event-matching test for the ARTEX Sigma rules (../../sigma/). check.sh
# installs a pinned pySigma inside a container and runs this script with the rule
# tree mounted read-only at /sigma and this directory at /src.
#
# WHAT THIS PROVES, AND WHY IT IS DIFFERENT FROM THE sigma/ SUITE
# --------------------------------------------------------------
# The sigma/ suite proves each rule is structurally valid and COMPILES to a
# backend query, and that its indicator strings survive into that query. It does
# NOT prove the rule actually fires on a matching event, or stays quiet on a
# benign one: a field renamed to something the log never carries, a wildcard that
# silently dropped, or an over-broad token would all still compile cleanly. The
# README's own principle is that "a detection you cannot run is only a claim," and
# the Suricata suite already backs its network rule with a real pcap replay
# (fires on the probe UA, silent on a benign browser). This suite closes the same
# gap for the host/log-layer Sigma rules on two levels:
# - ATOMIC rules (../../sigma/*.yml): for each rule a representative malicious
# event MATCHES and a benign event DOES NOT.
# - CORRELATION rules (../../sigma/correlation/*.yml): for each rule a positive
# timeline (threshold met, inside the window, within one group) FIRES and
# negative timelines (below threshold, threshold met but spread beyond the
# window, split across groups, or missing a leg) stay QUIET.
#
# HOW IT MATCHES (trust model)
# ----------------------------
# It does not hand-parse the YAML or re-implement Sigma's modifier logic. pySigma
# parses each rule and compiles its modifiers and condition into a tree:
# `|contains` becomes a wildcard-wrapped value, `|all` becomes an AND over values,
# `1 of selection_*` becomes an OR over the selection groups. This script only
# walks that compiled tree (AND / OR / NOT / field-equals / keyword) and tests
# each leaf against the event, so the authoritative parsing stays in pySigma. A
# leaf value or condition node this script does not explicitly support raises
# rather than passing silently (fail-closed), so a future rule using an
# unsupported construct surfaces loudly here instead of being waved through.
#
# For a correlation rule, pySigma likewise parses the aggregation spec — type
# (event_count / value_count / temporal), group-by fields, timespan, the
# threshold condition, and the resolved references to the atomic base rules. This
# script walks that parsed spec and applies it to a timeline, deciding which
# events feed each referenced rule with the very same atomic matcher above, so the
# Sigma logic again stays in pySigma; only the windowed aggregation is applied
# here. The correlation rules reference their atomics by id, so each is parsed in a
# collection that also holds every atomic rule (pySigma resolves the reference).
#
# SCOPE AND HONESTY (read before trusting a green run)
# ----------------------------------------------------
# - The CORRELATION window is the standard sliding-window interpretation: a
# window of `timespan` seconds anchored at each matching event, with inclusive
# bounds. Each timeline event carries an integer `ts` in relative seconds. A
# real SIEM's windowing (tumbling vs sliding, bound inclusivity, late arrival)
# may differ; this is a regression test for the rule's group-by / timespan /
# threshold logic — that it fires when they are satisfied and not when they are
# not — rather than a bit-exact model of any one backend's correlation engine.
# - Matching is CASE-INSENSITIVE. This mirrors the default of the splunk backend
# the sigma/ suite targets, and the destructive rule's own false-positive note
# assumes it (it warns that lowercase coreutils `truncate` shares the uppercase
# `TRUNCATE ` token and must be allow-listed). Your SIEM's case handling and
# field normalisation may differ; this is a regression test for the rules'
# field/value/condition logic, not a substitute for validating in your stack.
# - Keyword matching (the audit-framing rule) is modelled as a full-text
# substring search across all event field values, the common interpretation of
# an unbound Sigma keyword.
#
# Exits non-zero on any failure. Standard library only beyond pySigma.
import glob
import json
import os
import re
import sys
from sigma.collection import SigmaCollection
from sigma.conditions import (
ConditionAND,
ConditionFieldEqualsValueExpression,
ConditionNOT,
ConditionOR,
ConditionValueExpression,
)
from sigma.types import (
SigmaNull,
SigmaNumber,
SigmaRegularExpression,
SigmaString,
SpecialChars,
)
SIGMA_DIR = os.environ.get("SIGMA_DIR", "/sigma")
EVENTS_DIR = os.environ.get("EVENTS_DIR", "/src/events")
CORR_DIR = os.path.join(SIGMA_DIR, "correlation")
CORR_EVENTS_DIR = os.path.join(EVENTS_DIR, "correlation")
fail = 0
def note(msg):
print(f" {msg}")
def passed(msg):
note(f"PASS {msg}")
def bad(msg):
global fail
note(f"FAIL {msg}")
fail = 1
# --- matcher -------------------------------------------------------------------
def sigmastring_to_regex(value):
"""Compile a pySigma SigmaString (literal text plus wildcards) to an anchored,
case-insensitive regex. `|contains` already wrapped the value in multi
wildcards upstream, so a plain string compiles to an exact match and a
contains-value compiles to a substring match — exactly the Sigma semantics."""
parts = []
for part in value.s:
if part == SpecialChars.WILDCARD_MULTI:
parts.append(".*")
elif part == SpecialChars.WILDCARD_SINGLE:
parts.append(".")
elif isinstance(part, str):
parts.append(re.escape(part))
else:
raise ValueError(f"unsupported SigmaString part: {part!r}")
return re.compile("^" + "".join(parts) + "$", re.DOTALL | re.IGNORECASE)
def field_match(field, value, event):
if field not in event:
return False
observed = str(event[field])
if isinstance(value, SigmaString):
return sigmastring_to_regex(value).search(observed) is not None
if isinstance(value, SigmaNumber):
return observed == str(value.number)
if isinstance(value, SigmaNull):
return event.get(field) is None
if isinstance(value, SigmaRegularExpression):
return re.search(value.regexp, observed) is not None
raise ValueError(f"unsupported field value type: {type(value).__name__}")
def keyword_match(value, event):
"""Unbound keyword: full-text substring search across all field values."""
if not isinstance(value, SigmaString):
raise ValueError("unsupported keyword value type")
if any(not isinstance(p, str) for p in value.s):
raise ValueError("wildcard in keyword is not supported by this matcher")
token = "".join(value.s)
haystack = " ".join(str(v) for v in event.values())
return token.lower() in haystack.lower()
def evaluate(node, event):
if isinstance(node, ConditionAND):
return all(evaluate(a, event) for a in node.args)
if isinstance(node, ConditionOR):
return any(evaluate(a, event) for a in node.args)
if isinstance(node, ConditionNOT):
return not evaluate(node.args[0], event)
if isinstance(node, ConditionFieldEqualsValueExpression):
return field_match(node.field, node.value, event)
if isinstance(node, ConditionValueExpression):
return keyword_match(node.value, event)
raise ValueError(f"unsupported condition node: {type(node).__name__}")
def rule_matches(rule, event):
return any(evaluate(c.parsed, event) for c in rule.detection.parsed_condition)
# --- correlation evaluator -----------------------------------------------------
def group_key(event, fields):
if any(f not in event for f in fields):
return None
return tuple(event[f] for f in fields)
def correlation_fires(corr, timeline):
"""Apply a parsed SigmaCorrelationRule's aggregation to a timeline of events
(each carrying an integer `ts` in seconds). pySigma has parsed the rule into a
type, group-by fields, a timespan, a threshold condition, and resolved rule
references; this walks that parsed structure. Membership in a referenced rule
is decided by the same rule_matches the atomic suite uses, so the Sigma
detection logic stays in pySigma. The window is the standard sliding window:
`timespan` seconds anchored at each matching event, inclusive bounds."""
ctype = str(corr.type)
span = corr.timespan.seconds
group_by = corr.group_by or []
refs = [ref.rule for ref in corr.rules]
if ctype in ("event_count", "value_count"):
# A count correlation may reference several base rules; an event feeds the
# count if it matches ANY of them — the same union the temporal branch
# applies below. Looking at refs[0] alone would silently drop events
# matching the other referenced rules, a fail-open this suite's header
# forbids. With a single reference this reduces to the one-rule case, so
# the existing rules (each referencing one base rule) are unchanged.
matched = [e for e in timeline if any(rule_matches(r, e) for r in refs)]
groups = {}
for e in matched:
key = group_key(e, group_by)
if key is None:
continue
groups.setdefault(key, []).append(e)
threshold = corr.condition.count
fieldref = corr.condition.fieldref
for members in groups.values():
members = sorted(members, key=lambda e: e["ts"])
for anchor in members:
window = [
e for e in members if anchor["ts"] <= e["ts"] <= anchor["ts"] + span
]
if ctype == "event_count":
if len(window) >= threshold:
return True
else:
distinct = {e[fieldref] for e in window if fieldref in e}
if len(distinct) >= threshold:
return True
return False
if ctype == "temporal":
groups = {}
for e in timeline:
key = group_key(e, group_by)
if key is None:
continue
groups.setdefault(key, []).append(e)
for members in groups.values():
members = sorted(members, key=lambda e: e["ts"])
for anchor in members:
window = [
e for e in members if anchor["ts"] <= e["ts"] <= anchor["ts"] + span
]
if all(any(rule_matches(r, e) for e in window) for r in refs):
return True
return False
raise ValueError(f"unsupported correlation type: {ctype}")
def require_ts(events, stem, label):
for e in events:
if not isinstance(e.get("ts"), int):
raise ValueError(
f"{stem} ({label}): every timeline event needs an integer 'ts' "
f"(seconds); got {e!r}"
)
# --- loaders -------------------------------------------------------------------
def load_atomic_rules():
rules = {}
for path in sorted(glob.glob(os.path.join(SIGMA_DIR, "*.yml"))):
stem = os.path.splitext(os.path.basename(path))[0]
collection = SigmaCollection.from_yaml(open(path, encoding="utf-8").read())
for rule in collection.rules:
# Only plain atomic rules; correlation rules carry a `.type` and are
# handled separately below.
if type(rule).__name__ != "SigmaRule":
continue
rules[stem] = rule
return rules
def load_correlation_rules():
"""A correlation rule references its atomic base rules by id, so it must be
parsed in a collection that also contains those atomics. For each correlation
file, merge every atomic YAML with that one correlation YAML, parse the
collection (pySigma resolves the reference), and key the resulting
SigmaCorrelationRule by filename stem so it pairs with
events/correlation/<stem>.json."""
atomic_docs = [
open(p, encoding="utf-8").read()
for p in sorted(glob.glob(os.path.join(SIGMA_DIR, "*.yml")))
]
corrs = {}
for path in sorted(glob.glob(os.path.join(CORR_DIR, "*.yml"))):
stem = os.path.splitext(os.path.basename(path))[0]
merged = "\n---\n".join(atomic_docs + [open(path, encoding="utf-8").read()])
collection = SigmaCollection.from_yaml(merged)
found = [r for r in collection.rules if type(r).__name__ == "SigmaCorrelationRule"]
if len(found) != 1:
raise ValueError(f"{stem}: expected exactly 1 correlation rule, got {len(found)}")
corrs[stem] = found[0]
return corrs
def load_events(directory):
events = {}
for path in sorted(glob.glob(os.path.join(directory, "*.json"))):
stem = os.path.splitext(os.path.basename(path))[0]
events[stem] = json.load(open(path, encoding="utf-8"))
return events
def main():
rules = load_atomic_rules()
events = load_events(EVENTS_DIR)
print("== atomic 1/3 every atomic rule is paired with a sample-event file ==")
rule_stems = set(rules)
event_stems = set(events)
orphan_rules = sorted(rule_stems - event_stems)
orphan_events = sorted(event_stems - rule_stems)
if orphan_rules:
bad(f"atomic rules with no events/<name>.json: {orphan_rules}")
if orphan_events:
bad(f"event files with no matching atomic rule: {orphan_events}")
if not orphan_rules and not orphan_events:
passed(
f"rule/sample pairing: {len(rules)} atomic rules, "
f"{len(events)} event files, no orphans"
)
print("== atomic 2/3 each rule matches its malicious sample events (true positives) ==")
for stem in sorted(rule_stems & event_stems):
rule = rules[stem]
positives = events[stem].get("positive", [])
if not positives:
bad(f"{stem}: no positive sample events")
continue
missed = [e for e in positives if not rule_matches(rule, e)]
if missed:
bad(f"{stem}: {len(missed)}/{len(positives)} positive events did NOT match")
for e in missed:
note(f" unmatched: {json.dumps(e, ensure_ascii=False)}")
else:
passed(f"{stem}: {len(positives)}/{len(positives)} positive events matched")
print("== atomic 3/3 each rule rejects its benign sample events (true negatives) ==")
for stem in sorted(rule_stems & event_stems):
rule = rules[stem]
negatives = events[stem].get("negative", [])
if not negatives:
bad(f"{stem}: no negative sample events")
continue
fired = [e for e in negatives if rule_matches(rule, e)]
if fired:
bad(f"{stem}: {len(fired)}/{len(negatives)} benign events WRONGLY matched")
for e in fired:
note(f" wrongly matched: {json.dumps(e, ensure_ascii=False)}")
else:
passed(
f"{stem}: {len(negatives)}/{len(negatives)} benign events correctly "
"not matched"
)
corr_rules = load_correlation_rules()
corr_events = load_events(CORR_EVENTS_DIR)
print("== correlation 1/3 every correlation rule is paired with a timeline file ==")
corr_stems = set(corr_rules)
ce_stems = set(corr_events)
orphan_corr = sorted(corr_stems - ce_stems)
orphan_tl = sorted(ce_stems - corr_stems)
if orphan_corr:
bad(f"correlation rules with no events/correlation/<name>.json: {orphan_corr}")
if orphan_tl:
bad(f"timeline files with no matching correlation rule: {orphan_tl}")
if not orphan_corr and not orphan_tl:
passed(
f"rule/timeline pairing: {len(corr_rules)} correlation rules, "
f"{len(corr_events)} timeline files, no orphans"
)
print("== correlation 2/3 each rule fires on its positive timelines (true positives) ==")
for stem in sorted(corr_stems & ce_stems):
corr = corr_rules[stem]
positives = corr_events[stem].get("positive", [])
if not positives:
bad(f"{stem}: no positive timelines")
continue
for tl in positives:
require_ts(tl["events"], stem, tl["label"])
if correlation_fires(corr, tl["events"]):
passed(f"{stem}: fired — {tl['label']}")
else:
bad(f"{stem}: did NOT fire on a positive timeline — {tl['label']}")
print("== correlation 3/3 each rule stays quiet on its negative timelines (true negatives) ==")
for stem in sorted(corr_stems & ce_stems):
corr = corr_rules[stem]
negatives = corr_events[stem].get("negative", [])
if not negatives:
bad(f"{stem}: no negative timelines")
continue
for tl in negatives:
require_ts(tl["events"], stem, tl["label"])
if correlation_fires(corr, tl["events"]):
bad(f"{stem}: WRONGLY fired on a benign timeline — {tl['label']}")
else:
passed(f"{stem}: quiet — {tl['label']}")
print()
try:
import importlib.metadata as md
print(f"reference: pySigma {md.version('pysigma')}, atomic + correlation rules")
except Exception:
pass
print("RESULT: PASS" if fail == 0 else "RESULT: FAIL")
sys.exit(fail)
if __name__ == "__main__":
main()
+20
View File
@@ -0,0 +1,20 @@
#!/bin/sh
#
# In-container half of the ARTEX Sigma live event-matching test. run.sh launches
# this inside a Python container with the Sigma rule tree (atomic rules and the
# correlation/ subtree) mounted read-only at /sigma and this directory at /src. It
# installs a pinned pySigma, then hands off to check.py, which asserts that every
# atomic rule matches its malicious sample events and stays quiet on its benign
# ones, and that every correlation rule fires on its positive timeline and stays
# quiet on its negative ones (see check.py's header for the trust model and
# scope). pySigma does the parsing; check.py walks the compiled condition tree and
# aggregation spec and tests each sample event or timeline against it.
#
# POSIX sh (the slim image ships dash). Exits non-zero if any assertion fails.
set -eu
VERSION="${PYSIGMA_VERSION:-2.0.0}"
pip install --quiet --disable-pip-version-check "pysigma==${VERSION}" >/dev/null 2>&1
exec python3 /src/check.py
@@ -0,0 +1,10 @@
{
"note": "webserver access log. The rule matches cs-user-agent EXACTLY equal to 'artex-enrich/1.0' (enrich/enrich.go:233). A browser UA, and the same UA with a trailing suffix, must not match.",
"positive": [
{"cs-method": "GET", "cs-uri-stem": "/", "cs-user-agent": "artex-enrich/1.0", "c-ip": "203.0.113.7"}
],
"negative": [
{"cs-method": "GET", "cs-uri-stem": "/", "cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0", "c-ip": "203.0.113.8"},
{"cs-method": "GET", "cs-uri-stem": "/robots.txt", "cs-user-agent": "artex-enrich/1.0 (proxied)", "c-ip": "203.0.113.9"}
]
}
@@ -0,0 +1,9 @@
{
"note": "application log line. The rule is an unbound keyword matching the guard's audit-control framing marker (guard/guard.go). It should match wherever the marker appears in the message, and stay quiet on an ordinary log line.",
"positive": [
{"message": "2026-10-07T03:11:09Z guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"}
],
"negative": [
{"message": "2026-10-07T03:11:09Z auth: user login ok uid=42 ip=203.0.113.8"}
]
}
@@ -0,0 +1,10 @@
{
"note": "file-creation (file_event) telemetry. The rule needs TargetFilename to contain BOTH '_ca' AND 'mitmproxy-ca-cert.pem' (|all), which narrows it to ARTEX's '<dir>/_ca/mitmproxy-ca-cert.pem' layout (traffic/traffic.go). A standalone mitmproxy cert under .mitmproxy/ has the filename but not the '_ca' directory, so it must NOT match — that is the specificity the |all modifier buys.",
"positive": [
{"TargetFilename": "/home/ubuntu/.local/share/artex/data/_ca/mitmproxy-ca-cert.pem", "Image": "/opt/artex/artex"}
],
"negative": [
{"TargetFilename": "/home/ubuntu/.mitmproxy/mitmproxy-ca-cert.pem", "Image": "/usr/bin/mitmproxy"},
{"TargetFilename": "/etc/ssl/certs/ca-certificates.crt", "Image": "/usr/sbin/update-ca-certificates"}
]
}
@@ -0,0 +1,9 @@
{
"note": "forward-proxy egress log. The rule matches c-useragent EXACTLY equal to 'artex-selfupdate' (selfupdate/github.go), the UA ARTEX sets when it fetches its own release from GitHub. A generic client UA must not match.",
"positive": [
{"c-useragent": "artex-selfupdate", "cs-host": "github.com", "cs-uri-stem": "/Autumn-27/ARTEX/releases/latest"}
],
"negative": [
{"c-useragent": "curl/8.5.0", "cs-host": "github.com", "cs-uri-stem": "/"}
]
}
@@ -0,0 +1,864 @@
{
"note": "webserver access log timeline for the ARTEX Enrichment Fan-Out correlation (value_count of DISTINCT cs-host >= 20, grouped by c-ip, within a 10-minute window). 'ts' is relative seconds. Breadth — distinct hosts touched, not request volume — is the signal, so a high-volume/low-breadth burst must stay quiet.",
"positive": [
{
"label": "20 distinct hosts from one source within the 10-minute window",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 25,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 50,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 75,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.9",
"cs-host": "host04.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 125,
"c-ip": "10.0.0.9",
"cs-host": "host05.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 150,
"c-ip": "10.0.0.9",
"cs-host": "host06.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 175,
"c-ip": "10.0.0.9",
"cs-host": "host07.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "host08.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "host09.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 250,
"c-ip": "10.0.0.9",
"cs-host": "host10.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 275,
"c-ip": "10.0.0.9",
"cs-host": "host11.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 300,
"c-ip": "10.0.0.9",
"cs-host": "host12.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 325,
"c-ip": "10.0.0.9",
"cs-host": "host13.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 350,
"c-ip": "10.0.0.9",
"cs-host": "host14.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 375,
"c-ip": "10.0.0.9",
"cs-host": "host15.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 400,
"c-ip": "10.0.0.9",
"cs-host": "host16.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 425,
"c-ip": "10.0.0.9",
"cs-host": "host17.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 450,
"c-ip": "10.0.0.9",
"cs-host": "host18.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 475,
"c-ip": "10.0.0.9",
"cs-host": "host19.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
}
],
"negative": [
{
"label": "below the breadth threshold: only 19 distinct hosts",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 25,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 50,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 75,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.9",
"cs-host": "host04.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 125,
"c-ip": "10.0.0.9",
"cs-host": "host05.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 150,
"c-ip": "10.0.0.9",
"cs-host": "host06.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 175,
"c-ip": "10.0.0.9",
"cs-host": "host07.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "host08.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "host09.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 250,
"c-ip": "10.0.0.9",
"cs-host": "host10.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 275,
"c-ip": "10.0.0.9",
"cs-host": "host11.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 300,
"c-ip": "10.0.0.9",
"cs-host": "host12.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 325,
"c-ip": "10.0.0.9",
"cs-host": "host13.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 350,
"c-ip": "10.0.0.9",
"cs-host": "host14.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 375,
"c-ip": "10.0.0.9",
"cs-host": "host15.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 400,
"c-ip": "10.0.0.9",
"cs-host": "host16.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 425,
"c-ip": "10.0.0.9",
"cs-host": "host17.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 450,
"c-ip": "10.0.0.9",
"cs-host": "host18.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
},
{
"label": "20 distinct hosts but spread over ~13 minutes, so no single 10-minute window sees 20",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 40,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 80,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 120,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 160,
"c-ip": "10.0.0.9",
"cs-host": "host04.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "host05.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 240,
"c-ip": "10.0.0.9",
"cs-host": "host06.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 280,
"c-ip": "10.0.0.9",
"cs-host": "host07.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 320,
"c-ip": "10.0.0.9",
"cs-host": "host08.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 360,
"c-ip": "10.0.0.9",
"cs-host": "host09.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 400,
"c-ip": "10.0.0.9",
"cs-host": "host10.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 440,
"c-ip": "10.0.0.9",
"cs-host": "host11.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 480,
"c-ip": "10.0.0.9",
"cs-host": "host12.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 520,
"c-ip": "10.0.0.9",
"cs-host": "host13.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 560,
"c-ip": "10.0.0.9",
"cs-host": "host14.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 600,
"c-ip": "10.0.0.9",
"cs-host": "host15.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 640,
"c-ip": "10.0.0.9",
"cs-host": "host16.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 680,
"c-ip": "10.0.0.9",
"cs-host": "host17.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 720,
"c-ip": "10.0.0.9",
"cs-host": "host18.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 760,
"c-ip": "10.0.0.9",
"cs-host": "host19.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
},
{
"label": "high volume, low breadth: 25 requests from one source but only 4 distinct hosts",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 20,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 40,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 60,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 80,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 120,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 140,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 160,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 180,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 220,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 240,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 260,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 280,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 300,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 320,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 340,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 360,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 380,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 400,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 420,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 440,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 460,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 480,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
},
{
"label": "breadth split across two sources: 10 distinct hosts each, neither source reaches 20",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "host00.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 25,
"c-ip": "10.0.0.9",
"cs-host": "host01.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 50,
"c-ip": "10.0.0.9",
"cs-host": "host02.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 75,
"c-ip": "10.0.0.9",
"cs-host": "host03.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.9",
"cs-host": "host04.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 125,
"c-ip": "10.0.0.9",
"cs-host": "host05.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 150,
"c-ip": "10.0.0.9",
"cs-host": "host06.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 175,
"c-ip": "10.0.0.9",
"cs-host": "host07.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "host08.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "host09.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 0,
"c-ip": "10.0.0.10",
"cs-host": "host10.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 25,
"c-ip": "10.0.0.10",
"cs-host": "host11.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 50,
"c-ip": "10.0.0.10",
"cs-host": "host12.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 75,
"c-ip": "10.0.0.10",
"cs-host": "host13.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.10",
"cs-host": "host14.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 125,
"c-ip": "10.0.0.10",
"cs-host": "host15.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 150,
"c-ip": "10.0.0.10",
"cs-host": "host16.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 175,
"c-ip": "10.0.0.10",
"cs-host": "host17.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.10",
"cs-host": "host18.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.10",
"cs-host": "host19.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
}
]
}
@@ -0,0 +1,979 @@
{
"note": "webserver access log timeline for the ARTEX Enrichment Scan Velocity correlation (event_count >= 30 probes from one c-ip within a 5-minute window). 'ts' is relative seconds. Velocity (density in time), not total count, is the signal.",
"positive": [
{
"label": "30 enrichment probes from one source inside the 5-minute window",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 9,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 18,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 27,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 36,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 45,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 54,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 63,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 72,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 81,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 90,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 99,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 108,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 117,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 126,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 135,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 144,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 153,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 162,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 171,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 180,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 189,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 198,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 207,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 216,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 234,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 243,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 252,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 261,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
}
],
"negative": [
{
"label": "below the threshold: only 29 probes",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 9,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 18,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 27,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 36,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 45,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 54,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 63,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 72,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 81,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 90,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 99,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 108,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 117,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 126,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 135,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 144,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 153,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 162,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 171,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 180,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 189,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 198,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 207,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 216,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 234,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 243,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 252,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
},
{
"label": "30 probes total but spread over ~10 minutes, so no 5-minute window reaches 30",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 20,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 40,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 60,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 80,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 100,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 120,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 140,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 160,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 180,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 200,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 220,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 240,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 260,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 280,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 300,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 320,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 340,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 360,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 380,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 400,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 420,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 440,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 460,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 480,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 500,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 520,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 540,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 560,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
},
{
"ts": 580,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "artex-enrich/1.0"
}
]
},
{
"label": "30 requests in the window but the User-Agent is a normal browser (base rule does not match)",
"events": [
{
"ts": 0,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 9,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 18,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 27,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 36,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 45,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 54,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 63,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 72,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 81,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 90,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 99,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 108,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 117,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 126,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 135,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 144,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 153,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 162,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 171,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 180,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 189,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 198,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 207,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 216,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 225,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 234,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 243,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 252,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
},
{
"ts": 261,
"c-ip": "10.0.0.9",
"cs-host": "assets.example.test",
"cs-method": "GET",
"cs-uri-stem": "/",
"cs-user-agent": "Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
}
]
}
]
}
@@ -0,0 +1,152 @@
{
"note": "application/audit log timeline for the ARTEX Guard-Block Burst correlation (event_count >= 5 guard-control markers on one host within a 10-minute window). 'ts' is relative seconds. A single marker can be a quoted string; a burst on one host indicates an actively engaged ARTEX run.",
"positive": [
{
"label": "5 guard-control markers on one host inside the 10-minute window",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 60,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 120,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 180,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 240,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
}
]
}
],
"negative": [
{
"label": "below the threshold: only 4 markers",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 60,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 120,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 180,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
}
]
},
{
"label": "5 markers but spread over ~13 minutes, so no 10-minute window holds 5",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 200,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 400,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 600,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 800,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
}
]
},
{
"label": "markers split across two hosts: 3 and 2, neither host reaches 5",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 60,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 120,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 0,
"host": "web02",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 60,
"host": "web02",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
}
]
},
{
"label": "ordinary log lines on one host, no guard marker (base rule does not match)",
"events": [
{
"ts": 0,
"host": "web01",
"message": "auth: user login ok uid=42 ip=203.0.113.8"
},
{
"ts": 60,
"host": "web01",
"message": "auth: user login ok uid=42 ip=203.0.113.8"
},
{
"ts": 120,
"host": "web01",
"message": "auth: user login ok uid=42 ip=203.0.113.8"
},
{
"ts": 180,
"host": "web01",
"message": "auth: user login ok uid=42 ip=203.0.113.8"
},
{
"ts": 240,
"host": "web01",
"message": "auth: user login ok uid=42 ip=203.0.113.8"
}
]
}
]
}
@@ -0,0 +1,81 @@
{
"note": "host timeline joining application/audit logs and process-creation logs for the ARTEX Guard Marker With Destructive Command temporal correlation (both referenced rules must fire on the SAME host within a 30-minute window). 'ts' is relative seconds.",
"positive": [
{
"label": "guard marker then a destructive command on the same host within 30 minutes",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 600,
"host": "web01",
"CommandLine": "rm -rf / --no-preserve-root",
"Image": "/usr/bin/rm"
}
]
}
],
"negative": [
{
"label": "only the guard marker, no destructive command",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 120,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
}
]
},
{
"label": "only a destructive command, no guard marker",
"events": [
{
"ts": 0,
"host": "web01",
"CommandLine": "rm -rf / --no-preserve-root",
"Image": "/usr/bin/rm"
}
]
},
{
"label": "both present but ~60 minutes apart, outside the 30-minute window",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 3600,
"host": "web01",
"CommandLine": "rm -rf / --no-preserve-root",
"Image": "/usr/bin/rm"
}
]
},
{
"label": "the two legs on different hosts",
"events": [
{
"ts": 0,
"host": "web01",
"message": "guard 【ARTEX 平台管控·非目标防御】此操作被禁止: rm -rf /"
},
{
"ts": 600,
"host": "db02",
"CommandLine": "rm -rf / --no-preserve-root",
"Image": "/usr/bin/rm"
}
]
}
]
}
@@ -0,0 +1,13 @@
{
"note": "process_creation telemetry (CommandLine). The rule hunts destructive shell/DB/availability commands via '1 of selection_*', so one representative command from each of the three selection groups must match. Benign commands — including a plain 'rm' without the recursive/force flags — must not. (The rule's own false-positive note documents that lowercase coreutils 'truncate' DOES share the TRUNCATE token and must be allow-listed, so it is intentionally not used here as a negative.)",
"positive": [
{"CommandLine": "rm -rf /var/www/html", "Image": "/usr/bin/rm"},
{"CommandLine": "mysql -u root -e 'DROP TABLE customers'", "Image": "/usr/bin/mysql"},
{"CommandLine": "iptables -F", "Image": "/usr/sbin/iptables"}
],
"negative": [
{"CommandLine": "ls -la /var/www/html", "Image": "/usr/bin/ls"},
{"CommandLine": "rm /tmp/scratch.txt", "Image": "/usr/bin/rm"},
{"CommandLine": "git status", "Image": "/usr/bin/git"}
]
}
+53
View File
@@ -0,0 +1,53 @@
#!/usr/bin/env bash
#
# Reproducible live event-matching test for the ARTEX Sigma rules — both the
# atomic rules (../../sigma/*.yml) and the correlation rules
# (../../sigma/correlation/*.yml). The sibling sigma/ suite proves those rules are
# valid and COMPILE to a backend query; this suite proves they actually FIRE on a
# matching event (or timeline) and stay quiet on a benign one — the same "a
# detection you cannot run is only a claim" guarantee the suricata/ suite already
# gives the network rule with a pcap replay.
#
# It proves six properties with no host dependency beyond Docker (pySigma runs in
# a container, nothing is installed on the host and nothing is written to the repo
# tree):
#
# atomic 1 rule/sample pairing every atomic rule has an events/<name>.json and
# every events file maps to a rule (no orphans)
# atomic 2 true positives each rule matches all of its malicious events
# atomic 3 true negatives each rule matches none of its benign events
# corr 1 rule/timeline pairing every correlation rule has an
# events/correlation/<name>.json (no orphans)
# corr 2 true positives each rule FIRES on its positive timeline
# (threshold met, inside the window, one group)
# corr 3 true negatives each rule stays QUIET on its negative timelines
# (below threshold, window exceeded, split group,
# or a missing leg)
#
# pySigma parses each rule — for an atomic rule its condition tree, for a
# correlation rule its aggregation spec (type, group-by, timespan, threshold, and
# the resolved references to the atomic base rules) — and check.py only walks that
# parsed structure, so the authoritative Sigma logic stays in pySigma (see
# check.py's header). The correlation window is the standard sliding-window model
# and matching is case-insensitive; see check.py for the full scope and honesty
# notes.
#
# Usage: detections/tests/sigma_match/run.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
# PYSIGMA_VERSION (default 2.0.0 — the pinned reference version)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
SIGMA_DIR="$REPO/detections/sigma"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
PYSIGMA_VERSION="${PYSIGMA_VERSION:-2.0.0}"
# Everything runs inside the container: check.sh installs the pinned pySigma and
# runs check.py, which asserts the three properties and exits non-zero on any
# failure. The rule tree and this directory are mounted read-only.
docker run --rm \
-v "$SIGMA_DIR:/sigma:ro" \
-v "$HERE:/src:ro" \
-e PYSIGMA_VERSION="$PYSIGMA_VERSION" \
"$PYTHON_IMAGE" sh /src/check.sh
+5
View File
@@ -0,0 +1,5 @@
# scratch captures created by run.sh (binary pcaps); never committed.
# Match both files and directories (.scratch.* , not .scratch.*/) so that any
# leftover .scratch.<name> is ignored even when an interrupted run (SIGKILL,
# power loss) skips the EXIT cleanup, and so `git check-ignore` reports it.
.scratch.*
+109
View File
@@ -0,0 +1,109 @@
#!/usr/bin/env python3
"""Deterministic pcap generator for the ARTEX Suricata rule tests.
Synthesizes N independent plaintext HTTP request/response flows from a single
source, each carrying a chosen User-Agent, so `suricata -r` can be run offline
to prove the rules in ../../suricata/artex.rules fire (or stay silent) exactly
as documented. Output is regenerated on every run and is never committed -- the
test ships as source, not as a binary capture.
Usage:
gen_pcap.py <out.pcap> <user-agent> [num_flows] [interval_seconds]
The capture is fully deterministic: fixed addresses, ports derived from the
flow index, a fixed base timestamp, and flows spaced `interval_seconds` apart.
Nothing here sends a packet or touches a network -- it only writes a file.
"""
import sys
from scapy.all import Ether, IP, TCP, Raw, wrpcap
# Fixed, private, non-routable endpoints. One source so Suricata's
# `detection_filter ... track by_src` on sid 1000002 counts per source.
SRC_MAC = "02:00:00:00:00:01"
DST_MAC = "02:00:00:00:00:02"
SRC_IP = "10.10.10.9"
DST_IP = "10.10.10.80"
DST_PORT = 80
BASE_EPOCH = 1_760_000_000.0 # fixed so timestamps never depend on wall clock
CLIENT_ISN = 1000
SERVER_ISN = 2000
def http_request(user_agent: str) -> bytes:
return (
"GET /products?category=all HTTP/1.1\r\n"
"Host: shop.example.test\r\n"
f"User-Agent: {user_agent}\r\n"
"Accept: */*\r\n"
"Connection: close\r\n"
"\r\n"
).encode()
HTTP_RESPONSE = (
"HTTP/1.1 200 OK\r\n"
"Content-Type: text/html\r\n"
"Content-Length: 13\r\n"
"Connection: close\r\n"
"\r\n"
"<html></html>"
).encode()
def flow(index: int, user_agent: str, t0: float):
"""One complete TCP+HTTP conversation; returns a list of timestamped packets."""
sport = 40000 + index
eth_c = Ether(src=SRC_MAC, dst=DST_MAC)
eth_s = Ether(src=DST_MAC, dst=SRC_MAC)
ip_c = IP(src=SRC_IP, dst=DST_IP)
ip_s = IP(src=DST_IP, dst=SRC_IP)
req = http_request(user_agent)
rlen = len(req)
slen = len(HTTP_RESPONSE)
pkts = []
def add(pkt, offset):
pkt.time = t0 + offset
pkts.append(pkt)
# Handshake
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="S", seq=CLIENT_ISN), 0.000)
add(eth_s / ip_s / TCP(sport=DST_PORT, dport=sport, flags="SA", seq=SERVER_ISN, ack=CLIENT_ISN + 1), 0.001)
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="A", seq=CLIENT_ISN + 1, ack=SERVER_ISN + 1), 0.002)
# Request
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="PA", seq=CLIENT_ISN + 1, ack=SERVER_ISN + 1) / Raw(req), 0.003)
add(eth_s / ip_s / TCP(sport=DST_PORT, dport=sport, flags="A", seq=SERVER_ISN + 1, ack=CLIENT_ISN + 1 + rlen), 0.004)
# Response
add(eth_s / ip_s / TCP(sport=DST_PORT, dport=sport, flags="PA", seq=SERVER_ISN + 1, ack=CLIENT_ISN + 1 + rlen) / Raw(HTTP_RESPONSE), 0.005)
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="A", seq=CLIENT_ISN + 1 + rlen, ack=SERVER_ISN + 1 + slen), 0.006)
# Teardown
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="FA", seq=CLIENT_ISN + 1 + rlen, ack=SERVER_ISN + 1 + slen), 0.007)
add(eth_s / ip_s / TCP(sport=DST_PORT, dport=sport, flags="A", seq=SERVER_ISN + 1 + slen, ack=CLIENT_ISN + 2 + rlen), 0.008)
add(eth_s / ip_s / TCP(sport=DST_PORT, dport=sport, flags="FA", seq=SERVER_ISN + 1 + slen, ack=CLIENT_ISN + 2 + rlen), 0.009)
add(eth_c / ip_c / TCP(sport=sport, dport=DST_PORT, flags="A", seq=CLIENT_ISN + 2 + rlen, ack=SERVER_ISN + 2 + slen), 0.010)
return pkts
def main() -> int:
if len(sys.argv) < 3:
print(__doc__)
return 2
out = sys.argv[1]
user_agent = sys.argv[2]
num_flows = int(sys.argv[3]) if len(sys.argv) > 3 else 35
interval = float(sys.argv[4]) if len(sys.argv) > 4 else 1.0
packets = []
for i in range(num_flows):
packets.extend(flow(i, user_agent, BASE_EPOCH + i * interval))
wrpcap(out, packets)
print(f"wrote {len(packets)} packets across {num_flows} flows to {out} (UA={user_agent!r})")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+145
View File
@@ -0,0 +1,145 @@
#!/usr/bin/env bash
#
# Reproducible regression test for the ARTEX Suricata rules
# (../../suricata/artex.rules). It proves four properties with no committed
# binary capture and no host dependencies beyond Docker:
#
# 1. the whole rules file loads with zero errors (validity)
# `suricata -T --init-errors-fatal`; a rule that fails to parse or
# initialise is fatal even when no capture below exercises it
# 2. sid 1000001 fires exactly once per enrich probe (presence)
# 3. sid 1000002 fires once the 30-in-300s rate is hit (velocity)
# 4. sid 1000003 fires once per norma WebFetch request (presence)
# and the enrich sids stay silent on that capture (specificity)
# 5. an identical capture with a benign browser UA (specificity)
# produces zero alerts
#
# Everything runs in containers: `suricata -T` validates the ruleset, scapy
# synthesizes a deterministic pcap, then `suricata -r` reads it offline. The
# pcap is generated into a scratch dir that is removed on exit and is never
# committed.
#
# Usage: detections/tests/suricata/run.sh
# Env: SURICATA_IMAGE (default jasonish/suricata:latest)
# PYTHON_IMAGE (default python:3.12-slim)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../../.." && pwd)"
RULES_DIR="$REPO/detections/suricata"
SURICATA_IMAGE="${SURICATA_IMAGE:-jasonish/suricata:latest}"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
NUM_FLOWS=35
ENRICH_UA="artex-enrich/1.0"
# norma SDK WebFetch tool, hardcoded in github.com/Autumn-27/norma/tool/webfetch.go
# (literal "norma/0.4", verified in the go.sum-pinned v0.4.3 module source). Fewer flows
# than NUM_FLOWS because sid 1000003 is a single-hit presence rule with no rate component.
NORMA_UA="norma/0.4"
NORMA_FLOWS=8
BENIGN_UA="Mozilla/5.0 (X11; Linux x86_64; rv:128.0) Gecko/20100101 Firefox/128.0"
# Scratch must live under the repo tree so Docker Desktop (macOS) can bind-mount
# it; /tmp and $TMPDIR are not shared by default. It is git-ignored and removed
# on exit.
SCRATCH="$(mktemp -d "$HERE/.scratch.XXXXXX")"
cleanup() { rm -rf "$SCRATCH"; }
trap cleanup EXIT
fail=0
note() { printf ' %s\n' "$1"; }
echo "== 1/5 validate the full ruleset loads (suricata -T) =="
# `suricata -T` loads the whole rules file in test mode and exits; --init-errors-fatal
# makes any rule that fails to parse or initialise a hard error. This catches a broken
# rule even when no capture below exercises it: plain `suricata -r` skips such a rule and
# still exits 0, so the firing checks would stay green while a signature silently fails to
# load. This is the Suricata analogue of the Sigma suite's `sigma check` validity assertion.
if docker run --rm -v "$RULES_DIR:/r:ro" "$SURICATA_IMAGE" \
suricata -T -S /r/artex.rules -l /tmp --init-errors-fatal >/dev/null 2>&1; then
note "PASS ruleset loads with zero parse/init errors (suricata -T)"
else
note "FAIL ruleset loads with zero parse/init errors (suricata -T)"
fail=1
fi
echo "== 2/5 synthesize deterministic captures (scapy) =="
docker run --rm -v "$SCRATCH:/out" -v "$HERE:/src:ro" "$PYTHON_IMAGE" sh -c "
pip install --quiet --disable-pip-version-check scapy >/dev/null 2>&1 &&
python /src/gen_pcap.py /out/enrich.pcap '$ENRICH_UA' $NUM_FLOWS &&
python /src/gen_pcap.py /out/norma.pcap '$NORMA_UA' $NORMA_FLOWS &&
python /src/gen_pcap.py /out/benign.pcap '$BENIGN_UA' $NUM_FLOWS
"
run_suricata() { # $1 = capture basename
local name="$1"
mkdir -p "$SCRATCH/$name-out"
# -k none: crafted packets carry no valid checksums; do not drop on them.
docker run --rm -v "$SCRATCH:/data" -v "$RULES_DIR:/r:ro" "$SURICATA_IMAGE" \
suricata -r "/data/$name.pcap" -S /r/artex.rules -k none -l "/data/$name-out" \
>/dev/null 2>&1
}
alerts() { # $1 = capture basename, $2 = sid (or "any")
python3 - "$SCRATCH/$1-out/eve.json" "$2" <<'PY'
import json, sys
path, sid = sys.argv[1], sys.argv[2]
n = 0
with open(path) as f:
for line in f:
line = line.strip()
if not line:
continue
try:
e = json.loads(line)
except ValueError:
continue
if e.get("event_type") != "alert":
continue
if sid == "any" or e.get("alert", {}).get("signature_id") == int(sid):
n += 1
print(n)
PY
}
expect() { # $1 label, $2 actual, $3 op (eq|ge), $4 expected
local label="$1" actual="$2" op="$3" expected="$4" ok
case "$op" in
eq) [ "$actual" -eq "$expected" ] && ok=1 || ok=0 ;;
ge) [ "$actual" -ge "$expected" ] && ok=1 || ok=0 ;;
esac
if [ "$ok" -eq 1 ]; then
note "PASS $label (got $actual, want $op $expected)"
else
note "FAIL $label (got $actual, want $op $expected)"
fail=1
fi
}
echo "== 3/5 run Suricata offline over the enrich capture =="
run_suricata enrich
e1="$(alerts enrich 1000001)"
e2="$(alerts enrich 1000002)"
expect "sid 1000001 presence: one alert per probe" "$e1" eq "$NUM_FLOWS"
expect "sid 1000002 velocity: fires past 30-in-300s" "$e2" ge 1
note "reference (Suricata 8.0.7): sid 1000002 = 5 (flows 31-35)"
echo "== 4/5 run Suricata offline over the norma WebFetch capture =="
run_suricata norma
n3="$(alerts norma 1000003)"
nenrich="$(( $(alerts norma 1000001) + $(alerts norma 1000002) ))"
expect "sid 1000003 presence: one alert per WebFetch request" "$n3" eq "$NORMA_FLOWS"
expect "enrich sids stay silent on norma traffic (specificity)" "$nenrich" eq 0
echo "== 5/5 run Suricata offline over the benign capture =="
run_suricata benign
b="$(alerts benign any)"
expect "benign browser UA produces no ARTEX alerts" "$b" eq 0
echo
if [ "$fail" -eq 0 ]; then
echo "RESULT: PASS"
else
echo "RESULT: FAIL"
fi
exit "$fail"
+31
View File
@@ -0,0 +1,31 @@
#!/usr/bin/env bash
#
# Non-suite gate: runs the host-triage tool's built-in --self-test in a container.
#
# The triage tool (detections/triage/artex_host_triage.py) is a responder helper,
# not a detection rule, so it is deliberately NOT one of the detections/tests/<x>/
# run.sh suites — that keeps the harness-sync registry exactly the eight rule
# suites (check-harness-sync.py counts only SUITES entries, `detections/tests/<x>/
# run.sh` CI steps, and directories carrying a run.sh; this file is none of them).
# It is a gate like check-harness-sync.sh: both run-all.sh and the detections CI
# workflow call it, so a broken triage check fails the same merge gate as the
# rule suites. A detection you cannot run is only a claim.
#
# The self-test builds its own synthetic host in a temporary directory, asserts
# every check fires on it and that a clean host produces zero findings, and exits
# non-zero on any failure. No host dependency beyond Docker: the tool is pure
# Python standard library and runs in a container with only detections/triage
# mounted read-only; nothing is installed on the host and nothing is written to
# the repo tree.
#
# Usage: detections/tests/triage-selftest.sh
# Env: PYTHON_IMAGE (default python:3.12-slim)
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
REPO="$(cd "$HERE/../.." && pwd)"
PYTHON_IMAGE="${PYTHON_IMAGE:-python:3.12-slim}"
docker run --rm \
-v "$REPO/detections/triage:/triage:ro" \
"$PYTHON_IMAGE" python3 /triage/artex_host_triage.py --self-test
+106
View File
@@ -0,0 +1,106 @@
# ARTEX 호스트 분류(triage)
한국어 · [English](README.md)
[`artex_host_triage.py`](artex_host_triage.py) 는 ARTEX 가 실행된 정황이 의심되는 **호스트 한 대에서 직접**
돌리는 읽기 전용 분류(triage) 스크립트입니다. 이 디렉터리의 나머지 자료는 SIEM([Sigma](../sigma/))·네트워크
센서([Suricata](../suricata/))·위협 인텔리전스 플랫폼([침해지표](../indicators/))을 운용하는 방어자를 위한
것입니다. 이 스크립트는 그와 다른 대응자, 곧 SIEM 없이 의심 호스트의 셸 앞에 서서 로컬 상태만으로 "여기서
ARTEX 가 돌았는가"를 빠르고 근거 있게 답해야 하는 사람을 위한 것입니다.
이 스크립트는 디렉터리의 다른 자료가 담은 지문을 그대로 점검하고, 여기에 더해 **[침해지표
목록](../indicators/artex_indicators.csv)이 의도적으로 Sigma 규칙 없이 둔 세 가지 호스트·DB 지표**까지
점검합니다. 그 세 지표는 로그나 네트워크로 관측되지 않아 호스트에서 직접 확인할 수밖에 없는 것들입니다
(`server-listen-port`, `recording-proxy-endpoint`, `postgres-exploration-schema`).
## 무엇을 점검하는가
모든 점검 항목은 이 저장소 소스에서 확인한 문자열이나 경로에 근거하며, 각 발견에는 대응하는 Sigma 규칙이나
침해지표 행과 동일한 한계를 함께 적습니다.
- **리슨 포트**: `:8787`(관리 UI)과 `127.0.0.1:8788`(기록 프록시)을 확인합니다. 두 값은
[`cmd/artex/main.go`](../../cmd/artex/main.go) 의 `--addr`·`--proxy` 플래그 기본값입니다. 실행 중인
호스트에서는 `ss`·`netstat`·`lsof` 출력을 파싱하고, `--ports-from` 으로 넘긴 파일에서 읽을 수도 있습니다.
- **기록 프록시 아티팩트**: 기록기가 첫 실행 때 만드는 중간자(MITM) CA 파일
`<데이터 디렉터리>/traffic/_ca/mitmproxy-ca-cert.pem` 과, 그 옆의 `_index/index.sqlite`·`_blobs/` 를
확인합니다([`traffic/traffic.go`](../../traffic/traffic.go). 데이터 디렉터리 기본값은 실행 파일 옆의
`data/` 입니다). 이 CA 는 오가는 HTTP(S) 를 복호화해 기록하는 중간자 트래픽 기록기의 신뢰 앵커입니다
(MITRE ATT&CK T1557).
- **로그 마커**: 로그 파일에서 보강 프로브의 User-Agent `artex-enrich/1.0`
([`enrich/enrich.go`](../../enrich/enrich.go)), 자체 업데이트 송신의 User-Agent `artex-selfupdate`
([`selfupdate/github.go`](../../selfupdate/github.go)), 플랫폼 가드 감사 마커
([`guard/guard.go`](../../guard/guard.go))를 찾습니다. 가드 마커는 비(非)ASCII 프레이밍까지 원문 그대로
두어 grep 이 실제로 일치하도록 했습니다. 로그 회전으로 `.gz`·`.bz2`·`.xz` 로 압축된 과거 로그도 풀어서
함께 검사하므로 호스트의 로그 이력까지 포괄합니다. 다만 파이썬 표준 라이브러리에 코덱이 없는 형식
(`.zst`·`.lz4`)은 검사하지 않고 **건너뛴 파일로 보고**합니다. 조용히 깨끗하다고 처리하지 않으니, 그런
파일은 먼저 압축을 풀거나 손으로 `grep` 해서 따로 확인하십시오.
- **PostgreSQL 탐색 스키마**: ARTEX 저장소의 이중 그래프 테이블(`exploration_nodes`·`_edges`·`_anchors` 와
`assets`·`companies`·`activity`, 그리고 `agent_prompts` 시드)을 확인합니다
([`db/schema.sql`](../../db/schema.sql)). DSN 을 주면 `psql` 로 조회하고, `psql` 이 없으면 손으로 돌릴 수
있는 읽기 전용 쿼리를 그대로 출력합니다.
- **실행 중 프로세스의 환경변수 주입**: 프록시 변수(`HTTP_PROXY`·`HTTPS_PROXY`·`ALL_PROXY`)와, mitmproxy
CA(`mitmproxy-ca-cert.pem`)를 가리키는 툴체인 CA 신뢰 변수(`SSL_CERT_FILE`·`CURL_CA_BUNDLE`·
`REQUESTS_CA_BUNDLE`·`GIT_SSL_CAINFO`·`NODE_EXTRA_CA_CERTS`)를 **함께** 지닌 프로세스를 찾습니다. ARTEX 는
생성하는 모든 worker 도구에 바로 이 변수들을 주입합니다([`agent/worker.go`](../../agent/worker.go) 의
`proxyEnv`, [`agent/proxyenv_test.go`](../../agent/proxyenv_test.go) 가 단언). **변수 이름이 소스에 하드코딩**
이라(값만 바꿀 수 있음), 이 지문은 운영자가 바이너리 이름을 바꾸거나 포트를 바꿔도 남아 리슨 포트 하나보다
특이적입니다. 실행 중인 Linux 호스트에서는 `/proc` 를 읽고, 오프라인·포렌식 이미지에서는 `--proc-from` 으로
캡처한 환경변수 덤프를 읽습니다. 프록시·CA 가 함께면 높은 심각도, mitmproxy CA 하나 또는 ARTEX 기본 프록시
엔드포인트(`127.0.0.1:8788`) 하나만 있으면 중간 심각도로 보고하되, mitmproxy CA 가 없는 회사 프록시는
단서로 올리지 않습니다.
발견은 **분류를 위한 단서이지 단정이 아닙니다.** 또한 어떤 항목도 걸리지 않았다고 해서 안전하다는 뜻은
아닙니다. 운영자는 바이너리 이름을 바꾸거나, 데이터 디렉터리를 옮기거나, 포트를 바꿀 수 있기 때문입니다.
## 플랫폼 지원
이 스크립트는 순수 Python 3(표준 라이브러리만)이라서 Python 3 이 도는 곳이면 어디서나 동작하며, Linux(CI
자가 테스트)와 macOS 에서 실제로 돌려 확인했습니다. OS 에 의존하는 점검은 두 가지이고, 둘 다 실패하지 않고
깨끗하게 축소됩니다.
- **실행 중 포트 점검**: `ss` → `netstat` → `lsof` 순서로 시도해 출력이 나오는 첫 도구를 씁니다. Linux 에서는
`ss`·`netstat` 가 쓰이고, `ss` 가 없고 `netstat` 가 Linux 식 `-ltnp` 플래그를 받지 않는 macOS·BSD(이 경우
출력 없이 종료)에서는 `lsof -nP -iTCP -sTCP:LISTEN` 로 넘어가 같은 방식으로 파싱합니다. 실시간으로 훑는
대신 저장해 둔 목록을 읽으려면 `--ports-from` 을 쓰십시오.
- **실행 중 프로세스 환경변수 점검**: `/proc` 를 읽으므로 Linux 에서만 돕니다. `/proc` 가 없는 호스트
(macOS·BSD)에서는 깨끗함이 아니라 **건너뜀**으로 보고하므로, Linux 호스트에서 덤프를 떠 `--proc-from` 으로
넘기십시오(사용법 참조).
나머지 점검(기록 프록시 아티팩트·로그 마커·PostgreSQL 스키마)은 파일시스템·로그 파일·(DSN 이 있으면)
`psql` 를 읽으므로 OS 에 무관합니다.
## 사용법
```sh
# 호스트를 처음부터 끝까지 점검합니다
detections/triage/artex_host_triage.py \
--data-dir /opt/artex/data \
--log /var/log/syslog --log-dir /var/log/artex \
--pg-dsn "$ARTEX_PG_DSN"
# 기계가 읽는 형식으로 출력하고, 하나라도 걸리면 비-0 으로 종료합니다
detections/triage/artex_host_triage.py --data-dir /opt/artex/data --json --exit-code
# 오프라인·포렌식 이미지: 캡처한 프로세스 환경변수 덤프를 읽습니다
# 호스트에서 덤프를 만드는 법:
# for p in /proc/[0-9]*; do echo "# $p"; tr '\0' '\n' < "$p/environ"; echo; done > proc_env_dump.txt
detections/triage/artex_host_triage.py --proc-from proc_env_dump.txt
# 재현 가능한 픽스처 자가 테스트(호스트 상태를 건드리지 않습니다)
detections/triage/artex_host_triage.py --self-test
```
이 스크립트는 순수 Python 3 표준 라이브러리만 씁니다. 설치가 필요 없고, 네트워크를 쓰지 않으며,
`--self-test` 가 쓰는 자체 임시 디렉터리 말고는 아무 데도 쓰지 않습니다. 호스트 상태(열린 포트, 데이터
디렉터리, 로그 파일, 그리고 DSN 을 줄 때만 데이터베이스)를 읽어 발견한 내용을 출력합니다. 종료 코드는
기본적으로 `0` 입니다(게이트가 아니라 분류 용도입니다). `--exit-code` 를 주면 지표가 하나라도 걸렸을 때 `1`
로 종료합니다.
## 어떻게 정직함을 유지하는가
`--self-test` 는 합성 호스트를 만듭니다. 심어 둔 CA·색인·블롭 저장소가 있는 데이터 디렉터리, 각 마커가 든
로그, 포트 목록, 캡처한 프로세스 환경변수 덤프를 만든 뒤, 모든 점검이 그 위에서 발화하는지 단언하고, 이어서
깨끗한 호스트·정상 로그·회사 프록시 프로세스에서는 발견이 **0** 건인지(오탐이 없는지) 단언합니다. 이 자가 테스트는 [`detections` CI
워크플로](../../.github/workflows/detections.yml)에 연결되어 있고 [`detections/tests/run-all.sh`](../tests/run-all.sh)
가 다시 돌립니다. 그래서 어떤 점검이 깨지거나, 지표가 grep 하는 소스 문자열에서 어긋나면 머지 게이트에서
실패합니다. 돌려 볼 수 없는 탐지는 주장일 뿐이라는 원칙을 따릅니다.
+113
View File
@@ -0,0 +1,113 @@
# ARTEX host triage
English · [한국어](README.ko.md)
> 한국어: [`artex_host_triage.py`](artex_host_triage.py) 는 ARTEX 가 돌았다고 의심되는 **호스트 한 대에서 직접**
> 돌리는 읽기 전용 분류(triage) 스크립트입니다. SIEM(Sigma)·네트워크 센서(Suricata)·위협 인텔리전스
> 플랫폼(지표 CSV·MISP)을 쓰는 방어자 말고, SIEM 없이 의심 호스트의 셸 앞에 선 대응자를 위한 것입니다.
> 리슨 포트·기록 프록시 아티팩트·로그 마커·PostgreSQL 스키마를 저장소 소스에 근거해 점검하고, 각 발견에
> 같은 한계(포트는 바꿀 수 있음, CA 파일명은 단독 mitmproxy 와 공유됨 등)를 함께 적습니다. 자신이 소유하거나
> 서면 허가를 받은 호스트에만 사용하십시오. 한국어 전체 문서는 **[README.ko.md](README.ko.md)** 를 보십시오.
A read-only triage helper you run **on a single suspected host** to answer "did ARTEX run here?" from
local state. The rest of this directory serves defenders who run a SIEM ([Sigma](../sigma/)), a network
sensor ([Suricata](../suricata/)), or a threat-intelligence platform (the [indicators](../indicators/)).
This script serves the other responder: the one at a host's shell, with no SIEM, who needs a quick,
defensible answer from what is on the box.
It operationalizes the same fingerprints the rest of the directory ships, **plus the three host/DB
indicators the [indicator list](../indicators/artex_indicators.csv) deliberately carries without a Sigma
rule** because they are not log- or network-observable and can only be checked on the host itself
(`server-listen-port`, `recording-proxy-endpoint`, `postgres-exploration-schema`).
## What it checks
Every check is grounded in a string or path verified in this repository's source, and every finding
carries the same honest caveat as the matching Sigma rule or indicator row.
- **Listening ports** — `:8787` (admin UI) and `127.0.0.1:8788` (recording proxy), the defaults of the
`--addr` / `--proxy` flags in [`cmd/artex/main.go`](../../cmd/artex/main.go). Parsed from `ss`/`netstat`/`lsof`
on the live host, or from a file you pass with `--ports-from`.
- **Recording-proxy artifacts** — the MITM CA the recorder writes on first start,
`<data-dir>/traffic/_ca/mitmproxy-ca-cert.pem`, and the sibling `_index/index.sqlite` and `_blobs/`
([`traffic/traffic.go`](../../traffic/traffic.go); the data directory default is `data/` next to the binary).
The CA is the trust anchor of an adversary-in-the-middle traffic recorder (ATT&CK T1557).
- **Log markers** — the enrichment prober UA `artex-enrich/1.0` ([`enrich/enrich.go`](../../enrich/enrich.go)),
the self-update egress UA `artex-selfupdate` ([`selfupdate/github.go`](../../selfupdate/github.go)), and
the platform-guard audit marker ([`guard/guard.go`](../../guard/guard.go); kept verbatim, including the
non-ASCII framing, so the grep matches) in the log file(s) you point it at. Rotated logs compressed as
`.gz`/`.bz2`/`.xz` are decompressed and scanned too, so the host's log history is covered; a format with
no standard-library codec (`.zst`/`.lz4`) is reported as **skipped** rather than silently treated as
clean — decompress it first or `grep` it by hand.
- **PostgreSQL exploration schema** — the dual-graph tables (`exploration_nodes`/`_edges`/`_anchors` with
`assets`/`companies`/`activity` and the `agent_prompts` seed) in the ARTEX store
([`db/schema.sql`](../../db/schema.sql)). Run against a DSN with `psql` if available; otherwise the
script prints the exact read-only query for you to run by hand.
- **Process env injection** — a running process whose environment carries a proxy var (`HTTP_PROXY` /
`HTTPS_PROXY` / `ALL_PROXY`) **together with** a toolchain CA-trust var (`SSL_CERT_FILE` /
`CURL_CA_BUNDLE` / `REQUESTS_CA_BUNDLE` / `GIT_SSL_CAINFO` / `NODE_EXTRA_CA_CERTS`) pointing at a
`mitmproxy-ca-cert.pem`. ARTEX injects exactly these into every worker tool it spawns
([`agent/worker.go`](../../agent/worker.go) `proxyEnv`, asserted by
[`agent/proxyenv_test.go`](../../agent/proxyenv_test.go)). The variable **names are hard-coded** in the
source (only the values are configurable), so this tell survives an operator renaming the binary or
changing the ports — a stronger signal than the bare listen port. Read from `/proc` on the live Linux
host, or from a captured dump with `--proc-from`. A proxy and a mitmproxy CA together are reported high;
a mitmproxy CA alone, or the ARTEX default proxy endpoint (`127.0.0.1:8788`) alone, is medium; a
corporate proxy with no mitmproxy CA is deliberately not flagged.
A hit is a **triage lead, not an attribution**, and the absence of every finding is **not** a clean bill
of health: an operator can rename the binary, move the data directory, or change the ports.
## Platform support
The script is pure Python 3 (standard library only), so it runs wherever Python 3 does — verified on
Linux (the CI self-test) and macOS. Two checks are OS-specific, and both degrade cleanly rather than
failing:
- **Live port scan** — tries `ss`, then `netstat`, then `lsof`, and uses the first that produces output.
On Linux that is `ss`/`netstat`; on macOS/BSD, where `ss` is absent and `netstat` does not take the
Linux `-ltnp` flags (it exits with empty output), it falls through to `lsof -nP -iTCP -sTCP:LISTEN`,
parsed the same way. Pass `--ports-from` to read a saved listing instead of scanning live.
- **Live process-env scan** — reads `/proc`, so it runs only on Linux. On a host without `/proc`
(macOS/BSD) it is reported as **skipped**, not clean; capture a dump on the Linux host and pass it with
`--proc-from` (see Usage).
The remaining checks — recording-proxy artifacts, log markers, and the PostgreSQL schema — read the
filesystem, log files, and (with a DSN) `psql`, so they are OS-independent.
## Usage
```sh
# check a host end to end
detections/triage/artex_host_triage.py \
--data-dir /opt/artex/data \
--log /var/log/syslog --log-dir /var/log/artex \
--pg-dsn "$ARTEX_PG_DSN"
# machine-readable findings, and exit non-zero if anything fired
detections/triage/artex_host_triage.py --data-dir /opt/artex/data --json --exit-code
# offline / forensic image: read a captured process-environment dump
# make the dump on the host with:
# for p in /proc/[0-9]*; do echo "# $p"; tr '\0' '\n' < "$p/environ"; echo; done > proc_env_dump.txt
detections/triage/artex_host_triage.py --proc-from proc_env_dump.txt
# reproducible fixture test (no host state touched)
detections/triage/artex_host_triage.py --self-test
```
The script is pure Python 3 standard library: no install, no network, and it writes nothing anywhere
except the `--self-test`'s own temporary directory. It reads host state (open ports, a data directory,
log files, and — only if you pass a DSN — the database) and prints what it found. Exit code is `0` by
default (triage, not a gate); pass `--exit-code` to make it `1` when any indicator fired.
## How this stays honest
The `--self-test` builds a synthetic host — a data directory with a planted CA, index, and blob store; a
log containing each marker; a port listing; and a captured process-environment dump — and asserts every
check fires on it, then asserts a clean host, a benign log, and a corporate-proxy process produce **zero**
findings (no false positives). It is wired into the
[`detections` CI workflow](../../.github/workflows/detections.yml) and re-run by
[`detections/tests/run-all.sh`](../tests/run-all.sh), so a change that breaks a check, or that drifts an
indicator away from the source string it greps for, fails the merge gate. A detection you cannot run is
only a claim.
+911
View File
@@ -0,0 +1,911 @@
#!/usr/bin/env python3
#
# ARTEX host triage — a read-only responder helper for a suspected ARTEX host.
#
# The rest of detections/ serves defenders who run a SIEM (Sigma), a network
# sensor (Suricata), or a threat-intel platform (the MISP / CSV indicators). This
# script serves the other responder: the one standing at a single suspect host's
# shell, with no SIEM, who needs to answer "did ARTEX run here?" from local state.
# It operationalizes the same indicators the rest of the directory ships, plus the
# three host/DB indicators the indicator list deliberately carries WITHOUT a Sigma
# rule because they are not log- or network-observable and can only be checked on
# the box itself (see detections/indicators/artex_indicators.csv — the rows whose
# `rule` column is empty: server-listen-port, recording-proxy-endpoint,
# postgres-exploration-schema).
#
# It is a TRIAGE LEAD generator, not an alerting rule. Every check is grounded in
# a string or path verified in this repository's source, and every finding carries
# the same honest caveat the matching Sigma rule or indicator row carries: ports
# are configurable, the MITM CA filename is shared with standalone mitmproxy, the
# guard marker also appears in logs that merely quote this guide. A hit is a reason
# to look closer, never an attribution on its own, and the absence of every finding
# is NOT a clean bill of health — an operator can rename the binary, move the data
# directory, or change the ports.
#
# What it checks (each cites the source it is grounded in):
# 1. Listening ports :8787 (admin UI) and 127.0.0.1:8788 (recording proxy)
# — defaults of the --addr / --proxy flags in
# cmd/artex/main.go. Parsed from `ss`/`netstat`/`lsof`
# on the live host, or from --ports-from FILE.
# 2. Recording-proxy MITM <data-dir>/traffic/_ca/mitmproxy-ca-cert.pem and the
# CA + stores sibling _index/index.sqlite and _blobs/ the recorder
# writes on first start (traffic/traffic.go; the data
# dir default is data/ next to the binary — see
# cmd/artex/main.go). The CA is the trust anchor of an
# adversary-in-the-middle traffic recorder (ATT&CK
# T1557).
# 3. Log markers the enrichment prober UA `artex-enrich/1.0`
# (enrich/enrich.go), the self-update egress UA
# `artex-selfupdate` (selfupdate/github.go), and the
# platform-guard audit marker (guard/guard.go) in the
# log file(s) you point it at. Rotated logs that
# logrotate compressed as .gz/.bz2/.xz are read
# through their standard-library codec so their
# history is scanned too; a format with no stdlib
# codec (.zst/.lz4) is reported as skipped, never
# silently treated as clean.
# 4. PostgreSQL schema the dual-graph exploration tables (exploration_nodes /
# _edges / _anchors with assets / companies / activity
# and the agent_prompts seed) in the ARTEX store
# (db/schema.sql). Run against a DSN with `psql` if
# available; otherwise the script prints the exact
# read-only query for you to run by hand.
# 5. Process env injection a running process whose environment carries the
# recording proxy (HTTP_PROXY / HTTPS_PROXY / ALL_PROXY)
# together with a toolchain CA-trust var (SSL_CERT_FILE /
# CURL_CA_BUNDLE / REQUESTS_CA_BUNDLE / GIT_SSL_CAINFO /
# NODE_EXTRA_CA_CERTS) pointing at a mitmproxy-ca-cert.pem.
# ARTEX injects exactly these into every worker tool it
# spawns (agent/worker.go proxyEnv, asserted by
# agent/proxyenv_test.go). The variable NAMES are
# hard-coded in the source, so this tell survives an
# operator renaming the binary or changing the ports —
# a stronger signal than the bare listen port. Read from
# /proc on the live Linux host, or from --proc-from FILE.
#
# Safety: pure Python standard library, no network, no writes anywhere except the
# self-test's own temporary directory. It reads host state (open ports, a data
# directory, log files, the environments of running processes via /proc, and — only
# if you pass a DSN — the database) and prints what it found. Use it only on a host
# you own or are authorized in writing to inspect.
#
# Usage:
# detections/triage/artex_host_triage.py --data-dir /opt/artex/data \
# --log /var/log/syslog --log-dir /var/log/artex
# detections/triage/artex_host_triage.py --pg-dsn "$ARTEX_PG_DSN"
# detections/triage/artex_host_triage.py --proc-from proc_env_dump.txt # offline
# detections/triage/artex_host_triage.py --self-test # reproducible fixture test
# detections/triage/artex_host_triage.py --json # machine-readable findings
#
# Exit code: 0 by default (triage, not a gate). With --exit-code, exits 1 if any
# finding fired. --self-test exits non-zero on any self-test failure.
import argparse
import bz2
import gzip
import json
import lzma
import os
import re
import shutil
import subprocess
import sys
import tempfile
# --- grounded constants (every value is verified in this repository's source) ---
# Default listen / recording-proxy ports (cmd/artex/main.go --addr / --proxy).
SERVER_PORT = 8787
PROXY_HOST = "127.0.0.1"
PROXY_PORT = 8788
# Recording-proxy artifacts under <data-dir>/traffic/ (traffic/traffic.go;
# server/manager.go opens traffic.Open(filepath.Join(dir, "traffic"), ...)).
TRAFFIC_SUBDIR = "traffic"
CA_RELPATH = os.path.join("_ca", "mitmproxy-ca-cert.pem")
INDEX_RELPATH = os.path.join("_index", "index.sqlite")
BLOBS_RELDIR = "_blobs"
# Log markers. The guard marker is the original (untranslated) framing string the
# platform guard writes to the audit log on a blocked tool call (guard/guard.go);
# it is kept verbatim here because that is the exact byte sequence a responder
# greps for, and translating it would stop the match.
LOG_MARKERS = [
{
"value": "artex-enrich/1.0",
"title": "enrichment prober User-Agent",
"source": "enrich/enrich.go",
"severity": "high",
"caveat": "An operator can change the User-Agent; absence is not safety.",
},
{
"value": "artex-selfupdate",
"title": "self-update egress User-Agent",
"source": "selfupdate/github.go",
"severity": "medium",
"caveat": "Seen in outbound logs from a host running ARTEX; the string is configurable.",
},
{
"value": "【ARTEX 平台管控·非目标防御】",
"title": "platform-guard audit-log framing marker",
"source": "guard/guard.go",
"severity": "high",
"caveat": "Also appears in logs that merely quote this defense guide or the ARTEX source.",
},
]
# Dual-graph exploration schema fingerprint (db/schema.sql). Their presence
# together is the host-forensic tell; any one table name is generic.
SCHEMA_TABLES = [
"exploration_nodes",
"exploration_edges",
"exploration_anchors",
"assets",
"companies",
"activity",
"agent_prompts",
]
# Recording-proxy environment injection into spawned worker tools (agent/worker.go
# proxyEnv; asserted by agent/proxyenv_test.go). ARTEX routes every worker tool's
# traffic through the recording MITM proxy and, when a CA is present, makes the
# toolchain trust it — by setting these exact variables in the subprocess env. The
# variable NAMES are hard-coded in worker.go (only the values are configurable), so
# a tool process carrying a recording-proxy address in a proxy var AND a CA var
# pointing at a mitmproxy-ca-cert.pem is a far more specific tell than the bare
# listen port: it survives the operator renaming the binary or moving the data dir.
PROXY_ENV_VARS = (
"HTTP_PROXY", "HTTPS_PROXY", "http_proxy", "https_proxy", "ALL_PROXY", "all_proxy",
)
CA_ENV_VARS = (
"SSL_CERT_FILE", "CURL_CA_BUNDLE", "REQUESTS_CA_BUNDLE", "GIT_SSL_CAINFO", "NODE_EXTRA_CA_CERTS",
)
CA_BASENAME = "mitmproxy-ca-cert.pem" # basename of CA_RELPATH; the value a CA var points at
# The recording-proxy default endpoint (cmd/artex/main.go --proxy). An operator can
# point --proxy elsewhere, so the CA var is the anchor and this is only the fallback.
PROXY_DEFAULT_ENDPOINT = f"{PROXY_HOST}:{PROXY_PORT}" # 127.0.0.1:8788
class Finding:
def __init__(self, check, severity, title, detail, source, caveat):
self.check = check
self.severity = severity
self.title = title
self.detail = detail
self.source = source
self.caveat = caveat
def as_dict(self):
return {
"check": self.check,
"severity": self.severity,
"title": self.title,
"detail": self.detail,
"source": self.source,
"caveat": self.caveat,
}
# --- check 1: listening ports -------------------------------------------------
# Parse a port-listing produced by `ss -ltnp`, `netstat -ltnp`, or
# `lsof -nP -iTCP -sTCP:LISTEN`. Kept a pure function of its text input so the
# self-test can feed synthetic output without opening a real socket. Returns the
# set of (host, port) LISTEN endpoints it can parse out of any of those formats.
LISTEN_RE = re.compile(
r"(?P<host>\[?[0-9a-fA-F:.*]+\]?):(?P<port>\d{1,5})\b"
)
def parse_listen_endpoints(listing):
endpoints = set()
for line in listing.splitlines():
low = line.lower()
# ss/netstat lines for listeners contain the LISTEN state; lsof lines
# contain "(LISTEN)". Skip anything that is not a listening socket so a
# connected session to :8787 elsewhere is not misread as a local listener.
if "listen" not in low:
continue
for m in LISTEN_RE.finditer(line):
host = m.group("host").strip("[]")
try:
port = int(m.group("port"))
except ValueError:
continue
if 0 < port < 65536:
endpoints.add((host, port))
return endpoints
def gather_listen_listing():
"""Run the first available port tool; return its stdout, or '' if none work."""
for cmd in (
["ss", "-ltnp"],
["netstat", "-ltnp"],
["lsof", "-nP", "-iTCP", "-sTCP:LISTEN"],
):
if shutil.which(cmd[0]) is None:
continue
try:
out = subprocess.run(
cmd, capture_output=True, text=True, timeout=15, check=False
)
except (OSError, subprocess.SubprocessError):
continue
if out.stdout:
return out.stdout
return ""
def check_listening_ports(listing):
findings = []
endpoints = parse_listen_endpoints(listing)
for host, port in sorted(endpoints):
if port == SERVER_PORT:
findings.append(
Finding(
"listening-port",
"medium",
"ARTEX default admin-UI port is listening",
f"a process is listening on {host}:{port} (ARTEX --addr default :{SERVER_PORT})",
"cmd/artex/main.go",
"The port is configurable; confirm the process with `ss -ltnp` / `lsof`.",
)
)
if port == PROXY_PORT and (host == PROXY_HOST or host in ("*", "0.0.0.0", "::")):
findings.append(
Finding(
"listening-port",
"high",
"ARTEX recording-proxy loopback port is listening",
f"a process is listening on {host}:{port} (ARTEX --proxy default {PROXY_HOST}:{PROXY_PORT})",
"cmd/artex/main.go",
"Loopback-only and configurable; correlate with the MITM CA file under traffic/_ca/.",
)
)
return findings
# --- check 2: recording-proxy artifacts --------------------------------------
def check_recording_proxy_artifacts(data_dir):
findings = []
traffic = os.path.join(data_dir, TRAFFIC_SUBDIR)
ca = os.path.join(traffic, CA_RELPATH)
if os.path.isfile(ca):
findings.append(
Finding(
"recording-proxy-ca",
"medium",
"ARTEX recording-proxy MITM CA certificate present",
f"found {ca}",
"traffic/traffic.go",
"A bare mitmproxy-ca-cert.pem is shared with standalone mitmproxy; "
"the traffic/_ca/ layout narrows it to ARTEX.",
)
)
index = os.path.join(traffic, INDEX_RELPATH)
if os.path.isfile(index):
findings.append(
Finding(
"recording-proxy-index",
"medium",
"ARTEX recording-proxy traffic index store present",
f"found {index}",
"traffic/traffic.go",
"The recorder's SQLite index of captured HTTP(S) exchanges; a forensic artifact of a run.",
)
)
blobs = os.path.join(traffic, BLOBS_RELDIR)
if os.path.isdir(blobs):
findings.append(
Finding(
"recording-proxy-blobs",
"low",
"ARTEX recording-proxy body blob store present",
f"found {blobs}/",
"traffic/traffic.go",
"Spilled response bodies from the traffic recorder; corroborates the index/CA.",
)
)
return findings
# --- check 3: log markers -----------------------------------------------------
# logrotate (and journald) compress rotated logs. gzip is the historical default;
# bzip2 and xz show up when configured. Open those through their standard-library
# codec so the markers inside a rotated file are scanned too — a bare text open()
# would read the compressed bytes as UTF-8 and silently miss every marker in the
# host's log history, exactly the kind of "absence is not safety" gap this tool
# warns about. Formats with no stdlib codec (zstd, lz4) cannot be read here; they
# are reported as skipped so the responder decompresses them by hand rather than
# mistaking an unscanned file for a clean one.
STDLIB_LOG_OPENERS = {
".gz": gzip.open,
".bz2": bz2.open,
".xz": lzma.open,
".lzma": lzma.open,
}
UNSUPPORTED_COMPRESSED_EXTS = {".zst", ".zstd", ".lz4", ".lz", ".zip", ".7z", ".br"}
def open_log_stream(path):
"""Return a UTF-8 text stream for a log file, transparently decompressing a
gzip/bzip2/xz rotated log by extension. The caller uses it as a context
manager. Plaintext and anything unrecognized fall through to a plain open."""
opener = STDLIB_LOG_OPENERS.get(os.path.splitext(path)[1].lower())
if opener is not None:
return opener(path, "rt", encoding="utf-8", errors="replace")
return open(path, "r", encoding="utf-8", errors="replace")
def iter_log_files(logs, log_dirs):
seen = set()
for p in logs:
if os.path.isfile(p) and p not in seen:
seen.add(p)
yield p
for d in log_dirs:
if not os.path.isdir(d):
continue
for root, _dirs, files in os.walk(d):
for name in sorted(files):
p = os.path.join(root, name)
if p not in seen:
seen.add(p)
yield p
def scan_logs(logs, log_dirs):
"""Scan the log files/dirs for ARTEX markers. Returns (findings, skipped),
where skipped lists paths in a compressed format with no stdlib codec
(e.g. .zst/.lz4) that could not be read and so were NOT scanned."""
findings = []
skipped = []
for path in iter_log_files(logs, log_dirs):
if os.path.splitext(path)[1].lower() in UNSUPPORTED_COMPRESSED_EXTS:
# No stdlib codec: do not read gibberish and do not pretend it is
# clean — record it so run_checks can tell the responder to grep it
# by hand (zstdcat / lz4cat).
skipped.append(path)
continue
# Stream line by line instead of f.read(): the sanctioned log targets are
# whole syslogs (--log /var/log/syslog) that can be hundreds of MB, and all
# three markers live within a single line, so a line at a time keeps memory
# bounded to one line while matching exactly what a full read would.
# open_log_stream transparently decompresses a .gz/.bz2/.xz rotated log so
# its history is scanned too. Report each marker at most once per file (the
# full-read "value in text" did too), and stop early once all have fired.
fired = set()
try:
with open_log_stream(path) as f:
for line in f:
for marker in LOG_MARKERS:
if marker["value"] in fired:
continue
if marker["value"] in line:
fired.add(marker["value"])
findings.append(
Finding(
"log-marker",
marker["severity"],
f"ARTEX {marker['title']} in log",
f"{path!r} contains {marker['value']!r}",
marker["source"],
marker["caveat"],
)
)
if len(fired) == len(LOG_MARKERS):
break
except (OSError, EOFError, lzma.LZMAError):
# Unreadable or a corrupt/mislabeled compressed file (gzip.BadGzipFile
# and bz2 errors are OSError subclasses; lzma raises LZMAError). Skip it
# the same way the plain-read path always skipped an unreadable file.
continue
return findings, skipped
# --- check 4: PostgreSQL exploration schema -----------------------------------
# A single read-only query: how many of the dual-graph tables exist in the public
# schema. Printed for manual use when psql is unavailable or no DSN was given.
SCHEMA_QUERY = (
"SELECT count(*) FROM information_schema.tables "
"WHERE table_schema='public' AND table_name IN ("
+ ", ".join(f"'{t}'" for t in SCHEMA_TABLES)
+ ");"
)
def check_pg_schema(dsn):
findings = []
if not dsn:
return findings, (
"PostgreSQL schema check skipped (no --pg-dsn / ARTEX_PG_DSN). "
"To check by hand, run this read-only query against the suspected store:\n"
f" psql <DSN> -c \"{SCHEMA_QUERY}\"\n"
f" (a count at or near {len(SCHEMA_TABLES)} of these tables together is the dual-graph tell; db/schema.sql)"
)
if shutil.which("psql") is None:
return findings, (
"PostgreSQL schema check skipped (psql not found on PATH). "
f"Run by hand:\n psql <DSN> -c \"{SCHEMA_QUERY}\""
)
try:
out = subprocess.run(
["psql", dsn, "-tAc", SCHEMA_QUERY],
capture_output=True,
text=True,
timeout=30,
check=False,
)
except (OSError, subprocess.SubprocessError) as e:
return findings, f"PostgreSQL schema check could not run: {e}"
if out.returncode != 0:
return findings, (
"PostgreSQL schema check could not connect: "
+ (out.stderr.strip().splitlines()[-1] if out.stderr.strip() else "psql returned non-zero")
)
count = out.stdout.strip()
try:
n = int(count)
except ValueError:
return findings, f"PostgreSQL schema check returned an unexpected result: {count!r}"
if n >= 4:
findings.append(
Finding(
"postgres-schema",
"high" if n >= 6 else "medium",
"ARTEX dual-graph exploration schema present",
f"{n} of {len(SCHEMA_TABLES)} ARTEX exploration-graph tables found in the public schema",
"db/schema.sql",
"Inspect the database to confirm; a few table names overlap generic apps, the set does not.",
)
)
return findings, f"PostgreSQL schema check: {n} of {len(SCHEMA_TABLES)} ARTEX tables present."
# --- check 5: recording-proxy env injection in running processes --------------
def _proxy_env_hit(env):
"""First proxy var set to a non-empty value, as (var, value), else None."""
for var in PROXY_ENV_VARS:
val = env.get(var, "").strip()
if val:
return var, val
return None
def _ca_env_hit(env):
"""First CA-trust var pointing at a mitmproxy-ca-cert.pem, as (var, value), else None."""
for var in CA_ENV_VARS:
val = env.get(var, "").strip()
if val and os.path.basename(val) == CA_BASENAME:
return var, val
return None
def scan_process_env(label, env):
"""Findings for one process's environment dict. Pure function of its input so
the self-test can feed synthetic env without reading /proc. The strongest tell
is a proxy var AND a mitmproxy CA var together (the worker proxyEnv signature);
a mitmproxy CA alone, or the ARTEX default proxy endpoint alone, is a weaker
lead. A corporate proxy with no mitmproxy CA is deliberately not flagged."""
proxy = _proxy_env_hit(env)
ca = _ca_env_hit(env)
if ca and proxy:
pv, pval = proxy
cv, cval = ca
return [Finding(
"process-env-injection", "high",
"ARTEX recording-proxy env injection in a running process",
f"{label}: {pv}={pval} with {cv}={cval} — the worker proxyEnv signature "
"(routes through a proxy and trusts a mitmproxy CA)",
"agent/worker.go",
"A standalone mitmproxy or a MITM test harness can set these too; a proxy "
"together with a trusted mitmproxy-ca-cert.pem matches ARTEX's worker "
"injection. Capture can be disabled (--proxy ''), so absence is not safety.",
)]
if ca:
cv, cval = ca
return [Finding(
"process-env-injection", "medium",
"A running process is told to trust a mitmproxy CA",
f"{label}: {cv}={cval} points a toolchain CA-trust var at a mitmproxy-ca-cert.pem",
"agent/worker.go",
"The recording proxy injects this CA path into worker tools; the bare "
"filename is shared with standalone mitmproxy, so correlate with "
"traffic/_ca/ and the proxy port.",
)]
if proxy and PROXY_DEFAULT_ENDPOINT in proxy[1]:
pv, pval = proxy
return [Finding(
"process-env-injection", "medium",
"A running process routes through the ARTEX recording-proxy default endpoint",
f"{label}: {pv}={pval} (ARTEX --proxy default {PROXY_DEFAULT_ENDPOINT})",
"agent/worker.go",
"The endpoint is the --proxy default and is configurable; correlate with "
"the MITM CA under traffic/_ca/.",
)]
return []
def _parse_environ_bytes(raw):
"""Parse a NUL-separated /proc/<pid>/environ blob into a KEY->VALUE dict."""
env = {}
for tok in raw.split(b"\x00"):
if not tok:
continue
s = tok.decode("utf-8", "replace")
if "=" in s:
k, v = s.split("=", 1)
env[k] = v
return env
def parse_proc_dump(text):
r"""Parse a captured process-environment dump into [(label, env), ...]. Blocks
are separated by a blank line; a line starting with '#' sets the block label;
other entries are KEY=VALUE (NUL or newline separated). Produce such a dump on
the host with:
for p in /proc/[0-9]*; do echo "# $p"; tr '\0' '\n' < "$p/environ"; echo; done
"""
procs = []
for block in re.split(r"\n[ \t]*\n", text.replace("\x00", "\n")):
label = None
env = {}
for line in block.splitlines():
if not line.strip():
continue
if line.lstrip().startswith("#"):
label = line.lstrip()[1:].strip() or label
continue
if "=" in line:
k, v = line.split("=", 1)
env[k.strip()] = v
if env:
procs.append((label or "process", env))
return procs
def gather_process_envs():
"""(procs, note, unreadable): read each /proc/<pid>/environ on the live Linux
host. note is non-empty when /proc is unavailable (non-Linux) or some environs
were unreadable, so the caller never mistakes 'did not run' for 'clean'."""
if not sys.platform.startswith("linux") or not os.path.isdir("/proc"):
return [], (
"process-env check skipped (no /proc on this OS; run on the Linux host, "
"or pass --proc-from a captured env dump)."
), 0
procs = []
unreadable = 0
mypid = str(os.getpid())
try:
pids = os.listdir("/proc")
except OSError as e:
return [], f"process-env check could not list /proc: {e}", 0
for pid in pids:
if not pid.isdigit() or pid == mypid:
continue
try:
with open(os.path.join("/proc", pid, "environ"), "rb") as f:
raw = f.read()
except OSError:
unreadable += 1
continue
env = _parse_environ_bytes(raw)
if not env:
continue
comm = pid
try:
with open(os.path.join("/proc", pid, "comm"), "r", encoding="utf-8", errors="replace") as f:
comm = f.read().strip() or pid
except OSError:
pass
procs.append((f"pid {pid} ({comm})", env))
note = ""
if unreadable:
note = (
f"process-env check: {unreadable} process(es) had an unreadable "
"/proc/<pid>/environ — run as root to cover every process; absence is not safety."
)
return procs, note, unreadable
# --- reporting ----------------------------------------------------------------
SEVERITY_ORDER = {"high": 0, "medium": 1, "low": 2}
def run_checks(args):
findings = []
notes = []
if args.ports_from:
try:
with open(args.ports_from, "r", encoding="utf-8", errors="replace") as f:
listing = f.read()
except OSError as e:
listing = ""
notes.append(f"could not read --ports-from {args.ports_from}: {e}")
else:
listing = gather_listen_listing()
if not listing:
notes.append(
"listening-port check skipped (no ss/netstat/lsof output; "
"run on the host as a user that can see listeners, or pass --ports-from)."
)
findings += check_listening_ports(listing)
if args.data_dir:
for d in args.data_dir:
findings += check_recording_proxy_artifacts(d)
else:
notes.append(
"recording-proxy artifact check skipped (no --data-dir; "
"ARTEX's default is data/ next to the binary — cmd/artex/main.go)."
)
if args.log or args.log_dir:
log_findings, skipped = scan_logs(args.log, args.log_dir)
findings += log_findings
if skipped:
notes.append(
f"log-marker check could not read {len(skipped)} compressed log "
"file(s) with no standard-library codec (e.g. .zst/.lz4), so their "
"history was NOT scanned — decompress them first or grep them by "
"hand (e.g. `zstdcat FILE | grep -F artex-`): "
+ ", ".join(sorted(skipped))
)
else:
notes.append("log-marker check skipped (no --log / --log-dir).")
pg_findings, pg_note = check_pg_schema(args.pg_dsn or os.environ.get("ARTEX_PG_DSN"))
findings += pg_findings
if pg_note:
notes.append(pg_note)
if args.proc_from:
try:
with open(args.proc_from, "r", encoding="utf-8", errors="replace") as f:
dump = f.read()
except OSError as e:
dump = ""
notes.append(f"could not read --proc-from {args.proc_from}: {e}")
procs = parse_proc_dump(dump)
if not procs and dump.strip():
notes.append(
"process-env check: --proc-from file parsed no process blocks "
"(expected '# label' + KEY=VALUE lines, blocks split by a blank line)."
)
for label, env in procs:
findings += scan_process_env(label, env)
else:
procs, proc_note, _unreadable = gather_process_envs()
for label, env in procs:
findings += scan_process_env(label, env)
if proc_note:
notes.append(proc_note)
findings.sort(key=lambda f: (SEVERITY_ORDER.get(f.severity, 9), f.check, f.title))
return findings, notes
def print_report(findings, notes):
print("ARTEX host triage — read-only; findings are triage leads, not attribution.")
print("Use only on a host you own or are authorized in writing to inspect.\n")
if findings:
print(f"{len(findings)} indicator(s) fired:\n")
for f in findings:
print(f" [{f.severity.upper():6}] {f.title}")
print(f" {f.detail}")
print(f" grounded in: {f.source}")
print(f" caveat: {f.caveat}\n")
else:
print("No ARTEX indicators fired in the checks that ran.")
print("This is NOT a clean bill of health: an operator can rename the binary,")
print("move the data directory, or change the ports. Absence is not safety.\n")
if notes:
print("Notes:")
for n in notes:
print(" - " + n.replace("\n", "\n "))
# --- self-test ----------------------------------------------------------------
def self_test():
failures = []
def check(name, cond):
print(f" {'PASS' if cond else 'FAIL'} {name}")
if not cond:
failures.append(name)
# 1. parse_listen_endpoints across ss / netstat / lsof shapes.
ss_out = (
"State Recv-Q Send-Q Local Address:Port Peer Address:Port Process\n"
"LISTEN 0 4096 *:8787 *:* users:((\"artex\"))\n"
"LISTEN 0 4096 127.0.0.1:8788 0.0.0.0:* users:((\"artex\"))\n"
"LISTEN 0 128 127.0.0.1:5432 0.0.0.0:* users:((\"postgres\"))\n"
)
eps = parse_listen_endpoints(ss_out)
check("ports: ss output parses :8787 and 127.0.0.1:8788", ("*", 8787) in eps and ("127.0.0.1", 8788) in eps)
pf = check_listening_ports(ss_out)
checks_hit = {f.title for f in pf}
check("ports: both ARTEX listeners reported", len(pf) == 2)
check("ports: admin-UI listener reported", any("admin-UI" in t for t in checks_hit))
check("ports: recording-proxy listener reported", any("recording-proxy" in t for t in checks_hit))
lsof_out = "artex 42 root 7u IPv4 TCP 127.0.0.1:8788 (LISTEN)\n"
check("ports: lsof shape parses the proxy listener", ("127.0.0.1", 8788) in parse_listen_endpoints(lsof_out))
# a connected (non-LISTEN) session to :8787 must not be read as a local listener
estab = "ESTAB 0 0 10.0.0.5:51000 93.184.216.34:8787\n"
check("ports: a non-LISTEN session to :8787 is ignored", len(check_listening_ports(estab)) == 0)
with tempfile.TemporaryDirectory() as tmp:
# 2. recording-proxy artifacts under <data>/traffic/
data = os.path.join(tmp, "data")
traffic = os.path.join(data, TRAFFIC_SUBDIR)
os.makedirs(os.path.join(traffic, "_ca"))
os.makedirs(os.path.join(traffic, "_index"))
os.makedirs(os.path.join(traffic, "_blobs"))
open(os.path.join(traffic, CA_RELPATH), "w").close()
open(os.path.join(traffic, INDEX_RELPATH), "w").close()
af = check_recording_proxy_artifacts(data)
kinds = {f.check for f in af}
check("ca: CA + index + blobs all reported", kinds == {"recording-proxy-ca", "recording-proxy-index", "recording-proxy-blobs"})
# 3. log markers
logpath = os.path.join(tmp, "app.log")
with open(logpath, "w", encoding="utf-8") as f:
f.write("GET / HTTP/1.1 artex-enrich/1.0\n")
f.write("outbound artex-selfupdate to release host\n")
f.write("blocked: " + LOG_MARKERS[2]["value"] + " this operation is denied\n")
f.write("a normal line with no markers\n")
lf, _ = scan_logs([logpath], [])
check("logs: all three markers fire", len(lf) == 3)
# 3b. rotated (compressed) logs under a --log-dir are scanned too, not
# silently skipped. A responder pointing at /var/log/artex expects the
# rotated history to be covered; a plain read of the compressed bytes
# would miss every marker inside. Plant one marker per container: a
# plaintext current log, a .gz, a .bz2, and an .xz rotation.
rot = os.path.join(tmp, "rotated")
os.makedirs(rot)
with open(os.path.join(rot, "artex.log"), "w", encoding="utf-8") as f:
f.write("outbound artex-selfupdate to release host\n")
with gzip.open(os.path.join(rot, "artex.log.1.gz"), "wt", encoding="utf-8") as f:
f.write("GET / HTTP/1.1 artex-enrich/1.0\n")
with bz2.open(os.path.join(rot, "artex.log.2.bz2"), "wt", encoding="utf-8") as f:
f.write("blocked: " + LOG_MARKERS[2]["value"] + " this operation is denied\n")
with lzma.open(os.path.join(rot, "artex.log.3.xz"), "wt", encoding="utf-8") as f:
f.write("another GET / artex-enrich/1.0 probe\n")
rf, rskip = scan_logs([], [rot])
rtitles = [f.title for f in rf]
check("rotated: plaintext + .gz + .bz2 + .xz markers all fire via --log-dir", len(rf) == 4)
check("rotated: the .gz/.xz enrichment markers (missed by a plain read) are found",
sum("enrichment prober" in t for t in rtitles) == 2)
check("rotated: nothing is reported as skipped when every file has a stdlib codec", rskip == [])
# 3c. a format with no stdlib codec (.zst) is reported as skipped, never
# silently treated as clean.
zstpath = os.path.join(rot, "artex.log.4.zst")
with open(zstpath, "wb") as f:
f.write(b"\x28\xb5\x2f\xfd and bytes a plain read would mis-handle")
_rf2, rskip2 = scan_logs([], [rot])
check("rotated: a .zst log (no stdlib codec) is reported as skipped", zstpath in rskip2)
# 4. clean host: nothing fires, no false positives
clean = os.path.join(tmp, "clean")
os.makedirs(clean)
cleanlog = os.path.join(tmp, "clean.log")
with open(cleanlog, "w", encoding="utf-8") as f:
f.write("nothing to see here\nGET /health 200\n")
clean_lf, clean_skip = scan_logs([cleanlog], [])
check("clean: no artifact findings on an empty data dir", len(check_recording_proxy_artifacts(clean)) == 0)
check("clean: no log findings on a benign log", len(clean_lf) == 0 and clean_skip == [])
check("clean: no port findings on empty listing", len(check_listening_ports("")) == 0)
# 5. pg schema: skipped path returns a manual-query note, no finding
pgf, pgnote = check_pg_schema("")
check("pg: no DSN yields a manual-query note and no finding", len(pgf) == 0 and "psql" in pgnote and SCHEMA_TABLES[0] in SCHEMA_QUERY)
# 6. recording-proxy env injection in running processes (agent/worker.go proxyEnv)
dump = (
"# pid 101 (curl)\n"
"PATH=/usr/bin\n"
"HTTP_PROXY=127.0.0.1:8788\n"
"HTTPS_PROXY=127.0.0.1:8788\n"
"REQUESTS_CA_BUNDLE=/opt/artex/data/traffic/_ca/mitmproxy-ca-cert.pem\n"
"\n"
"# pid 202 (nginx)\n"
"PATH=/usr/sbin\n"
"HOME=/var/www\n"
"\n"
"# pid 303 (apt)\n"
"HTTP_PROXY=http://corp-proxy.local:3128\n"
"\n"
"# pid 404 (python)\n"
"REQUESTS_CA_BUNDLE=/opt/artex/data/traffic/_ca/mitmproxy-ca-cert.pem\n"
"\n"
"# pid 505 (wget)\n"
"https_proxy=127.0.0.1:8788\n"
)
procs = parse_proc_dump(dump)
check("procenv: dump parses five process blocks", len(procs) == 5)
by_label = {label: env for label, env in procs}
inj = scan_process_env("pid 101 (curl)", by_label.get("pid 101 (curl)", {}))
check("procenv: proxy + mitmproxy CA fires one HIGH injection finding",
len(inj) == 1 and inj[0].severity == "high" and inj[0].check == "process-env-injection")
benign = scan_process_env("pid 202 (nginx)", by_label.get("pid 202 (nginx)", {}))
check("procenv: a benign process fires nothing", len(benign) == 0)
corp = scan_process_env("pid 303 (apt)", by_label.get("pid 303 (apt)", {}))
check("procenv: a corporate proxy (not :8788, no mitm CA) is not a false positive", len(corp) == 0)
caonly = scan_process_env("pid 404 (python)", by_label.get("pid 404 (python)", {}))
check("procenv: a mitmproxy CA alone fires one MEDIUM finding",
len(caonly) == 1 and caonly[0].severity == "medium")
proxyonly = scan_process_env("pid 505 (wget)", by_label.get("pid 505 (wget)", {}))
check("procenv: the ARTEX default proxy endpoint alone fires one MEDIUM finding",
len(proxyonly) == 1 and proxyonly[0].severity == "medium")
print()
if failures:
print(f"RESULT: FAIL ({len(failures)} assertion(s) failed)")
return 1
print("RESULT: PASS")
return 0
def build_parser():
p = argparse.ArgumentParser(
description="Read-only host triage for a suspected ARTEX host (detections/triage).",
)
p.add_argument("--data-dir", action="append", default=[], metavar="PATH",
help="ARTEX data directory to check for recording-proxy artifacts (repeatable).")
p.add_argument("--log", action="append", default=[], metavar="PATH",
help="log file to scan for ARTEX markers (repeatable).")
p.add_argument("--log-dir", action="append", default=[], metavar="PATH",
help="directory of log files to scan recursively; rotated "
".gz/.bz2/.xz logs are decompressed and scanned too, while "
".zst/.lz4 (no stdlib codec) are reported as skipped "
"(repeatable).")
p.add_argument("--pg-dsn", default=None, metavar="DSN",
help="PostgreSQL DSN to check for the exploration schema (defaults to $ARTEX_PG_DSN).")
p.add_argument("--ports-from", default=None, metavar="FILE",
help="read a port listing from FILE instead of running ss/netstat/lsof.")
p.add_argument("--proc-from", default=None, metavar="FILE",
help="read a captured process-environment dump from FILE instead of "
"reading /proc on the live host (offline / forensic-image triage). "
"Format: '# label' + KEY=VALUE lines, process blocks split by a blank line.")
p.add_argument("--json", action="store_true", help="emit findings as JSON.")
p.add_argument("--exit-code", action="store_true",
help="exit 1 if any indicator fired (default: always exit 0).")
p.add_argument("--self-test", action="store_true",
help="run the built-in fixture test and exit.")
return p
def main(argv=None):
args = build_parser().parse_args(argv)
if args.self_test:
return self_test()
findings, notes = run_checks(args)
if args.json:
print(json.dumps(
{"findings": [f.as_dict() for f in findings], "notes": notes},
ensure_ascii=False, indent=2,
))
else:
print_report(findings, notes)
if args.exit_code and findings:
return 1
return 0
if __name__ == "__main__":
sys.exit(main())