rsigma-convert 0.24.0

Sigma rule conversion engine — convert rules to backend-native query strings
Documentation

rsigma-convert

CI

rsigma-convert is a Sigma rule conversion engine that transforms parsed Sigma rules into backend-native query strings (SQL, SPL, KQL, Lucene, etc.).

This library is part of rsigma.

Overview

The crate provides a generic conversion framework that any backend can plug into:

  • Backend trait with ~30 methods covering condition dispatch, detection item conversion, field/value escaping, regex, CIDR, comparison operators, field existence, field references, keywords, IN-list optimization, deferred expressions, and query finalization.
  • TextQueryConfig with ~90 configuration fields mirroring pySigma's TextQueryBackend class variables: precedence, boolean operators, wildcards, string/field quoting, match expressions (startswith/endswith/contains + case-sensitive variants), regex/CIDR templates, compare ops, IN-list optimization, unbound values, deferred parts, and query envelope.
  • Condition tree walker that recursively converts ConditionExpr nodes into query strings with selector/quantifier support.
  • Orchestrator via convert_collection(), which merges filters into the rules they target, applies pipelines, converts each rule, and collects results and errors. A backend without correlation support reports each correlation as an UnsupportedCorrelation error.
  • Deferred expressions through the DeferredExpression trait and DeferredTextExpression for backends that need post-query appendages (e.g. Splunk | regex, | where).
  • Test backend with TextQueryTestBackend and MandatoryPipelineTestBackend for backend-neutral foundation testing. Its output matches pySigma's TextQueryTestBackend: values are escaped and quoted the same way, the expression follows the value's wildcard shape, values of one field become in and contains-all lists, and CIDR matches render as cidrmatch('Field', "cidr").
  • PostgreSQL/TimescaleDB backend with native ILIKE, regex (~, ~* with |i), CIDR (inet/cidr), full-text search (tsvector/tsquery), JSONB field access, correlation via CTEs and window functions, and TimescaleDB-specific output formats (continuous aggregates, time_bucket queries, view generation).
  • LynxDB backend generating SPL2-compatible FROM <index> | search ... queries with glob wildcards and correct parenthesization for LynxDB's non-standard boolean precedence (NOT > OR > AND), and FROM <index> | where ... queries for rules whose values search cannot match exactly.

Backends

Backend Target names Description
Test test Backend-neutral text queries for foundation testing
PostgreSQL postgres, postgresql, pg Native PostgreSQL SQL with TimescaleDB support
LynxDB lynxdb SPL2-compatible search queries for LynxDB log analytics engine
Fibratus fibratus Fibratus rule YAML for the Apache-2.0 kernel-event detection and EDR engine

The optional sigma-cli feature (std-only, no extra dependencies) adds the sigma_cli module: discovery of an external sigma-cli, the sigma convert argument mapping (build_convert_args), and subprocess output classification (classify_output). It is the shared helper behind rsigma backend convert delegation and the MCP server's opt-in --allow-sigma-cli; the conversion API itself (convert_collection, the Backend trait) always converts natively and never spawns a subprocess.

Reverse conversion (query to Sigma)

The reverse module is the mirror of the Backend engine: it parses a SIEM query into the intermediate representation, raises a Sigma rule (via rsigma_ir::raise_rule), and emits YAML (via rsigma_parser::emit_rule_yaml). A Frontend trait plus a QueryDialect table drive a shared tokenizer and a precedence-climbing boolean parser (NOT > AND > OR); assemble_rule turns the boolean tree of leaves into named selections and a condition (AND-merged selections, same-field OR value lists, negated branches as filters), and reverse_collection converts a batch, collecting per-query errors. LuceneFrontend is the reference target, parsing the Lucene / Elasticsearch query_string subset (field:value with wildcards, quoted phrases, /regex/, [a TO b] and {a TO b} ranges, field:>=N comparison shorthand, field:(a OR b) value groups, _exists_, keyword terms, and AND/OR/NOT with grouping); constructs with no Sigma equivalent (boosting ^, fuzzy/proximity ~, non-numeric ranges) are rejected with a structured ConvertError. Adding a target is a new QueryDialect plus a Frontend::parse_atom. The CLI surface is rsigma rule reverse --from <dialect>.

Usage

Test backend

use rsigma_parser::parse_sigma_yaml;
use rsigma_convert::{convert_collection, Backend};
use rsigma_convert::backends::test::TextQueryTestBackend;

let yaml = r#"
title: Detect Whoami
logsource:
    category: process_creation
    product: windows
detection:
    selection:
        CommandLine|contains: 'whoami'
    condition: selection
level: medium
"#;

let collection = parse_sigma_yaml(yaml).unwrap();
let backend = TextQueryTestBackend::new();

let output = convert_collection(&backend, &collection, &[], "default").unwrap();
for result in &output.queries {
    for query in &result.queries {
        println!("{query}");
        // Output: CommandLine contains "whoami"
    }
}

PostgreSQL backend

use rsigma_parser::parse_sigma_yaml;
use rsigma_convert::{convert_collection, Backend};
use rsigma_convert::backends::postgres::PostgresBackend;

let yaml = r#"
title: Detect Whoami
logsource:
    category: process_creation
    product: windows
detection:
    selection:
        CommandLine|contains: 'whoami'
    condition: selection
level: medium
"#;

let collection = parse_sigma_yaml(yaml).unwrap();
let backend = PostgresBackend::new();

let output = convert_collection(&backend, &collection, &[], "default").unwrap();
for result in &output.queries {
    for query in &result.queries {
        println!("{query}");
        // Output: SELECT * FROM security_events WHERE "CommandLine" ILIKE '%whoami%'
    }
}

LynxDB backend

use rsigma_parser::parse_sigma_yaml;
use rsigma_convert::{convert_collection, Backend};
use rsigma_convert::backends::lynxdb::LynxDbBackend;

let yaml = r#"
title: Detect Whoami
logsource:
    category: process_creation
    product: windows
detection:
    selection:
        CommandLine|contains: 'whoami'
    condition: selection
level: medium
"#;

let collection = parse_sigma_yaml(yaml).unwrap();
let backend = LynxDbBackend::new();

let output = convert_collection(&backend, &collection, &[], "default").unwrap();
for result in &output.queries {
    for query in &result.queries {
        println!("{query}");
        // Output: FROM main | search CommandLine=*"whoami"*
    }
}

LynxDB output formats

Format Description
default Full query with index prefix: FROM main | search ...
minimal Search expression only (no FROM prefix), useful for the LynxDB API q parameter

LynxDB index selection

The target index defaults to main. Set it via pipeline state:

# In a pipeline YAML
transformations:
  - type: set_state
    key: index
    value: security_logs
rsigma backend convert -r rules/ -t lynxdb -p pipeline.yml
# Output: FROM security_logs | search ...

PostgreSQL output formats

Format Description
default Plain SELECT * FROM {table} WHERE ... queries
view CREATE OR REPLACE VIEW sigma_{id} AS SELECT ...
timescaledb Queries with time_bucket() for TimescaleDB optimization
continuous_aggregate CREATE MATERIALIZED VIEW ... WITH (timescaledb.continuous)
sliding_window Correlation queries using window functions for per-row sliding detection

SELECT column selection

When a Sigma rule specifies fields:, the backend emits SELECT field1, field2, ... instead of SELECT *. Function calls (e.g. count(*)) pass through unchanged, and field as alias is supported with both sides quoted independently.

CLI backend options

Backend configuration can be set via -O key=value flags on the CLI, which are wired through to PostgresBackend::from_options. Recognized keys: table, schema, database, timestamp_field, json_field, correlation_method, gap.

rsigma backend convert -r rules/ -t postgres -O table=security_logs -O schema=public -O timestamp_field=created_at

Custom table, schema, and database

The target table and schema can be set at three levels (highest precedence first):

  1. Rule-level custom_attributes: postgres.table, postgres.schema, postgres.database
  2. Pipeline state: set_state with key: table, key: schema
  3. CLI backend options: -O table=..., -O schema=..., -O database=...
  4. Backend defaults: PostgresBackend.table, .schema, .database

Example rule with custom attributes:

title: Process Creation
logsource:
    category: process_creation
detection:
    selection:
        CommandLine|contains: 'whoami'
    condition: selection
custom_attributes:
    postgres.table: process_events
    postgres.schema: siem

OCSF pipelines

Two OCSF processing pipelines are included:

Pipeline Description
pipelines/ocsf_postgres.yml Single-table: all events go to security_events
pipelines/ocsf_postgres_multi_table.yml Per-logsource routing: each category gets its own table (process_events, network_events, etc.)
# Single-table pipeline
rsigma backend convert -r rules/ -t postgres -p pipelines/ocsf_postgres.yml

# Multi-table pipeline (per-logsource routing)
rsigma backend convert -r rules/ -t postgres -p pipelines/ocsf_postgres_multi_table.yml

# With output format
rsigma backend convert -r rules/ -t postgres -p pipelines/ocsf_postgres.yml -f view
rsigma backend convert -r rules/ -t postgres -p pipelines/ocsf_postgres.yml -f continuous_aggregate

Multi-table temporal correlations

When a temporal correlation rule references detection rules that target different tables (via per-logsource pipeline routing or custom attributes), the backend automatically generates a UNION ALL CTE:

-- Rules targeting different tables produce UNION ALL
WITH matched AS (
    SELECT *, 'process_rule' AS rule_name FROM process_events
        WHERE time >= NOW() - INTERVAL '300 seconds'
    UNION ALL
    SELECT *, 'network_rule' AS rule_name FROM network_events
        WHERE time >= NOW() - INTERVAL '300 seconds'
)
SELECT "User", COUNT(DISTINCT rule_name) AS distinct_rules,
    MIN(time) AS first_seen, MAX(time) AS last_seen
FROM matched
GROUP BY "User"
HAVING COUNT(DISTINCT rule_name) >= 2

When all referenced rules share the same table, the simpler single-table approach is used instead.

Per-rule schemas are also tracked: if different detection rules set different schemas (via postgres.schema custom attribute or set_state key: schema in the pipeline), each leg of the UNION ALL uses the correct schema.table.

Important: The multi-table UNION ALL uses SELECT * in each leg, so PostgreSQL requires all referenced tables to have the same column count and compatible column types. This works well when tables share a normalized event schema. If your tables have different column layouts, either normalize them through pipeline field-mappings or use a single-table approach with a discriminator column (e.g. rule_name) instead.

Reference schema

A reference TimescaleDB schema is provided at schema/timescaledb_security_events.sql with hypertable setup, indexes (B-tree, GIN for full-text and JSONB), compression, retention policies, and an example continuous aggregate.

Backend Trait

Backends implement the Backend trait to produce query strings from a rule's intermediate representation (rsigma-ir). A rule is lowered to an IrRule once, and the trait's value leaves consume the faithful HIR (IrPattern, IrStrOp, IrNumber, resolved flags) rather than parser types, so a backend never touches rsigma-parser to emit a value match. The generic detection and condition walks live in the crate and call these leaves.

Key methods:

Method Description
convert_rule Convert a single SigmaRule into query strings
convert_ir_detection Walk an IrDetection (AllOf/AnyOf/Keywords/array match)
convert_ir_detection_item Convert a single IrDetectionItem (field + matcher)
convert_field_str String matching over an IrStrOp + wildcard-aware IrPattern
convert_field_regex Regex matching with explicit RegexFlags; RegexFlags::inline_prefix renders i, m, and s as an inline (?ims) group
convert_field_eq_cidr CIDR matching
convert_field_compare_op Numeric comparison via CompareOp (gt, gte, lt, lte)
convert_field_exists Field existence check
convert_keyword_str / convert_keyword_num Unbound/keyword value matching
convert_condition_and / convert_condition_or / convert_condition_not Combine sub-expressions
convert_condition_group Parenthesize a compound operand for its parent operator; the default follows NOT > AND > OR and always groups under NOT, as pySigma does. Text backends delegate to text_convert_condition_group so their precedence applies
finish_query Assemble final query with deferred parts
finalize_query Apply output format to a query
finalize_output Finalize the complete output

TextQueryConfig

For text-based query backends (the vast majority), create a TextQueryConfig with your backend's tokens and expressions, then delegate to the text_convert_* free functions:

Function Description
text_escape_and_quote_field Escape and optionally quote a field name
text_convert_ir_pattern Convert an IrPattern with escaping and quoting
text_convert_value_re Escape a regex pattern
text_convert_condition_and Join expressions with AND token
text_convert_condition_or Join expressions with OR token
text_convert_condition_not Negate an expression
text_convert_condition_group Precedence-aware grouping
text_convert_field_str_ir String match dispatch (contains/startswith/endswith/wildcard/exact)
text_finish_query Assemble query with deferred parts and state substitution

Implementing a Backend

  1. Define a TextQueryConfig constant with your backend's tokens and expressions.
  2. Create a struct that implements Backend, delegating most methods to the text_convert_* helpers.
  3. Override specific methods for backend-specific behavior (e.g. deferred regex for Splunk, SQL-specific CIDR handling for PostgreSQL).
  4. Register your backend in the CLI's try_native_backend() registry (a native backend takes precedence over sigma-cli delegation for that target).

See backends/test.rs for a complete reference implementation, backends/postgres.rs for a production backend with SQL-specific overrides, and backends/lynxdb/ for a TextQueryConfig-based backend with custom precedence handling that switches to a second config for where expressions.

PostgreSQL Backend Details

The PostgreSQL backend (PostgresBackend) leverages native PostgreSQL features that map cleanly to Sigma modifiers:

Sigma Modifier PostgreSQL SQL
contains ILIKE (case-insensitive)
startswith / endswith ILIKE
cased LIKE (case-sensitive)
re ~ (case-sensitive regex) or ~* (with i); (?w) prefix with m
cidr field::inet <<= 'value'::cidr
exists IS NOT NULL / IS NULL
keywords to_tsvector() @@ plainto_tsquery()

Correlation rules are converted to SQL using GROUP BY / HAVING for aggregation types (event_count, value_count, value_sum, value_avg, value_percentile, value_median) and CTEs for temporal correlation. Multi-table temporal correlations automatically generate UNION ALL CTEs when referenced rules target different tables. A collection conversion omits the standalone queries of referenced detection rules and referenced correlations unless at least one referencing correlation has top-level generate: true. The aggregate types embed the logic of referenced detection rules in a combined_events CTE, so their queries stay self-contained. temporal and temporal_ordered queries do not embed it: they filter on a rule_name column of the source table, so set generate: true when the standalone detection queries are still needed.

Non-temporal correlations support CTE-based pre-filtering: when the correlation references detection rules that were converted in the same collection, the backend wraps their queries in a WITH combined_events AS (q1 UNION ALL q2 ...) CTE so the aggregate only counts events matching the detection logic. In a collection conversion, a reference without a converted detection query (another correlation, or a detection rule that failed to convert) is an UnsupportedCorrelation error. Converting a single correlation without the collection scans the whole table within the timespan.

The sliding_window output format uses SQL window functions for event_count correlations, producing a per-row sliding window that emits every event crossing the threshold:

WITH combined_events AS (...),
event_counts AS (
    SELECT *, COUNT(*) OVER (
        PARTITION BY "User"
        ORDER BY time
        RANGE BETWEEN INTERVAL '300 seconds' PRECEDING AND CURRENT ROW
    ) AS correlation_event_count
    FROM combined_events
)
SELECT * FROM event_counts WHERE correlation_event_count >= 5

A correlation rule's window attribute selects the windowing strategy (independent of output format). An absent or sliding window keeps the SQL above unchanged; tumbling emits boundary-aligned buckets sized to the rule's timespan (time_bucket on TimescaleDB, date_bin on plain PostgreSQL); session emits a gaps-and-islands query (LAG + a running session_id) that honors the gap exactly and enforces the timespan cap as a HAVING filter. Backends can attach non-fatal diagnostics for such approximations: ConversionResult carries a warnings field, populated via Backend::convert_correlation_rule_with_warnings, and the CLI prints them to stderr.

Following pySigma's model, the windowing strategy can also be chosen at conversion time rather than taken from the rule. A backend advertises its strategies through Backend::correlation_methods (name/description pairs) and default_correlation_method; the PostgreSQL backend offers sliding/tumbling/session with sliding as the default. The CLI exposes the choice as rsigma backend convert -O correlation_method=NAME, which overrides a rule's own window hint for that conversion and is validated against the advertised methods. rsigma backend formats <target> lists them. For correlation_method=session over rules without a gap, -O gap=5m supplies the conversion-time default (a rule's own gap always wins).

Configuration

PostgresBackend fields:

Field Type Default Description
table String "security_events" Default table name (overridden by pipeline state or postgres.table custom attribute)
timestamp_field String "time" Timestamp column for time-windowed queries
json_field Option<String> None If set, fields are accessed via JSONB extraction (see JSONB field access)
schema Option<String> None PostgreSQL schema name (overridden by pipeline state or postgres.schema custom attribute)
database Option<String> None PostgreSQL database name (connection-level metadata)
timescaledb bool false Enable TimescaleDB-specific features
correlation_method Option<String> None Windowing strategy override for correlations (sliding/tumbling/session); None uses each rule's own window
session_gap_secs Option<u64> None Default session gap in seconds (from -O gap=5m), used when a session window is requested and the rule declares no gap

JSONB field access

When json_field is set (e.g. -O json_field=data), all Sigma field references are translated to PostgreSQL JSONB extraction operators instead of bare column names.

For top-level fields, the backend uses the ->> operator:

-- Sigma field: eventType
data->>'eventType'

For dotted field names (nested paths like securityContext.isProxy), the backend generates chained operators where intermediate segments use -> (returns jsonb) and the final segment uses ->> (returns text):

-- Sigma field: securityContext.isProxy
data->'securityContext'->>'isProxy'

-- Sigma field: actor.detail.alternateId
data->'actor'->'detail'->>'alternateId'

This matches the nested traversal behavior of the evaluation engine (rsigma-eval), which splits dotted field names on . and walks into nested JSON objects.

# Convert rules with JSONB field access against a "data" column
rsigma backend convert -r rules/ -t postgres -O table=okta_events -O json_field=data -O timestamp_field=time

LynxDB Backend Details

The LynxDB backend (LynxDbBackend) generates SPL2/Lynx Flow queries for the LynxDB log analytics engine. A rule renders as FROM <index> | search <predicates> when LynxDB's search matches every value in it exactly, and as FROM <index> | where <expression> otherwise.

Sigma Modifier search where
contains field=*"value"* match(field, "(?i)value")
startswith field="value"* match(field, "(?i)^value")
endswith field=*"value" match(field, "(?i)value$")
re match(field, "pattern")
cidr cidrmatch("cidr", field)
cased match(field, "^Value$")
lt, lte, gt, gte coalesce(tonumber(field)>1000, false)
wildcards (*, ?) field="value"* (leading or trailing * only) .* and . in the regex
exists field=* isnotnull(json_extract(_raw, "field"))
null isnull(json_extract(_raw, "field"))
keywords "value" (unbound search) match(_raw, "(?i)value")

search stays in use for strings of letters, digits, spaces, and . - _ : \ matched case-insensitively, numeric and boolean equality, and exists. Any other value puts the whole condition in where, so a regex nested under OR or NOT keeps its place in the condition.

Boolean precedence

LynxDB's search uses non-standard boolean operator precedence: NOT > OR > AND. This differs from most query languages where AND binds tighter than OR. The backend parenthesizes an AND nested under an OR and every compound operand of NOT, and leaves an OR nested under an AND bare because it already binds tighter:

Sigma: (A and B) or C        Query: (A AND B) OR C
Sigma: (A or B) and C        Query: A OR B AND C
Sigma: A and not (B or C)    Query: A AND NOT (B OR C)

where expressions use standard precedence:

FROM main | where coalesce(tonumber(status)=500, false) AND match(Path, "/api/.*")

License

MIT License.