bddkit
Acceptance testing for backend services — the API and everything behind it — written in Gherkin and run by a single Rust binary.
A backend test rarely ends at the HTTP response. The interesting question is usually what happened behind it: did the row change, did the balance move, did the state settle a second later. bddkit treats every system a scenario touches as a named resource — declared once in the config, reachable from any scenario — so seeding a row, calling the API, and asserting the row changed are three steps in one vocabulary instead of three tools with glue between them.
Scenario: registering a company charges the account
Given I have "accounts" with "email: <<unique(email)>>, balance: 100"
And the request body is:
"""
{"account_id": "<<last_insert_id_accounts>>", "name": "Acme"}
"""
When I request "/api/v1/companies" using HTTP POST
Then the response code is 201
And I expect the next assertion to pass within "10" seconds
And I should have "accounts" with "balance: 75"
What it solves
Extending the vocabulary shouldn't require a rebuild. Domain steps like
I login as user "..." are declared in YAML and loaded at startup, so
whoever writes the scenarios can also extend the language they're written in.
A typo shouldn't cost you three minutes of run time. Unknown steps,
ambiguous patterns, macro cycles, undeclared connections — all found in one
pass before the first request. Exit code 2 means "run not started", so CI
can tell a broken suite from a failing one. Pattern collisions fail loudly
instead of silently shadowing each other.
Checking the database shouldn't mean leaving the feature file. DB steps introspect the real schema — PostgreSQL, MySQL or MariaDB, the same step text on each: primary keys fill themselves the way the column declares them, values coerce to real column types, and a missing NOT NULL column fails by name at the step rather than as a driver error. Hawk signing, SRP-6a, and AES are steps too, so a login handshake stays declarative.
Async systems shouldn't need sleep loops. Prefix any assertion with
I expect the next assertion to pass within "10" seconds and it retries —
re-sending the request or re-running the query — until it holds.
Test data shouldn't collide, and should be removable afterwards. Every <<unique()>> value in a run shares one prefix drawn once per process, so uniqueness is guaranteed rather than probable — and <<run_id>> writes that prefix down, so cleanup afterwards is one step: I delete "companies" where "slug~: u<<run_id>>%".
A failure should explain itself. Every failed step prints the full last exchange — method, URL, headers, bodies, status — with no debug flag and no re-run. JSON mismatches point at the path that differs.
Quick start
run m4k2p9x7q3b1
✓ examples/features/methods.feature — scenarios: 8
✓ examples/features/json_matchers.feature — scenarios: 4
✓ examples/features/variables.feature — scenarios: 5
✓ examples/features/macros.feature — scenarios: 6
✓ examples/features/content_types.feature — scenarios: 4
✓ examples/features/eventual.feature — scenarios: 1
run m4k2p9x7q3b1
files: 6, scenarios: 28, failed: 0
The example suite talks to a local Smocker
instance seeded from examples/mocks/api-server.yaml, so it runs offline and
its responses are fixed by a file in this repo. Smocker's web UI is on
http://localhost:8081.
The database examples are a second suite with its own config and its own container:
examples/README.md covers both suites, what each feature file demonstrates,
and how to narrow a run to one file or one tag.
bddkit run flags: --config (required), positional paths to override the config's, --tag (repeatable), --env to pick a .env.<name> layer, --fail-fast. Exit codes: 0 passed, 1 a scenario failed, 2 the run never started.
Checking a suite before running it
bddkit doctor runs every check a run makes before its first request — config parse, ${VAR} expansion, ambiguous default_*, the plugin lock file and ABI versions, macro conflicts and cycles, undefined steps, malformed @priority/@serial tags, a paths that selects nothing, and a DSN the driver cannot parse — and reports all of them at once instead of stopping at the first:
$ bddkit doctor --config suite.yaml
config: suite.yaml
APP_ENV: dev
✓ config parsed, ${VAR} expanded, defaults resolved
✓ plugins none installed
✓ macros loaded, no conflict or cycle
✗ steps
features/auth.feature:12
unknown step: "I frobnicate"
✓ scheduling 4 chain(s)
✓ api main https://api.example.test (20s timeout)
- api main live probe skipped
✓ db primary postgres
- db primary live probe skipped
1 problem(s)
static checks only — pass --live to also probe the resources
A bare doctor opens no socket: it is fast, offline and deterministic, so it works on a train and a base_url pointing at a closed port is not a finding. --live adds the one class of check a run does not have — whether the resources the config names actually answer: a GET at each base_url (any status means reachable), a real connection to each database, and each declared plugin instance asked to probe its own resource, every one named individually so one dead DSN says which. A plugin that exports no probe is reported as skipped, never as a failure — a check that never ran has proved nothing.
bddkit doctor flags: --config (required), --env to pick a .env.<name> layer, --live, --json. Exit codes: 0 clean, 1 anything was reported — never 2, so a script's rule is simply "0 or fix something".
The header states which APP_ENV layer was selected, which is the first thing that is wrong when a suite passes locally and fails in CI. --json emits the same report machine-readably, one object per check with a probe flag separating a resource's static row from its live one.
doctor never fixes what it finds, and every static stage is a function run itself calls — if the two ever disagree about whether a config is valid, that is a bug in doctor.
What steps exist
bddkit steps list prints the whole vocabulary as templates, grouped by resource, with no config and no regex:
$ bddkit steps list db
db:
I use "<name>" connection
I have "<table>" where:
I have "<table>" with "<pairs>"
I have:
I update "<table>" with "<pairs>" where "<condition>"
...
Narrow it with a positional resource (api, db, srp, vars, debug, general, or a plugin's group) and with --filter <text>, a case-insensitive substring match. -v adds a one-line description under each step — filtering down to one step and adding -v is the help-for-one-step path, so there is no separate describe command:
$ bddkit steps list --filter "response code" -v
api:
the response code is <code>
asserts the HTTP status code of the last response
--json emits the same listing machine-readably, with the raw pattern included. --config <file> also loads that suite's plugins, so their steps appear under their own groups. Descriptions follow --lang, else $BDDKIT_LANG, else English; ru and lv ship with the binary, and an untranslated step falls back to English rather than to a blank line.
What a resource's config takes
bddkit resource fields prints the keys each kind of resources entry accepts, which key is mandatory, what value it takes, and what it is for:
$ bddkit resource fields db
db:
dsn string required connection string; its scheme selects the engine (postgres://, mysql://)
search_path nonscalar Postgres schema search path; refused on MySQL and MariaDB, which have no session equivalent
options nonscalar polling.timeout_secs / polling.interval_ms for eventual assertions against this connection
The type column is the answer to "can I set this with a flag": string, boolean and number can, and nonscalar — a map or a list — is what resource add --json is for. A plugin's group is described by the plugin itself, and a field whose manifest declares no type is a string.
Narrow it with a positional kind (api, db, srp, or a plugin's group). --config <file> also loads that suite's plugins, so the groups they serve are described too — by the plugins themselves, since only a plugin knows what its own group takes. A plugin whose manifest describes nothing says so and is not an error. --json emits the same listing machine-readably.
bddkit resource add
bddkit resource add <group> <name> --config <path> [--env <name>] [--json '{...}'] [--no-check] [--<field> <value>]...
$ bddkit resource add api staging --config suite.yaml --base_url http://staging.local --timeout_secs 5
$ bddkit resource add s3 main --config suite.yaml --no-check --bucket photos --endpoint https://s3.local
Writes one new resource into --config's YAML and reports whether it happened. A --<field> <value> becomes a body key; the value always arrives as a string, and it is the group's own field table (the host's for api/db/srp, a plugin's manifest for its own group) that converts it — timeout_secs becomes a YAML number, and a field the table calls a map or a list refuses the flag and names --json instead. --json '{...}' supplies the whole body at once and is applied first; any --<field> on the same command line then overrides that key, one key at a time, leaving the rest of the JSON body untouched. A name the group's table does not know is refused rather than silently added. The command's own flags (--config, --env, --json, --no-check) must be typed before any --<field> value — clap treats everything after the first flag it does not itself recognize as more field arguments, so one of these typed later would otherwise vanish into the resource body or be rejected as an unknown field. Before writing anything, the exact prospective file text is assembled, re-parsed and put through the checks doctor makes about a config and its resources — the suite's feature files, steps, macros and scheduling tags are deliberately not among them, since a broken .feature elsewhere is no reason to refuse a resource. A ${VAR} placeholder in a value is written out literally and only expanded for that check. The live probe that doctor --live runs against this one new resource is also run here, by default; --no-check skips only that probe, the shape validation still happens. A resource that already exists under that name is refused rather than replaced, because rewriting the surrounding YAML block is how hand-written comments in it get lost — insertion only ever splices in a new block next to the existing text. Exit code is 0 when the file was written and 1 when it was not (a malformed flag, a rejected field, a failed check, an existing resource), matching doctor's "0, or fix something" convention; the command itself never answers 2, though a usage error the argument parser catches before it starts still does.
Config
concurrency: 8 # files in flight; a file is one tokio task
macro_paths:
paths:
resources:
api:
review:
base_url: http://review.local
default_headers:
db:
default:
dsn: postgres://${DB_USER}:${DB_PASS}@db:5432/review
search_path:
options:
polling: # defaults for eventual assertions
${VAR} expands docker-compose style (:-, :?, :+ and friends) from the
real environment and a .env / .env.local / .env.<APP_ENV> layer stack
next to the config. With one resource of a kind its default_* is inferred;
with several it must be explicit. options set at the root cascades into
every resource and can be overridden per resource.
Variables live for the whole feature file, so a later scenario can read
what an earlier one produced. Request state and the current connection reset
per scenario. Files run in parallel; @serial(name) chains files that
contend, @priority(N) moves one up the queue.
Databases
PostgreSQL, MySQL 8 and MariaDB, from MariaDB 10.5 — the release that adds RETURNING, which bddkit uses to read a generated key back. Nothing refuses an older MariaDB at startup; it connects, and the first insert into a table with a primary key fails with a syntax error from the server. The engine comes from the DSN scheme (postgres://, mysql://), and because MySQL and MariaDB share one scheme they are told apart by a single SELECT VERSION() when the pool is built. Every DB step is spelled the same on all three — examples/db-features-mysql/db.feature is one file that passes unchanged against both MySQL and MariaDB.
What does not carry over:
- Sequences. MySQL has none, so
I get next value of sequencefails naming the engine. Postgres and MariaDB both have them. search_path. On MySQL and MariaDB a schema is a database, so there is no session-level search path. The key is refused at startup rather than quietly ignored: name the database in the DSN, or qualify a table asdatabase.tablein the step.binaryandvarbinarycolumns,binary(16)UUIDs included. This layer binds and compares everything as text, andUUID_TO_BIN'sswap_flag— which decides the byte order — is recorded nowhere ininformation_schema, so bddkit would have to guess and could write bytes the service under test disagrees with. It refuses the column instead. Store the UUID aschar(36)and bddkit fills it client-side.- A primary key filled by a server-side
DEFAULTthat is notAUTO_INCREMENT, on MySQL. There is noRETURNINGthere to read the value back with, so the step fails naming the column before it inserts anything — give the value explicitly. - The text a value reads back as is the engine's own.
I extractis portable as a step and not as a value:now()reads back as2026-08-29 15:11:50.884052+00on Postgres and2026-08-29 15:11:50on MySQL and MariaDB, a boolean astrueversus1, a barenumericas0wheredecimal(12,2)is0.00. SoI extractfollowed byvariable "x" should be equal to "..."is portable overvarchar-ish columns and nowhere else. Awhereagainst those same columns is unaffected — the engine coerces the bind back into its own type; it is only the extracted value that carries the engine's spelling.
examples/README.md has the same ground as a per-engine table, with the workaround for each row.
Cleaning up what a run created
Every <<unique()>> token is u followed by the run's own 12-character prefix and a counter, and <<run_id>> is that prefix. A ~ on the column name asks for SQL LIKE instead of =, so one step removes exactly the rows this run wrote — examples/db-features/cleanup.feature is it, working:
@serial(demo) @priority(-100)
Feature: cleaning up what this run created
Scenario: remove the companies this run created
When I delete "companies" where "slug~: u<<run_id>>%"
Four things to know about it:
- The operator lives on the column, never in the value.
slug: 20%is still exact equality against the text20%; onlyslug~:is a pattern. That direction is deliberate: a%that arrives through a variable — from an API response, say — must never be able to widen aDELETEon its own. - The condition grammar is four operators, and they work in every step that takes a condition:
col:is=,col!:is<>,col~:isLIKE,col!~:isNOT LIKE. With<<null>>:col:readsIS NULLandcol!:readsIS NOT NULL. Under a~the value is an SQLLIKEpattern in full —_matches any single character too, and\%/\_are the literal characters. - A negation with a value never matches a NULL column.
name!: keeperis SQL'sname <> 'keeper', and that is unknown — not true — for a row whosenameis NULL, so "everything that isn'tkeeper" quietly leaves those rows behind. It is ordinary three-valued logic and the engine is right; it is just rarely what the sentence in your head meant. Addname: <<null>>as a second condition if you want them too. <<run_id>>is now a reserved name. Like<<null>>,<<uuid()>>and<<unique()>>, it is answered before your own variables are looked at, so a variable you set calledrun_id— plausible, if you extract one from a column — is shadowed rather than read.- Only
<<unique()>>carries the run prefix — thetokenkind and itsemail,slug,urlaliases.<<unique(number)>>is microseconds plus a counter, with no prefix in it, so aLIKEon<<run_id>>never reaches a column filled that way: delete those rows by a column that does carry a token, or by their key. @priority(-100)makes the file last in its own chain, not last in the run. Chains run in parallel, so the cleanup file and the files whose data it removes belong in one@serial(<name>)chain — otherwise it can delete rows another chain is still using.
Plugins
api, db and srp are the resource kinds built into the binary. Any other key under resources: is a group served by a plugin — a shared library bddkit loads at startup — so reaching an object store, a queue or a mailbox is the same move as reaching a second database:
resources:
api:
review:
s3: # served by a plugin, not by bddkit itself
backups:
bucket: acme-backups
endpoint: http://minio:9000
archive:
bucket: acme-archive
default_s3: backups
The plugin brings its own steps, which read like any other step — a tester cannot tell a built-in from a plugin step, and a macro can call one. Instances are selected the same way as an API or a connection:
Given I use "archive" s3
When I upload file "report.pdf"
The selection resets to default_<group> at every scenario boundary, exactly like the current API and the current connection. With one instance in a group its default_<group> is inferred; with several it must be spelled out.
Which plugins are installed is machine state, not test config. It lives in .bddkit/plugins.yaml next to your config file — a list of {name, path} — and it does not belong in the repository with the suite: the config describes the system under test, a path to a .so describes one laptop or one CI runner.
# .bddkit/plugins.yaml
plugin:
- name: s3
path: /opt/bddkit/libbddkit_s3.so
A plugin runs inside the bddkit process with full privileges and there is no sandbox — installing one is the same trust decision as installing any other binary.
Writing one: docs/plugin-authoring.md is the complete contract, tests/fixtures/echo-plugin/ is a minimal plugin to copy, and bddkit-s3 is a real one to read — it serves the s3 group in the example above.
Where to look next
| For | Look at |
|---|---|
| Every step, authoritative | BUILTIN_STEPS in src/steps/mod.rs |
| How to run the examples | examples/README.md |
| A runnable HTTP example | examples/api.yaml, examples/features/ |
| Every HTTP method, 404 included | examples/features/methods.feature |
| JSON matchers and paths | examples/features/json_matchers.feature |
| Variables: set, extract, reuse | examples/features/variables.feature |
| Macros, nesting, Scenario Outline | examples/features/macros.feature, examples/macros/posts.yaml |
| Non-JSON responses, form login | examples/features/content_types.feature |
| Polling an assertion until it passes | examples/features/eventual.feature |
| The mock API behind all of it | examples/mocks/api-server.yaml |
| Every DB step, worked through | examples/db-features/db.feature |
| Cleaning up a run's data | examples/db-features/cleanup.feature |
| The same on MySQL and MariaDB | examples/db-features-mysql/db.feature |
| SRP handshake, Hawk signing | tests/features/ |
| Config schema | src/config.rs |
| Writing a plugin | docs/plugin-authoring.md, tests/fixtures/echo-plugin/ |
Design notes
Why not cucumber-rs. Three axes diverge: steps register at compile time
via proc-macros (here they load from YAML at run time), World is recreated
per scenario (here variables are file-scoped), and concurrency is a flat pool
over scenarios (here it's chains of files). Working around all three leaves
only its run loop while breaking its own step diagnostics — which is the part
worth strengthening. The gherkin crate underneath it is reused directly.
Why no transactional rollback. The service under test reads over its own connection and would never see uncommitted rows, so rolling back per scenario would break any test where a step writes and the API reads. Isolation comes from unique data instead.
Where this is going. api, db, and srp are the resource kinds that
ship, not the ceiling — any other key under resources: is a capability group
a plugin serves, so reaching an object store or a mailbox is the same move as
reaching a second database. The seams that made that possible were there from
the start: options cascade per instance, I use "<name>" <kind> is one step
shape, and dispatch returns passed | not yet | fatal so eventual assertions
work without knowing what they retry. What is still missing is the
bddkit plugin install side of it — today .bddkit/plugins.yaml is written
by hand.
Development
Some tests need a database; docker-compose.yml brings up all three — PostgreSQL on :5433 (schema from examples/db/init.sql), MySQL on :3307 and MariaDB on :3308 (both from examples/db/init-mysql.sql). The DB suite runs against Postgres by default; BDDKIT_TEST_ENGINE=mysql or =mariadb points it at the other two, and CI runs all three. The same file brings up Smocker as the HTTP example's mock API — it seeds examples/mocks/api-server.yaml at startup and serves it on localhost:8080 (web UI on localhost:8081).
License
Apache-2.0. See LICENSE.