bddkit
Acceptance testing for backend services — the API and everything behind it — written in Gherkin and run by a single Rust binary.
A backend test rarely ends at the HTTP response. The interesting question is usually what happened behind it: did the row change, did the balance move, did the state settle a second later. bddkit treats every system a scenario touches as a named resource — declared once in the config, reachable from any scenario — so seeding a row, calling the API, and asserting the row changed are three steps in one vocabulary instead of three tools with glue between them.
Scenario: registering a company charges the account
Given I have "accounts" with "email: <<unique(email)>>, balance: 100"
And the request body is:
"""
{"account_id": "<<last_insert_id_accounts>>", "name": "Acme"}
"""
When I request "/api/v1/companies" using HTTP POST
Then the response code is 201
And I expect the next assertion to pass within "10" seconds
And I should have "accounts" with "balance: 75"
What it solves
Extending the vocabulary shouldn't require a rebuild. Domain steps like
I login as user "..." are declared in YAML and loaded at startup, so
whoever writes the scenarios can also extend the language they're written in.
A typo shouldn't cost you three minutes of run time. Unknown steps,
ambiguous patterns, macro cycles, undeclared connections — all found in one
pass before the first request. Exit code 2 means "run not started", so CI
can tell a broken suite from a failing one. Pattern collisions fail loudly
instead of silently shadowing each other.
Checking the database shouldn't mean leaving the feature file. DB steps
introspect the real schema: primary keys fill themselves the way the column
declares them, values coerce to real column types, and a missing NOT NULL
column fails by name at the step rather than as a driver error. Hawk signing,
SRP-6a, and AES are steps too, so a login handshake stays declarative.
Async systems shouldn't need sleep loops. Prefix any assertion with
I expect the next assertion to pass within "10" seconds and it retries —
re-sending the request or re-running the query — until it holds.
Test data shouldn't collide, and should be removable afterwards. Every
<<unique()>> value in a run shares one prefix drawn once per process, so
uniqueness is guaranteed rather than probable, and cleanup afterwards is one
DELETE FROM t WHERE col LIKE '%<run_id>%'.
A failure should explain itself. Every failed step prints the full last exchange — method, URL, headers, bodies, status — with no debug flag and no re-run. JSON mismatches point at the path that differs.
Quick start
run m4k2p9x7q3b1
✓ examples/features/demo.feature — scenarios: 2
run m4k2p9x7q3b1
files: 1, scenarios: 2, failed: 0
Flags: --config (required), positional paths to override the config's,
--tag (repeatable), --env to pick a .env.<name> layer, --fail-fast.
Exit codes: 0 passed, 1 a scenario failed, 2 the run never started.
Config
concurrency: 8 # files in flight; a file is one tokio task
macro_paths:
paths:
resources:
api:
review:
base_url: http://review.local
default_headers:
db:
default:
dsn: postgres://${DB_USER}:${DB_PASS}@db:5432/review
search_path:
options:
polling: # defaults for eventual assertions
${VAR} expands docker-compose style (:-, :?, :+ and friends) from the
real environment and a .env / .env.local / .env.<APP_ENV> layer stack
next to the config. With one resource of a kind its default_* is inferred;
with several it must be explicit. options set at the root cascades into
every resource and can be overridden per resource.
Variables live for the whole feature file, so a later scenario can read
what an earlier one produced. Request state and the current connection reset
per scenario. Files run in parallel; @serial(name) chains files that
contend, @priority(N) moves one up the queue.
Where to look next
| For | Look at |
|---|---|
| Every step, authoritative | BUILTIN_STEPS in src/steps/mod.rs |
| A runnable HTTP example | examples/api.yaml, examples/features/ |
| Every DB step, worked through | examples/db-features/db.feature |
| Macros with parameters and exports | examples/macros/, tests/macros/ |
| SRP handshake, Hawk signing | tests/features/ |
| Config schema | src/config.rs |
Design notes
Why not cucumber-rs. Three axes diverge: steps register at compile time
via proc-macros (here they load from YAML at run time), World is recreated
per scenario (here variables are file-scoped), and concurrency is a flat pool
over scenarios (here it's chains of files). Working around all three leaves
only its run loop while breaking its own step diagnostics — which is the part
worth strengthening. The gherkin crate underneath it is reused directly.
Why no transactional rollback. The service under test reads over its own connection and would never see uncommitted rows, so rolling back per scenario would break any test where a step writes and the API reads. Isolation comes from unique data instead.
Where this is going. api, db, and srp are the resource kinds that
ship, not the ceiling — a key under resources: is meant to be a capability
group, so reaching an object store or a mailbox becomes the same move as
reaching a second database. The seams exist (options cascade per instance,
I use "<name>" <kind> is one step shape, dispatch returns
passed | not yet | fatal so eventual assertions work without knowing what
they retry). What's missing is loading steps and resource kinds from outside
the binary.
Development
Some tests need PostgreSQL; docker-compose.yml brings one up and
examples/db/init.sql creates the schema.
License
Apache-2.0. See LICENSE.