EzRaft
A beginner-friendly Raft consensus framework built on OpenRaft. EzRaft handles all Raft complexity internally - users only provide business logic and storage persistence.
Overview
Run your application on several machines at once, all holding the same state, so the
service survives losing some of them. That is what Raft is
for, and EzRaft reduces it to two traits - EzApp holds your state and applies
requests to it, EzStorage puts bytes on disk. Elections, replication,
membership, snapshots and the transport between machines are handled internally.
- Minimal user API: 4 methods total (3 storage + 1 app) vs 21+ in OpenRaft
- Smart defaults: 10/12 Raft types pre-configured, users specify only Request/Response
- Built-in networking: HTTP layer included, no user code needed
- Type-safe: Works directly with your types, not byte vectors
Status
Experimental. EzRaft is primarily an API design laboratory for exploring intuitive interface patterns. The APIs may change until the crate stabilizes. Production applications are not the primary audience.
Next phase: Stable API. Once the design exploration matures, EzRaft will provide a stable API with well-considered abstractions—exposing what users need while hiding unnecessary complexity.
Run a cluster
Three nodes, three terminals, from the repository root. The first node creates the cluster; every other node joins through a node already in it:
Each prints the node id the cluster gave it. Drive them from a fourth terminal - the
request bodies are the example's Request and ReadRequest enums as JSON:
# A write goes through the log: replicated and committed before it answers.
# {"value":null} - Set answers with the value it replaced, if there was one
# A read is answered from that node's memory: no consensus round, no log entry.
# {"hello":"world"} - from a node you never wrote to
# Any node accepts a write; a follower forwards it to the leader.
# {"value":"world"}
Now kill any one of the three. The two survivors are a majority, so they elect a new
leader within a few heartbeats and keep serving writes. Start the dead one again with the
exact same command line and it rejoins with its old id and its old state: each node keeps
that under ./data/127.0.0.1-8080/, relative to where you started it.
examples/kvstore.rs is all of the above, and the node prints these commands on startup
with its own address filled in.
What just happened
A cluster is a set of nodes, each one a process running your application. Every node
keeps the same list of requests - the log - in the same order. A request is committed
once a majority of the nodes have stored it, which is the point past which it can no
longer be lost. Each node then applies committed requests to its own copy of your state,
in log order, which is why all of them end up with the same state. Your apply method is
that step.
One node is elected leader and is the only one that may accept writes; the rest follow it, and forward writes to it. Since every node holds the whole state, any node can answer a read out of memory without asking anyone. A snapshot is that state serialized: it brings a node that fell too far behind back in one transfer instead of a log replay, and it lets old log entries be deleted. EzRaft builds and installs snapshots on its own.
The rest is five facts about EzRaft in particular:
One node creates, the rest join. EzRaft::create starts a cluster whose only member
is that node; every other node calls EzRaft::join with the address of a node already in
it. Calling create on two nodes gives two one-node clusters that never merge.
Ids are assigned, not configured. The cluster gives a joining node its id and the node persists it. Ids come from a log index, so they are unique but not consecutive - a three-node cluster might end up as 0, 7 and 2. On restart the persisted id is reused and the seed is not contacted again, so a seed address that has since left is harmless.
Joining and serving is all it takes to add fault tolerance. A new node arrives as a
learner - it receives the log but does not count towards a majority - and EzRaft::join
asks the cluster to promote it to voter. That promotion completes only once the node has
caught up, which it can only do through its own HTTP server, so EzRaft::serve is where
the promotion is awaited and where a failed one is reported. A joining node must call it.
Three voters tolerate one failure and five tolerate two. An even count buys nothing: four
tolerate one failure, the same as three.
A node can stay a learner. EzRaft::join_as_learner joins without asking to be
promoted: the node replicates the whole log and answers reads, but never votes, so it adds
read capacity without adding a node that writes must wait for. EzRaft::promote, called on
the leader, makes it a voter whenever you decide to.
Writes go through the log, reads do not. EzRaft::write returns once the entry is
committed and applied, and only the leader accepts one - POST /api/write is what takes a
write on any node and forwards it.
EzRaft::read runs a closure over local applied state - cheap, and possibly a little
behind. When a read must see everything acknowledged before it, call
EzRaft::linearizable first.
Restart with the same call. A node that was created starts with create again, a node
that joined starts with join again; both read their state back from storage. The example
does exactly this - the command line for a restart is the one that started the node.
Write your own service
This is a complete program - one request type, one response type, one state struct, one method of business logic:
use BTreeMap;
use io;
use async_trait;
use EzApp;
use EzConfig;
use EzRaft;
use FileStorage;
use Deserialize;
use Serialize;
/// What a client asks the cluster to do. It travels the network and goes into
/// the log, hence serde; openraft prints it in its logs, hence Display.
/// What `apply` hands back to whoever issued the write.
/// The application *is* the replicated state. A snapshot is this struct
/// serialized, so serde is all it takes.
async
User Traits
EzApp
The application itself: the implementing type is the replicated state, and
apply is the business logic. Snapshots are derived from the state via serde,
so there is nothing to implement for them.
Framework handles: Sequential application, snapshot build/restore and scheduling
You handle: Business logic
Because a snapshot is the whole state serialized, the state has to fit in memory. That suits the coordination/metadata class of application (ZooKeeper snapshots the same way); an application whose snapshot is a streamed checkpoint of something larger builds on openraft directly.
EzStorage
Handles persistence of Raft state (metadata, logs, snapshots).
Framework handles: Raft logic, when to persist, what to persist
You handle: Serialization and I/O
Raft's safety rests on three assumptions the implementation must keep: persist returns
only once the data would survive a machine crash, operations become durable in the order
they arrive, and read_logs reflects everything persist has written, deletions included.
FileStorage implements this over JSON files in a directory, so a first cluster
runs without writing storage code. It does not fsync, so a machine crash can
lose a vote or a log entry the node already acknowledged - implement EzStorage
yourself for anything past a demo. Its source is the worked example of doing so.
The EzRaft API
| Method | What it does |
|---|---|
write(req) |
Runs req through the log on this node, returns what apply produced. Leader only: on a follower it fails, naming the leader. POST /api/write is the forwarding front door. |
read(|app| ...) |
Runs a closure over this node's applied state. No consensus round, no log entry. |
linearizable() |
Leader only: confirms leadership with a quorum round-trip and waits for local state to catch up, so a read after it sees every acknowledged write. |
metrics(), is_leader(), node_id(), addr() |
Current node and cluster state. |
promote(node_id) |
Makes that learner a voter. Returns once it is one, so the call lasts as long as bringing it up to date takes. Fails while another membership change is in flight - a cluster admits one at a time - and on an id that is not a learner. |
demote(node_id) |
Makes that voter a learner. It keeps receiving the log, it stops being counted in the quorum. Refuses to empty the voter set. |
remove_node(node_id) |
Takes a node out of the cluster, voter or learner. Does not stop the node's own process. |
serve() |
Runs the HTTP server, and collects the promotion a join started - so a joining node must call it, and a promotion that fails is returned from here. Blocks, so spawn it - tokio::spawn(raft.clone().serve()) - and start it early: peers reach a node only through this server. |
inner() |
The openraft Raft underneath, for anything EzRaft does not wrap. |
Only a leader accepts a write or a membership change, and none of these methods go looking
for one: on a follower they fail, naming the leader. Reaching it is the server's job, not a
node's - POST /api/write forwards, and the admin endpoints answer with the leader's
address. That keeps EzRaft a node, with no idea how to talk to another one.
Every fallible method returns std::io::Error, including the ones that fail for reasons
that have nothing to do with I/O: a caller cannot usefully branch on the difference, since
write already finds the leader on its own. Code that does need typed errors reaches them
through inner().
Configuration
EzConfig provides sensible defaults for Raft timing parameters:
Election timeout is automatically calculated as 3-6x the heartbeat interval,
and the log-purge distance is derived from snapshot_interval.
Most users can use EzConfig::default().
HTTP API
EzRaft includes built-in HTTP endpoints:
- Raft RPC (
/raft/*): Internal consensus communication - Admin API (
/api/node_id,/api/membership,/api/metrics):node_idhands out an id,membershipadds a node, changes its role or removes it - tagged byop, so{"op": "Add", "node_id": 2, "addr": "..."},{"op": "SetRole", "node_id": 2, "role": "Voter"}or{"op": "Remove", "node_id": 2}. Joining is an id then two membership changes: enter as a learner, then ask for the voter role. One endpoint for all three because they are one decision made in stages - whether a node is a member, and whether it counts towards a quorum. Each answers either the result or the address of the leader to ask instead. - Application API (
POST /api/write): Propose client requests, using your request type as JSON - Application read (
POST /api/read): A read answered by the app'sreadmethod from local memory, without a consensus round. Body and answer are the app's own read types, so a read is not confined to what a key-and-value write API can phrase -examples/kvstore.rsserves both a keyed lookup and a prefix scan through it. What "nothing found" means is the app's to define; the framework does not interpret the answer.
ezraft::admin::AdminClient is the admin API from Rust rather than from curl: it follows the redirect to the leader and retries what a cluster that is still starting up fails at, which is how EzRaft::join joins.
The two application endpoints are a sample, not the API a real service exposes. Copy EzServer and replace them with your own routes, built on EzRaft::write and EzRaft::read. Nothing here is authenticated or encrypted, so keep the addresses on a trusted network.
Troubleshooting
Two one-node clusters instead of one. Every node after the first must join. Check
GET /api/metrics: membership_config.membership.configs must list every node.
A node comes back as a stranger. Its data directory holds its id - delete it and the node asks for a new one, or, if it was the founding node, starts a fresh single-node cluster of its own. Wipe data directories only when starting a whole cluster over.
One directory per node. Two nodes sharing one overwrite each other's log.
Writes fail after ~10 seconds. No leader: a cluster that has lost its majority cannot elect one. Bring enough nodes back.
Nothing in the logs. Raft reports through tracing, which goes nowhere until a
subscriber is installed - the example installs one, then RUST_LOG=info (or debug) shows
what Raft is doing.
Why EzRaft exists
API design exploration. EzRaft turns abstract ideas about "intuitive APIs" into concrete code. By testing different patterns—parameter organization, naming conventions, simplicity vs extensibility trade-offs—we discover what truly matches user intuition. These insights will guide future OpenRaft improvements.
Fast prototyping. As a secondary benefit, EzRaft lets beginners build working prototypes without understanding Raft internals or OpenRaft's architecture.
| Aspect | OpenRaft | EzRaft |
|---|---|---|
| Required traits | 7+ (RaftLogStorage, RaftStateMachine, etc.) | 2 (EzStorage, EzApp) |
| Required methods | 21+ | 4 |
| User-defined types | 12 (all generic parameters) | 2 (Request, Response) |
| Network code | User implements (~100 lines) | Built-in (0 lines) |
| Example complexity | ~400 lines | ~145 lines, plus 40 of startup help |
Build on OpenRaft directly when the state outgrows memory, when the transport has to be
something other than plain HTTP, or when the types EzRaft fixes - u64 node ids, an
address per node, standard one-leader-per-term voting - are not the ones you need.
License
Licensed under either of Apache-2.0 or MIT, at your option.