1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
//! Core traits and data types shared by the `r2l` workspace.
//!
//! `r2l-core` is the contracts crate. It defines the small set of interfaces
//! that environments, samplers, policies, agents, learners, and tensor
//! backends agree on. Backend-specific implementations live in crates such as
//! `r2l-burn` and `r2l-candle`; concrete algorithms and builders live outside
//! this crate as well.
//!
//! Most downstream code should start with the prelude:
//!
//! ```
//! use r2l_core::prelude::*;
//! ```
//!
//! The main extension points are:
//!
//! - [`Env`] and [`EnvBuilder`] for environment integrations.
//! - [`R2lTensor`] for tensor types used by environments
//! and learning code.
//! - [`Actor`], [`Policy`], [`ValueFunction`], and [`Learner`] for model
//! and optimizer components.
//! - [`TrajectoryBuffer`] and [`TrajectoryView`] for rollout storage.
//! - [`Agent`], [`Sampler`], and [`OnPolicyAlgorithm`] for on-policy training
//! loops.
//!
//! [`Actor`]: crate::models::Actor
//! [`Agent`]: crate::on_policy::algorithm::Agent
//! [`Env`]: crate::env::Env
//! [`EnvBuilder`]: crate::env::EnvBuilder
//! [`Learner`]: crate::models::Learner
//! [`OnPolicyAlgorithm`]: crate::on_policy::algorithm::OnPolicyAlgorithm
//! [`Policy`]: crate::models::Policy
//! [`R2lTensor`]: crate::tensor::R2lTensor
//! [`Sampler`]: crate::on_policy::algorithm::Sampler
//! [`TrajectoryBuffer`]: crate::buffers::buffer::TrajectoryBuffer
//! [`TrajectoryView`]: crate::buffers::buffer::TrajectoryView
//! [`ValueFunction`]: crate::models::ValueFunction
/// Rollout transition and trajectory storage.
/// Environment traits and space descriptions.
/// Error types
/// Actor, policy, value-function, and learner traits.
/// Shared interfaces for on-policy training loops.
/// Reproducible random-number utilities.
/// Online mean and variance estimators.
/// Backend-neutral tensor interfaces and adapters.
pub use ActorWrapper;
/// Control-flow result returned by training hooks.
///
/// Hook implementations use this to signal whether the surrounding training
/// loop should continue or stop at the current hook boundary.
/// Breaks out of the surrounding loop when a hook requests [`HookResult::Break`].
/// Returns `Ok(())` from the surrounding function when a hook requests
/// [`HookResult::Break`].
/// Common imports for implementing environments, policies, agents, samplers,
/// and learners.