1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
use Arc;
use async_trait;
use Mutex;
/// A trait that defines how batches of inference requests are processed by a model.
///
/// The `BatchHandler` trait abstracts the core operations required for batched inference
/// processing across different inference patterns (such as feedforward and autoregressive).
/// Implementations handle request batching, model execution, and output distribution,
/// while managing the state of active requests.
///
/// # Type Parameters
///
/// * `Request` - The type representing a single inference request.
/// * `ModelInput` - The type representing the batched input to the model, must be cloneable.
/// * `ModelOutput` - The type representing the output from the model, must be cloneable.
///