1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
//! Sparse Attention Masks
//!
//! This module implements prompt-aware attention head masking to skip
//! computation for heads that don't contribute to the output.
//!
//! # Key Insight
//!
//! Not all attention heads are equally important for every prompt. Portrait
//! prompts activate face-focused heads while landscape prompts activate
//! background/composition heads. By predicting which heads matter, we can
//! skip 50-70% of attention computation.
//!
//! # Architecture
//!
//! ```text
//! ┌─────────────────────────────────────────────────────────────────┐
//! │ Sparse Attention │
//! ├─────────────────────────────────────────────────────────────────┤
//! │ │
//! │ Standard Attention (32 heads × 64 layers = 2048 computations) │
//! │ ════════════════════════════════════════════════════════════ │
//! │ [████████████████████████████████] 100% compute │
//! │ │
//! │ Sparse Attention (prompt-aware masking) │
//! │ ════════════════════════════════════════════════════════════ │
//! │ "Portrait of a woman" │
//! │ [████████░░░░░░░░████░░░░░░░░░░░░] 35% compute │
//! │ ↑ face ↑ skip ↑ style │
//! │ │
//! │ "Mountain landscape at sunset" │
//! │ [░░░░░░░░████████████████████░░░░] 45% compute │
//! │ ↑ skip ↑ background/lighting ↑ skip │
//! └─────────────────────────────────────────────────────────────────┘
//! ```
pub use ;
pub use ;
pub use ;
pub use ;
pub use ;
pub use ;
/// Default sparsity target (fraction of heads to skip)
pub const DEFAULT_SPARSITY: f32 = 0.5;
/// Minimum heads to keep active per layer
pub const MIN_ACTIVE_HEADS: usize = 4;
/// Maximum quality degradation allowed
pub const MAX_QUALITY_LOSS: f32 = 0.02;
/// Prelude for common imports