Expand description
Mixtral Model V2 - Clean implementation using solid abstractions
This implements the Mixtral architecture which features:
- Mixture of Experts (MoE) with 8 experts, top-2 routing
- Sliding window attention (from Mistral)
- Grouped Query Attention (GQA)
- Uses unified Tensor type from tensor_core
- Implements Model trait from model_core
Structsยง
- Mixtral
Attention - Mixtral attention mechanism with sliding window
- Mixtral
Config - Mixtral
Expert - Single expert in Mixtral MoE
- Mixtral
Layer - Mixtral transformer layer with MoE
- Mixtral
MoE - Mixtral Mixture of Experts layer
- Mixtral
Model V2 - Main Mixtral model implementation