virtio-accel-coreml
A host-native virtio_accel_core::Accelerator
implementation for Apple's Core ML runtime and Neural Engine.
The backend requires macOS 14 or newer and refuses construction when Core ML does not report an
accessible MLNeuralEngineComputeDevice. Models are loaded with
MLComputeUnitsCPUAndNeuralEngine, which gives Core ML access to the ANE without permitting GPU
placement. Apple still decides placement per operation; a model containing unsupported ANE
operations can fall back to the CPU.
Portability tier: host-native — the real implementation is macOS-only. Other targets compile
a placeholder constructor that returns InitError::UnsupportedPlatform, keeping workspace and
cross-target dependency checks intact.
Model artifacts
CoreMlAccelerator::new(model_root) establishes a host-controlled root. A CoreMlArtifact names a
.mlmodel file, .mlpackage directory, or .mlmodelc directory beneath that root and maps every
model input/output feature to a virtio-accel slot. Absolute paths, parent traversal, symlink escape,
unmapped features, optional features, non-MLMultiArray features, and incompatible aliased layouts
are rejected.
#
#
Artifact ABI v1 supports nonoptional MLMultiArray features with Float16, Float32, Float64,
or Int32 elements. Each feature uses the model constraint's declared default shape; alternate
flexible shapes are not selected by the artifact. Image, sequence, dictionary, scalar, optional,
and Core ML state features are rejected at model load rather than failing after admission. The
model's default function is used for multi-function assets.
Source .mlmodel files and .mlpackage directories are compiled synchronously during
load_program; compiled .mlmodelc directories load directly and are recommended for predictable
startup latency. Fixed-shape models receive Core ML's infrequent-reshape hint on macOS 14.4+ and the
fast-prediction specialization strategy on macOS 15+. Core ML does not publish a finite
model-residency ceiling, so artifacts must declare REQUIRED_RESIDENT_BYTES (u64::MAX). This
deliberately forces the device's aggregate resident-program policy to opt into one Core ML model
instead of pretending an unverifiable smaller charge is exact.
Direct buffers and events
The backend advertises host and shared memory. Both are page-aligned provider allocations. Program
bindings wrap the exact bound range in MLMultiArray, and outputs use MLPredictionOptions output
backings. Completion verifies the returned output's data pointer, element type, shape, and strides;
a different Objective-C wrapper over the same exact storage remains valid, while a provider-side
result allocation fails with BackendError::Incompatible. Binding offsets must be aligned for the
model's scalar type.
Prediction uses Core ML's asynchronous completion API. Events retain every Rust allocation until
the native callback reaches a terminal state. Separate predictions may reuse a read-only input
allocation concurrently; any output or read-write binding retains exclusive native access. Host
transfers return BackendError::Busy while either access mode is active. Event cancellation is not
advertised because Core ML does not expose cancellation for an admitted prediction.
Submission sorts its bounded binding metadata once (O(b log b)) and the native bridge then performs
a linear validation and wrapping pass. It never copies tensor contents. Cumulative direct-binding
admissions and explicit-transfer bytes are available through direct_binding_admissions() and
explicit_transfer_bytes().
The FFI and allocation invariants are documented in SAFETY.md. Run the native end-to-end and conformance tests on an ANE-capable Mac with:
For local warm-path latency evidence, run the ignored release-mode measurement:
License
Licensed under either of Apache License, Version 2.0 or MIT license at your option.