Alvaro Bartolome provides a Rust-based implementation of the System One compatible API, designed specifically for open decision models such as Laya. The project features dynamic token-based batching and supports hardware acceleration via CPU, CUDA, and Metal (MPS). It is built using modern asynchronous frameworks like tokio and axum to provide high performance for model queries.
- Achieves approximately 14ms latency per query on an NVIDIA RTX Pro 6000.
- Includes support for ModernBert with custom decision heads for Laya models.
- Utilizes the Candle machine learning framework by Hugging Face.
- Supports multiple installation features via cargo, including specific flags for metal or cuda.