ML in Rust Best Practices
Reproducible, fast, memory-safe ML inference and serving in Rust.
Busca en todas las páginas de la documentación
Reproducible, fast, memory-safe ML inference and serving in Rust.
model_version field. Clients know when behavior changed.max_tokens and context length server-side. Client params are hints, not policy.spawn_blocking or dedicated pool for forwards. Never block Tokio worker threads on matmul./predict and generation routes. GPUs are expensive to share with the internet.Running synchronous model forward inside async Axum handlers without spawn_blocking, stalling unrelated routes.
At least one fixed input per model output head - classification logits, embedding vector, or first-token logit.
Fine for small models and research; large LLM training usually stays Python with Rust owning inference.
ort for exported ONNX from any framework; candle for HF Rust ports and tight Rust-only stacks.
VRAM at load, per-request KV growth, and max concurrent sessions in runbook tied to hardware SKU.
Treat tokenizer.json changes like model changes - rerun golden tests and bump model_version.
Version chunking policy and embedding model together; changing one without the other hurts recall.
CPU EP only in PR pipeline; nightly GPU job for perf regression on larger fixtures.
Result through blocking tasks; map to HTTP status without leaking tensor shape errors to clients.
See Serving Models with Axum for handler layout.
Stack versions: This page was written for Rust 1.97.0 (edition 2024), Tokio 1.x, Axum 0.8, serde 1.0, sqlx 0.8, clap 4, and Polars 0.46+.
Revisado por Chris St. John·Última actualización: 16 jul 2026