candle
Hugging Face's minimalist deep-learning framework in Rust - tensors, nn modules, and transformer inference without Python.
Search across all documentation pages
Hugging Face's minimalist deep-learning framework in Rust - tensors, nn modules, and transformer inference without Python.
use candle_core::{Device, Tensor};
use candle_nn::{linear, Linear, Module, VarBuilder};
fn mlp(vb: VarBuilder) -> candle_core::Result<Linear> {
linear(768, 128, vb.pp("proj"))
}
fn forward(x: &Tensor, layer: &Linear) -> candle_core::Result<Tensor> {
layer.forward(x)
}When to reach for this:
use candle_core::{Device, Tensor};
use candle_nn::{linear, Linear, Module, VarBuilder, VarMap};
fn main() -> candle_core::Result<()> {
let device = Device::Cpu;
let varmap = VarMap::new();
let vb = VarBuilder::from_varmap(&varmap, candle_core::DType::F32, &device);
let layer = linear(4, 2, vb.pp("fc"))?;
let x = Tensor::new(&[[1.0f32, 0.5, -0.2, 3.0]], &device)?;
let y = layer.forward(&x)?;
println!("output shape: {:?}", y.shape());
Ok(())
}What this demonstrates:
VarMap + VarBuilder for learnable parameterscandle-core provides Tensor, Device, and autodiff (feature-gated).candle-nn bundles layers, optimizers, and loss helpers.candle-transformers hosts model definitions matching HF configs.| Crate | Role |
|---|---|
candle-core | Tensor math, devices |
candle-nn | Layers, training utilities |
candle-transformers | BERT, Llama, etc. |
Cargo.lock and CI against HF examples.F32 unless model requires BF16.[batch, seq, hidden] confusion - attention fails silently with broadcast bugs. Fix: assert ranks at module boundaries.Device::cuda_if_available() at startup and log choice.| Alternative | Use When | Don't Use When |
|---|---|---|
| burn | Backend-agnostic training research | You want HF-maintained transformer ports |
| ort | Exported ONNX from any framework | You need training in Rust |
| tch (PyTorch bindings) | Must run arbitrary PyTorch graphs | You want pure Rust deps |
| Python HF | Largest model zoo and trainers | Single-binary edge deploy |
Possible for small/medium models; large LLM training still usually happens in Python, then weights convert for candle inference.
Use candle-transformers model builders plus safetensors files from the model repo - match config.json architecture names.
Enable cuda feature on candle-core and use Device::new_cuda(0)? when hardware exists.
GGUF and quantized checkpoints are common for LLM inference paths - follow HF candle examples for each model family.
Enable gradient clipping, lower learning rate, and inspect tensor stats after each layer in training loops.
Training in candle may still export via separate tooling; many teams train elsewhere and infer in candle from published weights.
candle is HF-centric for transformers; burn is general-purpose with pluggable backends - see burn.
Fine for research-scale training; verify numerics against PyTorch on a golden batch before trusting loss curves.
Pad token sequences, build attention masks, run one forward - pair with Tokenizers.
Share read-only weights with Arc; per-request state (KV cache) should not race across threads without synchronization.
Stack versions: This page was written for Rust 1.97.0 (edition 2024), Tokio 1.x, Axum 0.8, serde 1.0, sqlx 0.8, clap 4, and Polars 0.46+.
Reviewed by Chris St. John·Last updated Jul 16, 2026