ONNX Runtime (ort)
Run exported ONNX models in Rust with ONNX Runtime - the standard path from Python training to native inference.
Search across all documentation pages
Run exported ONNX models in Rust with ONNX Runtime - the standard path from Python training to native inference.
use ndarray::Array;
use ort::{inputs, session::Session, value::Tensor};
fn predict(session: &Session, flat: Vec<f32>) -> ort::Result<f32> {
let input = Array::from_shape_vec((1, flat.len()), flat)?;
let outputs = session.run(inputs![Tensor::from_array(input)?])?;
let view = outputs[0].try_extract_tensor::<f32>()?;
Ok(view[0])
}When to reach for this:
use ndarray::Array;
use ort::{inputs, session::Session, value::Tensor};
fn main() -> ort::Result<()> {
let session = Session::builder()?.commit_from_file("model.onnx")?;
let features = vec![0.1f32, 0.2, 0.3, 0.4];
let input = Array::from_shape_vec((1, 4), features)?;
let outputs = session.run(inputs![Tensor::from_array(input)?])?;
let tensor = outputs[0].try_extract_tensor::<f32>()?;
println!("score: {}", tensor[0]);
Ok(())
}What this demonstrates:
Session::run maps Rust tensors to ORT memory, executes providers, returns outputs.| Step | Action |
|---|---|
| Train | PyTorch/sklearn/etc. |
| Export | torch.onnx.export or framework tool |
| Validate | onnx.checker + sample inference in Rust |
input_ids not input. Fix: print session.inputs and match strings.Arc<Session> singleton at process start.| Alternative | Use When | Don't Use When |
|---|---|---|
| candle | Native Rust transformer without ONNX | Model only exists as ONNX today |
| burn import | Stay in burn ecosystem | Graph uses unsupported ONNX ops |
| Python microservice | Rapid iteration | Latency/ops cost of sidecar |
| TensorRT directly | NVIDIA-only max perf | Need portable CPU fallback |
Build ort with CUDA features and register CUDA execution provider in Session::builder() - driver must exist in container.
Yes - one Session per model, share across threads with Arc if ORT thread safety docs allow for your EP.
Set batch dimension in input tensor shape at runtime if export used dynamic axis annotations.
Most ONNX classifiers expect numeric tensors - tokenize text before ORT, not inside generic ORT APIs.
ORT profiling flags, smaller graph via onnx-simplifier, and GPU EP when batch size warrants it.
Match ort supported opset; re-export from training framework if import fails.
Stack rows in batch dimension when graph supports dynamic batch; pad to max length for sequences.
Treat ONNX like code - verify checksums and sign artifacts in object storage.
Check in tiny ONNX fixture and assert output tensor within epsilon of Python reference.
See Serving Models with Axum for HTTP wiring.
Stack versions: This page was written for Rust 1.97.0 (edition 2024), Tokio 1.x, Axum 0.8, serde 1.0, sqlx 0.8, clap 4, and Polars 0.46+.
Reviewed by Chris St. John·Last updated Jul 19, 2026