Performance Best Practices
Optimize the measured hot path only. Rust gives zero-cost abstractions, but discipline on allocations, profiling, and release settings still wins in production.
Search across all documentation pages
Optimize the measured hot path only. Rust gives zero-cost abstractions, but discipline on allocations, profiling, and release settings still wins in production.
&str, &[T], Cow at boundaries.Vec/String with with_capacity. Known sizes from input metadata.Box<dyn> and dynamic dispatch in hot paths. Use enums or generics.spawn_blocking for CPU/sync IO.tracing spans on handler boundaries. Find slow awaits in production.opt-level = 3, thin LTO, codegen-units = 1 for servers. Tune with data.target-cpu to deployment baseline. Not native unless uniform hardware.cargo bloat in release checklist.collect(). Polars 0.46+ vectorizes when possible.#[serde(borrow)] on &str fields.HashMap vs BTreeMap deliberately. Cache and ordering tradeoffs.Known anti-patterns (clone per row in tight loop) without profile; still verify after fix.
Yes: before/after numbers or "no hot path change" checkbox.
Author of change unless platform provides profiling support.
Optional on dedicated runner; store baselines for critical crates.
Use Polars for analytics; drop to Rust loops only for custom ops Polars lacks.
Ticket with profile attached; prioritize by SLO impact.
Document choice; servers often trade MB for latency if RAM is cheap.
Yes for user-facing APIs; Rust speed does not fix architectural N+1.
Mirror prod data volume anonymized; synthetic micro-bench insufficient alone.
When SLO met with headroom; invest in observability instead of micro-opts.
Stack versions: This page was written for Rust 1.97.0 (edition 2024), Tokio 1.x, Axum 0.8, serde 1.0, sqlx 0.8, clap 4, and Polars 0.46+.
Reviewed by Chris St. John·Last updated Jul 16, 2026