UTF-8 & Indexing
Rust str is UTF-8. Indexing s[i] is byte-based and panics on invalid char boundaries.
Busca en todas las páginas de la documentación
Rust str is UTF-8. Indexing s[i] is byte-based and panics on invalid char boundaries.
fn safe_prefix(s: &str, max_bytes: usize) -> &str {
if s.len() <= max_bytes { return s; }
let mut end = max_bytes;
while !s.is_char_boundary(end) { end -= 1; }
&s[..end]
}When to reach for this: Truncating UI strings, parsing protocols, any substring by length.
fn main() {
let s = "Hello 🦀";
// let bad = &s[0..7]; // may panic
for (i, ch) in s.char_indices() {
println!("{i}: {ch}");
}
}What this demonstrates:
char_indices gives byte index at char boundaryis_char_boundarylen() returns bytes. Use chars(), bytes(), or graphemes (unicode-segmentation crate) for user-perceived characters.
s.chars().nth(n).is_char_boundary.| Alternative | Use When | Don't Use When |
|---|---|---|
chars().nth | Char by index | Need O(1) random char |
unicode-segmentation | User graphemes | ASCII-only internal |
Vec<char> | Repeated char index | Memory overhead ok |
Stack versions: This page was written for Rust 1.97.0 (edition 2024), Tokio 1.x, Axum 0.8, serde 1.0, sqlx 0.8, clap 4, and Polars 0.46+.
Revisado por Chris St. John·Última actualización: 16 jul 2026