CSV/Parquet I/O
Read and write CSV and Parquet in Rust with predictable schemas, compression, and scan performance.
Busca en todas las páginas de la documentación
Read and write CSV and Parquet in Rust with predictable schemas, compression, and scan performance.
Quick-reference recipe card - copy-paste ready.
use polars::prelude::*;
fn csv_to_parquet(csv_path: &str, parquet_path: &str) -> PolarsResult<()> {
let mut df = CsvReadOptions::default()
.with_has_header(true)
.try_into_reader_with_file_path(Some(csv_path.into()))?
.finish()?;
ParquetWriter::new(std::fs::File::create(parquet_path)?)
.with_compression(ParquetCompression::Zstd(None))
.finish(&mut df)?;
Ok(())
}When to reach for this:
use polars::prelude::*;
fn main() -> PolarsResult<()> {
let schema = Schema::from_iter([
Field::new("order_id".into(), DataType::Int64),
Field::new("region".into(), DataType::String),
Field::new("amount".into(), DataType::Float64),
]);
let df = CsvReadOptions::default()
.with_has_header(true)
.with_schema(Some(Arc::new(schema)))
.try_into_reader_with_file_path(Some("orders.csv".into()))?
.finish()?;
let filtered = df
.lazy()
.filter(col("amount").gt(lit(0.0)))
.collect()?;
let mut out = filtered.clone();
ParquetWriter::new(std::fs::File::create("orders_clean.parquet")?)
.with_row_group_size(Some(128_000))
.finish(&mut out)?;
Ok(())
}What this demonstrates:
| Format | Strength | Watch out for |
|---|---|---|
| CSV | Universal export | No schema embedded; parsing cost |
| Parquet | Analytics, compression | Not human-readable; schema evolution discipline |
// Lazy scan avoids loading full CSV:
LazyCsvReader::new("big.csv".into()).with_has_header(true).finish()?;
// Read subset of Parquet columns:
LazyFrame::scan_parquet("big.parquet", ScanArgsParquet::default())?
.select([col("user_id"), col("event_time")]);1,5 vs 1.5 breaks inference. Fix: normalize files or set decimal separator in options.utf8-lossy preprocessing.| Alternative | Use When | Don't Use When |
|---|---|---|
| JSON Lines | Semi-structured event logs | Heavy numeric aggregation |
| Arrow IPC | In-process zero-copy handoff | Long-term archival |
| Avro | Schema evolution in Kafka pipelines | Interactive BI queries |
| SQLite | Small transactional slices | TB-scale columnar scans |
ZSTD balances ratio and speed for analytics. Snappy is faster, larger files. Benchmark on your hardware and read patterns.
Use with_ignore_errors cautiously, log rejects to a dead-letter file, and count dropped rows in metrics.
Yes - lazy scan_parquet with column projection avoids decoding unused fields.
Write region=west/data.parquet hive-style directories. Consumers filter paths by partition keys before scan.
Store timestamps as UTC with metadata notes. Convert on read for display time zones.
Use chunked readers beyond RAM size - see Streaming & Large Data.
Gzip CSV is fine for archive; prefer Parquet for anything queried more than once.
Assert dtypes on the DataFrame before ParquetWriter::finish and reject writes that drift from a versioned schema registry.
Yes - Parquet is cross-language. Keep Arrow-compatible logical types for smooth handoff.
Convert to CSV in a preprocessing step or use a dedicated XLSX crate - not ideal for production analytics pipelines.
Stack versions: This page was written for Rust 1.97.0 (edition 2024), Tokio 1.x, Axum 0.8, serde 1.0, sqlx 0.8, clap 4, and Polars 0.46+.
Revisado por Chris St. John·Última actualización: 16 jul 2026