Data Basics
9 examples to get you started with high-performance data in Rust - 7 basic and 2 intermediate.
Search across all documentation pages
9 examples to get you started with high-performance data in Rust - 7 basic and 2 intermediate.
cargo new data-playground && cd data-playground
cargo add polars --features lazy,csv,parquet
cargo add arrow
cargo add clap --features deriveTooling: Examples target Rust 1.97.0 (edition 2024) and Polars 0.46+.
Load a small file and inspect schema before any transforms.
use polars::prelude::*;
fn main() -> PolarsResult<()> {
let df = CsvReadOptions::default()
.with_has_header(true)
.try_into_reader_with_file_path(Some("sales.csv".into()))?
.finish()?;
println!("{}", df.head(Some(3)));
Ok(())
}CsvReadOptions centralizes delimiter, dtypes, and null handling.Related: CSV/Parquet I/O - format-specific tuning
Push predicates down so Polars reads less data from disk.
use polars::prelude::*;
fn main() -> PolarsResult<()> {
let q = LazyCsvReader::new("events.csv".into())
.with_has_header(true)
.finish()?
.filter(col("amount").gt(lit(100)))
.select([col("user_id"), col("amount")]);
let df = q.collect()?;
println!("rows: {}", df.height());
Ok(())
}collect() is the boundary where memory is allocated for the result.Related: Polars - lazy vs eager trade-offs
Summarize metrics without manual hash maps.
use polars::prelude::*;
fn revenue_by_region(df: LazyFrame) -> PolarsResult<LazyFrame> {
Ok(df
.group_by([col("region")])
.agg([col("revenue").sum().alias("total_revenue")])
.sort(["total_revenue"], SortMultipleOptions::default()))
}group_by + agg maps cleanly to SQL GROUP BY.Relational-style merges without leaving Rust.
use polars::prelude::*;
fn join_users(orders: LazyFrame, users: LazyFrame) -> LazyFrame {
orders.join(
users,
[col("user_id")],
[col("id")],
JoinArgs::new(JoinType::Inner),
)
}Columnar output compresses well and preserves types.
use polars::prelude::*;
fn write_parquet(df: &mut DataFrame, path: &str) -> PolarsResult<()> {
let file = std::fs::File::create(path)?;
ParquetWriter::new(file).finish(df)?;
Ok(())
}Related: CSV/Parquet I/O - compression and schema
Bound memory when the file exceeds RAM.
use polars::prelude::*;
use std::fs::File;
fn process_chunks(path: &str, chunk_rows: usize) -> PolarsResult<()> {
let file = File::open(path)?;
let reader = CsvReader::new(file).with_chunk_size(chunk_rows);
for chunk in reader.into_iter() {
let df = chunk?;
println!("chunk rows: {}", df.height());
}
Ok(())
}Related: Streaming & Large Data - backpressure patterns
Ship ETL as a single binary with typed flags.
use clap::Parser;
use polars::prelude::*;
#[derive(Parser)]
struct Cli {
#[arg(long)]
input: String,
#[arg(long, default_value = "out.parquet")]
output: String,
}
fn main() -> PolarsResult<()> {
let cli = Cli::parse();
let mut df = CsvReadOptions::default()
.try_into_reader_with_file_path(Some(cli.input.into()))?
.finish()?;
ParquetWriter::new(std::fs::File::create(&cli.output)?).finish(&mut df)?;
Ok(())
}clap derive gives help text and validation for free.PolarsResult from main for readable error chains.--filter or --columns flags as your pipeline grows.Related: Building Data CLIs & Services - production CLI layout
Share columnar memory without serializing row-by-row JSON.
use arrow::array::{Int32Array, StringArray};
use arrow::record_batch::RecordBatch;
use std::sync::Arc;
fn sample_batch() -> RecordBatch {
let ids = Int32Array::from(vec![1, 2, 3]);
let names = StringArray::from(vec!["alpha", "beta", "gamma"]);
RecordBatch::try_from_iter(vec![
("id", Arc::new(ids) as _),
("name", Arc::new(names) as _),
])
.expect("valid batch")
}Related: Apache Arrow - zero-copy semantics
Query in-memory Arrow tables without a separate database server.
use datafusion::prelude::*;
use datafusion::error::Result;
#[tokio::main]
async fn main() -> Result<()> {
let ctx = SessionContext::new();
ctx.register_csv("orders", "orders.csv", CsvReadOptions::default()).await?;
let df = ctx.sql("SELECT region, SUM(amount) AS total FROM orders GROUP BY region").await?;
df.show().await?;
Ok(())
}Related: DataFusion - custom UDFs and optimizers
Stack versions: This page was written for Rust 1.97.0 (edition 2024), Tokio 1.x, Axum 0.8, serde 1.0, sqlx 0.8, clap 4, and Polars 0.46+.
Reviewed by Chris St. John·Last updated Jul 16, 2026