Polars
The fast DataFrame library for Rust - lazy query optimization, vectorized execution, and Arrow-backed memory layout.
Busca en todas las páginas de la documentación
The fast DataFrame library for Rust - lazy query optimization, vectorized execution, and Arrow-backed memory layout.
Quick-reference recipe card - copy-paste ready.
use polars::prelude::*;
fn top_spenders(path: &str) -> PolarsResult<DataFrame> {
LazyCsvReader::new(path.into())
.with_has_header(true)
.finish()?
.group_by([col("customer_id")])
.agg([col("amount").sum().alias("total")])
.sort(["total"], SortMultipleOptions::default())
.limit(10)
.collect()
}When to reach for this:
use polars::prelude::*;
fn main() -> PolarsResult<()> {
let df = df! {
"region" => ["west", "east", "west", "north"],
"amount" => [120i64, 80, 200, 50],
}?;
let summary = df
.lazy()
.filter(col("amount").gt(lit(60)))
.group_by([col("region")])
.agg([
col("amount").sum().alias("total"),
col("amount").mean().alias("avg"),
])
.sort(["total"], SortMultipleOptions::default())
.collect()?;
println!("{}", summary);
Ok(())
}What this demonstrates:
.lazy() for optimized executioncollect() executes the plan - that is your memory and CPU spike boundary.DataFrame methods run immediately - simpler mental model, fewer optimizations.| Mode | API | Best for |
|---|---|---|
| Eager | DataFrame::filter, group_by | Unit tests, tiny tables |
| Lazy | LazyFrame + collect() | Large files, multi-step pipelines |
// Prefer explicit aliases in agg - unnamed exprs are hard to spot in explain plans.
.col("amount").sum().alias("total")
// Use .explain() on lazy plans during development.
println!("{}", lf.explain(true)?);collect() after every step defeats lazy optimization. Fix: chain transforms, collect once at the end.with_dtypes or cast after load.i32 vs i64 keys silently drop rows. Fix: cast both sides to the same type before join.id column when you need stable row identity.select * on 500 columns loads everything. Fix: project only needed columns in the lazy scan.| Alternative | Use When | Don't Use When |
|---|---|---|
| DataFusion SQL | Team thinks in SQL over registered tables | You want a Rust expression DSL in application code |
| pandas (Python) | Notebook exploration with a huge ecosystem | Production ETL needing single-binary deploy |
| DuckDB embedded | SQL-first analytics on files | You are standardizing on pure Rust dependencies |
| ndarray + manual loops | Tiny numeric kernels | Relational transforms and I/O heavy pipelines |
Default to lazy for file-backed pipelines. Use eager for tests, REPL-style exploration, and frames that already fit in memory.
Call .explain(true) on the LazyFrame before collect() to see predicate pushdown and projection pruning.
Yes. Many operations parallelize across columns and row groups. Control thread count with POLARS_MAX_THREADS when colocating services.
Use scan_parquet / scan_csv with compatible paths and feature flags for cloud backends, or stage files locally for simpler ops.
Use fill_null, drop_nulls, or coalesce expressions. Nullable dtypes must match schema expectations on write.
Stay on the expression API (col, lit, when) documented for 0.46+. Older lazy::dsl patterns may differ - pin versions in CI.
Convert to Arrow RecordBatch and hand off via PyArrow - see Interop with Python/pandas.
For bounded memory, combine chunked CSV reads or scan Parquet row groups - see Streaming & Large Data.
Use ParquetWriter or CsvWriter on eager frames after the final collect().
Register Polars output as Arrow tables in DataFusion for SQL layers - see DataFusion.
Stack versions: This page was written for Rust 1.97.0 (edition 2024), Tokio 1.x, Axum 0.8, serde 1.0, sqlx 0.8, clap 4, and Polars 0.46+.
Revisado por Chris St. John·Última actualización: 16 jul 2026