Interop with Python/pandas
Move tables between Rust and Python with Apache Arrow - avoid slow JSON serialization and preserve dtypes.
Search across all documentation pages
Move tables between Rust and Python with Apache Arrow - avoid slow JSON serialization and preserve dtypes.
Quick-reference recipe card - copy-paste ready.
use polars::prelude::*;
use std::fs::File;
fn export_arrow_ipc(df: &mut DataFrame, path: &str) -> PolarsResult<()> {
let mut file = File::create(path)?;
IpcWriter::new(&mut file).finish(df)?;
Ok(())
}import pyarrow as pa
import pyarrow.ipc as ipc
with pa.memory_map("/tmp/frame.arrow", "r") as source:
table = ipc.open_file(source).read_all()
df = table.to_pandas()When to reach for this:
Rust side writes IPC; Python reads into pandas without parsing CSV.
use polars::prelude::*;
fn main() -> PolarsResult<()> {
let mut df = df! {
"user_id" => [1i64, 2, 3],
"score" => [0.9, 0.4, 0.7],
}?;
let mut f = std::fs::File::create("/tmp/users.arrow")?;
IpcWriter::new(&mut f).finish(&mut df)?;
Ok(())
}import pandas as pd
import pyarrow.ipc as ipc
with open("/tmp/users.arrow", "rb") as f:
reader = ipc.open_file(f)
table = reader.read_all()
pdf = table.to_pandas(types_mapper=pd.ArrowDtype)
assert pdf["user_id"].dtype.name.startswith("int")What this demonstrates:
IpcWriter producing a file Python understandsPyArrow objects in-process (zero IPC file).pd.ArrowDtype for nullable integers.| Path | Copies | Best for |
|---|---|---|
| Arrow IPC file | Minimal with mmap | Batch handoff between jobs |
| PyO3 + pyarrow | In-process | Hot loops in Python calling Rust |
| Parquet on disk | Decode cost | Durable shared lake between teams |
| CSV | Full parse | Debugging only |
// When exposing PyO3, use maturin to build wheels:
// #[pyfunction] fn transform(py: Python, table: Bound<PyAny>) -> PyResult<PyObject> { ... }py.allow_threads around Rust compute.Cargo.lock and requirements.txt.| Alternative | Use When | Don't Use When |
|---|---|---|
| Parquet lake only | Async batch between teams | Interactive sub-second notebook loops |
| gRPC + Arrow Flight | Service-to-service streaming | Simple local CLI handoff |
| DuckDB in both | SQL contract over files | Heavy custom Rust kernels |
| Pure Python Polars | Team stays in Python API | Rust ownership of performance-critical path |
With IPC mmap and compatible buffers, often yes. In-process PyO3 may still copy depending on array ownership - benchmark your path.
Yes - write Arrow IPC or Parquet from pandas via PyArrow, read with Polars scan_ipc / read_parquet.
No for file/batch pipelines. PyO3 helps when Python orchestrates and Rust accelerates tight loops in one process.
Arrow bridges to NumPy for numeric columns without object dtype boxes - prefer Arrow table as the contract.
Use maturin develop locally and maturin build for wheels in CI - pair with a pyproject.toml pointing at your crate.
Yes - write IPC to /tmp, read in the next cell. For large data, mmap keeps RAM stable.
Print table.schema on both sides before transforms. Fail fast in Rust with explicit Schema on write.
Polars Python shares the same Rust core - IPC between Rust Polars and Python Polars is especially smooth.
Use Arrow int64 with null bitmap; avoid pandas float NaN sentinel for IDs.
See Apache Arrow for RecordBatch fundamentals underlying this handoff.
Stack versions: This page was written for Rust 1.97.0 (edition 2024), Tokio 1.x, Axum 0.8, serde 1.0, sqlx 0.8, clap 4, and Polars 0.46+.
Reviewed by Chris St. John·Last updated Jul 16, 2026