Python Polars Cheat Sheet: Fast DataFrames for Busy Engineers

Copy-paste Polars cheat sheet: filter by number, group, join, pl.lit, dates, and Pandas to Polars. Real syntax from production data pipelines.

4 min read

Python Polars Cheat Sheet: Fast DataFrames for Busy Engineers cover

Polars hits the sweet spot between Pandas’ ease and Spark’s scale. If you’ve ever waited on a groupby or cursed a memory error, this cheat sheet is for you. I’ve pulled the patterns that save time in real pipelines, not just toy examples. Bookmark this before your next ETL run.

Setup and Basics

First, get Polars and a dataset. The lazy API is the default now, so you’ll rarely need to call .lazy() explicitly. Start with a CSV or Parquet file, or create a DataFrame from scratch.

  • pip install polars pyarrow

  • import polars as pl

  • df = pl.read_csv('data.csv') # or pl.read_parquet()

  • df = pl.DataFrame({'a': [1, 2], 'b': ['x', 'y']})

Selecting and Filtering

Polars uses expressions, not strings. This feels odd at first but pays off when you chain operations. The syntax is consistent: every column is an expression you can transform, filter, or aggregate.

  • df.select(['a', 'b']) # columns by name

  • df.select(pl.col('a').alias('renamed'))

  • df.filter(pl.col('a') > 10)

  • df.filter(pl.col('b').is_in(['x', 'z']))

  • df.filter(pl.col('a').is_null())

Transforming Data

Polars expressions are composable. You can nest them, reuse them, and even store them in variables. This is where the library shines over Pandas.

  • df.with_columns(pl.col('a').cast(pl.Float64))

  • df.with_columns(pl.col('a').fill_null(0))

  • df.with_columns((pl.col('a') * 2).alias('a_doubled'))

  • df.with_columns(pl.col('b').str.to_uppercase())

  • df.with_columns(pl.col('a').is_between(10, 20))

Grouping and Aggregations

Groupbys in Polars are lazy by default. This means you can stack multiple aggregations without materializing intermediate results. The syntax is clean, but watch out for the order of operations.

  • df.group_by('b').agg(pl.col('a').sum())

  • df.group_by('b').agg([pl.col('a').mean(), pl.col('a').max()])

  • df.group_by('b').agg(pl.col('a').quantile(0.9))

  • df.group_by_dynamic('timestamp', every='1d').agg(pl.col('a').sum())

Joins and Concatenation

Joins in Polars are explicit. You’ll specify the join type and the columns to join on. Concatenation is straightforward, but remember that Polars is strict about schema matching.

  • df.join(other, on='key', how='inner')

  • df.join(other, on='key', how='left')

  • df.hstack([other]) # column-wise

  • df.vstack([other]) # row-wise

  • pl.concat([df1, df2], how='diagonal') # schema-safe union

Performance Tips

Polars is fast, but you can make it faster. The lazy API is your friend. Use it for any pipeline longer than a few operations. Also, avoid Python loops. Polars expressions are vectorized, so let the engine do the work.

  • Use .lazy() and .collect() for pipelines with multiple steps.

  • Prefer Parquet over CSV for I/O. It’s faster and smaller.

  • Use .with_columns() instead of multiple .select() calls.

  • Avoid .apply() unless absolutely necessary. Use expressions first.

  • Set pl.Config.set_fmt_str_lengths(100) to see full strings in debug output.

Debugging and Inspection

Polars has great tools for debugging. The .explain() method shows the query plan, which is invaluable for optimizing lazy pipelines. For quick checks, use .head() or .sample().

  • df.head(5) # first 5 rows

  • df.sample(5) # random 5 rows

  • df.describe() # summary stats

  • df.schema # column names and types

  • lazy_df.explain() # query plan for lazy DataFrames

Literals and constants with pl.lit

pl.lit wraps a raw Python value as a Polars expression, so Polars treats it as a constant instead of a column name. Reach for it when you need a fixed column, a comparison value, or a default inside select, filter, or with_columns.

  • df.with_columns(pl.lit('prod').alias('env')) # constant column

  • df.with_columns(pl.lit(0).alias('score'))

  • df.filter(pl.col('status') == pl.lit('active'))

  • df.select((pl.col('price') * pl.lit(1.2)).alias('with_tax'))

  • df.with_columns(pl.when(pl.col('a') > 0).then(pl.lit('pos')).otherwise(pl.lit('neg')).alias('sign'))

Filter by number, string, or null

Every filter is an expression. For numbers, compare the column directly and combine conditions with & and |, wrapping each in parentheses. For ranges, is_between reads cleaner than two comparisons.

  • df.filter(pl.col('amount') > 100) # filter by number

  • df.filter(pl.col('amount').is_between(10, 50)) # numeric range

  • df.filter((pl.col('amount') > 100) & (pl.col('country') == 'US'))

  • df.filter(pl.col('name').str.contains('adil')) # string match

  • df.filter(pl.col('email').is_null()) # or .is_not_null()

  • df.filter(pl.col('id').is_in([1, 2, 3]))

Dates: parsing and finding non-dates

Parse strings to dates with str.to_date and pass the format explicitly for speed. To select rows where a value is not a valid date, parse with strict=False so bad values become null, then filter on is_null.

  • df.with_columns(pl.col('d').str.to_date('%Y-%m-%d'))

  • df.with_columns(pl.col('d').str.to_datetime('%Y-%m-%d %H:%M:%S'))

  • # rows where the value is NOT a valid date format:

  • df.with_columns(pl.col('d').str.to_date('%Y-%m-%d', strict=False).alias('parsed')).filter(pl.col('parsed').is_null())

  • df.filter(pl.col('date') > pl.date(2026, 1, 1))

  • df.with_columns(pl.col('date').dt.year().alias('year'))

Coming from Pandas, dicts, or pyreadstat (SPSS/SAS)

Polars converts directly from Pandas, dicts, and Arrow. For SPSS or SAS files read with pyreadstat, read into a Pandas frame first, then hand it to pl.from_pandas. That path keeps your column names and types intact.

  • pl.from_pandas(pandas_df)

  • pl.from_dict({'a': [1, 2], 'b': ['x', 'y']})

  • pl.from_arrow(arrow_table)

  • # pyreadstat (.sav SPSS / .sas7bdat SAS): read with pandas, then convert

  • import pyreadstat; pdf, meta = pyreadstat.read_sav('file.sav'); df = pl.from_pandas(pdf)

  • df.to_pandas() # convert back when a library needs Pandas

FAQ

What does pl.lit do in Polars? It turns a raw Python value into an expression, so Polars uses it as a constant rather than a column name. Use it for fixed values inside select, filter, and with_columns.

How do I filter by a number in Polars? Compare the column expression directly, for example df.filter(pl.col('amount') > 100), and combine conditions with & and |, each wrapped in parentheses.

How do I convert a Pandas or pyreadstat DataFrame to Polars? Use pl.from_pandas(df). For SPSS or SAS files, read them with pyreadstat into a Pandas frame first, then pass that frame to pl.from_pandas.

How do I select rows where a value is not a valid date? Parse the column with str.to_date(..., strict=False) so invalid values become null, then filter with is_null to keep only the rows that are not valid dates.

Polars won’t replace Pandas for every task, but it’s the right tool for most data pipelines. The syntax takes a day to learn and a week to master. Once it clicks, you’ll write faster, cleaner code. Keep this cheat sheet handy until the patterns stick.

More Polars guides

Each of these goes deeper on one job: rename columns, create DataFrames from dict, NumPy, Pandas, or pyreadstat, filter rows by number, string, or date, read and write CSV, Parquet, and JSON, and view, inspect, and stack frames.

Building something with AI? Let's talk.

I design and ship production AI and full-stack products for US teams. See how I can help.

View all services

Join the newsletter

Be the first to read our articles.

Python Polars Cheat Sheet: Fast DataFrames for Busy Engineers