30 Matching Annotations
  1. Sep 2026
    1. Pandas Should Go Extinct

      Summary: Pandas Should Go Extinct

      Core Thesis

      • Pandas forces teams into expensive, complex distributed systems (such as Spark or Snowflake) prematurely due to its memory inefficiency and single-threaded execution.
      • Most organizations do not actually possess "Big Data"; instead, they face "Medium Data" problems that modern single-machine tools can easily handle.

      The "Big Data" Myth vs. Reality

      • Telemetry from Amazon Redshift's fleet indicates that true Big Data is rare for typical analytics:
        • Approximately 94.7% of tables contain under 100 GB of data.
        • Around 86.9% of queries scan 80 GB or less and execute in under a second.
      • The "Pandas cliff" occurs around tens of gigabytes, leaving an unserved gap up to ~100 GB that single-node engines are ideally suited to bridge.

      Architectural Limitations of Pandas

      • Operates eagerly and sequentially, reading full datasets into memory before filtering or aggregation.
      • Lacks native query optimization, resulting in poor CPU core utilization, high memory footprint (USS), and severe OS swapping.
      • Features an increasingly baroque and counterintuitive API that complicates maintenance.

      The Alternatives: Polars and DuckDB

      • Polars:
        • Rust-based DataFrame library with familiar syntax.
        • Implements lazy evaluation, optimized query plans, predicate pushdown, and streaming execution across all available CPU cores.
      • DuckDB:
        • Embedded in-process analytical (OLAP) database, operating as an "SQLite for analytics."
        • Uses vectorized SQL queries, parallel processing, and can query in-memory Python objects and Parquet/CSV files directly.

      Benchmark Insights (1 Billion Row Challenge & NYC Taxi Data)

      • On high-spec cloud hardware (32 cores, 128 GB RAM):
        • Pandas finished the 1BRC aggregation in ~4m 28s, bottlenecked at ~113% CPU with 38 GB memory usage.
        • Polars completed in ~5.04s and DuckDB in ~5.19s, both achieving over 3,100% CPU utilization.
      • On local developer hardware (M1 MBA):
        • Polars and DuckDB finished within 39–47 seconds with negligible swap usage.
        • Pandas required over 12 minutes and triggered 21 GB of disk swapping.
      • Key built-in benefits include out-of-the-box multithreading, memory streaming, disk spilling, and frictionless migration via Apache Arrow zero-copy memory sharing.

      Hacker News Discussion

      • The Literal "Panda" Bait-and-Switch:
        • Multiple commenters humorously admitted clicking expecting a biology/ecology debate about giant panda conservation and captive breeding, expressing amused appreciation for the double-meaning title.
      • Ecosystem Inertia and Educational Use:
        • Critics noted that 99%+ of scripts, student projects, and quick data checks operate on tiny datasets where Pandas' performance is negligible.
        • Extensive documentation, StackOverflow answers, and existing pipelines provide strong inertia favoring Pandas, though others warned this inertia fuels stagnation.
      • API Ergonomics and Code Maintainability:
        • Strong agreement emerged regarding Pandas' awkward API; minor query requirement changes (such as filtering within a group) often necessitate rewriting entire pipelines or resorting to inefficient lambdas.
        • Polars and R's Tidyverse were highlighted as having significantly cleaner, more predictable expression-based semantics.
      • Small-Data Performance Trade-offs:
        • Users pointed out that Polars can occasionally run slower than Pandas on very small dataframes (e.g., thousands of rows) due to thread synchronization overhead, emphasizing that thread counts should be tuned (POLARS_MAX_THREADS=1) and profiled.
      • Advancements in the Data Science Landscape:
        • Participants debated whether data science tooling has stagnated amid AI hype, with many emphasizing that DuckDB and Polars represent the most impactful and practical innovations in local data analysis in recent years.
  2. Dec 2023
    1. Technique #2: Sampling

      How do you load only a subset of the rows?

      When you load your data, you can specify a skiprows function that will randomly decide whether to load that row or not:

      ```

      from random import random

      def sample(row_number): ... if row_number == 0: ... # Never drop the row with column names: ... return False ... # random() returns uniform numbers between 0 and 1: ... return random() > 0.001 ... sampled = pd.read_csv("/tmp/voting.csv", skiprows=sample) len(sampled) 973 ```

  3. Jul 2023
    1. The parameter by specifies the columns, and ascending takes a list to define the sorting direction per each column. In this case, we're sorting by Country name in descending order first (in lexicographical order), and by number of Employees in ascending order second.

      Pandas DataFrame allows for multiple sorting

    1. Bollinger bands are just a simple visualization/analysis technique that creates two bands, one "roof" and one "floor" of some "support" for a given time series. The reasoning is that, if the time series is "below" the "floor", it's a historic low, and if it's "above" the "roof", it's a historic high. In terms of stock prices and other financial instruments, when the price crosses a band, it's said to be too cheap or too expensive.

      How to display Bollinger bands with Pandas.

  4. May 2023
    1. Return the first 5 rows of the DataFrame

      5 تا ردیف اول را برات بر میگردونه. یه ورودی هم شاید بگیره که در واقع تعداد ردیف هایی است که میخواد برگردونه

  5. Apr 2023
    1. ff = ef['x','y']

      Máscaras em Pandas são uma maneira de selecionar um subconjunto de dados de um DataFrame, Series ou outro objeto de dados baseado em uma condição booleana.

      O código que deve ser adicionado no lugar de # a fazer é:

      ff = ef[['x', 'y']]

      Isso irá selecionar apenas as colunas 'x' e 'y' do DataFrame ef, que é o resultado da máscara m. A máscara m seleciona apenas as linhas onde o valor da coluna 'z' é False, e então, ef contém apenas essas linhas. Finalmente, ff é criado selecionando as colunas 'x' e 'y' do DataFrame ef.

  6. Dec 2021
  7. Nov 2021
  8. Sep 2021
  9. Aug 2020
  10. Mar 2020
    1. It’s just that it often makes sense to write code in the order JOIN / WHERE / GROUP BY / HAVING. (I’ll often put a WHERE first to improve performance though, and I think most database engines will also do a WHERE first in practice)

      Pandas usually writes code in this syntax:

      1. JOIN
      2. WHERE
      3. GROUP BY
      4. HAVING

      Example:

      1. df = thing1.join(thing2) # like a JOIN
      2. df = df[df.created_at > 1000] # like a WHERE
      3. df = df.groupby('something', num_yes = ('yes', 'sum')) # like a GROUP BY
      4. df = df[df.num_yes > 2] # like a HAVING, filtering on the result of a GROUP BY
      5. df = df[['num_yes', 'something1', 'something']] # pick the columns I want to display, like a SELECT
      6. df.sort_values('sometthing', ascending=True)[:30] # ORDER BY and LIMIT
      7. df[:30]
  11. Nov 2019
  12. Oct 2019
    1. Indicate number of NA values placed in non-numeric columns.

      This is only true when using the Python parsing engine.

      Filled 3 NA values in column name
      

      If using the C parsing engine you get something like the following output:

      Tokenization took: 0.01 ms
      Type conversion took: 0.70 ms
      Parser memory cleanup took: 0.01 ms
      
  13. Feb 2019
  14. Jun 2018
  15. May 2018
  16. Apr 2018
  17. Mar 2018
    1. I'll skip the inefficient method I used before with the custom groupby aggregationm, and go for some neat trick using the mighty transform method.

      a more constrained. and thus more efficient way to do transformations on groupbys than the apply method. You can do very cool stuff with it. For those of you who know splunk - this has the neat "streamstats" and "eventstats" capabilities

  18. Dec 2017