{
  "id": 248968,
  "title": "Difference in standard deviation calc between packages",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/248968",
  "author_name": "Mark Tenenholtz",
  "post_date": "2021-06-25T21:55:04.782000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Since there's a few high-profile public kernels that are doing some aggregations on the target statistic, I thought a PSA/friendly reminder was in order: the way that pandas and numpy calculate standard deviation by default are different! Pandas uses the unbiased estimator for the standard deviation by dividing by N-1, while numpy does not. If you give the argument <code>ddof=1</code> to numpy or, alternatively, the argument <code>ddof=0</code> to pandas, they will replicate each other's behavior. </p>",
  "messages": [
    {
      "id": 1365566,
      "postDate": "2021-06-25T21:55:04.783Z",
      "content": "<p>Since there's a few high-profile public kernels that are doing some aggregations on the target statistic, I thought a PSA/friendly reminder was in order: the way that pandas and numpy calculate standard deviation by default are different! Pandas uses the unbiased estimator for the standard deviation by dividing by N-1, while numpy does not. If you give the argument <code>ddof=1</code> to numpy or, alternatively, the argument <code>ddof=0</code> to pandas, they will replicate each other's behavior. </p>",
      "rawMarkdown": "Since there's a few high-profile public kernels that are doing some aggregations on the target statistic, I thought a PSA/friendly reminder was in order: the way that pandas and numpy calculate standard deviation by default are different! Pandas uses the unbiased estimator for the standard deviation by dividing by N-1, while numpy does not. If you give the argument `ddof=1` to numpy or, alternatively, the argument `ddof=0` to pandas, they will replicate each other's behavior. ",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1365566": "Since there's a few high-profile public kernels that are doing some aggregations on the target statistic, I thought a PSA/friendly reminder was in order: the way that pandas and numpy calculate standard deviation by default are different! Pandas uses the unbiased estimator for the standard deviation by dividing by N-1, while numpy does not. If you give the argument `ddof=1` to numpy or, alternatively, the argument `ddof=0` to pandas, they will replicate each other's behavior. "
  }
}