{
  "id": 344846,
  "title": "pandas noob",
  "url": "/competitions/amex-default-prediction/discussion/344846",
  "author_name": "",
  "post_date": "2022-08-16T19:58:37.648095800Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Maybe it's just that time of the night:</p>\n<pre><code>z = pd.DataFrame({\n    'cid':[0,0,0,0,0,0,1,1,1,1,1,1],\n    'b':  [5,1,2,2,3,3,4,4,4,4,4,1],\n})\nz.groupby('cid').b.std()\n</code></pre>\n<p>The above displays:</p>\n<pre><code>cid\n0    1.366260\n1    1.224745\nName: b, dtype: float64\n</code></pre>\n<p>But if I do something like this, I get the following results:</p>\n<pre><code>np.std([5,1,2,2,3,3]) # 1.247219128924647\nnp.std([4,4,4,4,4,1]) # 1.118033988749895\n</code></pre>\n<p>So, what's going on pandas?</p>",
  "messages": [
    {
      "id": "1901675",
      "postDate": "08/16/2022 19:58:37",
      "content": "<p>Maybe it's just that time of the night:</p>\n<pre><code>z = pd.DataFrame({\n    'cid':[0,0,0,0,0,0,1,1,1,1,1,1],\n    'b':  [5,1,2,2,3,3,4,4,4,4,4,1],\n})\nz.groupby('cid').b.std()\n</code></pre>\n<p>The above displays:</p>\n<pre><code>cid\n0    1.366260\n1    1.224745\nName: b, dtype: float64\n</code></pre>\n<p>But if I do something like this, I get the following results:</p>\n<pre><code>np.std([5,1,2,2,3,3]) # 1.247219128924647\nnp.std([4,4,4,4,4,1]) # 1.118033988749895\n</code></pre>\n<p>So, what's going on pandas?</p>",
      "rawMarkdown": "Maybe it's just that time of the night:\n\n```\nz = pd.DataFrame({\n    'cid':[0,0,0,0,0,0,1,1,1,1,1,1],\n    'b':  [5,1,2,2,3,3,4,4,4,4,4,1],\n})\nz.groupby('cid').b.std()\n```\n\nThe above displays:\n\n```\ncid\n0    1.366260\n1    1.224745\nName: b, dtype: float64\n```\n\nBut if I do something like this, I get the following results:\n\n```\nnp.std([5,1,2,2,3,3]) # 1.247219128924647\nnp.std([4,4,4,4,4,1]) # 1.118033988749895\n```\n\nSo, what's going on pandas?",
      "votes": null
    },
    {
      "id": "1901689",
      "postDate": "08/16/2022 20:13:00",
      "content": "<p>read the docs:<br>\n<a href=\"https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.std.html\" target=\"_blank\">https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.std.html</a></p>\n<p><code>To have the same behaviour as numpy.std, use ddof=0 (instead of the default ddof=1)</code></p>\n<p>pandas applies different degrees of freedom by default in the stdev formula</p>",
      "rawMarkdown": "read the docs:\nhttps://pandas.pydata.org/docs/reference/api/pandas.DataFrame.std.html\n\n`To have the same behaviour as numpy.std, use ddof=0 (instead of the default ddof=1)`\n\npandas applies different degrees of freedom by default in the stdev formula",
      "votes": null
    },
    {
      "id": "1901691",
      "postDate": "08/16/2022 20:13:13",
      "content": "<p>.          </p>",
      "rawMarkdown": ".",
      "votes": null
    },
    {
      "id": "1901700",
      "postDate": "08/16/2022 20:16:53",
      "content": "<p>Numpy divides the sum of squares by the length of the array, pandas divides by the length minus one.</p>\n<pre><code>s = pd.Series([1, 2, 3])\nprint(s.std(), s.values.std())\nprint(np.sqrt(np.square(s - s.mean()).sum() / (len(s)-1)),\n      np.sqrt(np.square(s - s.mean()).sum() / len(s)))\n</code></pre>\n<pre><code>1.0 0.816496580927726\n1.0 0.816496580927726\n</code></pre>\n<p>Compare the default value for <code>ddof</code> <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.std.html\" target=\"_blank\">here</a> and <a href=\"https://numpy.org/doc/stable/reference/generated/numpy.std.html\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "Numpy divides the sum of squares by the length of the array, pandas divides by the length minus one.\n\n```\ns = pd.Series([1, 2, 3])\nprint(s.std(), s.values.std())\nprint(np.sqrt(np.square(s - s.mean()).sum() / (len(s)-1)),\n      np.sqrt(np.square(s - s.mean()).sum() / len(s)))\n```\n\n```\n1.0 0.816496580927726\n1.0 0.816496580927726\n```\n\nCompare the default value for `ddof` [here](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.std.html) and [here](https://numpy.org/doc/stable/reference/generated/numpy.std.html).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1901689,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "08/16/2022 20:13:00",
      "content": "<p>read the docs:<br>\n<a href=\"https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.std.html\" target=\"_blank\">https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.std.html</a></p>\n<p><code>To have the same behaviour as numpy.std, use ddof=0 (instead of the default ddof=1)</code></p>\n<p>pandas applies different degrees of freedom by default in the stdev formula</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1901691,
      "author_name": "authman",
      "author_url": "",
      "post_date": "08/16/2022 20:13:13",
      "content": "<p>.          </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1901700,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "08/16/2022 20:16:53",
      "content": "<p>Numpy divides the sum of squares by the length of the array, pandas divides by the length minus one.</p>\n<pre><code>s = pd.Series([1, 2, 3])\nprint(s.std(), s.values.std())\nprint(np.sqrt(np.square(s - s.mean()).sum() / (len(s)-1)),\n      np.sqrt(np.square(s - s.mean()).sum() / len(s)))\n</code></pre>\n<pre><code>1.0 0.816496580927726\n1.0 0.816496580927726\n</code></pre>\n<p>Compare the default value for <code>ddof</code> <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.std.html\" target=\"_blank\">here</a> and <a href=\"https://numpy.org/doc/stable/reference/generated/numpy.std.html\" target=\"_blank\">here</a>.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1901675": "Maybe it's just that time of the night:\n\n```\nz = pd.DataFrame({\n    'cid':[0,0,0,0,0,0,1,1,1,1,1,1],\n    'b':  [5,1,2,2,3,3,4,4,4,4,4,1],\n})\nz.groupby('cid').b.std()\n```\n\nThe above displays:\n\n```\ncid\n0    1.366260\n1    1.224745\nName: b, dtype: float64\n```\n\nBut if I do something like this, I get the following results:\n\n```\nnp.std([5,1,2,2,3,3]) # 1.247219128924647\nnp.std([4,4,4,4,4,1]) # 1.118033988749895\n```\n\nSo, what's going on pandas?",
    "1901689": "read the docs:\nhttps://pandas.pydata.org/docs/reference/api/pandas.DataFrame.std.html\n\n`To have the same behaviour as numpy.std, use ddof=0 (instead of the default ddof=1)`\n\npandas applies different degrees of freedom by default in the stdev formula",
    "1901691": ".",
    "1901700": "Numpy divides the sum of squares by the length of the array, pandas divides by the length minus one.\n\n```\ns = pd.Series([1, 2, 3])\nprint(s.std(), s.values.std())\nprint(np.sqrt(np.square(s - s.mean()).sum() / (len(s)-1)),\n      np.sqrt(np.square(s - s.mean()).sum() / len(s)))\n```\n\n```\n1.0 0.816496580927726\n1.0 0.816496580927726\n```\n\nCompare the default value for `ddof` [here](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.std.html) and [here](https://numpy.org/doc/stable/reference/generated/numpy.std.html)."
  },
  "source": "meta"
}