{
  "id": 195032,
  "title": "Which is correct? pandas or cuDF",
  "url": "/competitions/riiid-test-answer-prediction/discussion/195032",
  "author_name": "tereka",
  "post_date": "2020-11-03T06:46:11.548000",
  "votes": 32,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hi.  <br>\nI found the different results between pandas and cuDF.   <br>\nI calculate to mean \"prior_question_elapsed_time\". result is as follows.</p>\n<p>Pandas: 13005.0810546875<br>\ncuDF:25423.810042960275</p>\n<p>I have two questions.</p>\n<ol>\n<li>Why differences?  <br>\nI guess that result is loss of digits.</li>\n<li>How do we get same result?</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/tereka/pandas-cudf-meanresult-difference\" target=\"_blank\">https://www.kaggle.com/tereka/pandas-cudf-meanresult-difference</a></p>",
  "messages": [
    {
      "id": 1068188,
      "postDate": "2020-11-03T06:46:11.547Z",
      "content": "<p>Hi.  <br>\nI found the different results between pandas and cuDF.   <br>\nI calculate to mean \"prior_question_elapsed_time\". result is as follows.</p>\n<p>Pandas: 13005.0810546875<br>\ncuDF:25423.810042960275</p>\n<p>I have two questions.</p>\n<ol>\n<li>Why differences?  <br>\nI guess that result is loss of digits.</li>\n<li>How do we get same result?</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/tereka/pandas-cudf-meanresult-difference\" target=\"_blank\">https://www.kaggle.com/tereka/pandas-cudf-meanresult-difference</a></p>",
      "rawMarkdown": "Hi.  \nI found the different results between pandas and cuDF.   \nI calculate to mean \"prior_question_elapsed_time\". result is as follows.\n\nPandas: 13005.0810546875\ncuDF:25423.810042960275\n\nI have two questions.\n1. Why differences?  \n  I guess that result is loss of digits.\n2. How do we get same result?\n\n\nhttps://www.kaggle.com/tereka/pandas-cudf-meanresult-difference",
      "votes": 31
    },
    {
      "id": 1068930,
      "postDate": "2020-11-03T22:08:39.070Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F474194%2F57cd260ec30a153fa3920ef5efea35f3%2FScreenshot%202020-11-03%20at%2023.05.06.png?generation=1604441307458282&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F474194%2F57cd260ec30a153fa3920ef5efea35f3%2FScreenshot%202020-11-03%20at%2023.05.06.png?generation=1604441307458282&alt=media)",
      "votes": 9
    },
    {
      "id": 1068268,
      "postDate": "2020-11-03T08:15:16.573Z",
      "content": "<p>I've faced this issue before. From what I remember (a while back though), pandas uses <a href=\"https://pandas.pydata.org/pandas-docs/stable/getting_started/install.html#recommended-dependencies\" target=\"_blank\">bottleneck</a> which speeds up computations on large datasets but uses some approximations that can give particularly poor results when using 32-bit dtypes.</p>\n<p>Highly recommended to avoid using 32-bit types for this dataset with pandas.<br>\nNot sure if it still holds, but <a href=\"https://stackoverflow.com/a/53144736\" target=\"_blank\">this</a> is a really good explanation of the internal workings of bottleneck.</p>",
      "rawMarkdown": "I've faced this issue before. From what I remember (a while back though), pandas uses [bottleneck](https://pandas.pydata.org/pandas-docs/stable/getting_started/install.html#recommended-dependencies) which speeds up computations on large datasets but uses some approximations that can give particularly poor results when using 32-bit dtypes.\n\nHighly recommended to avoid using 32-bit types for this dataset with pandas.\nNot sure if it still holds, but [this](https://stackoverflow.com/a/53144736) is a really good explanation of the internal workings of bottleneck.",
      "votes": 7,
      "replies": [
        {
          "id": 1068816,
          "postDate": "2020-11-03T19:02:19.170Z",
          "content": "<p>This is why my local results are different. I do not have <code>bottleneck</code> installed.</p>",
          "rawMarkdown": "This is why my local results are different. I do not have `bottleneck` installed.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1068264,
      "postDate": "2020-11-03T08:09:38.367Z",
      "content": "<p>I think cuDF is correct, because the result of BigQuery shows 25423.810042961624.<br>\nSomething wrong with pandas.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1878632%2Fea3b9031dd763e3cdde08958e9f7d493%2F2020-11-03%2017.07.16.png?generation=1604390887398189&amp;alt=media\" alt=\"BQ result\"></p>",
      "rawMarkdown": "I think cuDF is correct, because the result of BigQuery shows 25423.810042961624.\nSomething wrong with pandas.\n\n![BQ result](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1878632%2Fea3b9031dd763e3cdde08958e9f7d493%2F2020-11-03%2017.07.16.png?generation=1604390887398189&alt=media)",
      "votes": 4
    },
    {
      "id": 1068195,
      "postDate": "2020-11-03T06:52:46.850Z",
      "content": "<p>For what it's worth, using Pandas locally I get pretty much the same answer as cuDF. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F150865%2F959b6b4a8720396a1d59704324411e4b%2FScreenshot%20from%202020-11-02%2022-48-50.png?generation=1604386159894507&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "For what it's worth, using Pandas locally I get pretty much the same answer as cuDF. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F150865%2F959b6b4a8720396a1d59704324411e4b%2FScreenshot%20from%202020-11-02%2022-48-50.png?generation=1604386159894507&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 1068197,
          "postDate": "2020-11-03T06:56:07.717Z",
          "content": "<p>It's very interesting.<br>\npandas version is the same as '1.1.3' in the notebook.<br>\nthat's mystery…</p>",
          "rawMarkdown": "It's very interesting.\npandas version is the same as '1.1.3' in the notebook.\nthat's mystery..."
        },
        {
          "id": 1068209,
          "postDate": "2020-11-03T07:16:00.393Z",
          "content": "<p>my local pc.(pandas)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F282250%2F13596e7c24cf17f7ac36bebd546666db%2F2020-11-03%2016.15.11.png?generation=1604387735882764&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "my local pc.(pandas)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F282250%2F13596e7c24cf17f7ac36bebd546666db%2F2020-11-03%2016.15.11.png?generation=1604387735882764&alt=media)"
        },
        {
          "id": 1068214,
          "postDate": "2020-11-03T07:21:51.137Z",
          "content": "<p>Converting to float64 in the Notebook fixes it. I still don't understand why it works on my machine without converting it.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F150865%2F521180612dda21b23bb8faeb27a3cc9e%2FScreenshot%20from%202020-11-02%2023-20-17.png?generation=1604388090189519&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Converting to float64 in the Notebook fixes it. I still don't understand why it works on my machine without converting it.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F150865%2F521180612dda21b23bb8faeb27a3cc9e%2FScreenshot%20from%202020-11-02%2023-20-17.png?generation=1604388090189519&alt=media)",
          "votes": 5
        },
        {
          "id": 1068219,
          "postDate": "2020-11-03T07:26:00.240Z",
          "content": "<p>Thanks, After I convert to float64, fixed that problem.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F282250%2Fbdf2bd8fa4f6a8713c703dc4e5c3e118%2F2020-11-03%2016.24.41.png?generation=1604388344514607&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thanks, After I convert to float64, fixed that problem.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F282250%2Fbdf2bd8fa4f6a8713c703dc4e5c3e118%2F2020-11-03%2016.24.41.png?generation=1604388344514607&alt=media)"
        },
        {
          "id": 1068226,
          "postDate": "2020-11-03T07:34:34.690Z",
          "content": "<p>Something to do with the python version I guess? The Notebooks use 3.7.6.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F150865%2F1b550f5813450572d7f0039101fce649%2FScreenshot%20from%202020-11-02%2023-32-27.png?generation=1604388798924659&amp;alt=media\" alt=\"\"></p>\n<p>[Edit] Nevermind, my pandas version for 3.6 is different than the one for 3.7.9 so it's not a fair comparison</p>",
          "rawMarkdown": "Something to do with the python version I guess? The Notebooks use 3.7.6.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F150865%2F1b550f5813450572d7f0039101fce649%2FScreenshot%20from%202020-11-02%2023-32-27.png?generation=1604388798924659&alt=media)\n\n[Edit] Nevermind, my pandas version for 3.6 is different than the one for 3.7.9 so it's not a fair comparison"
        }
      ]
    },
    {
      "id": 1070130,
      "postDate": "2020-11-05T11:46:43.440Z",
      "content": "<p>Mine are ok with float32 condition… I've never faced with 13005 values </p>",
      "rawMarkdown": "Mine are ok with float32 condition... I've never faced with 13005 values "
    },
    {
      "id": 1068340,
      "postDate": "2020-11-03T09:57:12.593Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1068225,
      "postDate": "2020-11-03T07:34:13.257Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1068229,
          "postDate": "2020-11-03T07:35:52.190Z",
          "content": "<p>yes. v1.1.4(my local) have 13005  <br>\nIt is a difference between cuDF and pandas.</p>",
          "rawMarkdown": "yes. v1.1.4(my local) have 13005  \nIt is a difference between cuDF and pandas."
        },
        {
          "id": 1068245,
          "postDate": "2020-11-03T07:49:11.550Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1068252,
          "postDate": "2020-11-03T07:58:53.890Z",
          "content": "<p>sum/cnt is same, but mean is difference… this is so strange.</p>",
          "rawMarkdown": "sum/cnt is same, but mean is difference... this is so strange."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1068930,
      "author_name": "Rafi Hai",
      "author_url": "",
      "post_date": "2020-11-03T22:08:39.070000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F474194%2F57cd260ec30a153fa3920ef5efea35f3%2FScreenshot%202020-11-03%20at%2023.05.06.png?generation=1604441307458282&amp;alt=media\" alt=\"\"></p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 1068268,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2020-11-03T08:15:16.573000",
      "content": "<p>I've faced this issue before. From what I remember (a while back though), pandas uses <a href=\"https://pandas.pydata.org/pandas-docs/stable/getting_started/install.html#recommended-dependencies\" target=\"_blank\">bottleneck</a> which speeds up computations on large datasets but uses some approximations that can give particularly poor results when using 32-bit dtypes.</p>\n<p>Highly recommended to avoid using 32-bit types for this dataset with pandas.<br>\nNot sure if it still holds, but <a href=\"https://stackoverflow.com/a/53144736\" target=\"_blank\">this</a> is a really good explanation of the internal workings of bottleneck.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1068816,
          "author_name": "Branden Murray",
          "author_url": "",
          "post_date": "2020-11-03T19:02:19.170000",
          "content": "<p>This is why my local results are different. I do not have <code>bottleneck</code> installed.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1068264,
      "author_name": "Leo0523",
      "author_url": "",
      "post_date": "2020-11-03T08:09:38.367000",
      "content": "<p>I think cuDF is correct, because the result of BigQuery shows 25423.810042961624.<br>\nSomething wrong with pandas.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1878632%2Fea3b9031dd763e3cdde08958e9f7d493%2F2020-11-03%2017.07.16.png?generation=1604390887398189&amp;alt=media\" alt=\"BQ result\"></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1068195,
      "author_name": "Branden Murray",
      "author_url": "",
      "post_date": "2020-11-03T06:52:46.850000",
      "content": "<p>For what it's worth, using Pandas locally I get pretty much the same answer as cuDF. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F150865%2F959b6b4a8720396a1d59704324411e4b%2FScreenshot%20from%202020-11-02%2022-48-50.png?generation=1604386159894507&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1068197,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "2020-11-03T06:56:07.717000",
          "content": "<p>It's very interesting.<br>\npandas version is the same as '1.1.3' in the notebook.<br>\nthat's mystery…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1068209,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "2020-11-03T07:16:00.393000",
          "content": "<p>my local pc.(pandas)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F282250%2F13596e7c24cf17f7ac36bebd546666db%2F2020-11-03%2016.15.11.png?generation=1604387735882764&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1068214,
          "author_name": "Branden Murray",
          "author_url": "",
          "post_date": "2020-11-03T07:21:51.137000",
          "content": "<p>Converting to float64 in the Notebook fixes it. I still don't understand why it works on my machine without converting it.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F150865%2F521180612dda21b23bb8faeb27a3cc9e%2FScreenshot%20from%202020-11-02%2023-20-17.png?generation=1604388090189519&amp;alt=media\" alt=\"\"></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1068219,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "2020-11-03T07:26:00.240000",
          "content": "<p>Thanks, After I convert to float64, fixed that problem.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F282250%2Fbdf2bd8fa4f6a8713c703dc4e5c3e118%2F2020-11-03%2016.24.41.png?generation=1604388344514607&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1068226,
          "author_name": "Branden Murray",
          "author_url": "",
          "post_date": "2020-11-03T07:34:34.690000",
          "content": "<p>Something to do with the python version I guess? The Notebooks use 3.7.6.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F150865%2F1b550f5813450572d7f0039101fce649%2FScreenshot%20from%202020-11-02%2023-32-27.png?generation=1604388798924659&amp;alt=media\" alt=\"\"></p>\n<p>[Edit] Nevermind, my pandas version for 3.6 is different than the one for 3.7.9 so it's not a fair comparison</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1070130,
      "author_name": "jwc",
      "author_url": "",
      "post_date": "2020-11-05T11:46:43.440000",
      "content": "<p>Mine are ok with float32 condition… I've never faced with 13005 values </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1068340,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-03T09:57:12.593000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1068225,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-03T07:34:13.257000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1068229,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "2020-11-03T07:35:52.190000",
          "content": "<p>yes. v1.1.4(my local) have 13005  <br>\nIt is a difference between cuDF and pandas.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1068245,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-03T07:49:11.550000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1068252,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "2020-11-03T07:58:53.890000",
          "content": "<p>sum/cnt is same, but mean is difference… this is so strange.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1068188": "Hi.  \nI found the different results between pandas and cuDF.   \nI calculate to mean \"prior_question_elapsed_time\". result is as follows.\n\nPandas: 13005.0810546875\ncuDF:25423.810042960275\n\nI have two questions.\n1. Why differences?  \n  I guess that result is loss of digits.\n2. How do we get same result?\n\n\nhttps://www.kaggle.com/tereka/pandas-cudf-meanresult-difference",
    "1068930": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F474194%2F57cd260ec30a153fa3920ef5efea35f3%2FScreenshot%202020-11-03%20at%2023.05.06.png?generation=1604441307458282&alt=media)",
    "1068268": "I've faced this issue before. From what I remember (a while back though), pandas uses [bottleneck](https://pandas.pydata.org/pandas-docs/stable/getting_started/install.html#recommended-dependencies) which speeds up computations on large datasets but uses some approximations that can give particularly poor results when using 32-bit dtypes.\n\nHighly recommended to avoid using 32-bit types for this dataset with pandas.\nNot sure if it still holds, but [this](https://stackoverflow.com/a/53144736) is a really good explanation of the internal workings of bottleneck.",
    "1068264": "I think cuDF is correct, because the result of BigQuery shows 25423.810042961624.\nSomething wrong with pandas.\n\n![BQ result](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1878632%2Fea3b9031dd763e3cdde08958e9f7d493%2F2020-11-03%2017.07.16.png?generation=1604390887398189&alt=media)",
    "1068195": "For what it's worth, using Pandas locally I get pretty much the same answer as cuDF. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F150865%2F959b6b4a8720396a1d59704324411e4b%2FScreenshot%20from%202020-11-02%2022-48-50.png?generation=1604386159894507&alt=media)",
    "1070130": "Mine are ok with float32 condition... I've never faced with 13005 values ",
    "1068340": "",
    "1068225": ""
  }
}