{
  "id": 196474,
  "title": "Faster than pandas.merge ?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/196474",
  "author_name": "",
  "post_date": "2020-11-11T11:33:16.561324200Z",
  "votes": 5,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I think pandas.merge is too slow during submission (takes 3 hours on average), and I am looking for something faster, any ideas ?<br>\nbetter methods using python or numpy, or other libraries faster than pandas ?</p>",
  "messages": [
    {
      "id": "1075091",
      "postDate": "11/11/2020 11:33:16",
      "content": "<p>I think pandas.merge is too slow during submission (takes 3 hours on average), and I am looking for something faster, any ideas ?<br>\nbetter methods using python or numpy, or other libraries faster than pandas ?</p>",
      "rawMarkdown": "I think pandas.merge is too slow during submission (takes 3 hours on average), and I am looking for something faster, any ideas ?\nbetter methods using python or numpy, or other libraries faster than pandas ?",
      "votes": null
    },
    {
      "id": "1075629",
      "postDate": "11/11/2020 20:41:42",
      "content": "<p>people reported using concat, you can try that</p>",
      "rawMarkdown": "people reported using concat, you can try that",
      "votes": null
    },
    {
      "id": "1075634",
      "postDate": "11/11/2020 20:44:24",
      "content": "<p>check this  - <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009</a> </p>",
      "rawMarkdown": "check this  - https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009",
      "votes": null
    },
    {
      "id": "1075655",
      "postDate": "11/11/2020 21:19:44",
      "content": "<p>Doesn't concat require storing two copies of the Dataframe in memory?<br>\nHow would you work around that with such a large dataset?</p>",
      "rawMarkdown": "Doesn't concat require storing two copies of the Dataframe in memory?\nHow would you work around that with such a large dataset?",
      "votes": null
    },
    {
      "id": "1075819",
      "postDate": "11/12/2020 02:19:14",
      "content": "<p>From my experiment because each iter in submission consist of small total rows its better to just use loop and dict</p>",
      "rawMarkdown": "From my experiment because each iter in submission consist of small total rows its better to just use loop and dict",
      "votes": null
    },
    {
      "id": "1077210",
      "postDate": "11/13/2020 10:20:33",
      "content": "<p>that's very useful, thanks</p>",
      "rawMarkdown": "that's very useful, thanks",
      "votes": null
    },
    {
      "id": "1077211",
      "postDate": "11/13/2020 10:21:43",
      "content": "<p>I believe that too,<br>\nmerge seems good with big datasets, but with small chunks, I thinks it is slow</p>",
      "rawMarkdown": "I believe that too,\nmerge seems good with big datasets, but with small chunks, I thinks it is slow",
      "votes": null
    },
    {
      "id": "1077314",
      "postDate": "11/13/2020 13:12:46",
      "content": "<p>use join to replace merge ,can save 50% time</p>",
      "rawMarkdown": "use join to replace merge ,can save 50% time",
      "votes": null
    },
    {
      "id": "1077535",
      "postDate": "11/13/2020 17:30:32",
      "content": "<p>How to measure the memory cost for the step? </p>",
      "rawMarkdown": "How to measure the memory cost for the step?",
      "votes": null
    },
    {
      "id": "1077850",
      "postDate": "11/14/2020 02:13:00",
      "content": "<p><a href=\"https://www.kaggle.com/tkm2261\" target=\"_blank\">@tkm2261</a> has shared his wonderful methods in the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197023\" target=\"_blank\">discussion</a> and <a href=\"https://www.kaggle.com/tkm2261/fast-pandas-left-join-357x-faster-than-pd-merge\" target=\"_blank\">notebook</a>.</p>",
      "rawMarkdown": "tkm2261 has shared his wonderful methods in the [discussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197023) and [notebook](https://www.kaggle.com/tkm2261/fast-pandas-left-join-357x-faster-than-pd-merge).",
      "votes": null
    },
    {
      "id": "1077969",
      "postDate": "11/14/2020 06:47:06",
      "content": "<p>correct, I tried join, and it saved me about 50%<br>\nthanks for that.</p>\n<p>but, now I am trying faster idea that I saw in below discussion, it might be even faster than join.</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197023\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197023</a></p>",
      "rawMarkdown": "correct, I tried join, and it saved me about 50%\nthanks for that.\n\nbut, now I am trying faster idea that I saw in below discussion, it might be even faster than join.\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197023",
      "votes": null
    },
    {
      "id": "1077970",
      "postDate": "11/14/2020 06:47:58",
      "content": "<p>I checked it, its very interesting, I am working on it now.<br>\nthanks for notifying me about it.</p>",
      "rawMarkdown": "I checked it, its very interesting, I am working on it now.\nthanks for notifying me about it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1075629,
      "author_name": "chenmingml",
      "author_url": "",
      "post_date": "11/11/2020 20:41:42",
      "content": "<p>people reported using concat, you can try that</p>",
      "votes": null,
      "replies": [
        {
          "id": 1075655,
          "author_name": "cshorten30",
          "author_url": "",
          "post_date": "11/11/2020 21:19:44",
          "content": "<p>Doesn't concat require storing two copies of the Dataframe in memory?<br>\nHow would you work around that with such a large dataset?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1077535,
          "author_name": "chenmingml",
          "author_url": "",
          "post_date": "11/13/2020 17:30:32",
          "content": "<p>How to measure the memory cost for the step? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1075634,
      "author_name": "rashmibanthia",
      "author_url": "",
      "post_date": "11/11/2020 20:44:24",
      "content": "<p>check this  - <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1077210,
          "author_name": "mohamadnawfal",
          "author_url": "",
          "post_date": "11/13/2020 10:20:33",
          "content": "<p>that's very useful, thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1075819,
      "author_name": "marcellosusanto",
      "author_url": "",
      "post_date": "11/12/2020 02:19:14",
      "content": "<p>From my experiment because each iter in submission consist of small total rows its better to just use loop and dict</p>",
      "votes": null,
      "replies": [
        {
          "id": 1077211,
          "author_name": "mohamadnawfal",
          "author_url": "",
          "post_date": "11/13/2020 10:21:43",
          "content": "<p>I believe that too,<br>\nmerge seems good with big datasets, but with small chunks, I thinks it is slow</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1077314,
      "author_name": "pp2file",
      "author_url": "",
      "post_date": "11/13/2020 13:12:46",
      "content": "<p>use join to replace merge ,can save 50% time</p>",
      "votes": null,
      "replies": [
        {
          "id": 1077969,
          "author_name": "mohamadnawfal",
          "author_url": "",
          "post_date": "11/14/2020 06:47:06",
          "content": "<p>correct, I tried join, and it saved me about 50%<br>\nthanks for that.</p>\n<p>but, now I am trying faster idea that I saw in below discussion, it might be even faster than join.</p>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197023\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197023</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1077850,
      "author_name": "yutsumura",
      "author_url": "",
      "post_date": "11/14/2020 02:13:00",
      "content": "<p><a href=\"https://www.kaggle.com/tkm2261\" target=\"_blank\">@tkm2261</a> has shared his wonderful methods in the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197023\" target=\"_blank\">discussion</a> and <a href=\"https://www.kaggle.com/tkm2261/fast-pandas-left-join-357x-faster-than-pd-merge\" target=\"_blank\">notebook</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1077970,
          "author_name": "mohamadnawfal",
          "author_url": "",
          "post_date": "11/14/2020 06:47:58",
          "content": "<p>I checked it, its very interesting, I am working on it now.<br>\nthanks for notifying me about it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1075091": "I think pandas.merge is too slow during submission (takes 3 hours on average), and I am looking for something faster, any ideas ?\nbetter methods using python or numpy, or other libraries faster than pandas ?",
    "1075629": "people reported using concat, you can try that",
    "1075634": "check this  - https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009",
    "1075655": "Doesn't concat require storing two copies of the Dataframe in memory?\nHow would you work around that with such a large dataset?",
    "1075819": "From my experiment because each iter in submission consist of small total rows its better to just use loop and dict",
    "1077210": "that's very useful, thanks",
    "1077211": "I believe that too,\nmerge seems good with big datasets, but with small chunks, I thinks it is slow",
    "1077314": "use join to replace merge ,can save 50% time",
    "1077535": "How to measure the memory cost for the step?",
    "1077850": "tkm2261 has shared his wonderful methods in the [discussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197023) and [notebook](https://www.kaggle.com/tkm2261/fast-pandas-left-join-357x-faster-than-pd-merge).",
    "1077969": "correct, I tried join, and it saved me about 50%\nthanks for that.\n\nbut, now I am trying faster idea that I saw in below discussion, it might be even faster than join.\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/197023",
    "1077970": "I checked it, its very interesting, I am working on it now.\nthanks for notifying me about it."
  },
  "source": "meta"
}