{
  "id": 197023,
  "title": "357x faster pandas left join",
  "url": "/competitions/riiid-test-answer-prediction/discussion/197023",
  "author_name": "",
  "post_date": "2020-11-13T21:38:03.236237400Z",
  "votes": 50,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I reimplemented the left outer join with <code>pd.reindex</code> and <code>pd.concat</code>. It is around 357 times faster <code>pd.merge</code> as shown in the below kernel.  </p>\n<p><a href=\"https://www.kaggle.com/tkm2261/fast-pandas-left-join-357x-faster\" target=\"_blank\">https://www.kaggle.com/tkm2261/fast-pandas-left-join-357x-faster</a></p>\n<p>I was very surprised even though my method is only applicable to a particular case: The left table is small, the right table is large, the join key is single, and the index of the right table is the join key. Do you know some reasons why so slow?</p>",
  "messages": [
    {
      "id": "1077731",
      "postDate": "11/13/2020 21:38:03",
      "content": "<p>I reimplemented the left outer join with <code>pd.reindex</code> and <code>pd.concat</code>. It is around 357 times faster <code>pd.merge</code> as shown in the below kernel.  </p>\n<p><a href=\"https://www.kaggle.com/tkm2261/fast-pandas-left-join-357x-faster\" target=\"_blank\">https://www.kaggle.com/tkm2261/fast-pandas-left-join-357x-faster</a></p>\n<p>I was very surprised even though my method is only applicable to a particular case: The left table is small, the right table is large, the join key is single, and the index of the right table is the join key. Do you know some reasons why so slow?</p>",
      "rawMarkdown": "I reimplemented the left outer join with `pd.reindex` and `pd.concat`. It is around 357 times faster `pd.merge` as shown in the below kernel.  \n\nhttps://www.kaggle.com/tkm2261/fast-pandas-left-join-357x-faster\n\nI was very surprised even though my method is only applicable to a particular case: The left table is small, the right table is large, the join key is single, and the index of the right table is the join key. Do you know some reasons why so slow?",
      "votes": null
    },
    {
      "id": "1077747",
      "postDate": "11/13/2020 22:12:24",
      "content": "<p>You can make merge 10x faster by changing it to use index - change<br>\n<code>df_test.merge(df_user, how='left', on='user_id')</code><br>\nto<br>\n<code>df_test.merge(df_user, how='left', left_on='user_id', right_index=True)</code><br>\nbut it still will be slower than concat.</p>\n<p>I'm currently using loops myself. Will compare my loops with your concat method - maybe your approach is faster.</p>",
      "rawMarkdown": "You can make merge 10x faster by changing it to use index - change\n`df_test.merge(df_user, how='left', on='user_id')`\nto\n`df_test.merge(df_user, how='left', left_on='user_id', right_index=True)`\nbut it still will be slower than concat.\n\nI'm currently using loops myself. Will compare my loops with your concat method - maybe your approach is faster.",
      "votes": null
    },
    {
      "id": "1077750",
      "postDate": "11/13/2020 22:18:34",
      "content": "<p>Thanks! I will add the <code>right_index=True</code> results. I am also thinking of doing loops, i.e., everything is done in numpy. But, it needs heavy engineering efforts. Plz let us know if your method outperforms :-). </p>",
      "rawMarkdown": "Thanks! I will add the `right_index=True` results. I am also thinking of doing loops, i.e., everything is done in numpy. But, it needs heavy engineering efforts. Plz let us know if your method outperforms :-).",
      "votes": null
    },
    {
      "id": "1077786",
      "postDate": "11/13/2020 23:25:08",
      "content": "<p>For me the bottleneck was frequent data manipulation as concat worked pretty fast for final merging. So I moved to loops and it's been much easier, customizable &amp; faster… but yeah creating the inference pipeline was a little tricky. </p>",
      "rawMarkdown": "For me the bottleneck was frequent data manipulation as concat worked pretty fast for final merging. So I moved to loops and it's been much easier, customizable & faster... but yeah creating the inference pipeline was a little tricky.",
      "votes": null
    },
    {
      "id": "1077952",
      "postDate": "11/14/2020 06:02:02",
      "content": "<p>Another solution that I have been using is to find the present users and use loc to index just those ids. Doing this can make the merge take 65 ms.</p>\n<pre><code>df_test.merge(df_user.loc[u_present], how='left', on='user_id')\n</code></pre>\n<p>Combining this with what <a href=\"https://www.kaggle.com/alijs1\" target=\"_blank\">@alijs1</a> suggested can speed it up slightly to 58 ms </p>\n<pre><code>df_test.merge(df_user.loc[u_present], how='left', left_on='user_id', right_index=True)\n</code></pre>\n<p>Concat is still by far the fastest approach. Thanks for sharing!</p>\n<p>Edit: <code>reindex</code> is a superior approach to using LOC since it fills in NANs automatically, which I didn't know until now. reindex and concat is definitely the way to go :)</p>",
      "rawMarkdown": "Another solution that I have been using is to find the present users and use loc to index just those ids. Doing this can make the merge take 65 ms.\n\n```\ndf_test.merge(df_user.loc[u_present], how='left', on='user_id')\n```\n\nCombining this with what @alijs1 suggested can speed it up slightly to 58 ms \n\n```\ndf_test.merge(df_user.loc[u_present], how='left', left_on='user_id', right_index=True)\n```\n\nConcat is still by far the fastest approach. Thanks for sharing!\n\nEdit: `reindex` is a superior approach to using LOC since it fills in NANs automatically, which I didn't know until now. reindex and concat is definitely the way to go :)",
      "votes": null
    },
    {
      "id": "1077955",
      "postDate": "11/14/2020 06:16:19",
      "content": "<p>Thank you for your sharing! I will also add the result to the kernel.</p>",
      "rawMarkdown": "Thank you for your sharing! I will also add the result to the kernel.",
      "votes": null
    },
    {
      "id": "1077996",
      "postDate": "11/14/2020 07:42:01",
      "content": "<p>simple, yet brilliant idea !<br>\nI applied it and It saved me a lot of wasted submit time 😊<br>\nthanks for sharing.</p>",
      "rawMarkdown": "simple, yet brilliant idea !\nI applied it and It saved me a lot of wasted submit time 😊\nthanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1077747,
      "author_name": "alijs1",
      "author_url": "",
      "post_date": "11/13/2020 22:12:24",
      "content": "<p>You can make merge 10x faster by changing it to use index - change<br>\n<code>df_test.merge(df_user, how='left', on='user_id')</code><br>\nto<br>\n<code>df_test.merge(df_user, how='left', left_on='user_id', right_index=True)</code><br>\nbut it still will be slower than concat.</p>\n<p>I'm currently using loops myself. Will compare my loops with your concat method - maybe your approach is faster.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1077750,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "11/13/2020 22:18:34",
          "content": "<p>Thanks! I will add the <code>right_index=True</code> results. I am also thinking of doing loops, i.e., everything is done in numpy. But, it needs heavy engineering efforts. Plz let us know if your method outperforms :-). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1077786,
          "author_name": "dexarsal",
          "author_url": "",
          "post_date": "11/13/2020 23:25:08",
          "content": "<p>For me the bottleneck was frequent data manipulation as concat worked pretty fast for final merging. So I moved to loops and it's been much easier, customizable &amp; faster… but yeah creating the inference pipeline was a little tricky. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1077952,
      "author_name": "doctorkael",
      "author_url": "",
      "post_date": "11/14/2020 06:02:02",
      "content": "<p>Another solution that I have been using is to find the present users and use loc to index just those ids. Doing this can make the merge take 65 ms.</p>\n<pre><code>df_test.merge(df_user.loc[u_present], how='left', on='user_id')\n</code></pre>\n<p>Combining this with what <a href=\"https://www.kaggle.com/alijs1\" target=\"_blank\">@alijs1</a> suggested can speed it up slightly to 58 ms </p>\n<pre><code>df_test.merge(df_user.loc[u_present], how='left', left_on='user_id', right_index=True)\n</code></pre>\n<p>Concat is still by far the fastest approach. Thanks for sharing!</p>\n<p>Edit: <code>reindex</code> is a superior approach to using LOC since it fills in NANs automatically, which I didn't know until now. reindex and concat is definitely the way to go :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1077955,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "11/14/2020 06:16:19",
          "content": "<p>Thank you for your sharing! I will also add the result to the kernel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1077996,
      "author_name": "mohamadnawfal",
      "author_url": "",
      "post_date": "11/14/2020 07:42:01",
      "content": "<p>simple, yet brilliant idea !<br>\nI applied it and It saved me a lot of wasted submit time 😊<br>\nthanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1077731": "I reimplemented the left outer join with `pd.reindex` and `pd.concat`. It is around 357 times faster `pd.merge` as shown in the below kernel.  \n\nhttps://www.kaggle.com/tkm2261/fast-pandas-left-join-357x-faster\n\nI was very surprised even though my method is only applicable to a particular case: The left table is small, the right table is large, the join key is single, and the index of the right table is the join key. Do you know some reasons why so slow?",
    "1077747": "You can make merge 10x faster by changing it to use index - change\n`df_test.merge(df_user, how='left', on='user_id')`\nto\n`df_test.merge(df_user, how='left', left_on='user_id', right_index=True)`\nbut it still will be slower than concat.\n\nI'm currently using loops myself. Will compare my loops with your concat method - maybe your approach is faster.",
    "1077750": "Thanks! I will add the `right_index=True` results. I am also thinking of doing loops, i.e., everything is done in numpy. But, it needs heavy engineering efforts. Plz let us know if your method outperforms :-).",
    "1077786": "For me the bottleneck was frequent data manipulation as concat worked pretty fast for final merging. So I moved to loops and it's been much easier, customizable & faster... but yeah creating the inference pipeline was a little tricky.",
    "1077952": "Another solution that I have been using is to find the present users and use loc to index just those ids. Doing this can make the merge take 65 ms.\n\n```\ndf_test.merge(df_user.loc[u_present], how='left', on='user_id')\n```\n\nCombining this with what @alijs1 suggested can speed it up slightly to 58 ms \n\n```\ndf_test.merge(df_user.loc[u_present], how='left', left_on='user_id', right_index=True)\n```\n\nConcat is still by far the fastest approach. Thanks for sharing!\n\nEdit: `reindex` is a superior approach to using LOC since it fills in NANs automatically, which I didn't know until now. reindex and concat is definitely the way to go :)",
    "1077955": "Thank you for your sharing! I will also add the result to the kernel.",
    "1077996": "simple, yet brilliant idea !\nI applied it and It saved me a lot of wasted submit time 😊\nthanks for sharing."
  },
  "source": "meta"
}