{
  "id": 192779,
  "title": "Reduce submission time",
  "url": "/competitions/riiid-test-answer-prediction/discussion/192779",
  "author_name": "Alex",
  "post_date": "2020-10-23T08:46:07.386000",
  "votes": 7,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi I'm trying to speed-up the feature engineering both in training and inference portions. I've heard that pd.merge() could be slow and after googling I've found that join() could work better.<br>\nMerge:</p>\n<pre><code>%%time\nT2 = pd.merge(X, results_user, on=['user_id'], how=\"left\")\nCPU times: user 977 ms, sys: 489 ms, total: 1.47 s\nWall time: 1.46 s\n</code></pre>\n<p>Join:</p>\n<pre><code>%%time\nX.set_index('user_id', inplace=True)\nT1 = X.join(results_user, how=\"left\")\nX.reset_index(inplace=True)\nCPU times: user 2.03 s, sys: 322 ms, total: 2.35 s\nWall time: 2.34 s\n</code></pre>\n<p>So apparently merge is much faster than join here.. What could be the cause? I'm also trying to improve groupby operations which appear to be really really slow but I'm running out of solution.<br>\nHow to you handle such cases ? </p>",
  "messages": [
    {
      "id": 1058011,
      "postDate": "2020-10-23T08:46:07.387Z",
      "content": "<p>Hi I'm trying to speed-up the feature engineering both in training and inference portions. I've heard that pd.merge() could be slow and after googling I've found that join() could work better.<br>\nMerge:</p>\n<pre><code>%%time\nT2 = pd.merge(X, results_user, on=['user_id'], how=\"left\")\nCPU times: user 977 ms, sys: 489 ms, total: 1.47 s\nWall time: 1.46 s\n</code></pre>\n<p>Join:</p>\n<pre><code>%%time\nX.set_index('user_id', inplace=True)\nT1 = X.join(results_user, how=\"left\")\nX.reset_index(inplace=True)\nCPU times: user 2.03 s, sys: 322 ms, total: 2.35 s\nWall time: 2.34 s\n</code></pre>\n<p>So apparently merge is much faster than join here.. What could be the cause? I'm also trying to improve groupby operations which appear to be really really slow but I'm running out of solution.<br>\nHow to you handle such cases ? </p>",
      "rawMarkdown": "Hi I'm trying to speed-up the feature engineering both in training and inference portions. I've heard that pd.merge() could be slow and after googling I've found that join() could work better.\nMerge:\n```\n%%time\nT2 = pd.merge(X, results_user, on=['user_id'], how=\"left\")\nCPU times: user 977 ms, sys: 489 ms, total: 1.47 s\nWall time: 1.46 s\n```\n\nJoin:\n```\n%%time\nX.set_index('user_id', inplace=True)\nT1 = X.join(results_user, how=\"left\")\nX.reset_index(inplace=True)\nCPU times: user 2.03 s, sys: 322 ms, total: 2.35 s\nWall time: 2.34 s\n```\nSo apparently merge is much faster than join here.. What could be the cause? I'm also trying to improve groupby operations which appear to be really really slow but I'm running out of solution.\nHow to you handle such cases ? \n",
      "votes": 7
    },
    {
      "id": 1058802,
      "postDate": "2020-10-24T09:55:05.587Z",
      "content": "<p>You could try out <code>datatable</code> for merging (<a href=\"https://github.com/h2oai/datatable)\" target=\"_blank\">https://github.com/h2oai/datatable)</a>. On my machine, the example below takes about 26 seconds using pandas vs. ~6 seconds using datatable.</p>\n<pre><code>import  pandas as pd\nfrom tqdm import tqdm\nimport datatable as dt\nfrom datatable import join\nfrom time import time\n\ntrain = pd.read_csv(\"./data/riiid-test-answer-prediction/train.csv\")\ntest = pd.read_csv(\"./data/riiid-test-answer-prediction/example_test.csv\")\nlectures = pd.read_csv(\"./data/riiid-test-answer-prediction/lectures.csv\")\n\n# Replicate the test data 100 times to get a better idea of timing\ndf_list = []\nfor i, _ in tqdm(enumerate(range(100))):\n    i = 3*i + 1\n    tmp = test.copy()\n    tmp['group_num'] = ( tmp['group_num'] + 1 )* i\n    df_list.append(tmp)\n\ntest = pd.concat(df_list)\n\n# Create summary data to merge with\nuser_lectures_bin = train.groupby('user_id')['content_type_id'].max().reset_index()\nuser_lectures_bin.columns = ['user_id','content_type_id_max']\n\ntest_iter = test.groupby('group_num')\n\n\npandas_merge_list = []\nstart_time = time()\nfor i, grp in tqdm(test_iter):\n    grp = grp.merge(user_lectures_bin, how='left', on=['user_id'])\n    pandas_merge_list.append(grp)\nprint(time() - start_time)\n# ~26 seconds\npandas_merge = pd.concat(pandas_merge_list)\n\n\n# Convert pandas frame to datatable\nuser_lectures_bin_dt = dt.Frame(user_lectures_bin)\n\ndatatable_merge_list = []\nstart_time = time()\n#include setting the key here because it takes a little extra time\nuser_lectures_bin_dt.key = 'user_id'\nfor i, grp in tqdm(test_iter):\n    grp = dt.Frame(grp)\n    grp = grp[:,:,join(user_lectures_bin_dt)]\n    grp = grp.to_pandas()\n    datatable_merge_list.append(grp)\nprint(time() - start_time)\n# ~6 seconds\ndatatable_merge = pd.concat(datatable_merge_list)\n\n# Make sure the results are the same\npandas_merge.equals(datatable_merge)\n#True\n</code></pre>",
      "rawMarkdown": "You could try out `datatable` for merging (https://github.com/h2oai/datatable). On my machine, the example below takes about 26 seconds using pandas vs. ~6 seconds using datatable.\n\n```\nimport  pandas as pd\nfrom tqdm import tqdm\nimport datatable as dt\nfrom datatable import join\nfrom time import time\n\ntrain = pd.read_csv(\"./data/riiid-test-answer-prediction/train.csv\")\ntest = pd.read_csv(\"./data/riiid-test-answer-prediction/example_test.csv\")\nlectures = pd.read_csv(\"./data/riiid-test-answer-prediction/lectures.csv\")\n\n# Replicate the test data 100 times to get a better idea of timing\ndf_list = []\nfor i, _ in tqdm(enumerate(range(100))):\n    i = 3*i + 1\n    tmp = test.copy()\n    tmp['group_num'] = ( tmp['group_num'] + 1 )* i\n    df_list.append(tmp)\n\ntest = pd.concat(df_list)\n\n# Create summary data to merge with\nuser_lectures_bin = train.groupby('user_id')['content_type_id'].max().reset_index()\nuser_lectures_bin.columns = ['user_id','content_type_id_max']\n\ntest_iter = test.groupby('group_num')\n\n\npandas_merge_list = []\nstart_time = time()\nfor i, grp in tqdm(test_iter):\n    grp = grp.merge(user_lectures_bin, how='left', on=['user_id'])\n    pandas_merge_list.append(grp)\nprint(time() - start_time)\n# ~26 seconds\npandas_merge = pd.concat(pandas_merge_list)\n\n\n# Convert pandas frame to datatable\nuser_lectures_bin_dt = dt.Frame(user_lectures_bin)\n\ndatatable_merge_list = []\nstart_time = time()\n#include setting the key here because it takes a little extra time\nuser_lectures_bin_dt.key = 'user_id'\nfor i, grp in tqdm(test_iter):\n    grp = dt.Frame(grp)\n    grp = grp[:,:,join(user_lectures_bin_dt)]\n    grp = grp.to_pandas()\n    datatable_merge_list.append(grp)\nprint(time() - start_time)\n# ~6 seconds\ndatatable_merge = pd.concat(datatable_merge_list)\n\n# Make sure the results are the same\npandas_merge.equals(datatable_merge)\n#True\n```\n\n",
      "votes": 3,
      "replies": [
        {
          "id": 1062219,
          "postDate": "2020-10-27T17:02:39.440Z",
          "content": "<p>Thanks for the hint ! </p>",
          "rawMarkdown": "Thanks for the hint ! "
        }
      ]
    },
    {
      "id": 1058493,
      "postDate": "2020-10-23T19:18:15.870Z",
      "content": "<p>I'm at the stage where I'm timing and over-engineering every single line in my feature engineering code to even finish within the 9 hour limit. What I find the trickiest bit is to make the feature engineering be efficient for really large batches (think of preprocessing train data before training a model) but also really small batches (which are apparently quite common in the private test set).</p>\n<p>Some computationally more intensive parts of feature engineering are now written in such a way that a different code logic is used depending on whether the size is large or small, or depending on whether the user IDs contained in it are unique or not. Which can make quite a difference I find.</p>\n<p>Still, I can easily process all train data in a single notebook (or could, if the output file size wasn't limited to 5 GB) but it's reaaaaally close and just about finishes within 9 hours when I submit it. Guess I have to find more ways to scrape some time here and there before I can introduce new features. Or drop not-so-awesome ones.</p>",
      "rawMarkdown": "I'm at the stage where I'm timing and over-engineering every single line in my feature engineering code to even finish within the 9 hour limit. What I find the trickiest bit is to make the feature engineering be efficient for really large batches (think of preprocessing train data before training a model) but also really small batches (which are apparently quite common in the private test set).\n\nSome computationally more intensive parts of feature engineering are now written in such a way that a different code logic is used depending on whether the size is large or small, or depending on whether the user IDs contained in it are unique or not. Which can make quite a difference I find.\n\nStill, I can easily process all train data in a single notebook (or could, if the output file size wasn't limited to 5 GB) but it's reaaaaally close and just about finishes within 9 hours when I submit it. Guess I have to find more ways to scrape some time here and there before I can introduce new features. Or drop not-so-awesome ones.",
      "votes": 2,
      "replies": [
        {
          "id": 1058520,
          "postDate": "2020-10-23T20:10:45.220Z",
          "content": "<blockquote>\n  <p>Still, I can easily process all train data in a single notebook (or could, if the output file size wasn't limited to 5 GB) but it's reaaaaally close and just about finishes within 9 hours when I submit it. Guess I have to find more ways to scrape some time here and there before I can introduce new features. Or drop not-so-awesome ones.</p>\n</blockquote>\n<p>For the training data, you could always just split it into multiple notebooks / datasets as a hack workaround, right? My view is that the test processing is the real FE bottleneck - anything derived statically from train seems like it should be relatively easy to work with, especially since there aren't really that many distinct users at the end of the day. Your size-dependent handling code sounds like a great idea, since squeezing every last bit you can at test time might become the big focus once the low-hanging-fruit / most important features are worked out.</p>",
          "rawMarkdown": "> Still, I can easily process all train data in a single notebook (or could, if the output file size wasn't limited to 5 GB) but it's reaaaaally close and just about finishes within 9 hours when I submit it. Guess I have to find more ways to scrape some time here and there before I can introduce new features. Or drop not-so-awesome ones.\n\nFor the training data, you could always just split it into multiple notebooks / datasets as a hack workaround, right? My view is that the test processing is the real FE bottleneck - anything derived statically from train seems like it should be relatively easy to work with, especially since there aren't really that many distinct users at the end of the day. Your size-dependent handling code sounds like a great idea, since squeezing every last bit you can at test time might become the big focus once the low-hanging-fruit / most important features are worked out."
        },
        {
          "id": 1058524,
          "postDate": "2020-10-23T20:22:57.813Z",
          "content": "<p>Yeah that's what I'm doing, processing the train data in like 3-5 notebooks (which are all done within like an hour) using the same code for preprocessing as in the inference kernel (where it then takes nearly 9 hours on a dataset about 2% of the size of train…).<br>\nEfficient test preprocessing is definitely key here and will eventually limit the features we can use.</p>",
          "rawMarkdown": "Yeah that's what I'm doing, processing the train data in like 3-5 notebooks (which are all done within like an hour) using the same code for preprocessing as in the inference kernel (where it then takes nearly 9 hours on a dataset about 2% of the size of train...).\nEfficient test preprocessing is definitely key here and will eventually limit the features we can use."
        },
        {
          "id": 1058573,
          "postDate": "2020-10-23T22:29:13.520Z",
          "content": "<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> , would you mind to share why it's so slow when doing inference  at test time? I am just curious, personally, I haven't done any work on FE.</p>",
          "rawMarkdown": "@spacelx , would you mind to share why it's so slow when doing inference  at test time? I am just curious, personally, I haven't done any work on FE."
        },
        {
          "id": 1058680,
          "postDate": "2020-10-24T05:06:15.283Z",
          "content": "<p>I think the difference is that for the train set I can process many thousands of rows in a batch (10000 rows take a second or two maybe) whereas the test data appears to be provided in many small batches of like 100 or less rows each on average - processing this will always be less efficient.</p>",
          "rawMarkdown": "I think the difference is that for the train set I can process many thousands of rows in a batch (10000 rows take a second or two maybe) whereas the test data appears to be provided in many small batches of like 100 or less rows each on average - processing this will always be less efficient.",
          "votes": 3
        },
        {
          "id": 1058692,
          "postDate": "2020-10-24T05:45:39.787Z",
          "content": "<p>It might be a good idea to use ~ 100 rows or less per batch in your validation data as well and the runtime to validate 2.5M such rows is so far pretty close to the time taken for my models to score on private test set.</p>\n<p>Maybe not for every experiment since it might take a long time to validate but certainly once you decide to submit a notebook for scoring.</p>",
          "rawMarkdown": "It might be a good idea to use ~ 100 rows or less per batch in your validation data as well and the runtime to validate 2.5M such rows is so far pretty close to the time taken for my models to score on private test set.\n\nMaybe not for every experiment since it might take a long time to validate but certainly once you decide to submit a notebook for scoring.",
          "votes": 3
        },
        {
          "id": 1058702,
          "postDate": "2020-10-24T05:59:53.560Z",
          "content": "<blockquote>\n  <p>the test data appears to be provided in many small batches of like 100 or less rows each on average</p>\n</blockquote>\n<p>Can you share, How did we come to this conclusion?</p>",
          "rawMarkdown": ">the test data appears to be provided in many small batches of like 100 or less rows each on average\n\nCan you share, How did we come to this conclusion?"
        },
        {
          "id": 1058712,
          "postDate": "2020-10-24T06:24:50.457Z",
          "content": "<p>I think it's from <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1056282\" target=\"_blank\">here</a> and while I don't know if 10 or 100 or 1000 or whatever is the right number of not, I certainly feel there is a good chunk of smaller-sized batches (&lt;= 500).</p>\n<p>In general processing 2x size of data once is faster than processing 1x size of data twice. Hence the size of batch and total runtime of scoring are inversely proportional.</p>",
          "rawMarkdown": "I think it's from [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1056282) and while I don't know if 10 or 100 or 1000 or whatever is the right number of not, I certainly feel there is a good chunk of smaller-sized batches (<= 500).\n\nIn general processing 2x size of data once is faster than processing 1x size of data twice. Hence the size of batch and total runtime of scoring are inversely proportional.",
          "votes": 2
        },
        {
          "id": 1058724,
          "postDate": "2020-10-24T06:44:30.740Z",
          "content": "<p>Yep, that's it. Not sure about actual numbers but optimizing for small batches is definitely key.</p>",
          "rawMarkdown": "Yep, that's it. Not sure about actual numbers but optimizing for small batches is definitely key."
        },
        {
          "id": 1058780,
          "postDate": "2020-10-24T08:48:52.403Z",
          "content": "<p>Yes, I am also worried about transformer model won't finish in time. But not tested yet</p>",
          "rawMarkdown": "Yes, I am also worried about transformer model won't finish in time. But not tested yet"
        }
      ]
    },
    {
      "id": 1058147,
      "postDate": "2020-10-23T11:51:58.733Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1058802,
      "author_name": "Branden Murray",
      "author_url": "",
      "post_date": "2020-10-24T09:55:05.587000",
      "content": "<p>You could try out <code>datatable</code> for merging (<a href=\"https://github.com/h2oai/datatable)\" target=\"_blank\">https://github.com/h2oai/datatable)</a>. On my machine, the example below takes about 26 seconds using pandas vs. ~6 seconds using datatable.</p>\n<pre><code>import  pandas as pd\nfrom tqdm import tqdm\nimport datatable as dt\nfrom datatable import join\nfrom time import time\n\ntrain = pd.read_csv(\"./data/riiid-test-answer-prediction/train.csv\")\ntest = pd.read_csv(\"./data/riiid-test-answer-prediction/example_test.csv\")\nlectures = pd.read_csv(\"./data/riiid-test-answer-prediction/lectures.csv\")\n\n# Replicate the test data 100 times to get a better idea of timing\ndf_list = []\nfor i, _ in tqdm(enumerate(range(100))):\n    i = 3*i + 1\n    tmp = test.copy()\n    tmp['group_num'] = ( tmp['group_num'] + 1 )* i\n    df_list.append(tmp)\n\ntest = pd.concat(df_list)\n\n# Create summary data to merge with\nuser_lectures_bin = train.groupby('user_id')['content_type_id'].max().reset_index()\nuser_lectures_bin.columns = ['user_id','content_type_id_max']\n\ntest_iter = test.groupby('group_num')\n\n\npandas_merge_list = []\nstart_time = time()\nfor i, grp in tqdm(test_iter):\n    grp = grp.merge(user_lectures_bin, how='left', on=['user_id'])\n    pandas_merge_list.append(grp)\nprint(time() - start_time)\n# ~26 seconds\npandas_merge = pd.concat(pandas_merge_list)\n\n\n# Convert pandas frame to datatable\nuser_lectures_bin_dt = dt.Frame(user_lectures_bin)\n\ndatatable_merge_list = []\nstart_time = time()\n#include setting the key here because it takes a little extra time\nuser_lectures_bin_dt.key = 'user_id'\nfor i, grp in tqdm(test_iter):\n    grp = dt.Frame(grp)\n    grp = grp[:,:,join(user_lectures_bin_dt)]\n    grp = grp.to_pandas()\n    datatable_merge_list.append(grp)\nprint(time() - start_time)\n# ~6 seconds\ndatatable_merge = pd.concat(datatable_merge_list)\n\n# Make sure the results are the same\npandas_merge.equals(datatable_merge)\n#True\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 1062219,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-10-27T17:02:39.440000",
          "content": "<p>Thanks for the hint ! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1058493,
      "author_name": "Alex Bader",
      "author_url": "",
      "post_date": "2020-10-23T19:18:15.870000",
      "content": "<p>I'm at the stage where I'm timing and over-engineering every single line in my feature engineering code to even finish within the 9 hour limit. What I find the trickiest bit is to make the feature engineering be efficient for really large batches (think of preprocessing train data before training a model) but also really small batches (which are apparently quite common in the private test set).</p>\n<p>Some computationally more intensive parts of feature engineering are now written in such a way that a different code logic is used depending on whether the size is large or small, or depending on whether the user IDs contained in it are unique or not. Which can make quite a difference I find.</p>\n<p>Still, I can easily process all train data in a single notebook (or could, if the output file size wasn't limited to 5 GB) but it's reaaaaally close and just about finishes within 9 hours when I submit it. Guess I have to find more ways to scrape some time here and there before I can introduce new features. Or drop not-so-awesome ones.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1058520,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2020-10-23T20:10:45.220000",
          "content": "<blockquote>\n  <p>Still, I can easily process all train data in a single notebook (or could, if the output file size wasn't limited to 5 GB) but it's reaaaaally close and just about finishes within 9 hours when I submit it. Guess I have to find more ways to scrape some time here and there before I can introduce new features. Or drop not-so-awesome ones.</p>\n</blockquote>\n<p>For the training data, you could always just split it into multiple notebooks / datasets as a hack workaround, right? My view is that the test processing is the real FE bottleneck - anything derived statically from train seems like it should be relatively easy to work with, especially since there aren't really that many distinct users at the end of the day. Your size-dependent handling code sounds like a great idea, since squeezing every last bit you can at test time might become the big focus once the low-hanging-fruit / most important features are worked out.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058524,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-23T20:22:57.813000",
          "content": "<p>Yeah that's what I'm doing, processing the train data in like 3-5 notebooks (which are all done within like an hour) using the same code for preprocessing as in the inference kernel (where it then takes nearly 9 hours on a dataset about 2% of the size of train…).<br>\nEfficient test preprocessing is definitely key here and will eventually limit the features we can use.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058573,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-10-23T22:29:13.520000",
          "content": "<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> , would you mind to share why it's so slow when doing inference  at test time? I am just curious, personally, I haven't done any work on FE.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058680,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-24T05:06:15.283000",
          "content": "<p>I think the difference is that for the train set I can process many thousands of rows in a batch (10000 rows take a second or two maybe) whereas the test data appears to be provided in many small batches of like 100 or less rows each on average - processing this will always be less efficient.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1058692,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-24T05:45:39.787000",
          "content": "<p>It might be a good idea to use ~ 100 rows or less per batch in your validation data as well and the runtime to validate 2.5M such rows is so far pretty close to the time taken for my models to score on private test set.</p>\n<p>Maybe not for every experiment since it might take a long time to validate but certainly once you decide to submit a notebook for scoring.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1058702,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-24T05:59:53.560000",
          "content": "<blockquote>\n  <p>the test data appears to be provided in many small batches of like 100 or less rows each on average</p>\n</blockquote>\n<p>Can you share, How did we come to this conclusion?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058712,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-24T06:24:50.457000",
          "content": "<p>I think it's from <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1056282\" target=\"_blank\">here</a> and while I don't know if 10 or 100 or 1000 or whatever is the right number of not, I certainly feel there is a good chunk of smaller-sized batches (&lt;= 500).</p>\n<p>In general processing 2x size of data once is faster than processing 1x size of data twice. Hence the size of batch and total runtime of scoring are inversely proportional.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1058724,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-24T06:44:30.740000",
          "content": "<p>Yep, that's it. Not sure about actual numbers but optimizing for small batches is definitely key.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058780,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-10-24T08:48:52.403000",
          "content": "<p>Yes, I am also worried about transformer model won't finish in time. But not tested yet</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1058147,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-23T11:51:58.733000",
      "content": "",
      "votes": -3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1058011": "Hi I'm trying to speed-up the feature engineering both in training and inference portions. I've heard that pd.merge() could be slow and after googling I've found that join() could work better.\nMerge:\n```\n%%time\nT2 = pd.merge(X, results_user, on=['user_id'], how=\"left\")\nCPU times: user 977 ms, sys: 489 ms, total: 1.47 s\nWall time: 1.46 s\n```\n\nJoin:\n```\n%%time\nX.set_index('user_id', inplace=True)\nT1 = X.join(results_user, how=\"left\")\nX.reset_index(inplace=True)\nCPU times: user 2.03 s, sys: 322 ms, total: 2.35 s\nWall time: 2.34 s\n```\nSo apparently merge is much faster than join here.. What could be the cause? I'm also trying to improve groupby operations which appear to be really really slow but I'm running out of solution.\nHow to you handle such cases ? \n",
    "1058802": "You could try out `datatable` for merging (https://github.com/h2oai/datatable). On my machine, the example below takes about 26 seconds using pandas vs. ~6 seconds using datatable.\n\n```\nimport  pandas as pd\nfrom tqdm import tqdm\nimport datatable as dt\nfrom datatable import join\nfrom time import time\n\ntrain = pd.read_csv(\"./data/riiid-test-answer-prediction/train.csv\")\ntest = pd.read_csv(\"./data/riiid-test-answer-prediction/example_test.csv\")\nlectures = pd.read_csv(\"./data/riiid-test-answer-prediction/lectures.csv\")\n\n# Replicate the test data 100 times to get a better idea of timing\ndf_list = []\nfor i, _ in tqdm(enumerate(range(100))):\n    i = 3*i + 1\n    tmp = test.copy()\n    tmp['group_num'] = ( tmp['group_num'] + 1 )* i\n    df_list.append(tmp)\n\ntest = pd.concat(df_list)\n\n# Create summary data to merge with\nuser_lectures_bin = train.groupby('user_id')['content_type_id'].max().reset_index()\nuser_lectures_bin.columns = ['user_id','content_type_id_max']\n\ntest_iter = test.groupby('group_num')\n\n\npandas_merge_list = []\nstart_time = time()\nfor i, grp in tqdm(test_iter):\n    grp = grp.merge(user_lectures_bin, how='left', on=['user_id'])\n    pandas_merge_list.append(grp)\nprint(time() - start_time)\n# ~26 seconds\npandas_merge = pd.concat(pandas_merge_list)\n\n\n# Convert pandas frame to datatable\nuser_lectures_bin_dt = dt.Frame(user_lectures_bin)\n\ndatatable_merge_list = []\nstart_time = time()\n#include setting the key here because it takes a little extra time\nuser_lectures_bin_dt.key = 'user_id'\nfor i, grp in tqdm(test_iter):\n    grp = dt.Frame(grp)\n    grp = grp[:,:,join(user_lectures_bin_dt)]\n    grp = grp.to_pandas()\n    datatable_merge_list.append(grp)\nprint(time() - start_time)\n# ~6 seconds\ndatatable_merge = pd.concat(datatable_merge_list)\n\n# Make sure the results are the same\npandas_merge.equals(datatable_merge)\n#True\n```\n\n",
    "1058493": "I'm at the stage where I'm timing and over-engineering every single line in my feature engineering code to even finish within the 9 hour limit. What I find the trickiest bit is to make the feature engineering be efficient for really large batches (think of preprocessing train data before training a model) but also really small batches (which are apparently quite common in the private test set).\n\nSome computationally more intensive parts of feature engineering are now written in such a way that a different code logic is used depending on whether the size is large or small, or depending on whether the user IDs contained in it are unique or not. Which can make quite a difference I find.\n\nStill, I can easily process all train data in a single notebook (or could, if the output file size wasn't limited to 5 GB) but it's reaaaaally close and just about finishes within 9 hours when I submit it. Guess I have to find more ways to scrape some time here and there before I can introduce new features. Or drop not-so-awesome ones.",
    "1058147": ""
  }
}