{
  "id": 196009,
  "title": "Ditch Pandas - from 9 hours submission time to 20 minutes.",
  "url": "/competitions/riiid-test-answer-prediction/discussion/196009",
  "author_name": "",
  "post_date": "2020-11-08T17:29:00.951113700Z",
  "votes": 72,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I have seen many calling to ditch Pandas, I thought I'd make it official in a discussion.</p>\n<p>My submission time took almost 9 hours, ditching Pandas made scoring take only 20 minutes. Through my findings, for small batches (&lt;30 Rows), looping on numpy arrays and dealing with each row independently is way faster than trying to vectorize the operations. Maybe it'd come obvious for some, but it wasn't obvious for me.</p>",
  "messages": [
    {
      "id": "1072792",
      "postDate": "11/08/2020 17:29:00",
      "content": "<p>I have seen many calling to ditch Pandas, I thought I'd make it official in a discussion.</p>\n<p>My submission time took almost 9 hours, ditching Pandas made scoring take only 20 minutes. Through my findings, for small batches (&lt;30 Rows), looping on numpy arrays and dealing with each row independently is way faster than trying to vectorize the operations. Maybe it'd come obvious for some, but it wasn't obvious for me.</p>",
      "rawMarkdown": "I have seen many calling to ditch Pandas, I thought I'd make it official in a discussion.\n\nMy submission time took almost 9 hours, ditching Pandas made scoring take only 20 minutes. Through my findings, for small batches (<30 Rows), looping on numpy arrays and dealing with each row independently is way faster than trying to vectorize the operations. Maybe it'd come obvious for some, but it wasn't obvious for me.",
      "votes": null
    },
    {
      "id": "1073751",
      "postDate": "11/09/2020 22:17:24",
      "content": "<p>How do you do the merging? You check every key you would need in a pd.merge for every row? </p>",
      "rawMarkdown": "How do you do the merging? You check every key you would need in a pd.merge for every row?",
      "votes": null
    },
    {
      "id": "1073776",
      "postDate": "11/09/2020 22:50:13",
      "content": "<p>I stored the data that I used to use pd.merge on inside python dictionaries, thus lookup time is constant. So while looping, I just retrieve data from those dictionaries and use them to populate my input array. </p>",
      "rawMarkdown": "I stored the data that I used to use pd.merge on inside python dictionaries, thus lookup time is constant. So while looping, I just retrieve data from those dictionaries and use them to populate my input array.",
      "votes": null
    },
    {
      "id": "1074761",
      "postDate": "11/11/2020 03:02:45",
      "content": "<p>9h -&gt; 20min? That's amazing!<br>\nChange DataFrame to ndarray,  loop a dict,  concat by hand, and you saved 8h?!<br>\nMan, that's unbelievable<br>\nI would like to try</p>",
      "rawMarkdown": "9h -> 20min? That's amazing!\nChange DataFrame to ndarray,  loop a dict,  concat by hand, and you saved 8h?!\nMan, that's unbelievable\nI would like to try",
      "votes": null
    },
    {
      "id": "1074928",
      "postDate": "11/11/2020 08:49:56",
      "content": "<p>In terms of feature engineering, do you still use pandas?</p>",
      "rawMarkdown": "In terms of feature engineering, do you still use pandas?",
      "votes": null
    },
    {
      "id": "1075316",
      "postDate": "11/11/2020 15:19:36",
      "content": "<p>Do you mean something like a dictionary with user_id as keys and an array with all the features as value? And then a loop over each line of the test_df and concat? I still can't see it. Could you share a couple of lines of code to see how this is done?</p>",
      "rawMarkdown": "Do you mean something like a dictionary with user_id as keys and an array with all the features as value? And then a loop over each line of the test_df and concat? I still can't see it. Could you share a couple of lines of code to see how this is done?",
      "votes": null
    },
    {
      "id": "1075359",
      "postDate": "11/11/2020 16:31:34",
      "content": "<p>The offline features yes, but during inference no.</p>",
      "rawMarkdown": "The offline features yes, but during inference no.",
      "votes": null
    },
    {
      "id": "1075371",
      "postDate": "11/11/2020 16:44:45",
      "content": "<p>It's pretty straightforward. As you said, I have a dictionary with user_id as a key with dictionaries with their features as values; although I have more than three dictionaries of that sort.</p>\n<ol>\n<li>Create an empty numpy array. Shape: (number of rows, number of feature)</li>\n<li>Loop through every row</li>\n<li>Populate it with data from these dictionaries and input batch</li>\n<li>Proceed to inference and scaling</li>\n</ol>",
      "rawMarkdown": "It's pretty straightforward. As you said, I have a dictionary with user_id as a key with dictionaries with their features as values; although I have more than three dictionaries of that sort.\n1. Create an empty numpy array. Shape: (number of rows, number of feature)\n2. Loop through every row\n3. Populate it with data from these dictionaries and input batch\n4. Proceed to inference and scaling",
      "votes": null
    },
    {
      "id": "1075447",
      "postDate": "11/11/2020 18:19:37",
      "content": "<p>this maybe true for small datasets (I haven't tried myself), but for larger datasets we may need new dataframe handling library which is faster than pandas?? Tell if you know any??</p>",
      "rawMarkdown": "this maybe true for small datasets (I haven't tried myself), but for larger datasets we may need new dataframe handling library which is faster than pandas?? Tell if you know any??",
      "votes": null
    },
    {
      "id": "1075452",
      "postDate": "11/11/2020 18:25:07",
      "content": "<p>I guess during the inference phase the dataset are quite small.</p>",
      "rawMarkdown": "I guess during the inference phase the dataset are quite small.",
      "votes": null
    },
    {
      "id": "1075482",
      "postDate": "11/11/2020 18:40:43",
      "content": "<p>You can use CUDF from Rapids, it uses GPU, speeds up alot</p>",
      "rawMarkdown": "You can use CUDF from Rapids, it uses GPU, speeds up alot",
      "votes": null
    },
    {
      "id": "1075752",
      "postDate": "11/11/2020 23:43:58",
      "content": "<p>xarray might be worth a try.</p>",
      "rawMarkdown": "xarray might be worth a try.",
      "votes": null
    },
    {
      "id": "1075770",
      "postDate": "11/12/2020 00:21:19",
      "content": "<p>Seems interesting, does kaggle support it? Or is there a offline dataset to add in the kernel?</p>",
      "rawMarkdown": "Seems interesting, does kaggle support it? Or is there a offline dataset to add in the kernel?",
      "votes": null
    },
    {
      "id": "1077110",
      "postDate": "11/13/2020 08:46:16",
      "content": "<p>I applied same idea, but got small improvement in the execution time<br>\nfrom 3:00 hrs ---&gt; it became 2:30 hours (0:30 minutes improvement).<br>\nbelow is how I implemented it:</p>\n<pre><code>env = riiideducation.make_env()\niter_test = env.iter_test()\nfor (test_df, sample_prediction_df) in iter_test:\n    np_test_df = test_df[['user_id', 'content_id']].to_numpy()\n    X_submit = np.empty((len(np_test_df), len(model_cols)))\n    for i in np.arange(len(X_submit)):\n        X_submit[i][0] = users_dict[np_test_df[i][0]]\n        X_submit[i][1] = questions_dict[np_test_df[i][1]]\n    test_df['answered_correctly'] = model.predict(X_submit)\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n</code></pre>",
      "rawMarkdown": "I applied same idea, but got small improvement in the execution time\nfrom 3:00 hrs ---> it became 2:30 hours (0:30 minutes improvement).\nbelow is how I implemented it:\n```\nenv = riiideducation.make_env()\niter_test = env.iter_test()\nfor (test_df, sample_prediction_df) in iter_test:\n    np_test_df = test_df[['user_id', 'content_id']].to_numpy()\n    X_submit = np.empty((len(np_test_df), len(model_cols)))\n    for i in np.arange(len(X_submit)):\n        X_submit[i][0] = users_dict[np_test_df[i][0]]\n        X_submit[i][1] = questions_dict[np_test_df[i][1]]\n    test_df['answered_correctly'] = model.predict(X_submit)\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n```",
      "votes": null
    },
    {
      "id": "1077121",
      "postDate": "11/13/2020 08:55:12",
      "content": "<p><a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a><br>\nI did quick search about xarray, and found below article says that xarray is slow when dealing with  small chunks of data. what is your opinion about that?<br>\n<a href=\"https://github.com/pydata/xarray/issues/2799\" target=\"_blank\">https://github.com/pydata/xarray/issues/2799</a></p>",
      "rawMarkdown": "jpmiller\nI did quick search about xarray, and found below article says that xarray is slow when dealing with  small chunks of data. what is your opinion about that?\nhttps://github.com/pydata/xarray/issues/2799",
      "votes": null
    },
    {
      "id": "1077437",
      "postDate": "11/13/2020 16:03:58",
      "content": "<p>Nice find. I thought that xarray was speedy but clearly not for these functions. In that case it sounds like Numpy is the way to go. Plain old python dictionaries might also work in some cases.</p>",
      "rawMarkdown": "Nice find. I thought that xarray was speedy but clearly not for these functions. In that case it sounds like Numpy is the way to go. Plain old python dictionaries might also work in some cases.",
      "votes": null
    },
    {
      "id": "1078414",
      "postDate": "11/14/2020 18:42:54",
      "content": "<p>Pandas is great for exploring the dataset as well but for operations numpy is better because its datastructure (arrays) are optmized for compute. </p>",
      "rawMarkdown": "Pandas is great for exploring the dataset as well but for operations numpy is better because its datastructure (arrays) are optmized for compute.",
      "votes": null
    },
    {
      "id": "1105323",
      "postDate": "12/07/2020 19:23:38",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/abdessalemboukil\" target=\"_blank\">@abdessalemboukil</a>, I wanted to ask, if I'll be using numpy array everywhere instead of dataframe, then how will I be able to keep track of the column name, especially while merging the data of different dataframes, as I require the name of the column on which I can perform a merge operation.</p>",
      "rawMarkdown": "Hi @abdessalemboukil, I wanted to ask, if I'll be using numpy array everywhere instead of dataframe, then how will I be able to keep track of the column name, especially while merging the data of different dataframes, as I require the name of the column on which I can perform a merge operation.",
      "votes": null
    },
    {
      "id": "1105671",
      "postDate": "12/08/2020 04:32:51",
      "content": "<p>you don't need names in a matrix, keep track of column index rather.</p>",
      "rawMarkdown": "you don't need names in a matrix, keep track of column index rather.",
      "votes": null
    },
    {
      "id": "1140847",
      "postDate": "01/06/2021 10:10:50",
      "content": "<p>If my submission does not exceed 9h limit in the current public leaderboard, can it be sure that it will not exceed the 9h limit later in the private leaderboard? Thanks.</p>",
      "rawMarkdown": "If my submission does not exceed 9h limit in the current public leaderboard, can it be sure that it will not exceed the 9h limit later in the private leaderboard? Thanks.",
      "votes": null
    },
    {
      "id": "1140884",
      "postDate": "01/06/2021 10:46:58",
      "content": "<p>Your score on the private leaderboard is computed at the same time as your public score, so you have nothing to worry about if your submission already succeeded. </p>",
      "rawMarkdown": "Your score on the private leaderboard is computed at the same time as your public score, so you have nothing to worry about if your submission already succeeded.",
      "votes": null
    },
    {
      "id": "1141099",
      "postDate": "01/06/2021 13:53:17",
      "content": "<p>Thank you very much</p>",
      "rawMarkdown": "Thank you very much",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1073751,
      "author_name": "iuryck",
      "author_url": "",
      "post_date": "11/09/2020 22:17:24",
      "content": "<p>How do you do the merging? You check every key you would need in a pd.merge for every row? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1073776,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "11/09/2020 22:50:13",
          "content": "<p>I stored the data that I used to use pd.merge on inside python dictionaries, thus lookup time is constant. So while looping, I just retrieve data from those dictionaries and use them to populate my input array. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075316,
          "author_name": "jcesquiveld",
          "author_url": "",
          "post_date": "11/11/2020 15:19:36",
          "content": "<p>Do you mean something like a dictionary with user_id as keys and an array with all the features as value? And then a loop over each line of the test_df and concat? I still can't see it. Could you share a couple of lines of code to see how this is done?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075371,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "11/11/2020 16:44:45",
          "content": "<p>It's pretty straightforward. As you said, I have a dictionary with user_id as a key with dictionaries with their features as values; although I have more than three dictionaries of that sort.</p>\n<ol>\n<li>Create an empty numpy array. Shape: (number of rows, number of feature)</li>\n<li>Loop through every row</li>\n<li>Populate it with data from these dictionaries and input batch</li>\n<li>Proceed to inference and scaling</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1074761,
      "author_name": "solomonxian",
      "author_url": "",
      "post_date": "11/11/2020 03:02:45",
      "content": "<p>9h -&gt; 20min? That's amazing!<br>\nChange DataFrame to ndarray,  loop a dict,  concat by hand, and you saved 8h?!<br>\nMan, that's unbelievable<br>\nI would like to try</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1074928,
      "author_name": "sparked",
      "author_url": "",
      "post_date": "11/11/2020 08:49:56",
      "content": "<p>In terms of feature engineering, do you still use pandas?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1075359,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "11/11/2020 16:31:34",
          "content": "<p>The offline features yes, but during inference no.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1078414,
          "author_name": "luukhofman",
          "author_url": "",
          "post_date": "11/14/2020 18:42:54",
          "content": "<p>Pandas is great for exploring the dataset as well but for operations numpy is better because its datastructure (arrays) are optmized for compute. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1075447,
      "author_name": "ujjwalsood07",
      "author_url": "",
      "post_date": "11/11/2020 18:19:37",
      "content": "<p>this maybe true for small datasets (I haven't tried myself), but for larger datasets we may need new dataframe handling library which is faster than pandas?? Tell if you know any??</p>",
      "votes": null,
      "replies": [
        {
          "id": 1075452,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "11/11/2020 18:25:07",
          "content": "<p>I guess during the inference phase the dataset are quite small.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075482,
          "author_name": "iuryck",
          "author_url": "",
          "post_date": "11/11/2020 18:40:43",
          "content": "<p>You can use CUDF from Rapids, it uses GPU, speeds up alot</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075752,
          "author_name": "jpmiller",
          "author_url": "",
          "post_date": "11/11/2020 23:43:58",
          "content": "<p>xarray might be worth a try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075770,
          "author_name": "iuryck",
          "author_url": "",
          "post_date": "11/12/2020 00:21:19",
          "content": "<p>Seems interesting, does kaggle support it? Or is there a offline dataset to add in the kernel?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1077121,
          "author_name": "mohamadnawfal",
          "author_url": "",
          "post_date": "11/13/2020 08:55:12",
          "content": "<p><a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a><br>\nI did quick search about xarray, and found below article says that xarray is slow when dealing with  small chunks of data. what is your opinion about that?<br>\n<a href=\"https://github.com/pydata/xarray/issues/2799\" target=\"_blank\">https://github.com/pydata/xarray/issues/2799</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1077437,
          "author_name": "jpmiller",
          "author_url": "",
          "post_date": "11/13/2020 16:03:58",
          "content": "<p>Nice find. I thought that xarray was speedy but clearly not for these functions. In that case it sounds like Numpy is the way to go. Plain old python dictionaries might also work in some cases.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1077110,
      "author_name": "mohamadnawfal",
      "author_url": "",
      "post_date": "11/13/2020 08:46:16",
      "content": "<p>I applied same idea, but got small improvement in the execution time<br>\nfrom 3:00 hrs ---&gt; it became 2:30 hours (0:30 minutes improvement).<br>\nbelow is how I implemented it:</p>\n<pre><code>env = riiideducation.make_env()\niter_test = env.iter_test()\nfor (test_df, sample_prediction_df) in iter_test:\n    np_test_df = test_df[['user_id', 'content_id']].to_numpy()\n    X_submit = np.empty((len(np_test_df), len(model_cols)))\n    for i in np.arange(len(X_submit)):\n        X_submit[i][0] = users_dict[np_test_df[i][0]]\n        X_submit[i][1] = questions_dict[np_test_df[i][1]]\n    test_df['answered_correctly'] = model.predict(X_submit)\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1105323,
      "author_name": "smritisingh1997",
      "author_url": "",
      "post_date": "12/07/2020 19:23:38",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/abdessalemboukil\" target=\"_blank\">@abdessalemboukil</a>, I wanted to ask, if I'll be using numpy array everywhere instead of dataframe, then how will I be able to keep track of the column name, especially while merging the data of different dataframes, as I require the name of the column on which I can perform a merge operation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1105671,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "12/08/2020 04:32:51",
          "content": "<p>you don't need names in a matrix, keep track of column index rather.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1140847,
      "author_name": "",
      "author_url": "",
      "post_date": "01/06/2021 10:10:50",
      "content": "<p>If my submission does not exceed 9h limit in the current public leaderboard, can it be sure that it will not exceed the 9h limit later in the private leaderboard? Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1140884,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "01/06/2021 10:46:58",
          "content": "<p>Your score on the private leaderboard is computed at the same time as your public score, so you have nothing to worry about if your submission already succeeded. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1141099,
      "author_name": "",
      "author_url": "",
      "post_date": "01/06/2021 13:53:17",
      "content": "<p>Thank you very much</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1072792": "I have seen many calling to ditch Pandas, I thought I'd make it official in a discussion.\n\nMy submission time took almost 9 hours, ditching Pandas made scoring take only 20 minutes. Through my findings, for small batches (<30 Rows), looping on numpy arrays and dealing with each row independently is way faster than trying to vectorize the operations. Maybe it'd come obvious for some, but it wasn't obvious for me.",
    "1073751": "How do you do the merging? You check every key you would need in a pd.merge for every row?",
    "1073776": "I stored the data that I used to use pd.merge on inside python dictionaries, thus lookup time is constant. So while looping, I just retrieve data from those dictionaries and use them to populate my input array.",
    "1074761": "9h -> 20min? That's amazing!\nChange DataFrame to ndarray,  loop a dict,  concat by hand, and you saved 8h?!\nMan, that's unbelievable\nI would like to try",
    "1074928": "In terms of feature engineering, do you still use pandas?",
    "1075316": "Do you mean something like a dictionary with user_id as keys and an array with all the features as value? And then a loop over each line of the test_df and concat? I still can't see it. Could you share a couple of lines of code to see how this is done?",
    "1075359": "The offline features yes, but during inference no.",
    "1075371": "It's pretty straightforward. As you said, I have a dictionary with user_id as a key with dictionaries with their features as values; although I have more than three dictionaries of that sort.\n1. Create an empty numpy array. Shape: (number of rows, number of feature)\n2. Loop through every row\n3. Populate it with data from these dictionaries and input batch\n4. Proceed to inference and scaling",
    "1075447": "this maybe true for small datasets (I haven't tried myself), but for larger datasets we may need new dataframe handling library which is faster than pandas?? Tell if you know any??",
    "1075452": "I guess during the inference phase the dataset are quite small.",
    "1075482": "You can use CUDF from Rapids, it uses GPU, speeds up alot",
    "1075752": "xarray might be worth a try.",
    "1075770": "Seems interesting, does kaggle support it? Or is there a offline dataset to add in the kernel?",
    "1077110": "I applied same idea, but got small improvement in the execution time\nfrom 3:00 hrs ---> it became 2:30 hours (0:30 minutes improvement).\nbelow is how I implemented it:\n```\nenv = riiideducation.make_env()\niter_test = env.iter_test()\nfor (test_df, sample_prediction_df) in iter_test:\n    np_test_df = test_df[['user_id', 'content_id']].to_numpy()\n    X_submit = np.empty((len(np_test_df), len(model_cols)))\n    for i in np.arange(len(X_submit)):\n        X_submit[i][0] = users_dict[np_test_df[i][0]]\n        X_submit[i][1] = questions_dict[np_test_df[i][1]]\n    test_df['answered_correctly'] = model.predict(X_submit)\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n```",
    "1077121": "jpmiller\nI did quick search about xarray, and found below article says that xarray is slow when dealing with  small chunks of data. what is your opinion about that?\nhttps://github.com/pydata/xarray/issues/2799",
    "1077437": "Nice find. I thought that xarray was speedy but clearly not for these functions. In that case it sounds like Numpy is the way to go. Plain old python dictionaries might also work in some cases.",
    "1078414": "Pandas is great for exploring the dataset as well but for operations numpy is better because its datastructure (arrays) are optmized for compute.",
    "1105323": "Hi @abdessalemboukil, I wanted to ask, if I'll be using numpy array everywhere instead of dataframe, then how will I be able to keep track of the column name, especially while merging the data of different dataframes, as I require the name of the column on which I can perform a merge operation.",
    "1105671": "you don't need names in a matrix, keep track of column index rather.",
    "1140847": "If my submission does not exceed 9h limit in the current public leaderboard, can it be sure that it will not exceed the 9h limit later in the private leaderboard? Thanks.",
    "1140884": "Your score on the private leaderboard is computed at the same time as your public score, so you have nothing to worry about if your submission already succeeded.",
    "1141099": "Thank you very much"
  },
  "source": "meta"
}