{
  "id": 394837,
  "title": "How to avoid Timeout in inference part?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/394837",
  "author_name": "",
  "post_date": "2023-03-15T02:05:28.399092800Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I'm stuck in timeout problem recently. I test nearly 9k features in training and select 1k-2k features for 3 level-group models respectively. However, in inference part, I still need to generate all the 9k features and then select the useful features by loading the importance dict from the training result. Seems this process will cost too much time. <br>\nIs there any way too avoid these redundant features generate again during inference?</p>",
  "messages": [
    {
      "id": "2182203",
      "postDate": "03/15/2023 02:05:28",
      "content": "<p>I'm stuck in timeout problem recently. I test nearly 9k features in training and select 1k-2k features for 3 level-group models respectively. However, in inference part, I still need to generate all the 9k features and then select the useful features by loading the importance dict from the training result. Seems this process will cost too much time. <br>\nIs there any way too avoid these redundant features generate again during inference?</p>",
      "rawMarkdown": "I'm stuck in timeout problem recently. I test nearly 9k features in training and select 1k-2k features for 3 level-group models respectively. However, in inference part, I still need to generate all the 9k features and then select the useful features by loading the importance dict from the training result. Seems this process will cost too much time. \nIs there any way too avoid these redundant features generate again during inference?",
      "votes": null
    },
    {
      "id": "2182253",
      "postDate": "03/15/2023 02:53:12",
      "content": "<p>hey, <br>\nIf we don't know how you create the features we can't tell you what you should change, <br>\nhowever, from my understanding of your problem, you have features that are created looping over rooms, fqid and other stuff and you <strong>don't want all of them but only specific ones</strong> correct ? if so, consider creating a new list that is kind of a subset of the original fqid or room list too loop over less items.<br>\nIf your problem is that all of your 9k features are useful but not for every model, a simple <code>if</code> statement would do the job where you would only create them for specific level_groups, which can easily be done since for <strong>inference they give us data level groups at a time.</strong></p>\n<p><em>Off topic</em>, but you said you only used 3 models, you seem to be doing fine but check out my discussion on the topic <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/393614\" target=\"_blank\">here</a>, maybe you'll try it out and get an even better score. I gave arguments to why you shoundn't consider your approach.</p>",
      "rawMarkdown": "hey, \nIf we don't know how you create the features we can't tell you what you should change, \nhowever, from my understanding of your problem, you have features that are created looping over rooms, fqid and other stuff and you **don't want all of them but only specific ones** correct ? if so, consider creating a new list that is kind of a subset of the original fqid or room list too loop over less items.\nIf your problem is that all of your 9k features are useful but not for every model, a simple `if` statement would do the job where you would only create them for specific level_groups, which can easily be done since for **inference they give us data level groups at a time.**\n\n*Off topic*, but you said you only used 3 models, you seem to be doing fine but check out my discussion on the topic [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/393614), maybe you'll try it out and get an even better score. I gave arguments to why you shoundn't consider your approach.",
      "votes": null
    },
    {
      "id": "2182254",
      "postDate": "03/15/2023 02:54:43",
      "content": "<p>One last thing, have you considered <strong>polars</strong>? it's a pandas's like library that has way faster computing speed. If not, give it a try.</p>",
      "rawMarkdown": "One last thing, have you considered **polars**? it's a pandas's like library that has way faster computing speed. If not, give it a try.",
      "votes": null
    },
    {
      "id": "2182721",
      "postDate": "03/15/2023 09:11:42",
      "content": "<p>I reuse the <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> xgb baseline to make a proof of concept on how to estimate inference time and how to debug bottlenecks in inference pipeline. you can find notebook  <a href=\"https://www.kaggle.com/code/steubk/xgboost-baseline-and-inference-time-estimation\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "I reuse the @cdeotte xgb baseline to make a proof of concept on how to estimate inference time and how to debug bottlenecks in inference pipeline. you can find notebook  [here](https://www.kaggle.com/code/steubk/xgboost-baseline-and-inference-time-estimation)",
      "votes": null
    },
    {
      "id": "2182892",
      "postDate": "03/15/2023 11:20:04",
      "content": "<p>Completely agree with <a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> on this one - using subsets of the column values you want to calculate, if statements to only include specific aggregates, and using polars are all great steps. </p>\n<p>It could also be worth making a submission (or running the excellent inference time estimation notebook that <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> shared) with just your feature engineering to see how long that takes and compare it to the model predictions.</p>\n<p>Out of interest, how are you selecting your features? Is it just using the top feature importances from the models with a different number of features for each level group, or are you doing something different like permutation importance or recursively removing features to find the optimal subset for each model? I think feature selection is probably the next key step for these models, and I'm still playing around with some different methods so would be really interested in hearing how you're selecting your features!</p>",
      "rawMarkdown": "Completely agree with @janmpia on this one - using subsets of the column values you want to calculate, if statements to only include specific aggregates, and using polars are all great steps. \n\nIt could also be worth making a submission (or running the excellent inference time estimation notebook that @steubk shared) with just your feature engineering to see how long that takes and compare it to the model predictions.\n\nOut of interest, how are you selecting your features? Is it just using the top feature importances from the models with a different number of features for each level group, or are you doing something different like permutation importance or recursively removing features to find the optimal subset for each model? I think feature selection is probably the next key step for these models, and I'm still playing around with some different methods so would be really interested in hearing how you're selecting your features!",
      "votes": null
    },
    {
      "id": "2182897",
      "postDate": "03/15/2023 11:22:10",
      "content": "<p>And speaking of polars, <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> made these great notebooks showing how we can use polars at inference time. I think they've got another notebook somewhere for training too :) <a href=\"https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference\" target=\"_blank\">https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference</a></p>",
      "rawMarkdown": "And speaking of polars, @carnozhao made these great notebooks showing how we can use polars at inference time. I think they've got another notebook somewhere for training too :) https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference",
      "votes": null
    },
    {
      "id": "2183452",
      "postDate": "03/15/2023 17:09:11",
      "content": "<p>Based on my experiments, if you can run your inference code within 8-9s with the sample test dataset, it will not be timeout when submitting</p>",
      "rawMarkdown": "Based on my experiments, if you can run your inference code within 8-9s with the sample test dataset, it will not be timeout when submitting",
      "votes": null
    },
    {
      "id": "2183850",
      "postDate": "03/16/2023 00:59:57",
      "content": "<p>Thanks a lot. It's really helpful for me!</p>",
      "rawMarkdown": "Thanks a lot. It's really helpful for me!",
      "votes": null
    },
    {
      "id": "2183852",
      "postDate": "03/16/2023 01:01:17",
      "content": "<p>Thanks for your reply. I have used polars. Maybe I should loop only specific items as you said.</p>",
      "rawMarkdown": "Thanks for your reply. I have used polars. Maybe I should loop only specific items as you said.",
      "votes": null
    },
    {
      "id": "2183858",
      "postDate": "03/16/2023 01:07:44",
      "content": "<p>Thanks for your reply, I will try as u said. I simply add the remove collinear variables method to the baseline. Here is my feature selection log. By the way, your \"Saving predictions from previous predictions\" helps me a lot :)<br>\n:)<br>\ndf1 shape:  (11779, 9344)<br>\ndf2 shape:  (11779, 9356)<br>\ndf3 shape:  (11779, 9352)</p>\n<h6>#</h6>\n<p>we need to drop number of columns with over 90% nan values in df1:  6788<br>\nwe need to drop number of columns with over 90% nan values in df2:  5263<br>\nwe need to drop number of columns with over 90% nan values in df3:  4705</p>\n<h6>#</h6>\n<p>then we need to drop number of columns with all the same value in df1:  211<br>\nthen we need to drop number of columns with all the same value in df2:  153<br>\nthen we need to drop number of columns with all the same value in df3:  112</p>\n<h6>#</h6>\n<p>then we need to drop number of Collinear Variables in df1. 1360<br>\nthen we need to drop number of Collinear Variables in df2. 2163<br>\nthen we need to drop number of Collinear Variables in df3. 2374</p>\n<h6>#</h6>\n<p>we need to drop total number of columns in df1:  8359<br>\nwe need to drop total number of columns in df2:  7579<br>\nwe need to drop total number of columns in df3:  7191</p>",
      "rawMarkdown": "Thanks for your reply, I will try as u said. I simply add the remove collinear variables method to the baseline. Here is my feature selection log. By the way, your \"Saving predictions from previous predictions\" helps me a lot :)\n:)\ndf1 shape:  (11779, 9344)\ndf2 shape:  (11779, 9356)\ndf3 shape:  (11779, 9352)\n#########################\nwe need to drop number of columns with over 90% nan values in df1:  6788\nwe need to drop number of columns with over 90% nan values in df2:  5263\nwe need to drop number of columns with over 90% nan values in df3:  4705\n#########################\nthen we need to drop number of columns with all the same value in df1:  211\nthen we need to drop number of columns with all the same value in df2:  153\nthen we need to drop number of columns with all the same value in df3:  112\n#########################\nthen we need to drop number of Collinear Variables in df1. 1360\nthen we need to drop number of Collinear Variables in df2. 2163\nthen we need to drop number of Collinear Variables in df3. 2374\n#########################\nwe need to drop total number of columns in df1:  8359\nwe need to drop total number of columns in df2:  7579\nwe need to drop total number of columns in df3:  7191",
      "votes": null
    },
    {
      "id": "2188800",
      "postDate": "03/20/2023 02:08:48",
      "content": "<p>One thing that I found really helpful especially for polars is that you can select features even at the feature generation process.<br>\nFor example,</p>\n<pre><code>pruned_cols = []\nfor col in cols:\n     if select_col(col):\n        pruned_cols.append(col)\ndf = df.groupby(\"session_id\", \"level_group).agg(pruned_cols)\n</code></pre>\n<p>The <code>select_col</code> function might check whether the column's name is in the list of features that you want. With this code snippet, you only generate the features that you need. </p>",
      "rawMarkdown": "One thing that I found really helpful especially for polars is that you can select features even at the feature generation process.\nFor example,\n\n```\npruned_cols = []\nfor col in cols:\n     if select_col(col):\n        pruned_cols.append(col)\ndf = df.groupby(\"session_id\", \"level_group).agg(pruned_cols)\n```\n\nThe `select_col` function might check whether the column's name is in the list of features that you want. With this code snippet, you only generate the features that you need.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2182253,
      "author_name": "janmpia",
      "author_url": "",
      "post_date": "03/15/2023 02:53:12",
      "content": "<p>hey, <br>\nIf we don't know how you create the features we can't tell you what you should change, <br>\nhowever, from my understanding of your problem, you have features that are created looping over rooms, fqid and other stuff and you <strong>don't want all of them but only specific ones</strong> correct ? if so, consider creating a new list that is kind of a subset of the original fqid or room list too loop over less items.<br>\nIf your problem is that all of your 9k features are useful but not for every model, a simple <code>if</code> statement would do the job where you would only create them for specific level_groups, which can easily be done since for <strong>inference they give us data level groups at a time.</strong></p>\n<p><em>Off topic</em>, but you said you only used 3 models, you seem to be doing fine but check out my discussion on the topic <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/393614\" target=\"_blank\">here</a>, maybe you'll try it out and get an even better score. I gave arguments to why you shoundn't consider your approach.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2182254,
          "author_name": "janmpia",
          "author_url": "",
          "post_date": "03/15/2023 02:54:43",
          "content": "<p>One last thing, have you considered <strong>polars</strong>? it's a pandas's like library that has way faster computing speed. If not, give it a try.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2183852,
              "author_name": "chris2042",
              "author_url": "",
              "post_date": "03/16/2023 01:01:17",
              "content": "<p>Thanks for your reply. I have used polars. Maybe I should loop only specific items as you said.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2182721,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "03/15/2023 09:11:42",
      "content": "<p>I reuse the <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> xgb baseline to make a proof of concept on how to estimate inference time and how to debug bottlenecks in inference pipeline. you can find notebook  <a href=\"https://www.kaggle.com/code/steubk/xgboost-baseline-and-inference-time-estimation\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2183850,
          "author_name": "chris2042",
          "author_url": "",
          "post_date": "03/16/2023 00:59:57",
          "content": "<p>Thanks a lot. It's really helpful for me!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2182892,
      "author_name": "judehunt23",
      "author_url": "",
      "post_date": "03/15/2023 11:20:04",
      "content": "<p>Completely agree with <a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> on this one - using subsets of the column values you want to calculate, if statements to only include specific aggregates, and using polars are all great steps. </p>\n<p>It could also be worth making a submission (or running the excellent inference time estimation notebook that <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> shared) with just your feature engineering to see how long that takes and compare it to the model predictions.</p>\n<p>Out of interest, how are you selecting your features? Is it just using the top feature importances from the models with a different number of features for each level group, or are you doing something different like permutation importance or recursively removing features to find the optimal subset for each model? I think feature selection is probably the next key step for these models, and I'm still playing around with some different methods so would be really interested in hearing how you're selecting your features!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2182897,
          "author_name": "judehunt23",
          "author_url": "",
          "post_date": "03/15/2023 11:22:10",
          "content": "<p>And speaking of polars, <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> made these great notebooks showing how we can use polars at inference time. I think they've got another notebook somewhere for training too :) <a href=\"https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference\" target=\"_blank\">https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2183858,
              "author_name": "chris2042",
              "author_url": "",
              "post_date": "03/16/2023 01:07:44",
              "content": "<p>Thanks for your reply, I will try as u said. I simply add the remove collinear variables method to the baseline. Here is my feature selection log. By the way, your \"Saving predictions from previous predictions\" helps me a lot :)<br>\n:)<br>\ndf1 shape:  (11779, 9344)<br>\ndf2 shape:  (11779, 9356)<br>\ndf3 shape:  (11779, 9352)</p>\n<h6>#</h6>\n<p>we need to drop number of columns with over 90% nan values in df1:  6788<br>\nwe need to drop number of columns with over 90% nan values in df2:  5263<br>\nwe need to drop number of columns with over 90% nan values in df3:  4705</p>\n<h6>#</h6>\n<p>then we need to drop number of columns with all the same value in df1:  211<br>\nthen we need to drop number of columns with all the same value in df2:  153<br>\nthen we need to drop number of columns with all the same value in df3:  112</p>\n<h6>#</h6>\n<p>then we need to drop number of Collinear Variables in df1. 1360<br>\nthen we need to drop number of Collinear Variables in df2. 2163<br>\nthen we need to drop number of Collinear Variables in df3. 2374</p>\n<h6>#</h6>\n<p>we need to drop total number of columns in df1:  8359<br>\nwe need to drop total number of columns in df2:  7579<br>\nwe need to drop total number of columns in df3:  7191</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2183452,
      "author_name": "minhtu123",
      "author_url": "",
      "post_date": "03/15/2023 17:09:11",
      "content": "<p>Based on my experiments, if you can run your inference code within 8-9s with the sample test dataset, it will not be timeout when submitting</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2188800,
      "author_name": "mrhantato",
      "author_url": "",
      "post_date": "03/20/2023 02:08:48",
      "content": "<p>One thing that I found really helpful especially for polars is that you can select features even at the feature generation process.<br>\nFor example,</p>\n<pre><code>pruned_cols = []\nfor col in cols:\n     if select_col(col):\n        pruned_cols.append(col)\ndf = df.groupby(\"session_id\", \"level_group).agg(pruned_cols)\n</code></pre>\n<p>The <code>select_col</code> function might check whether the column's name is in the list of features that you want. With this code snippet, you only generate the features that you need. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2182203": "I'm stuck in timeout problem recently. I test nearly 9k features in training and select 1k-2k features for 3 level-group models respectively. However, in inference part, I still need to generate all the 9k features and then select the useful features by loading the importance dict from the training result. Seems this process will cost too much time. \nIs there any way too avoid these redundant features generate again during inference?",
    "2182253": "hey, \nIf we don't know how you create the features we can't tell you what you should change, \nhowever, from my understanding of your problem, you have features that are created looping over rooms, fqid and other stuff and you **don't want all of them but only specific ones** correct ? if so, consider creating a new list that is kind of a subset of the original fqid or room list too loop over less items.\nIf your problem is that all of your 9k features are useful but not for every model, a simple `if` statement would do the job where you would only create them for specific level_groups, which can easily be done since for **inference they give us data level groups at a time.**\n\n*Off topic*, but you said you only used 3 models, you seem to be doing fine but check out my discussion on the topic [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/393614), maybe you'll try it out and get an even better score. I gave arguments to why you shoundn't consider your approach.",
    "2182254": "One last thing, have you considered **polars**? it's a pandas's like library that has way faster computing speed. If not, give it a try.",
    "2182721": "I reuse the @cdeotte xgb baseline to make a proof of concept on how to estimate inference time and how to debug bottlenecks in inference pipeline. you can find notebook  [here](https://www.kaggle.com/code/steubk/xgboost-baseline-and-inference-time-estimation)",
    "2182892": "Completely agree with @janmpia on this one - using subsets of the column values you want to calculate, if statements to only include specific aggregates, and using polars are all great steps. \n\nIt could also be worth making a submission (or running the excellent inference time estimation notebook that @steubk shared) with just your feature engineering to see how long that takes and compare it to the model predictions.\n\nOut of interest, how are you selecting your features? Is it just using the top feature importances from the models with a different number of features for each level group, or are you doing something different like permutation importance or recursively removing features to find the optimal subset for each model? I think feature selection is probably the next key step for these models, and I'm still playing around with some different methods so would be really interested in hearing how you're selecting your features!",
    "2182897": "And speaking of polars, @carnozhao made these great notebooks showing how we can use polars at inference time. I think they've got another notebook somewhere for training too :) https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference",
    "2183452": "Based on my experiments, if you can run your inference code within 8-9s with the sample test dataset, it will not be timeout when submitting",
    "2183850": "Thanks a lot. It's really helpful for me!",
    "2183852": "Thanks for your reply. I have used polars. Maybe I should loop only specific items as you said.",
    "2183858": "Thanks for your reply, I will try as u said. I simply add the remove collinear variables method to the baseline. Here is my feature selection log. By the way, your \"Saving predictions from previous predictions\" helps me a lot :)\n:)\ndf1 shape:  (11779, 9344)\ndf2 shape:  (11779, 9356)\ndf3 shape:  (11779, 9352)\n#########################\nwe need to drop number of columns with over 90% nan values in df1:  6788\nwe need to drop number of columns with over 90% nan values in df2:  5263\nwe need to drop number of columns with over 90% nan values in df3:  4705\n#########################\nthen we need to drop number of columns with all the same value in df1:  211\nthen we need to drop number of columns with all the same value in df2:  153\nthen we need to drop number of columns with all the same value in df3:  112\n#########################\nthen we need to drop number of Collinear Variables in df1. 1360\nthen we need to drop number of Collinear Variables in df2. 2163\nthen we need to drop number of Collinear Variables in df3. 2374\n#########################\nwe need to drop total number of columns in df1:  8359\nwe need to drop total number of columns in df2:  7579\nwe need to drop total number of columns in df3:  7191",
    "2188800": "One thing that I found really helpful especially for polars is that you can select features even at the feature generation process.\nFor example,\n\n```\npruned_cols = []\nfor col in cols:\n     if select_col(col):\n        pruned_cols.append(col)\ndf = df.groupby(\"session_id\", \"level_group).agg(pruned_cols)\n```\n\nThe `select_col` function might check whether the column's name is in the list of features that you want. With this code snippet, you only generate the features that you need."
  },
  "source": "meta"
}