{
  "id": 401564,
  "title": "What are your inference times?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/401564",
  "author_name": "",
  "post_date": "2023-04-13T20:58:03.164915500Z",
  "votes": 3,
  "comment_count": 38,
  "views": 0,
  "content": "<p>I think my feature engineering is super slow and my total inference time (i.e. from when I make a submission until I see my score) is:</p>\n<p>XGBOOST (pre-trained) | ~9 hours |  LB 0.695</p>\n<p>Can anybody share their inference times for reference? Thanks!</p>",
  "messages": [
    {
      "id": "2220989",
      "postDate": "04/13/2023 20:58:03",
      "content": "<p>I think my feature engineering is super slow and my total inference time (i.e. from when I make a submission until I see my score) is:</p>\n<p>XGBOOST (pre-trained) | ~9 hours |  LB 0.695</p>\n<p>Can anybody share their inference times for reference? Thanks!</p>",
      "rawMarkdown": "I think my feature engineering is super slow and my total inference time (i.e. from when I make a submission until I see my score) is:\n\nXGBOOST (pre-trained) | ~9 hours |  LB 0.695\n\n\nCan anybody share their inference times for reference? Thanks!",
      "votes": null
    },
    {
      "id": "2221106",
      "postDate": "04/14/2023 02:08:52",
      "content": "<p>I didn't have the exact number but it was like 1.5 hours or so for my xgboost models. But I should mention that my submission notebook only contains inference part. Models were trained, saved, and then uploaded to the dataset.</p>",
      "rawMarkdown": "I didn't have the exact number but it was like 1.5 hours or so for my xgboost models. But I should mention that my submission notebook only contains inference part. Models were trained, saved, and then uploaded to the dataset.",
      "votes": null
    },
    {
      "id": "2221362",
      "postDate": "04/14/2023 07:43:37",
      "content": "<p>Thanks for sharing! Yeah, my 9 hours is with loaded and trained xgboost model too, so pretty damn slow xax. If I may ask, how many features do you have?</p>",
      "rawMarkdown": "Thanks for sharing! Yeah, my 9 hours is with loaded and trained xgboost model too, so pretty damn slow xax. If I may ask, how many features do you have?",
      "votes": null
    },
    {
      "id": "2222057",
      "postDate": "04/14/2023 20:52:06",
      "content": "<p>I had ~2900 features in the best case.</p>",
      "rawMarkdown": "I had ~2900 features in the best case.",
      "votes": null
    },
    {
      "id": "2222062",
      "postDate": "04/14/2023 21:17:24",
      "content": "<p>cool, thanks for the info. I also have 2-3k features. I've tried to select top 500/600/1000 features using feature importance but the LB either dropped or was the same. Have you tried feature selection?</p>",
      "rawMarkdown": "cool, thanks for the info. I also have 2-3k features. I've tried to select top 500/600/1000 features using feature importance but the LB either dropped or was the same. Have you tried feature selection?",
      "votes": null
    },
    {
      "id": "2222075",
      "postDate": "04/14/2023 21:34:02",
      "content": "<p>I did, but I arrived at the same conclusion: my CV and LB both drops after retaining features selected by xgboost. So at the end, I keep all features. Perhaps I didn't do it right, but I'm a bit puzzled as why my F1 drops by keeping those \"important features\" only.</p>",
      "rawMarkdown": "I did, but I arrived at the same conclusion: my CV and LB both drops after retaining features selected by xgboost. So at the end, I keep all features. Perhaps I didn't do it right, but I'm a bit puzzled as why my F1 drops by keeping those \"important features\" only.",
      "votes": null
    },
    {
      "id": "2222083",
      "postDate": "04/14/2023 21:41:26",
      "content": "<p>I see. Btw what optimizations do you use to do faster feature engineering during submission? I only use polars..</p>",
      "rawMarkdown": "I see. Btw what optimizations do you use to do faster feature engineering during submission? I only use polars..",
      "votes": null
    },
    {
      "id": "2222086",
      "postDate": "04/14/2023 21:53:03",
      "content": "<p>I'm using polars too. I don't think I have any tricks for speed up, I just did groupby on <code>session_id</code> and <code>level_group</code> then aggregate <code>elapsed_time</code> difference conditioning on various combinations of original and extracted column features.</p>\n<p>By the way, if you want to load pretrained models, do import them outside the <code>iter_test</code> loop. Loading models <em>within</em> the loop is pretty slow.</p>",
      "rawMarkdown": "I'm using polars too. I don't think I have any tricks for speed up, I just did groupby on `session_id` and `level_group` then aggregate `elapsed_time` difference conditioning on various combinations of original and extracted column features.\n\nBy the way, if you want to load pretrained models, do import them outside the `iter_test` loop. Loading models _within_ the loop is pretty slow.",
      "votes": null
    },
    {
      "id": "2222103",
      "postDate": "04/14/2023 22:29:43",
      "content": "<p>That's good advice, very appreciated! </p>\n<p>I am a bit confused why my feature engineering is taking so much time during submission… When I do feature engineering on the train set it takes about 2 minutes, but submission time is hours. Do you experience the same?</p>",
      "rawMarkdown": "That's good advice, very appreciated! \n\nI am a bit confused why my feature engineering is taking so much time during submission... When I do feature engineering on the train set it takes about 2 minutes, but submission time is hours. Do you experience the same?",
      "votes": null
    },
    {
      "id": "2223094",
      "postDate": "04/15/2023 20:29:31",
      "content": "<p>Yeah, same here: feature engineering (FE) was way faster in the training phase than in the inference phase. Without details of the inference part, it seems difficult to diagnose slow FE. A trivial guess is that there is a <em>lot</em> more test data. But this can't explain our difference in inference FE time….</p>",
      "rawMarkdown": "Yeah, same here: feature engineering (FE) was way faster in the training phase than in the inference phase. Without details of the inference part, it seems difficult to diagnose slow FE. A trivial guess is that there is a _lot_ more test data. But this can't explain our difference in inference FE time....",
      "votes": null
    },
    {
      "id": "2223113",
      "postDate": "04/15/2023 21:20:05",
      "content": "<p>To be fair when we submit Kaggle  gives us just 2 CPU-s, and polars uses heavily parallelization. So polars might not be that useful for submission. But I am still suprised…</p>",
      "rawMarkdown": "To be fair when we submit Kaggle  gives us just 2 CPU-s, and polars uses heavily parallelization. So polars might not be that useful for submission. But I am still suprised...",
      "votes": null
    },
    {
      "id": "2224206",
      "postDate": "04/17/2023 05:20:33",
      "content": "<p>That's a good point about the number of available CPUs in the inference phase. I totally forgot about it.</p>",
      "rawMarkdown": "That's a good point about the number of available CPUs in the inference phase. I totally forgot about it.",
      "votes": null
    },
    {
      "id": "2224432",
      "postDate": "04/17/2023 10:02:02",
      "content": "<p>2 hrs with 500~1000 feats</p>",
      "rawMarkdown": "2 hrs with 500~1000 feats",
      "votes": null
    },
    {
      "id": "2224525",
      "postDate": "04/17/2023 12:01:34",
      "content": "<p>that's really good - thanks! Do you mind sharing how you've chosen the 500-1000 features? I suppose you had much more than that and did feature reduction right?</p>",
      "rawMarkdown": "that's really good - thanks! Do you mind sharing how you've chosen the 500-1000 features? I suppose you had much more than that and did feature reduction right?",
      "votes": null
    },
    {
      "id": "2224754",
      "postDate": "04/17/2023 16:19:55",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ngocuong\" target=\"_blank\">@ngocuong</a>. Are you using pandas or polars? Polars is faster.</p>",
      "rawMarkdown": "Hi @ngocuong. Are you using pandas or polars? Polars is faster.",
      "votes": null
    },
    {
      "id": "2224948",
      "postDate": "04/17/2023 19:10:13",
      "content": "<p>hi, I use polars.</p>",
      "rawMarkdown": "hi, I use polars.",
      "votes": null
    },
    {
      "id": "2225492",
      "postDate": "04/18/2023 07:51:18",
      "content": "<p>Permutation importance can help you to select features, but you also need to choose the features which is meaningful(err, it's hard to say). Because some new features are noise. Xgb importance may induce overfitting with data leakage. Kfold using for features selection is also induce overfitting.</p>",
      "rawMarkdown": "Permutation importance can help you to select features, but you also need to choose the features which is meaningful(err, it's hard to say). Because some new features are noise. Xgb importance may induce overfitting with data leakage. Kfold using for features selection is also induce overfitting.",
      "votes": null
    },
    {
      "id": "2225564",
      "postDate": "04/18/2023 08:53:17",
      "content": "<p>thanks for the advice! What CV improvement did you get from permutation importance, if I may ask?</p>",
      "rawMarkdown": "thanks for the advice! What CV improvement did you get from permutation importance, if I may ask?",
      "votes": null
    },
    {
      "id": "2227805",
      "postDate": "04/20/2023 03:05:07",
      "content": "<p>boost +0.003 both CV and LB with permutation importance, and you should do feature selection before model training. In other word, you should separete feature selection from model training.</p>",
      "rawMarkdown": "boost +0.003 both CV and LB with permutation importance, and you should do feature selection before model training. In other word, you should separete feature selection from model training.",
      "votes": null
    },
    {
      "id": "2227827",
      "postDate": "04/20/2023 03:46:00",
      "content": "<p>My inference time is about 40 minutes with aggregating 5-fold models. I used 200 features for 0-4, 500 features for 5-12, and 700 features for 13-22. </p>",
      "rawMarkdown": "My inference time is about 40 minutes with aggregating 5-fold models. I used 200 features for 0-4, 500 features for 5-12, and 700 features for 13-22.",
      "votes": null
    },
    {
      "id": "2227870",
      "postDate": "04/20/2023 04:47:02",
      "content": "<p>This is a 'today I learned' moment for me. Never knew this permutation importance exist.</p>",
      "rawMarkdown": "This is a 'today I learned' moment for me. Never knew this permutation importance exist.",
      "votes": null
    },
    {
      "id": "2228083",
      "postDate": "04/20/2023 08:59:47",
      "content": "<p><a href=\"https://www.kaggle.com/mengvision\" target=\"_blank\">@mengvision</a>  that is impressive! Do you mind sharing what scoring function did you use for permutation importance, was it R2?</p>",
      "rawMarkdown": "mengvision  that is impressive! Do you mind sharing what scoring function did you use for permutation importance, was it R2?",
      "votes": null
    },
    {
      "id": "2228678",
      "postDate": "04/20/2023 17:53:27",
      "content": "<p>If it is not a secret, how do you calculate permutation importance prior to training your model? I recon that you need a model for this. </p>",
      "rawMarkdown": "If it is not a secret, how do you calculate permutation importance prior to training your model? I recon that you need a model for this.",
      "votes": null
    },
    {
      "id": "2228688",
      "postDate": "04/20/2023 18:08:55",
      "content": "<p>I think he is doing it manually - i.e, before training:</p>\n<ol>\n<li>he permutes one col in his features data, makes predictions and compute score</li>\n<li>he does not permute anything, makes predictions and compute score.</li>\n</ol>\n<p>Finally he compares two scores, if score 1 &lt; score 2 - then feature is important, otherwise is not important.   </p>",
      "rawMarkdown": "I think he is doing it manually - i.e, before training:\n1. he permutes one col in his features data, makes predictions and compute score\n2. he does not permute anything, makes predictions and compute score.\n\nFinally he compares two scores, if score 1 < score 2 - then feature is important, otherwise is not important.",
      "votes": null
    },
    {
      "id": "2228741",
      "postDate": "04/20/2023 19:05:02",
      "content": "<p>Interesting approach!</p>",
      "rawMarkdown": "Interesting approach!",
      "votes": null
    },
    {
      "id": "2228842",
      "postDate": "04/20/2023 21:37:38",
      "content": "<p>Do you also train your <code>XGBoost</code> in your notebook or just run inference? </p>\n<p>It might be worth to train your model in a side notebook and then use the trained model in the inference notebook.</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Do you also train your `XGBoost` in your notebook or just run inference? \n\nIt might be worth to train your model in a side notebook and then use the trained model in the inference notebook.\n\nThe Devastator.",
      "votes": null
    },
    {
      "id": "2228970",
      "postDate": "04/21/2023 01:18:53",
      "content": "<p>Exactly!! And Don't use the same model for selecting features and training model, e.g. use Random Forest for feature selection and XGBoost for model training. </p>",
      "rawMarkdown": "Exactly!! And Don't use the same model for selecting features and training model, e.g. use Random Forest for feature selection and XGBoost for model training.",
      "votes": null
    },
    {
      "id": "2229208",
      "postDate": "04/21/2023 06:44:57",
      "content": "<p>40 mins is so efficient! How did you implement data prerpocess in infer part?</p>",
      "rawMarkdown": "40 mins is so efficient! How did you implement data prerpocess in infer part?",
      "votes": null
    },
    {
      "id": "2229316",
      "postDate": "04/21/2023 08:57:11",
      "content": "<p>quite impressive indeed.. Do you just use polars?</p>",
      "rawMarkdown": "quite impressive indeed.. Do you just use polars?",
      "votes": null
    },
    {
      "id": "2229317",
      "postDate": "04/21/2023 09:00:42",
      "content": "<p>very interesting insight <a href=\"https://www.kaggle.com/mengvision\" target=\"_blank\">@mengvision</a>! What is your reasoning/intuition to use different models for feature selection and XGBoost? My thinking is that a feature could be very good for one model and be very bad for another. So why would you use different model architectures for feature selection and modelling?</p>",
      "rawMarkdown": "very interesting insight @mengvision! What is your reasoning/intuition to use different models for feature selection and XGBoost? My thinking is that a feature could be very good for one model and be very bad for another. So why would you use different model architectures for feature selection and modelling?",
      "votes": null
    },
    {
      "id": "2229359",
      "postDate": "04/21/2023 09:56:53",
      "content": "<p>hey <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, first of all thank for all your contribution in Kaggle - you are great, I love your notebooks/discussions!</p>\n<p>No, I do not train XGBoost in my notebook, I do only inference (using predict_proba).</p>",
      "rawMarkdown": "hey @thedevastator, first of all thank for all your contribution in Kaggle - you are great, I love your notebooks/discussions!\n\n\nNo, I do not train XGBoost in my notebook, I do only inference (using predict_proba).",
      "votes": null
    },
    {
      "id": "2229514",
      "postDate": "04/21/2023 12:47:23",
      "content": "<p>I used polars but I only generate the features I need. For my feature engineering training code, I generate about 6k features. In my inference code, I only generate the features I need, so for each group, I generate a subset of them. Since I use features from previous level_group, I also need to merge two data frames. Merging two dataframes is very slow, so I use numpy instead. </p>\n<p><code>data = np.hstack([df_cache[session_id]['0-4'][cache_features].values, df[df_features].values])</code></p>\n<p>Hopefully, this is helpful for you guys. </p>",
      "rawMarkdown": "I used polars but I only generate the features I need. For my feature engineering training code, I generate about 6k features. In my inference code, I only generate the features I need, so for each group, I generate a subset of them. Since I use features from previous level_group, I also need to merge two data frames. Merging two dataframes is very slow, so I use numpy instead. \n\n`data = np.hstack([df_cache[session_id]['0-4'][cache_features].values, df[df_features].values])`\n\nHopefully, this is helpful for you guys.",
      "votes": null
    },
    {
      "id": "2230069",
      "postDate": "04/22/2023 01:35:01",
      "content": "<p>Two stage using the same model may induce data leakage. At least, it’s the situation for me.</p>",
      "rawMarkdown": "Two stage using the same model may induce data leakage. At least, it’s the situation for me.",
      "votes": null
    },
    {
      "id": "2237334",
      "postDate": "04/27/2023 14:41:50",
      "content": "<p>we got between 5-7 hours right now , based on number of features .</p>",
      "rawMarkdown": "we got between 5-7 hours right now , based on number of features .",
      "votes": null
    },
    {
      "id": "2237341",
      "postDate": "04/27/2023 14:50:32",
      "content": "<p>cool, do you mind sharing how many features do you use?</p>",
      "rawMarkdown": "cool, do you mind sharing how many features do you use?",
      "votes": null
    },
    {
      "id": "2238010",
      "postDate": "04/28/2023 06:35:08",
      "content": "<p>To follow up, I cut the number of features according to feature importance then permuted the remaining features. The inference time was reduced to about an hour from 1.5 hour previously. No magic here essentially, the time reduction was simply due to the reduced number of features….</p>",
      "rawMarkdown": "To follow up, I cut the number of features according to feature importance then permuted the remaining features. The inference time was reduced to about an hour from 1.5 hour previously. No magic here essentially, the time reduction was simply due to the reduced number of features....",
      "votes": null
    },
    {
      "id": "2239761",
      "postDate": "04/29/2023 20:16:58",
      "content": "<p>Lot of features in best one ,range in 1k we trying to prune now 😀. It’s a single model </p>",
      "rawMarkdown": "Lot of features in best one ,range in 1k we trying to prune now 😀. It’s a single model",
      "votes": null
    },
    {
      "id": "2240195",
      "postDate": "04/30/2023 10:21:57",
      "content": "<p>lol, good job 😀</p>",
      "rawMarkdown": "lol, good job 😀",
      "votes": null
    },
    {
      "id": "2241550",
      "postDate": "05/01/2023 15:39:24",
      "content": "<p>2 hour for my silver solution.</p>",
      "rawMarkdown": "2 hour for my silver solution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2221106,
      "author_name": "randwlkr",
      "author_url": "",
      "post_date": "04/14/2023 02:08:52",
      "content": "<p>I didn't have the exact number but it was like 1.5 hours or so for my xgboost models. But I should mention that my submission notebook only contains inference part. Models were trained, saved, and then uploaded to the dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2221362,
          "author_name": "ngocuong",
          "author_url": "",
          "post_date": "04/14/2023 07:43:37",
          "content": "<p>Thanks for sharing! Yeah, my 9 hours is with loaded and trained xgboost model too, so pretty damn slow xax. If I may ask, how many features do you have?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2222057,
              "author_name": "randwlkr",
              "author_url": "",
              "post_date": "04/14/2023 20:52:06",
              "content": "<p>I had ~2900 features in the best case.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2222062,
                  "author_name": "ngocuong",
                  "author_url": "",
                  "post_date": "04/14/2023 21:17:24",
                  "content": "<p>cool, thanks for the info. I also have 2-3k features. I've tried to select top 500/600/1000 features using feature importance but the LB either dropped or was the same. Have you tried feature selection?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2222075,
                      "author_name": "randwlkr",
                      "author_url": "",
                      "post_date": "04/14/2023 21:34:02",
                      "content": "<p>I did, but I arrived at the same conclusion: my CV and LB both drops after retaining features selected by xgboost. So at the end, I keep all features. Perhaps I didn't do it right, but I'm a bit puzzled as why my F1 drops by keeping those \"important features\" only.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2222083,
                          "author_name": "ngocuong",
                          "author_url": "",
                          "post_date": "04/14/2023 21:41:26",
                          "content": "<p>I see. Btw what optimizations do you use to do faster feature engineering during submission? I only use polars..</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2222086,
                              "author_name": "randwlkr",
                              "author_url": "",
                              "post_date": "04/14/2023 21:53:03",
                              "content": "<p>I'm using polars too. I don't think I have any tricks for speed up, I just did groupby on <code>session_id</code> and <code>level_group</code> then aggregate <code>elapsed_time</code> difference conditioning on various combinations of original and extracted column features.</p>\n<p>By the way, if you want to load pretrained models, do import them outside the <code>iter_test</code> loop. Loading models <em>within</em> the loop is pretty slow.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2222103,
                                  "author_name": "ngocuong",
                                  "author_url": "",
                                  "post_date": "04/14/2023 22:29:43",
                                  "content": "<p>That's good advice, very appreciated! </p>\n<p>I am a bit confused why my feature engineering is taking so much time during submission… When I do feature engineering on the train set it takes about 2 minutes, but submission time is hours. Do you experience the same?</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2223094,
                                      "author_name": "randwlkr",
                                      "author_url": "",
                                      "post_date": "04/15/2023 20:29:31",
                                      "content": "<p>Yeah, same here: feature engineering (FE) was way faster in the training phase than in the inference phase. Without details of the inference part, it seems difficult to diagnose slow FE. A trivial guess is that there is a <em>lot</em> more test data. But this can't explain our difference in inference FE time….</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 2223113,
                                          "author_name": "ngocuong",
                                          "author_url": "",
                                          "post_date": "04/15/2023 21:20:05",
                                          "content": "<p>To be fair when we submit Kaggle  gives us just 2 CPU-s, and polars uses heavily parallelization. So polars might not be that useful for submission. But I am still suprised…</p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 2224206,
                                              "author_name": "randwlkr",
                                              "author_url": "",
                                              "post_date": "04/17/2023 05:20:33",
                                              "content": "<p>That's a good point about the number of available CPUs in the inference phase. I totally forgot about it.</p>",
                                              "votes": null,
                                              "replies": []
                                            }
                                          ]
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2238010,
          "author_name": "randwlkr",
          "author_url": "",
          "post_date": "04/28/2023 06:35:08",
          "content": "<p>To follow up, I cut the number of features according to feature importance then permuted the remaining features. The inference time was reduced to about an hour from 1.5 hour previously. No magic here essentially, the time reduction was simply due to the reduced number of features….</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2224432,
      "author_name": "mengvision",
      "author_url": "",
      "post_date": "04/17/2023 10:02:02",
      "content": "<p>2 hrs with 500~1000 feats</p>",
      "votes": null,
      "replies": [
        {
          "id": 2224525,
          "author_name": "ngocuong",
          "author_url": "",
          "post_date": "04/17/2023 12:01:34",
          "content": "<p>that's really good - thanks! Do you mind sharing how you've chosen the 500-1000 features? I suppose you had much more than that and did feature reduction right?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2225492,
              "author_name": "mengvision",
              "author_url": "",
              "post_date": "04/18/2023 07:51:18",
              "content": "<p>Permutation importance can help you to select features, but you also need to choose the features which is meaningful(err, it's hard to say). Because some new features are noise. Xgb importance may induce overfitting with data leakage. Kfold using for features selection is also induce overfitting.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2225564,
                  "author_name": "ngocuong",
                  "author_url": "",
                  "post_date": "04/18/2023 08:53:17",
                  "content": "<p>thanks for the advice! What CV improvement did you get from permutation importance, if I may ask?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2227805,
                      "author_name": "mengvision",
                      "author_url": "",
                      "post_date": "04/20/2023 03:05:07",
                      "content": "<p>boost +0.003 both CV and LB with permutation importance, and you should do feature selection before model training. In other word, you should separete feature selection from model training.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2228083,
                          "author_name": "ngocuong",
                          "author_url": "",
                          "post_date": "04/20/2023 08:59:47",
                          "content": "<p><a href=\"https://www.kaggle.com/mengvision\" target=\"_blank\">@mengvision</a>  that is impressive! Do you mind sharing what scoring function did you use for permutation importance, was it R2?</p>",
                          "votes": null,
                          "replies": []
                        },
                        {
                          "id": 2228678,
                          "author_name": "woprime",
                          "author_url": "",
                          "post_date": "04/20/2023 17:53:27",
                          "content": "<p>If it is not a secret, how do you calculate permutation importance prior to training your model? I recon that you need a model for this. </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2228688,
                              "author_name": "ngocuong",
                              "author_url": "",
                              "post_date": "04/20/2023 18:08:55",
                              "content": "<p>I think he is doing it manually - i.e, before training:</p>\n<ol>\n<li>he permutes one col in his features data, makes predictions and compute score</li>\n<li>he does not permute anything, makes predictions and compute score.</li>\n</ol>\n<p>Finally he compares two scores, if score 1 &lt; score 2 - then feature is important, otherwise is not important.   </p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2228741,
                                  "author_name": "woprime",
                                  "author_url": "",
                                  "post_date": "04/20/2023 19:05:02",
                                  "content": "<p>Interesting approach!</p>",
                                  "votes": null,
                                  "replies": []
                                },
                                {
                                  "id": 2228970,
                                  "author_name": "mengvision",
                                  "author_url": "",
                                  "post_date": "04/21/2023 01:18:53",
                                  "content": "<p>Exactly!! And Don't use the same model for selecting features and training model, e.g. use Random Forest for feature selection and XGBoost for model training. </p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2229317,
                                      "author_name": "ngocuong",
                                      "author_url": "",
                                      "post_date": "04/21/2023 09:00:42",
                                      "content": "<p>very interesting insight <a href=\"https://www.kaggle.com/mengvision\" target=\"_blank\">@mengvision</a>! What is your reasoning/intuition to use different models for feature selection and XGBoost? My thinking is that a feature could be very good for one model and be very bad for another. So why would you use different model architectures for feature selection and modelling?</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 2230069,
                                          "author_name": "mengvision",
                                          "author_url": "",
                                          "post_date": "04/22/2023 01:35:01",
                                          "content": "<p>Two stage using the same model may induce data leakage. At least, it’s the situation for me.</p>",
                                          "votes": null,
                                          "replies": []
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                },
                {
                  "id": 2227870,
                  "author_name": "woprime",
                  "author_url": "",
                  "post_date": "04/20/2023 04:47:02",
                  "content": "<p>This is a 'today I learned' moment for me. Never knew this permutation importance exist.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2224754,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "04/17/2023 16:19:55",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ngocuong\" target=\"_blank\">@ngocuong</a>. Are you using pandas or polars? Polars is faster.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2224948,
          "author_name": "ngocuong",
          "author_url": "",
          "post_date": "04/17/2023 19:10:13",
          "content": "<p>hi, I use polars.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2227827,
      "author_name": "mrhantato",
      "author_url": "",
      "post_date": "04/20/2023 03:46:00",
      "content": "<p>My inference time is about 40 minutes with aggregating 5-fold models. I used 200 features for 0-4, 500 features for 5-12, and 700 features for 13-22. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2229208,
          "author_name": "mengvision",
          "author_url": "",
          "post_date": "04/21/2023 06:44:57",
          "content": "<p>40 mins is so efficient! How did you implement data prerpocess in infer part?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2229514,
              "author_name": "mrhantato",
              "author_url": "",
              "post_date": "04/21/2023 12:47:23",
              "content": "<p>I used polars but I only generate the features I need. For my feature engineering training code, I generate about 6k features. In my inference code, I only generate the features I need, so for each group, I generate a subset of them. Since I use features from previous level_group, I also need to merge two data frames. Merging two dataframes is very slow, so I use numpy instead. </p>\n<p><code>data = np.hstack([df_cache[session_id]['0-4'][cache_features].values, df[df_features].values])</code></p>\n<p>Hopefully, this is helpful for you guys. </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2229316,
          "author_name": "ngocuong",
          "author_url": "",
          "post_date": "04/21/2023 08:57:11",
          "content": "<p>quite impressive indeed.. Do you just use polars?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228842,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "04/20/2023 21:37:38",
      "content": "<p>Do you also train your <code>XGBoost</code> in your notebook or just run inference? </p>\n<p>It might be worth to train your model in a side notebook and then use the trained model in the inference notebook.</p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2229359,
          "author_name": "ngocuong",
          "author_url": "",
          "post_date": "04/21/2023 09:56:53",
          "content": "<p>hey <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>, first of all thank for all your contribution in Kaggle - you are great, I love your notebooks/discussions!</p>\n<p>No, I do not train XGBoost in my notebook, I do only inference (using predict_proba).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2237334,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "04/27/2023 14:41:50",
      "content": "<p>we got between 5-7 hours right now , based on number of features .</p>",
      "votes": null,
      "replies": [
        {
          "id": 2237341,
          "author_name": "ngocuong",
          "author_url": "",
          "post_date": "04/27/2023 14:50:32",
          "content": "<p>cool, do you mind sharing how many features do you use?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2239761,
              "author_name": "gauravbrills",
              "author_url": "",
              "post_date": "04/29/2023 20:16:58",
              "content": "<p>Lot of features in best one ,range in 1k we trying to prune now 😀. It’s a single model </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2240195,
                  "author_name": "ngocuong",
                  "author_url": "",
                  "post_date": "04/30/2023 10:21:57",
                  "content": "<p>lol, good job 😀</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2241550,
      "author_name": "littlstar123",
      "author_url": "",
      "post_date": "05/01/2023 15:39:24",
      "content": "<p>2 hour for my silver solution.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2220989": "I think my feature engineering is super slow and my total inference time (i.e. from when I make a submission until I see my score) is:\n\nXGBOOST (pre-trained) | ~9 hours |  LB 0.695\n\n\nCan anybody share their inference times for reference? Thanks!",
    "2221106": "I didn't have the exact number but it was like 1.5 hours or so for my xgboost models. But I should mention that my submission notebook only contains inference part. Models were trained, saved, and then uploaded to the dataset.",
    "2221362": "Thanks for sharing! Yeah, my 9 hours is with loaded and trained xgboost model too, so pretty damn slow xax. If I may ask, how many features do you have?",
    "2222057": "I had ~2900 features in the best case.",
    "2222062": "cool, thanks for the info. I also have 2-3k features. I've tried to select top 500/600/1000 features using feature importance but the LB either dropped or was the same. Have you tried feature selection?",
    "2222075": "I did, but I arrived at the same conclusion: my CV and LB both drops after retaining features selected by xgboost. So at the end, I keep all features. Perhaps I didn't do it right, but I'm a bit puzzled as why my F1 drops by keeping those \"important features\" only.",
    "2222083": "I see. Btw what optimizations do you use to do faster feature engineering during submission? I only use polars..",
    "2222086": "I'm using polars too. I don't think I have any tricks for speed up, I just did groupby on `session_id` and `level_group` then aggregate `elapsed_time` difference conditioning on various combinations of original and extracted column features.\n\nBy the way, if you want to load pretrained models, do import them outside the `iter_test` loop. Loading models _within_ the loop is pretty slow.",
    "2222103": "That's good advice, very appreciated! \n\nI am a bit confused why my feature engineering is taking so much time during submission... When I do feature engineering on the train set it takes about 2 minutes, but submission time is hours. Do you experience the same?",
    "2223094": "Yeah, same here: feature engineering (FE) was way faster in the training phase than in the inference phase. Without details of the inference part, it seems difficult to diagnose slow FE. A trivial guess is that there is a _lot_ more test data. But this can't explain our difference in inference FE time....",
    "2223113": "To be fair when we submit Kaggle  gives us just 2 CPU-s, and polars uses heavily parallelization. So polars might not be that useful for submission. But I am still suprised...",
    "2224206": "That's a good point about the number of available CPUs in the inference phase. I totally forgot about it.",
    "2224432": "2 hrs with 500~1000 feats",
    "2224525": "that's really good - thanks! Do you mind sharing how you've chosen the 500-1000 features? I suppose you had much more than that and did feature reduction right?",
    "2224754": "Hi @ngocuong. Are you using pandas or polars? Polars is faster.",
    "2224948": "hi, I use polars.",
    "2225492": "Permutation importance can help you to select features, but you also need to choose the features which is meaningful(err, it's hard to say). Because some new features are noise. Xgb importance may induce overfitting with data leakage. Kfold using for features selection is also induce overfitting.",
    "2225564": "thanks for the advice! What CV improvement did you get from permutation importance, if I may ask?",
    "2227805": "boost +0.003 both CV and LB with permutation importance, and you should do feature selection before model training. In other word, you should separete feature selection from model training.",
    "2227827": "My inference time is about 40 minutes with aggregating 5-fold models. I used 200 features for 0-4, 500 features for 5-12, and 700 features for 13-22.",
    "2227870": "This is a 'today I learned' moment for me. Never knew this permutation importance exist.",
    "2228083": "mengvision  that is impressive! Do you mind sharing what scoring function did you use for permutation importance, was it R2?",
    "2228678": "If it is not a secret, how do you calculate permutation importance prior to training your model? I recon that you need a model for this.",
    "2228688": "I think he is doing it manually - i.e, before training:\n1. he permutes one col in his features data, makes predictions and compute score\n2. he does not permute anything, makes predictions and compute score.\n\nFinally he compares two scores, if score 1 < score 2 - then feature is important, otherwise is not important.",
    "2228741": "Interesting approach!",
    "2228842": "Do you also train your `XGBoost` in your notebook or just run inference? \n\nIt might be worth to train your model in a side notebook and then use the trained model in the inference notebook.\n\nThe Devastator.",
    "2228970": "Exactly!! And Don't use the same model for selecting features and training model, e.g. use Random Forest for feature selection and XGBoost for model training.",
    "2229208": "40 mins is so efficient! How did you implement data prerpocess in infer part?",
    "2229316": "quite impressive indeed.. Do you just use polars?",
    "2229317": "very interesting insight @mengvision! What is your reasoning/intuition to use different models for feature selection and XGBoost? My thinking is that a feature could be very good for one model and be very bad for another. So why would you use different model architectures for feature selection and modelling?",
    "2229359": "hey @thedevastator, first of all thank for all your contribution in Kaggle - you are great, I love your notebooks/discussions!\n\n\nNo, I do not train XGBoost in my notebook, I do only inference (using predict_proba).",
    "2229514": "I used polars but I only generate the features I need. For my feature engineering training code, I generate about 6k features. In my inference code, I only generate the features I need, so for each group, I generate a subset of them. Since I use features from previous level_group, I also need to merge two data frames. Merging two dataframes is very slow, so I use numpy instead. \n\n`data = np.hstack([df_cache[session_id]['0-4'][cache_features].values, df[df_features].values])`\n\nHopefully, this is helpful for you guys.",
    "2230069": "Two stage using the same model may induce data leakage. At least, it’s the situation for me.",
    "2237334": "we got between 5-7 hours right now , based on number of features .",
    "2237341": "cool, do you mind sharing how many features do you use?",
    "2238010": "To follow up, I cut the number of features according to feature importance then permuted the remaining features. The inference time was reduced to about an hour from 1.5 hour previously. No magic here essentially, the time reduction was simply due to the reduced number of features....",
    "2239761": "Lot of features in best one ,range in 1k we trying to prune now 😀. It’s a single model",
    "2240195": "lol, good job 😀",
    "2241550": "2 hour for my silver solution."
  },
  "source": "meta"
}