{
  "id": 386218,
  "title": "Train GPU, Infer CPU, Boost CV LB",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/386218",
  "author_name": "Chris Deotte",
  "post_date": "2023-02-11T21:08:13.542000",
  "votes": 149,
  "comment_count": 70,
  "views": 0,
  "content": "<p>I would like to point out that we do <strong>not need</strong> to feature engineer on CPU <strong>nor</strong> train on CPU in this competition. Only inference needs to be done on CPU. We should train in a one notebook and save model. Then infer with another notebook and load model.</p>\n<h1>Use CPU Kaggle Notebook for Inference</h1>\n<p>During inference we must use CPU 8GB notebook, however there is plenty of memory because the Kaggle API uses a <code>for-loop</code>. In each iteration of the <code>for-loop</code> we are given two very small dataframes (<code>test.csv</code> and <code>sample_submission.csv</code>) that we process and infer. Each iteration we receive a new test.csv and sample_submission.csv. There are no memory issues because each test dataframe is a few hundred rows with 20 columns. Its size in memory is under 1MB. </p>\n<p>FYI, in each iteration of the <code>for-loop</code> we receive one chapter of one user (out of total of 3 chapters and 11k test users). A chapter is a few hundred rows of dataframe with one event per row (and 20 columns). When we receive  chapter one (i.e. <code>level_group = '0-4'</code>) then we use this data to predict questions <code>1-3</code> for the assoicated user. When we receive chapter two (i.e. <code>level_group = '5-12'</code>) then we use this data to predict questions <code>4-13</code>. When we receive chapter three (i.e. <code>level_group = '13-22'</code>) then we predict questions <code>14-18</code>.</p>\n<h1>Use GPU Kaggle Notebook for Train</h1>\n<p>During training, we must load and process <code>train.csv</code> which has 13 million rows and 20 columns. It contains all 3 chapters for 11k train users. This file is a few GB and may cause memory problems when we feature engineer and/or train. Furthermore feature engineering and training takes time. Therefore I suggest using GPU Kaggle notebook during train. This notebook has 16GB GPU VRAM and 32GB CPU RAM for a total of 48GB of RAM! This is plenty of memory and speed to process <code>train.csv</code> and train our models.</p>\n<h1>XGBoost Starter Notebook - LB 0.680</h1>\n<p>I published an XGBoost starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\" target=\"_blank\">here</a>. Currently feature engineer, train, and infer are in one CPU notebook. I suggest creating two notebooks. Put the <code>train.csv</code> feature engineering and model training in the first GPU notebook. Then after training XGB Classifier model use </p>\n<pre><code>model.save_model(f'XGB_question_{t}.xgb')\n</code></pre>\n<p>Upload these models to a Kaggle dataset. Then put the inference code in a second CPU notebook. And load model with</p>\n<pre><code>model = XGBClassifier()\nmodel.load_model(f'XGB_question_{t}.xgb')\n</code></pre>\n<p>Using GPU to train XGB will give us <strong>4x or more speed increase!</strong> for model training and <strong>10x or more speed increase!</strong> for feature engineering</p>\n<h1>RAPIDS cuDF Feature Engineer</h1>\n<p>In the first notebook to utilize cuDF, we can change </p>\n<pre><code>import pandas as pd\n</code></pre>\n<p>into the following</p>\n<pre><code>import cudf as pd\n</code></pre>\n<p>Then our feature engineering will be <strong>10x or more faster!</strong> There are two other changes we need to convert Pandas to RAPIDS cuDF. Use the two codes below</p>\n<pre><code>targets = pd.read_csv('train_labels.csv')\ntargets['session'] = pd.to_numeric( targets.session_id.str.split('_').list.get(0) )\ntargets['q'] = pd.to_numeric( targets.session_id.str.split('_q').list.get(1) )\n</code></pre>\n<p>And because GroupKFold wants Pandas:</p>\n<pre><code>P1 = df.iloc[:,0].to_pandas()\nP2 = df.index.to_pandas()\nfor i, (train_index, test_index) in enumerate(gkf.split(X=P1, groups=P2)):\n</code></pre>\n<h1>How To Improve CV and LB</h1>\n<p>To secret to improving CV and LB is lots of fast experiments. Using GPU cuDF and GPU XGB, we can engineer a few new features and train 5x18=90 models in a few minutes. If the CV score increases, then we keep the new features, otherwise we discard and try again. </p>\n<p>So far, using GPU, I have created and tested over 1000 features! My current best solution is XGB single model using my best 400 features. It achieves CV 0.703 and LB 0.702!</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2023/gpu.jpeg\" alt=\"\"></p>\n<h1>UPDATE - GPU Starter Notebook - LB 0.677</h1>\n<p>Shashwat has provided a RAPIDS cuDF GPU train/validate notebook <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\" target=\"_blank\">here</a> and CPU inference notebook <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference\" target=\"_blank\">here</a>. Thanks <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> . His starter notebook is faster than my original CPU starter notebook and achieves <code>+0.001</code> CV boost and <code>+0.001</code> LB boost, just using GPU instead of CPU (without changing the features or model). Hooray!</p>",
  "messages": [
    {
      "id": 2140569,
      "postDate": "2023-02-11T21:08:13.543Z",
      "content": "<p>I would like to point out that we do <strong>not need</strong> to feature engineer on CPU <strong>nor</strong> train on CPU in this competition. Only inference needs to be done on CPU. We should train in a one notebook and save model. Then infer with another notebook and load model.</p>\n<h1>Use CPU Kaggle Notebook for Inference</h1>\n<p>During inference we must use CPU 8GB notebook, however there is plenty of memory because the Kaggle API uses a <code>for-loop</code>. In each iteration of the <code>for-loop</code> we are given two very small dataframes (<code>test.csv</code> and <code>sample_submission.csv</code>) that we process and infer. Each iteration we receive a new test.csv and sample_submission.csv. There are no memory issues because each test dataframe is a few hundred rows with 20 columns. Its size in memory is under 1MB. </p>\n<p>FYI, in each iteration of the <code>for-loop</code> we receive one chapter of one user (out of total of 3 chapters and 11k test users). A chapter is a few hundred rows of dataframe with one event per row (and 20 columns). When we receive  chapter one (i.e. <code>level_group = '0-4'</code>) then we use this data to predict questions <code>1-3</code> for the assoicated user. When we receive chapter two (i.e. <code>level_group = '5-12'</code>) then we use this data to predict questions <code>4-13</code>. When we receive chapter three (i.e. <code>level_group = '13-22'</code>) then we predict questions <code>14-18</code>.</p>\n<h1>Use GPU Kaggle Notebook for Train</h1>\n<p>During training, we must load and process <code>train.csv</code> which has 13 million rows and 20 columns. It contains all 3 chapters for 11k train users. This file is a few GB and may cause memory problems when we feature engineer and/or train. Furthermore feature engineering and training takes time. Therefore I suggest using GPU Kaggle notebook during train. This notebook has 16GB GPU VRAM and 32GB CPU RAM for a total of 48GB of RAM! This is plenty of memory and speed to process <code>train.csv</code> and train our models.</p>\n<h1>XGBoost Starter Notebook - LB 0.680</h1>\n<p>I published an XGBoost starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\" target=\"_blank\">here</a>. Currently feature engineer, train, and infer are in one CPU notebook. I suggest creating two notebooks. Put the <code>train.csv</code> feature engineering and model training in the first GPU notebook. Then after training XGB Classifier model use </p>\n<pre><code>model.save_model(f'XGB_question_{t}.xgb')\n</code></pre>\n<p>Upload these models to a Kaggle dataset. Then put the inference code in a second CPU notebook. And load model with</p>\n<pre><code>model = XGBClassifier()\nmodel.load_model(f'XGB_question_{t}.xgb')\n</code></pre>\n<p>Using GPU to train XGB will give us <strong>4x or more speed increase!</strong> for model training and <strong>10x or more speed increase!</strong> for feature engineering</p>\n<h1>RAPIDS cuDF Feature Engineer</h1>\n<p>In the first notebook to utilize cuDF, we can change </p>\n<pre><code>import pandas as pd\n</code></pre>\n<p>into the following</p>\n<pre><code>import cudf as pd\n</code></pre>\n<p>Then our feature engineering will be <strong>10x or more faster!</strong> There are two other changes we need to convert Pandas to RAPIDS cuDF. Use the two codes below</p>\n<pre><code>targets = pd.read_csv('train_labels.csv')\ntargets['session'] = pd.to_numeric( targets.session_id.str.split('_').list.get(0) )\ntargets['q'] = pd.to_numeric( targets.session_id.str.split('_q').list.get(1) )\n</code></pre>\n<p>And because GroupKFold wants Pandas:</p>\n<pre><code>P1 = df.iloc[:,0].to_pandas()\nP2 = df.index.to_pandas()\nfor i, (train_index, test_index) in enumerate(gkf.split(X=P1, groups=P2)):\n</code></pre>\n<h1>How To Improve CV and LB</h1>\n<p>To secret to improving CV and LB is lots of fast experiments. Using GPU cuDF and GPU XGB, we can engineer a few new features and train 5x18=90 models in a few minutes. If the CV score increases, then we keep the new features, otherwise we discard and try again. </p>\n<p>So far, using GPU, I have created and tested over 1000 features! My current best solution is XGB single model using my best 400 features. It achieves CV 0.703 and LB 0.702!</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2023/gpu.jpeg\" alt=\"\"></p>\n<h1>UPDATE - GPU Starter Notebook - LB 0.677</h1>\n<p>Shashwat has provided a RAPIDS cuDF GPU train/validate notebook <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\" target=\"_blank\">here</a> and CPU inference notebook <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference\" target=\"_blank\">here</a>. Thanks <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> . His starter notebook is faster than my original CPU starter notebook and achieves <code>+0.001</code> CV boost and <code>+0.001</code> LB boost, just using GPU instead of CPU (without changing the features or model). Hooray!</p>",
      "rawMarkdown": "I would like to point out that we do **not need** to feature engineer on CPU **nor** train on CPU in this competition. Only inference needs to be done on CPU. We should train in a one notebook and save model. Then infer with another notebook and load model.\n\n# Use CPU Kaggle Notebook for Inference\nDuring inference we must use CPU 8GB notebook, however there is plenty of memory because the Kaggle API uses a `for-loop`. In each iteration of the `for-loop` we are given two very small dataframes (`test.csv` and `sample_submission.csv`) that we process and infer. Each iteration we receive a new test.csv and sample_submission.csv. There are no memory issues because each test dataframe is a few hundred rows with 20 columns. Its size in memory is under 1MB. \n\nFYI, in each iteration of the `for-loop` we receive one chapter of one user (out of total of 3 chapters and 11k test users). A chapter is a few hundred rows of dataframe with one event per row (and 20 columns). When we receive  chapter one (i.e. `level_group = '0-4'`) then we use this data to predict questions `1-3` for the assoicated user. When we receive chapter two (i.e. `level_group = '5-12'`) then we use this data to predict questions `4-13`. When we receive chapter three (i.e. `level_group = '13-22'`) then we predict questions `14-18`.\n\n# Use GPU Kaggle Notebook for Train\nDuring training, we must load and process `train.csv` which has 13 million rows and 20 columns. It contains all 3 chapters for 11k train users. This file is a few GB and may cause memory problems when we feature engineer and/or train. Furthermore feature engineering and training takes time. Therefore I suggest using GPU Kaggle notebook during train. This notebook has 16GB GPU VRAM and 32GB CPU RAM for a total of 48GB of RAM! This is plenty of memory and speed to process `train.csv` and train our models.\n\n# XGBoost Starter Notebook - LB 0.680\nI published an XGBoost starter notebook [here][1]. Currently feature engineer, train, and infer are in one CPU notebook. I suggest creating two notebooks. Put the `train.csv` feature engineering and model training in the first GPU notebook. Then after training XGB Classifier model use \n\n    model.save_model(f'XGB_question_{t}.xgb')\n\nUpload these models to a Kaggle dataset. Then put the inference code in a second CPU notebook. And load model with\n\n    model = XGBClassifier()\n    model.load_model(f'XGB_question_{t}.xgb')\n\nUsing GPU to train XGB will give us **4x or more speed increase!** for model training and **10x or more speed increase!** for feature engineering\n\n# RAPIDS cuDF Feature Engineer\nIn the first notebook to utilize cuDF, we can change \n\n    import pandas as pd\n\ninto the following\n\n    import cudf as pd\n\nThen our feature engineering will be **10x or more faster!** There are two other changes we need to convert Pandas to RAPIDS cuDF. Use the two codes below\n\n    targets = pd.read_csv('train_labels.csv')\n    targets['session'] = pd.to_numeric( targets.session_id.str.split('_').list.get(0) )\n    targets['q'] = pd.to_numeric( targets.session_id.str.split('_q').list.get(1) )\n\nAnd because GroupKFold wants Pandas:\n\n    P1 = df.iloc[:,0].to_pandas()\n    P2 = df.index.to_pandas()\n    for i, (train_index, test_index) in enumerate(gkf.split(X=P1, groups=P2)):\n\n# How To Improve CV and LB\nTo secret to improving CV and LB is lots of fast experiments. Using GPU cuDF and GPU XGB, we can engineer a few new features and train 5x18=90 models in a few minutes. If the CV score increases, then we keep the new features, otherwise we discard and try again. \n\nSo far, using GPU, I have created and tested over 1000 features! My current best solution is XGB single model using my best 400 features. It achieves CV 0.703 and LB 0.702!\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2023/gpu.jpeg)\n\n# UPDATE - GPU Starter Notebook - LB 0.677\nShashwat has provided a RAPIDS cuDF GPU train/validate notebook [here][2] and CPU inference notebook [here][3]. Thanks @shashwatraman . His starter notebook is faster than my original CPU starter notebook and achieves `+0.001` CV boost and `+0.001` LB boost, just using GPU instead of CPU (without changing the features or model). Hooray!\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\n[2]: https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\n[3]: https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference",
      "votes": 148
    },
    {
      "id": 2140625,
      "postDate": "2023-02-11T23:55:14.080Z",
      "content": "<p>GPU is key, as usual 😃.</p>",
      "rawMarkdown": "GPU is key, as usual 😃.",
      "votes": 2
    },
    {
      "id": 2222841,
      "postDate": "2023-04-15T15:54:55.730Z",
      "content": "<p>Thanks for all the info. I am learning quite a bit from your XGBoost notebook. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Thanks for all the info. I am learning quite a bit from your XGBoost notebook. @cdeotte ",
      "votes": 1
    },
    {
      "id": 2187315,
      "postDate": "2023-03-18T16:20:45.820Z",
      "content": "<p>I don't quite understand what kind of features can you add to the model to reach 400, I'm really confused about that part. Also thank you for you contribution!</p>",
      "rawMarkdown": "I don't quite understand what kind of features can you add to the model to reach 400, I'm really confused about that part. Also thank you for you contribution!",
      "votes": 1,
      "replies": [
        {
          "id": 2187349,
          "postDate": "2023-03-18T16:52:56.423Z",
          "content": "<p>There are dozens of levels, dozens of rooms, dozens of events. If you start making features by grouping by pairs of level, room, event then you can quickly reach 1000 features.</p>",
          "rawMarkdown": "There are dozens of levels, dozens of rooms, dozens of events. If you start making features by grouping by pairs of level, room, event then you can quickly reach 1000 features.",
          "votes": 2,
          "replies": [
            {
              "id": 2187356,
              "postDate": "2023-03-18T16:59:48.153Z",
              "content": "<p>ok I understand now, thank you very much!!</p>",
              "rawMarkdown": "ok I understand now, thank you very much!!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2170739,
      "postDate": "2023-03-06T08:52:08.313Z",
      "content": "<p>Thanks for sharing.　I now have a better understanding of this competition.👍</p>",
      "rawMarkdown": "Thanks for sharing.　I now have a better understanding of this competition.👍",
      "votes": 1
    },
    {
      "id": 2163243,
      "postDate": "2023-02-28T17:21:00.207Z",
      "content": "<p>Thanks Chris for all the info. I am learning quite a bit from your XGBoost notebook. I assume there is temporal information (subsequent sessions) for a player, so I am thinking to try a model that keeps track of this information (memory), for example a RNN. Any thoughts or recommendations for that?</p>",
      "rawMarkdown": "Thanks Chris for all the info. I am learning quite a bit from your XGBoost notebook. I assume there is temporal information (subsequent sessions) for a player, so I am thinking to try a model that keeps track of this information (memory), for example a RNN. Any thoughts or recommendations for that?",
      "votes": 1
    },
    {
      "id": 2162811,
      "postDate": "2023-02-28T12:21:14.503Z",
      "content": "<p>Thanks Chris ! BTW if anyone isnt able to import cudf, you can easily install this <code>pip install cupy-cuda11x</code> in the notebook console. </p>",
      "rawMarkdown": "Thanks Chris ! BTW if anyone isnt able to import cudf, you can easily install this `pip install cupy-cuda11x` in the notebook console. ",
      "votes": 1
    },
    {
      "id": 2150685,
      "postDate": "2023-02-19T14:01:30.880Z",
      "content": "<p><strong>UPDATE - GPU Starter Notebook - LB 0.677</strong><br>\nShashwat has provided a RAPIDS cuDF GPU train/validate notebook <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\" target=\"_blank\">here</a> and CPU inference notebook <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference\" target=\"_blank\">here</a>. Thanks <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> . His starter notebook is faster than my original CPU starter notebook and achieves <code>+0.001</code> CV boost and <code>+0.001</code> LB boost, just using GPU instead of CPU (without changing the features or model). Hooray!</p>",
      "rawMarkdown": "**UPDATE - GPU Starter Notebook - LB 0.677**\nShashwat has provided a RAPIDS cuDF GPU train/validate notebook [here][2] and CPU inference notebook [here][3]. Thanks @shashwatraman . His starter notebook is faster than my original CPU starter notebook and achieves `+0.001` CV boost and `+0.001` LB boost, just using GPU instead of CPU (without changing the features or model). Hooray!\n\n[2]: https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\n[3]: https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference",
      "votes": 1
    },
    {
      "id": 2145873,
      "postDate": "2023-02-15T12:20:35.997Z",
      "content": "<p>Nice, Amazing use of RAPIDS</p>",
      "rawMarkdown": "Nice, Amazing use of RAPIDS",
      "votes": 1
    },
    {
      "id": 2145567,
      "postDate": "2023-02-15T07:42:08.447Z",
      "content": "<p>Good job!!!</p>",
      "rawMarkdown": "Good job!!!",
      "votes": 1
    },
    {
      "id": 2144188,
      "postDate": "2023-02-14T19:55:19.290Z",
      "content": "<p>Amazing !!</p>",
      "rawMarkdown": "Amazing !!\n\n\n",
      "votes": 1
    },
    {
      "id": 2143874,
      "postDate": "2023-02-14T15:38:34.827Z",
      "content": "<p>Hmmm,  i think, what kaggle time limit for GPU too big for experiments(</p>",
      "rawMarkdown": "Hmmm,  i think, what kaggle time limit for GPU too big for experiments(",
      "votes": 1,
      "replies": [
        {
          "id": 2143895,
          "postDate": "2023-02-14T15:57:52.410Z",
          "content": "<p>There is no time limit for GPU. We can train / experiment as much as we want. The time limit only applies to our inference notebook which we submit to Kaggle.</p>",
          "rawMarkdown": "There is no time limit for GPU. We can train / experiment as much as we want. The time limit only applies to our inference notebook which we submit to Kaggle.",
          "votes": 1,
          "replies": [
            {
              "id": 2147864,
              "postDate": "2023-02-17T00:09:20.933Z",
              "content": "<p>He might be referring to the 30 hour weekly limit Kaggle imposes.</p>",
              "rawMarkdown": "He might be referring to the 30 hour weekly limit Kaggle imposes.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2143806,
      "postDate": "2023-02-14T14:25:50.323Z",
      "content": "<p>Thanks for your sharing and I have a question.Will you try deep neural network methods like transformer, training on GPU and infer on CPU?I think the deep model may have better performance since we have complex sequence data.Maybe you think there is no way to use deep model with the computation constraint? Looking forward to your reply.</p>",
      "rawMarkdown": "Thanks for your sharing and I have a question.Will you try deep neural network methods like transformer, training on GPU and infer on CPU?I think the deep model may have better performance since we have complex sequence data.Maybe you think there is no way to use deep model with the computation constraint? Looking forward to your reply.",
      "votes": 1,
      "replies": [
        {
          "id": 2143899,
          "postDate": "2023-02-14T15:59:41.460Z",
          "content": "<p>Yes I plan to try Transformer and RNN. If we build our own transformer from scratch like <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\">here</a>, then we can keep it simple and fast to train and infer. This is what i plan to do after GBT. (There is RNN example <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a>)</p>",
          "rawMarkdown": "Yes I plan to try Transformer and RNN. If we build our own transformer from scratch like [here][1], then we can keep it simple and fast to train and infer. This is what i plan to do after GBT. (There is RNN example [here][2])\n\n[1]: https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\n[2]: https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790",
          "votes": 4
        }
      ]
    },
    {
      "id": 2143600,
      "postDate": "2023-02-14T11:35:21.877Z",
      "content": "<p>For the people familar with polars thats also an option.<br>\nIn moment I create 430 features in under 1.5 minutes with polars. <a href=\"https://www.kaggle.com/datasets/zaemon1251hesty/polars01516\" target=\"_blank\">this dataset</a> lets you install polars offline.</p>",
      "rawMarkdown": "For the people familar with polars thats also an option.\nIn moment I create 430 features in under 1.5 minutes with polars. [this dataset](https://www.kaggle.com/datasets/zaemon1251hesty/polars01516) lets you install polars offline.",
      "votes": 1
    },
    {
      "id": 2142155,
      "postDate": "2023-02-13T10:56:23.723Z",
      "content": "<p>Amazing !!</p>",
      "rawMarkdown": "Amazing !!",
      "votes": 1
    },
    {
      "id": 2141696,
      "postDate": "2023-02-13T02:40:30.020Z",
      "content": "<p>that is amazing！</p>",
      "rawMarkdown": "that is amazing！",
      "votes": 1
    },
    {
      "id": 2140847,
      "postDate": "2023-02-12T07:55:26.263Z",
      "content": "<blockquote>\n  <p>Only inference needs to be done on CPU. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> my understanding, inference can be done on GPU for normal <strong>Leaderboard Prizes</strong>. Is it correct ?<br>\nI think inference must be done on CPU only for <strong>Efficiency Prizes</strong> .</p>",
      "rawMarkdown": "> Only inference needs to be done on CPU. \n\n@cdeotte my understanding, inference can be done on GPU for normal **Leaderboard Prizes**. Is it correct ?\nI think inference must be done on CPU only for **Efficiency Prizes** .",
      "votes": 1,
      "replies": [
        {
          "id": 2140885,
          "postDate": "2023-02-12T08:41:46.567Z",
          "content": "<p>The GPU is not officially provided for this competition, so I don't think the inference is done on the GPU. I think the idea of the notebook is to do a fast experiment with the GPU to create and filter new features, and if the features are effective we will save the features and then complete the inference on the CPU and submit it. Is it correct ? <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "rawMarkdown": "The GPU is not officially provided for this competition, so I don't think the inference is done on the GPU. I think the idea of the notebook is to do a fast experiment with the GPU to create and filter new features, and if the features are effective we will save the features and then complete the inference on the CPU and submit it. Is it correct ? @cdeotte ",
          "votes": 2
        },
        {
          "id": 2141074,
          "postDate": "2023-02-12T12:33:19.907Z",
          "content": "<p>Could not submit when GPU was turned on.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1985486%2F1b76a55bd776cca7b9e71037da1c8ba2%2FScreenshot_20230212-212951.png?generation=1676205041905134&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Could not submit when GPU was turned on.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1985486%2F1b76a55bd776cca7b9e71037da1c8ba2%2FScreenshot_20230212-212951.png?generation=1676205041905134&alt=media)",
          "votes": 2,
          "replies": [
            {
              "id": 2141153,
              "postDate": "2023-02-12T13:51:10.617Z",
              "content": "<p><a href=\"https://www.kaggle.com/luzhonghao\" target=\"_blank\">@luzhonghao</a> <a href=\"https://www.kaggle.com/zakopur0\" target=\"_blank\">@zakopur0</a> Oh! I misunderstood the regulation. Thank you for sharing !</p>",
              "rawMarkdown": "@luzhonghao @zakopur0 Oh! I misunderstood the regulation. Thank you for sharing !",
              "votes": 2
            },
            {
              "id": 2141344,
              "postDate": "2023-02-12T17:11:16.293Z",
              "content": "<p>On the top of code requirements page <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview/code-requirements\" target=\"_blank\">here</a> it says the paragraph below. And as others say above if you try to submit a GPU notebook, the competition will not let you. (Also note the CPU notebook is not the normal Kaggle CPU notebook, it only has 8GB RAM in this comp).</p>\n<blockquote>\n  <p>Note: this competition is aimed at producing models that are small and lightweight. We have introduced compute constraints to match - your VMs will have only 2 CPUs, 8GB of RAM, and no GPU available. You will still have a maximum of 9 hours to complete the task, but between the constraints and the efficiency prize there will be some interesting sub-problems to solve. Good luck!</p>\n</blockquote>",
              "rawMarkdown": "On the top of code requirements page [here][1] it says the paragraph below. And as others say above if you try to submit a GPU notebook, the competition will not let you. (Also note the CPU notebook is not the normal Kaggle CPU notebook, it only has 8GB RAM in this comp).\n\n>Note: this competition is aimed at producing models that are small and lightweight. We have introduced compute constraints to match - your VMs will have only 2 CPUs, 8GB of RAM, and no GPU available. You will still have a maximum of 9 hours to complete the task, but between the constraints and the efficiency prize there will be some interesting sub-problems to solve. Good luck!\n\n[1]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview/code-requirements",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2219917,
      "postDate": "2023-04-13T01:47:16.173Z",
      "content": "<p>Hi, I have a question.<br>\nFrom \" XGB single model using my best 400 features.\" so If I want to use 400 features out of 1000 features. Do I need to retrain my model?</p>",
      "rawMarkdown": "Hi, I have a question.\nFrom \" XGB single model using my best 400 features.\" so If I want to use 400 features out of 1000 features. Do I need to retrain my model?",
      "votes": 2,
      "replies": [
        {
          "id": 2219977,
          "postDate": "2023-04-13T03:40:49.740Z",
          "content": "<p>Yes. My model is only trained on 400 features. During experimentation, i add features and compute CV score. If CV score increases then i keep the features. If CV score decreases, then i discard the features. After my experiments, i use all 400 features that helped CV score and train my final model.</p>",
          "rawMarkdown": "Yes. My model is only trained on 400 features. During experimentation, i add features and compute CV score. If CV score increases then i keep the features. If CV score decreases, then i discard the features. After my experiments, i use all 400 features that helped CV score and train my final model.",
          "votes": 2,
          "replies": [
            {
              "id": 2222069,
              "postDate": "2023-04-14T21:23:31.697Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks, I have learnt so much from you! Do you mind sharing how you choose the top 400 features? Do you compute feature importance from one of the folds only?</p>",
              "rawMarkdown": "@cdeotte thanks, I have learnt so much from you! Do you mind sharing how you choose the top 400 features? Do you compute feature importance from one of the folds only?"
            },
            {
              "id": 2222091,
              "postDate": "2023-04-14T22:12:37.797Z",
              "content": "<p>I have done forward selection. I start with my public notebook. Then i add a dozen new features if they improve CV, then i keep them all. If they do not then i discard them all. Next i add another dozen new features etc etc</p>",
              "rawMarkdown": "I have done forward selection. I start with my public notebook. Then i add a dozen new features if they improve CV, then i keep them all. If they do not then i discard them all. Next i add another dozen new features etc etc",
              "votes": 2
            },
            {
              "id": 2234620,
              "postDate": "2023-04-25T11:05:32.730Z",
              "content": "<p>thanks - if it is not too much to ask, do you mind sharing roughly how much features you add on each iteration? Also what CV improvement you consider 'real' and what is just noise? </p>",
              "rawMarkdown": "thanks - if it is not too much to ask, do you mind sharing roughly how much features you add on each iteration? Also what CV improvement you consider 'real' and what is just noise? "
            }
          ]
        }
      ]
    },
    {
      "id": 2193819,
      "postDate": "2023-03-23T14:19:48.010Z",
      "content": "<p>Hi, i am a fan of you.I have a question.<br>\nAfter the operation of Kfold,actually we get 18*5 models,will we get more well score in LB if we use the 90 models that conbine each model,for example,for each question,we can get the label that appear most frequently in 5 model output。 <br>\nInstead，you retrain the model and keep one model for each question. why ?</p>",
      "rawMarkdown": "Hi, i am a fan of you.I have a question.\nAfter the operation of Kfold,actually we get 18*5 models,will we get more well score in LB if we use the 90 models that conbine each model,for example,for each question,we can get the label that appear most frequently in 5 model output。 \nInstead，you retrain the model and keep one model for each question. why ?",
      "votes": 2,
      "replies": [
        {
          "id": 2193827,
          "postDate": "2023-03-23T14:26:41.163Z",
          "content": "<p>We only keep 1 KFold model per question to speed up inference. If you want to improve CV and LB, there are 3 options</p>\n<ul>\n<li>keep all KFold models. Then during inference use either 18, 36, 54, 72, or 90 models</li>\n<li>After KFold finishes, train 18 new models that each use 100% of data (with number of trees found by KFold early stop). Then during inference use the eighteen 100% models.</li>\n<li>train with <code>K=20</code> KFold, then only keep 1 KFold model per question. Then each of these models train with 95% of data. (And will do better than K=5 where each KFold model only uses 80% of data).</li>\n</ul>",
          "rawMarkdown": "We only keep 1 KFold model per question to speed up inference. If you want to improve CV and LB, there are 3 options\n* keep all KFold models. Then during inference use either 18, 36, 54, 72, or 90 models\n* After KFold finishes, train 18 new models that each use 100% of data (with number of trees found by KFold early stop). Then during inference use the eighteen 100% models.\n* train with `K=20` KFold, then only keep 1 KFold model per question. Then each of these models train with 95% of data. (And will do better than K=5 where each KFold model only uses 80% of data).",
          "votes": 5,
          "replies": [
            {
              "id": 2193904,
              "postDate": "2023-03-23T15:31:01.467Z",
              "content": "<p>Thanks,i get it. Futher more,regarding the options you mentioned, can i understand it this way.We can't say which options is the best,and the optimal options is different  for different situations.So,we should try all solution.</p>",
              "rawMarkdown": "Thanks,i get it. Futher more,regarding the options you mentioned, can i understand it this way.We can't say which options is the best,and the optimal options is different  for different situations.So,we should try all solution.",
              "votes": 1
            },
            {
              "id": 2193915,
              "postDate": "2023-03-23T15:34:06.633Z",
              "content": "<p>They should all perform similar. I suggest number 2 or 3 since they will infer faster. Number 2 is the best since it will infer <strong>and</strong> train faster. Number 2 only requires that we train 6 = 5 + 1 models per question whereas number 3 trains 20 models per question.</p>",
              "rawMarkdown": "They should all perform similar. I suggest number 2 or 3 since they will infer faster. Number 2 is the best since it will infer **and** train faster. Number 2 only requires that we train 6 = 5 + 1 models per question whereas number 3 trains 20 models per question."
            },
            {
              "id": 2193920,
              "postDate": "2023-03-23T15:39:39.967Z",
              "content": "<p>Thanks a lot! </p>",
              "rawMarkdown": "Thanks a lot! ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2184165,
      "postDate": "2023-03-16T07:17:58.570Z",
      "content": "<p>Thanks a lot!<br>\nFor those who are wondering what <br>\n<code>targets['session'] = pd.to_numeric( targets.session_id.str.split('_').list.get(0) )</code><br>\ndoes (as I did), it is basically just the translation of<br>\n<code>targets['session'] = targets.session_id.apply(lambda x: int(x.split('_')[0]) )</code> to cudf.</p>\n<p>targets.session_id.str.split('_') gives a cudf series like<br>\n0          [20090312431273200, q1]<br>\n1          [20090312433251036, q1]</p>\n<p>and then .list.get(0) gives a cudf series like<br>\n0         20090312431273200<br>\n1         20090312433251036</p>\n<p>with dtype = object<br>\nand then pd.to_numeric is used to change the dtype to int</p>",
      "rawMarkdown": "Thanks a lot!\nFor those who are wondering what \n`targets['session'] = pd.to_numeric( targets.session_id.str.split('_').list.get(0) )`\ndoes (as I did), it is basically just the translation of\n`targets['session'] = targets.session_id.apply(lambda x: int(x.split('_')[0]) )` to cudf.\n\n\ntargets.session_id.str.split('_') gives a cudf series like\n0          [20090312431273200, q1]\n1          [20090312433251036, q1]\n\nand then .list.get(0) gives a cudf series like\n0         20090312431273200\n1         20090312433251036\n\nwith dtype = object\nand then pd.to_numeric is used to change the dtype to int",
      "votes": 2
    },
    {
      "id": 2166817,
      "postDate": "2023-03-03T03:33:27.220Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>,<br>\nI'm facing a problem.<br>\nAfter adding 80-90 features, a point comes where the score decreases with any new bunch of features added. <br>\nIs this because the new features are just not helpful, there is too much noise in the previous features or there is some other reason? <br>\nCan I do anything to fix this?<br>\nThanks.</p>",
      "rawMarkdown": "Hi @cdeotte,\nI'm facing a problem.\nAfter adding 80-90 features, a point comes where the score decreases with any new bunch of features added. \nIs this because the new features are just not helpful, there is too much noise in the previous features or there is some other reason? \nCan I do anything to fix this?\nThanks.",
      "votes": 2
    },
    {
      "id": 2155379,
      "postDate": "2023-02-22T15:55:17.987Z",
      "content": "<p>This is brilliant Chris, it's great to see how much fast experimentation you can do for feature engineering using a GPU - it's all the extra features that make the difference! Out of interest, how are you handling features from different level groups? </p>\n<p>I can imagine adding these would have quite a boost on the model rather than only including features where you group by session ID and level group, but saving all of the features for a previous level groups and concatenating these on seems excessive, inefficient, and might add noise to the dataset. On the other hand, adding the score a model predicted for a session ID's previous levels could help a lot, but that might miss some of the information the model could use to get there. How are you making features between different level groups, and how are you deciding which ones to use at inference?</p>",
      "rawMarkdown": "This is brilliant Chris, it's great to see how much fast experimentation you can do for feature engineering using a GPU - it's all the extra features that make the difference! Out of interest, how are you handling features from different level groups? \n\nI can imagine adding these would have quite a boost on the model rather than only including features where you group by session ID and level group, but saving all of the features for a previous level groups and concatenating these on seems excessive, inefficient, and might add noise to the dataset. On the other hand, adding the score a model predicted for a session ID's previous levels could help a lot, but that might miss some of the information the model could use to get there. How are you making features between different level groups, and how are you deciding which ones to use at inference?",
      "votes": 2,
      "replies": [
        {
          "id": 2155399,
          "postDate": "2023-02-22T16:06:30.920Z",
          "content": "<p>Great question Jude. I believe adding features from different level groups can boost our CV and LB by at least <code>+0.004</code>. The LB was stuck at 0.694 before Kaggle fixed the API. Then the LB jumped up to 0.698. The fix allowed us to use info from different level groups so I suspect we can get at least <code>+0.004</code> boost (using info from different level groups).</p>\n<p>I have not tried using information from other level groups yet (with the correct Kaggle API). I did before the fix and was able to get about <code>+0.015</code> using the backwards leak in time level groups. But i have not tried since the fix. Currently, i'm just exploring more and more features within the one level group. I think we can get at least LB 0.694 using within only and i'm not quite there yet but close. There are lots of features to try!</p>",
          "rawMarkdown": "Great question Jude. I believe adding features from different level groups can boost our CV and LB by at least `+0.004`. The LB was stuck at 0.694 before Kaggle fixed the API. Then the LB jumped up to 0.698. The fix allowed us to use info from different level groups so I suspect we can get at least `+0.004` boost (using info from different level groups).\n\nI have not tried using information from other level groups yet (with the correct Kaggle API). I did before the fix and was able to get about `+0.015` using the backwards leak in time level groups. But i have not tried since the fix. Currently, i'm just exploring more and more features within the one level group. I think we can get at least LB 0.694 using within only and i'm not quite there yet but close. There are lots of features to try!",
          "votes": 5,
          "replies": [
            {
              "id": 2155459,
              "postDate": "2023-02-22T16:34:23.283Z",
              "content": "<p>That's really interesting, I hadn't spotted the LB jump after the API was fixed. It's amazing that you're getting such a strong score just within the level groups - I'll get my head down and build some more features then! </p>",
              "rawMarkdown": "That's really interesting, I hadn't spotted the LB jump after the API was fixed. It's amazing that you're getting such a strong score just within the level groups - I'll get my head down and build some more features then! ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2155337,
      "postDate": "2023-02-22T15:23:55.290Z",
      "content": "<p>Hello sir,<br>\nDo you first create all the features possible and then perform feature selection, or do you add the new features one by one and keep them if they improve the score?<br>\nThanks.</p>",
      "rawMarkdown": "Hello sir,\nDo you first create all the features possible and then perform feature selection, or do you add the new features one by one and keep them if they improve the score?\nThanks.",
      "votes": 2,
      "replies": [
        {
          "id": 2155341,
          "postDate": "2023-02-22T15:28:36.943Z",
          "content": "<p>I add dozens of new (related) features then compute CV score. If CV score increases I keep them all otherwise I discard them. Sometimes, when a new dozen fail (and I believe that it should help), I will give it another try by adding just half of them (that i think are best). I have not tried doing more careful feature selection yet. Right now i'm just exploring tons of new feature ideas.</p>",
          "rawMarkdown": "I add dozens of new (related) features then compute CV score. If CV score increases I keep them all otherwise I discard them. Sometimes, when a new dozen fail (and I believe that it should help), I will give it another try by adding just half of them (that i think are best). I have not tried doing more careful feature selection yet. Right now i'm just exploring tons of new feature ideas.",
          "votes": 1,
          "replies": [
            {
              "id": 2155362,
              "postDate": "2023-02-22T15:47:54.450Z",
              "content": "<p>Thank you for the reply.<br>\nSo just to clarify, if the half dozen also fail, then I should discard those or try with the half of those again?<br>\nAnd sir, will you do any feature selection at the end?<br>\nThanks.</p>",
              "rawMarkdown": "Thank you for the reply.\nSo just to clarify, if the half dozen also fail, then I should discard those or try with the half of those again?\nAnd sir, will you do any feature selection at the end?\nThanks.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2142337,
      "postDate": "2023-02-13T13:12:02.887Z",
      "content": "<p>Hello sir,<br>\nThank you for guiding us with such valuable information. <br>\nI have a question about hyperparameter optimization. Should I optimize the hyperparameters for the model after every feature that I create? <br>\nThe thing I want to ask is, at which steps of the modelling process should I do hyperparamater optimization?</p>",
      "rawMarkdown": "Hello sir,\nThank you for guiding us with such valuable information. \nI have a question about hyperparameter optimization. Should I optimize the hyperparameters for the model after every feature that I create? \nThe thing I want to ask is, at which steps of the modelling process should I do hyperparamater optimization?",
      "votes": 2,
      "replies": [
        {
          "id": 2142468,
          "postDate": "2023-02-13T14:52:32.093Z",
          "content": "<p>Different Kagglers have different opinions. I <strong>do not</strong> optimize hyperparameters after adding each new feature. Furthermore, I do not use many hyperparameters so there isn't much to optimize anyway. For GBT models, I only use <code>max_depth</code>, <code>subsample</code>, and <code>colsample_bytree</code>. My current best submission (CV 704 LB 703) uses the hyperparameters in my public notebook.</p>\n<p>I optimze these 3 parameters in my first model which has dozens of initial features. Sometimes after a week goes by and i have hundreds more features, i will check these hyperparameters. I will also check how much different <code>learning_rate</code> affects CV and LB. But as you can see, I keep the hyperparameters simple and <strong>do not</strong> change them often. So far changing them (from my public notebook) has not improved my CV nor LB.</p>\n<p>I spend the majority of my time searching for new features. Engineering new features is the most important.</p>",
          "rawMarkdown": "Different Kagglers have different opinions. I **do not** optimize hyperparameters after adding each new feature. Furthermore, I do not use many hyperparameters so there isn't much to optimize anyway. For GBT models, I only use `max_depth`, `subsample`, and `colsample_bytree`. My current best submission (CV 704 LB 703) uses the hyperparameters in my public notebook.\n\nI optimze these 3 parameters in my first model which has dozens of initial features. Sometimes after a week goes by and i have hundreds more features, i will check these hyperparameters. I will also check how much different `learning_rate` affects CV and LB. But as you can see, I keep the hyperparameters simple and **do not** change them often. So far changing them (from my public notebook) has not improved my CV nor LB.\n\nI spend the majority of my time searching for new features. Engineering new features is the most important.",
          "votes": 13,
          "replies": [
            {
              "id": 2142489,
              "postDate": "2023-02-13T15:10:02.910Z",
              "content": "<p>Thank you so much, this will help me a lot. There's just one more thing. Do you optimize the hyperparameters manually or you use something like GridSearch/RandomSearch/BayesianOptimization?</p>",
              "rawMarkdown": "Thank you so much, this will help me a lot. There's just one more thing. Do you optimize the hyperparameters manually or you use something like GridSearch/RandomSearch/BayesianOptimization?",
              "votes": 2
            },
            {
              "id": 2142501,
              "postDate": "2023-02-13T15:15:08.847Z",
              "content": "<p>I do a manual grid search. Since I only adjust 3 parameters it is simple. I check different max depths of 3, 4, 5, 6 etc. Then my subsample is usually always 0.8. Then i check colsample_bytree by checking 0.2, 0.3, 0.4, 0.5, 0.6</p>",
              "rawMarkdown": "I do a manual grid search. Since I only adjust 3 parameters it is simple. I check different max depths of 3, 4, 5, 6 etc. Then my subsample is usually always 0.8. Then i check colsample_bytree by checking 0.2, 0.3, 0.4, 0.5, 0.6",
              "votes": 2
            },
            {
              "id": 2142509,
              "postDate": "2023-02-13T15:17:00.657Z",
              "content": "<p>The <code>max_depth</code> is the most important to check (and will affect CV LB the most). In most competitions, we can just use <code>subsample = 0.8</code> and <code>colsample_bytree = 0.5</code> and then check max depths from 3, 4, 5, 6, 7, 8. Sometimes adjusting colsample_bytree helps. Changing subsample rarely helps.</p>",
              "rawMarkdown": "The `max_depth` is the most important to check (and will affect CV LB the most). In most competitions, we can just use `subsample = 0.8` and `colsample_bytree = 0.5` and then check max depths from 3, 4, 5, 6, 7, 8. Sometimes adjusting colsample_bytree helps. Changing subsample rarely helps.",
              "votes": 3
            },
            {
              "id": 2142537,
              "postDate": "2023-02-13T15:32:46.220Z",
              "content": "<p>Thank you so muchh!!! 😁</p>",
              "rawMarkdown": "Thank you so muchh!!! 😁",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2141795,
      "postDate": "2023-02-13T05:11:30.633Z",
      "content": "<p>That's a great strategy to boost your CV and LB score! Training on GPU and using CPU for inference can definitely speed up the process. And utilizing RAPID cuDF for feature engineering can make it even faster. I'm impressed with your effort to engineer over 1000 features, that's a lot of experimentation.</p>\n<p>Thanks for sharing! 🎉</p>",
      "rawMarkdown": "That's a great strategy to boost your CV and LB score! Training on GPU and using CPU for inference can definitely speed up the process. And utilizing RAPID cuDF for feature engineering can make it even faster. I'm impressed with your effort to engineer over 1000 features, that's a lot of experimentation.\n\nThanks for sharing! 🎉",
      "votes": 2
    },
    {
      "id": 2140578,
      "postDate": "2023-02-11T21:16:56.723Z",
      "content": "<p>Great post as usual , I just want to ask how much time did the 400 features inference took ?</p>",
      "rawMarkdown": "Great post as usual , I just want to ask how much time did the 400 features inference took ?",
      "votes": 2,
      "replies": [
        {
          "id": 2140582,
          "postDate": "2023-02-11T21:24:25.080Z",
          "content": "<p>My current best submission takes 6 hours. (Note I have not done any CPU inference optimization. I just load question models, create features using 1 thread and infer using 1 thread).  There are 18 separate question models and each model uses a different subset of features. Each model uses approximately 250 features.</p>",
          "rawMarkdown": "My current best submission takes 6 hours. (Note I have not done any CPU inference optimization. I just load question models, create features using 1 thread and infer using 1 thread).  There are 18 separate question models and each model uses a different subset of features. Each model uses approximately 250 features.",
          "votes": 8,
          "replies": [
            {
              "id": 2140585,
              "postDate": "2023-02-11T21:29:12.890Z",
              "content": "<p>Thanks for your reply !</p>",
              "rawMarkdown": "Thanks for your reply !"
            }
          ]
        }
      ]
    },
    {
      "id": 2981950,
      "postDate": "2024-09-07T09:59:28.667Z",
      "content": "<p>The CPU is the base😉</p>",
      "rawMarkdown": "The CPU is the base😉"
    },
    {
      "id": 2282084,
      "postDate": "2023-05-31T11:02:53.340Z",
      "content": "<p>Thanks for the guidance. I am getting <code>No module named 'jo_wilder.competition'</code> while trying to run Shashwat's inference notebook. I tried setting the environment option to <code>Pin to the original environment</code> but that is also not working. Am I the only one facing this problem? Could somebody help me with this?</p>",
      "rawMarkdown": "Thanks for the guidance. I am getting `No module named 'jo_wilder.competition'` while trying to run Shashwat's inference notebook. I tried setting the environment option to `Pin to the original environment` but that is also not working. Am I the only one facing this problem? Could somebody help me with this?"
    },
    {
      "id": 2276314,
      "postDate": "2023-05-26T17:28:33.623Z",
      "content": "<p>I encountered some errors during the feature engineering process, which were specifically raised during the ranking time. However, I managed to address those errors successfully. Although I was able to fix the issues, I encountered a new problem of exceeding the time limit during the ranking phase. Initially, my code took about 300 seconds to run. After refactoring the code, I managed to reduce the runtime to 120 seconds. Despite the improvement, I'm still facing the frustrating challenge of exceeding the time limit without receiving any feedback or specific information about the cause of the failure.</p>",
      "rawMarkdown": "I encountered some errors during the feature engineering process, which were specifically raised during the ranking time. However, I managed to address those errors successfully. Although I was able to fix the issues, I encountered a new problem of exceeding the time limit during the ranking phase. Initially, my code took about 300 seconds to run. After refactoring the code, I managed to reduce the runtime to 120 seconds. Despite the improvement, I'm still facing the frustrating challenge of exceeding the time limit without receiving any feedback or specific information about the cause of the failure."
    },
    {
      "id": 2246109,
      "postDate": "2023-05-04T22:10:12.830Z",
      "content": "<p>Amazing! It's really helpful!</p>",
      "rawMarkdown": "Amazing! It's really helpful!"
    },
    {
      "id": 2219021,
      "postDate": "2023-04-12T07:52:50.353Z",
      "content": "<p>Hi, Do you have a code snippet of using my best 400 features out of your 1000 features? How can I do that?</p>",
      "rawMarkdown": "Hi, Do you have a code snippet of using my best 400 features out of your 1000 features? How can I do that?"
    },
    {
      "id": 2226125,
      "postDate": "2023-04-18T17:07:30.417Z",
      "content": "<p>Thanks a lot for the information you shared. I never knew that feature engineering can get us such good results.  I am definitely going to try this way of training on GPU and trying inference on CPU. </p>",
      "rawMarkdown": "Thanks a lot for the information you shared. I never knew that feature engineering can get us such good results.  I am definitely going to try this way of training on GPU and trying inference on CPU. ",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 2224576,
      "postDate": "2023-04-17T12:42:24.750Z",
      "content": "<p>I have a question about the second notebook. There is a model created in the first notebook in the input tab, but how do I upload it to the input tab?（<a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference\" target=\"_blank\">here</a>）</p>",
      "rawMarkdown": "I have a question about the second notebook. There is a model created in the first notebook in the input tab, but how do I upload it to the input tab?（[here](https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference)）",
      "isDeleted": true
    },
    {
      "id": 2224316,
      "postDate": "2023-04-17T07:57:33.263Z",
      "content": "<p>Thank you for providing the notebook. When I ran your notebook, I encountered a memory error. Do you think there might be any settings that need to be adjusted outside of the notebook?（<a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\" target=\"_blank\">here</a>）</p>",
      "rawMarkdown": "Thank you for providing the notebook. When I ran your notebook, I encountered a memory error. Do you think there might be any settings that need to be adjusted outside of the notebook?（[here](https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train)）",
      "isDeleted": true,
      "replies": [
        {
          "id": 2224326,
          "postDate": "2023-04-17T08:06:47.277Z",
          "content": "<p>I tried running it on two different GPU environments on Kaggle, but in both cases, the notebook crashes at the beginning of the Feature Engineering cell.</p>",
          "rawMarkdown": "I tried running it on two different GPU environments on Kaggle, but in both cases, the notebook crashes at the beginning of the Feature Engineering cell.",
          "isDeleted": true,
          "replies": [
            {
              "id": 2224574,
              "postDate": "2023-04-17T12:39:21.760Z",
              "content": "<p>I have found a solution. I set up the method to use the 36GB of memory that you recommended, and by deleting train1, train2, and train3 at the FeatureEngineer stage, it worked without any problems.</p>",
              "rawMarkdown": "I have found a solution. I set up the method to use the 36GB of memory that you recommended, and by deleting train1, train2, and train3 at the FeatureEngineer stage, it worked without any problems.",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2140613,
      "postDate": "2023-02-11T22:40:28.547Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2140579,
      "postDate": "2023-02-11T21:17:42.497Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 2218143,
      "postDate": "2023-04-11T13:14:23.013Z",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": 1
    },
    {
      "id": 2216385,
      "postDate": "2023-04-10T04:00:52.780Z",
      "content": "<p>That's great! Thanks for sharing.</p>",
      "rawMarkdown": "That's great! Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 2187292,
      "postDate": "2023-03-18T15:44:23.133Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 2146355,
      "postDate": "2023-02-15T20:00:05.450Z",
      "content": "<p>That's wow, thanks for sharing.</p>",
      "rawMarkdown": "That's wow, thanks for sharing.",
      "votes": 2
    },
    {
      "id": 3294817,
      "postDate": "2025-09-26T20:13:08.593Z",
      "content": "<p>Thanks for all the useful information. </p>",
      "rawMarkdown": "Thanks for all the useful information. "
    }
  ],
  "comments": [
    {
      "id": 2140625,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2023-02-11T23:55:14.080000",
      "content": "<p>GPU is key, as usual 😃.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2222841,
      "author_name": "Tazria Helal",
      "author_url": "",
      "post_date": "2023-04-15T15:54:55.730000",
      "content": "<p>Thanks for all the info. I am learning quite a bit from your XGBoost notebook. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2187315,
      "author_name": "Omar Dramé",
      "author_url": "",
      "post_date": "2023-03-18T16:20:45.820000",
      "content": "<p>I don't quite understand what kind of features can you add to the model to reach 400, I'm really confused about that part. Also thank you for you contribution!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2187349,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-03-18T16:52:56.423000",
          "content": "<p>There are dozens of levels, dozens of rooms, dozens of events. If you start making features by grouping by pairs of level, room, event then you can quickly reach 1000 features.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2187356,
              "author_name": "Omar Dramé",
              "author_url": "",
              "post_date": "2023-03-18T16:59:48.153000",
              "content": "<p>ok I understand now, thank you very much!!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2170739,
      "author_name": "NHopeT",
      "author_url": "",
      "post_date": "2023-03-06T08:52:08.313000",
      "content": "<p>Thanks for sharing.　I now have a better understanding of this competition.👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2163243,
      "author_name": "Ebc",
      "author_url": "",
      "post_date": "2023-02-28T17:21:00.207000",
      "content": "<p>Thanks Chris for all the info. I am learning quite a bit from your XGBoost notebook. I assume there is temporal information (subsequent sessions) for a player, so I am thinking to try a model that keeps track of this information (memory), for example a RNN. Any thoughts or recommendations for that?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2162811,
      "author_name": "gemammi",
      "author_url": "",
      "post_date": "2023-02-28T12:21:14.503000",
      "content": "<p>Thanks Chris ! BTW if anyone isnt able to import cudf, you can easily install this <code>pip install cupy-cuda11x</code> in the notebook console. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2150685,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-02-19T14:01:30.880000",
      "content": "<p><strong>UPDATE - GPU Starter Notebook - LB 0.677</strong><br>\nShashwat has provided a RAPIDS cuDF GPU train/validate notebook <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\" target=\"_blank\">here</a> and CPU inference notebook <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference\" target=\"_blank\">here</a>. Thanks <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> . His starter notebook is faster than my original CPU starter notebook and achieves <code>+0.001</code> CV boost and <code>+0.001</code> LB boost, just using GPU instead of CPU (without changing the features or model). Hooray!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2145873,
      "author_name": "Santha Kumar V",
      "author_url": "",
      "post_date": "2023-02-15T12:20:35.997000",
      "content": "<p>Nice, Amazing use of RAPIDS</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2145567,
      "author_name": "Alex_X",
      "author_url": "",
      "post_date": "2023-02-15T07:42:08.447000",
      "content": "<p>Good job!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2144188,
      "author_name": "ashwath selvam",
      "author_url": "",
      "post_date": "2023-02-14T19:55:19.290000",
      "content": "<p>Amazing !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2143874,
      "author_name": "MIkhail Donskoy",
      "author_url": "",
      "post_date": "2023-02-14T15:38:34.827000",
      "content": "<p>Hmmm,  i think, what kaggle time limit for GPU too big for experiments(</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2143895,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-02-14T15:57:52.410000",
          "content": "<p>There is no time limit for GPU. We can train / experiment as much as we want. The time limit only applies to our inference notebook which we submit to Kaggle.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2147864,
              "author_name": "Abhishek Shah",
              "author_url": "",
              "post_date": "2023-02-17T00:09:20.933000",
              "content": "<p>He might be referring to the 30 hour weekly limit Kaggle imposes.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2143806,
      "author_name": "KobeInMemory",
      "author_url": "",
      "post_date": "2023-02-14T14:25:50.323000",
      "content": "<p>Thanks for your sharing and I have a question.Will you try deep neural network methods like transformer, training on GPU and infer on CPU?I think the deep model may have better performance since we have complex sequence data.Maybe you think there is no way to use deep model with the computation constraint? Looking forward to your reply.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2143899,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-02-14T15:59:41.460000",
          "content": "<p>Yes I plan to try Transformer and RNN. If we build our own transformer from scratch like <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\">here</a>, then we can keep it simple and fast to train and infer. This is what i plan to do after GBT. (There is RNN example <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a>)</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2143600,
      "author_name": "Simon Veitner",
      "author_url": "",
      "post_date": "2023-02-14T11:35:21.877000",
      "content": "<p>For the people familar with polars thats also an option.<br>\nIn moment I create 430 features in under 1.5 minutes with polars. <a href=\"https://www.kaggle.com/datasets/zaemon1251hesty/polars01516\" target=\"_blank\">this dataset</a> lets you install polars offline.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2142155,
      "author_name": "Maxens D",
      "author_url": "",
      "post_date": "2023-02-13T10:56:23.723000",
      "content": "<p>Amazing !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2141696,
      "author_name": "huyidao",
      "author_url": "",
      "post_date": "2023-02-13T02:40:30.020000",
      "content": "<p>that is amazing！</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2140847,
      "author_name": "toshi_k",
      "author_url": "",
      "post_date": "2023-02-12T07:55:26.263000",
      "content": "<blockquote>\n  <p>Only inference needs to be done on CPU. </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> my understanding, inference can be done on GPU for normal <strong>Leaderboard Prizes</strong>. Is it correct ?<br>\nI think inference must be done on CPU only for <strong>Efficiency Prizes</strong> .</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2140885,
          "author_name": "暴躁鹿",
          "author_url": "",
          "post_date": "2023-02-12T08:41:46.567000",
          "content": "<p>The GPU is not officially provided for this competition, so I don't think the inference is done on the GPU. I think the idea of the notebook is to do a fast experiment with the GPU to create and filter new features, and if the features are effective we will save the features and then complete the inference on the CPU and submit it. Is it correct ? <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2141074,
          "author_name": "zakopuro",
          "author_url": "",
          "post_date": "2023-02-12T12:33:19.907000",
          "content": "<p>Could not submit when GPU was turned on.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1985486%2F1b76a55bd776cca7b9e71037da1c8ba2%2FScreenshot_20230212-212951.png?generation=1676205041905134&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2141153,
              "author_name": "toshi_k",
              "author_url": "",
              "post_date": "2023-02-12T13:51:10.617000",
              "content": "<p><a href=\"https://www.kaggle.com/luzhonghao\" target=\"_blank\">@luzhonghao</a> <a href=\"https://www.kaggle.com/zakopur0\" target=\"_blank\">@zakopur0</a> Oh! I misunderstood the regulation. Thank you for sharing !</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2141344,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-02-12T17:11:16.293000",
              "content": "<p>On the top of code requirements page <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview/code-requirements\" target=\"_blank\">here</a> it says the paragraph below. And as others say above if you try to submit a GPU notebook, the competition will not let you. (Also note the CPU notebook is not the normal Kaggle CPU notebook, it only has 8GB RAM in this comp).</p>\n<blockquote>\n  <p>Note: this competition is aimed at producing models that are small and lightweight. We have introduced compute constraints to match - your VMs will have only 2 CPUs, 8GB of RAM, and no GPU available. You will still have a maximum of 9 hours to complete the task, but between the constraints and the efficiency prize there will be some interesting sub-problems to solve. Good luck!</p>\n</blockquote>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2219917,
      "author_name": "THITIWAT RUANGSAKORN",
      "author_url": "",
      "post_date": "2023-04-13T01:47:16.173000",
      "content": "<p>Hi, I have a question.<br>\nFrom \" XGB single model using my best 400 features.\" so If I want to use 400 features out of 1000 features. Do I need to retrain my model?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2219977,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-04-13T03:40:49.740000",
          "content": "<p>Yes. My model is only trained on 400 features. During experimentation, i add features and compute CV score. If CV score increases then i keep the features. If CV score decreases, then i discard the features. After my experiments, i use all 400 features that helped CV score and train my final model.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2222069,
              "author_name": "durvorezbariq",
              "author_url": "",
              "post_date": "2023-04-14T21:23:31.697000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks, I have learnt so much from you! Do you mind sharing how you choose the top 400 features? Do you compute feature importance from one of the folds only?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2222091,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-04-14T22:12:37.797000",
              "content": "<p>I have done forward selection. I start with my public notebook. Then i add a dozen new features if they improve CV, then i keep them all. If they do not then i discard them all. Next i add another dozen new features etc etc</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2234620,
              "author_name": "durvorezbariq",
              "author_url": "",
              "post_date": "2023-04-25T11:05:32.730000",
              "content": "<p>thanks - if it is not too much to ask, do you mind sharing roughly how much features you add on each iteration? Also what CV improvement you consider 'real' and what is just noise? </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2193819,
      "author_name": "huangzeyi123",
      "author_url": "",
      "post_date": "2023-03-23T14:19:48.010000",
      "content": "<p>Hi, i am a fan of you.I have a question.<br>\nAfter the operation of Kfold,actually we get 18*5 models,will we get more well score in LB if we use the 90 models that conbine each model,for example,for each question,we can get the label that appear most frequently in 5 model output。 <br>\nInstead，you retrain the model and keep one model for each question. why ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2193827,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-03-23T14:26:41.163000",
          "content": "<p>We only keep 1 KFold model per question to speed up inference. If you want to improve CV and LB, there are 3 options</p>\n<ul>\n<li>keep all KFold models. Then during inference use either 18, 36, 54, 72, or 90 models</li>\n<li>After KFold finishes, train 18 new models that each use 100% of data (with number of trees found by KFold early stop). Then during inference use the eighteen 100% models.</li>\n<li>train with <code>K=20</code> KFold, then only keep 1 KFold model per question. Then each of these models train with 95% of data. (And will do better than K=5 where each KFold model only uses 80% of data).</li>\n</ul>",
          "votes": 5,
          "replies": [
            {
              "id": 2193904,
              "author_name": "huangzeyi123",
              "author_url": "",
              "post_date": "2023-03-23T15:31:01.467000",
              "content": "<p>Thanks,i get it. Futher more,regarding the options you mentioned, can i understand it this way.We can't say which options is the best,and the optimal options is different  for different situations.So,we should try all solution.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2193915,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-03-23T15:34:06.633000",
              "content": "<p>They should all perform similar. I suggest number 2 or 3 since they will infer faster. Number 2 is the best since it will infer <strong>and</strong> train faster. Number 2 only requires that we train 6 = 5 + 1 models per question whereas number 3 trains 20 models per question.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2193920,
              "author_name": "huangzeyi123",
              "author_url": "",
              "post_date": "2023-03-23T15:39:39.967000",
              "content": "<p>Thanks a lot! </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2184165,
      "author_name": "Cherry Tomato",
      "author_url": "",
      "post_date": "2023-03-16T07:17:58.570000",
      "content": "<p>Thanks a lot!<br>\nFor those who are wondering what <br>\n<code>targets['session'] = pd.to_numeric( targets.session_id.str.split('_').list.get(0) )</code><br>\ndoes (as I did), it is basically just the translation of<br>\n<code>targets['session'] = targets.session_id.apply(lambda x: int(x.split('_')[0]) )</code> to cudf.</p>\n<p>targets.session_id.str.split('_') gives a cudf series like<br>\n0          [20090312431273200, q1]<br>\n1          [20090312433251036, q1]</p>\n<p>and then .list.get(0) gives a cudf series like<br>\n0         20090312431273200<br>\n1         20090312433251036</p>\n<p>with dtype = object<br>\nand then pd.to_numeric is used to change the dtype to int</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2166817,
      "author_name": "Shashwat Raman",
      "author_url": "",
      "post_date": "2023-03-03T03:33:27.220000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>,<br>\nI'm facing a problem.<br>\nAfter adding 80-90 features, a point comes where the score decreases with any new bunch of features added. <br>\nIs this because the new features are just not helpful, there is too much noise in the previous features or there is some other reason? <br>\nCan I do anything to fix this?<br>\nThanks.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2155379,
      "author_name": "Jude Hunt",
      "author_url": "",
      "post_date": "2023-02-22T15:55:17.987000",
      "content": "<p>This is brilliant Chris, it's great to see how much fast experimentation you can do for feature engineering using a GPU - it's all the extra features that make the difference! Out of interest, how are you handling features from different level groups? </p>\n<p>I can imagine adding these would have quite a boost on the model rather than only including features where you group by session ID and level group, but saving all of the features for a previous level groups and concatenating these on seems excessive, inefficient, and might add noise to the dataset. On the other hand, adding the score a model predicted for a session ID's previous levels could help a lot, but that might miss some of the information the model could use to get there. How are you making features between different level groups, and how are you deciding which ones to use at inference?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2155399,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-02-22T16:06:30.920000",
          "content": "<p>Great question Jude. I believe adding features from different level groups can boost our CV and LB by at least <code>+0.004</code>. The LB was stuck at 0.694 before Kaggle fixed the API. Then the LB jumped up to 0.698. The fix allowed us to use info from different level groups so I suspect we can get at least <code>+0.004</code> boost (using info from different level groups).</p>\n<p>I have not tried using information from other level groups yet (with the correct Kaggle API). I did before the fix and was able to get about <code>+0.015</code> using the backwards leak in time level groups. But i have not tried since the fix. Currently, i'm just exploring more and more features within the one level group. I think we can get at least LB 0.694 using within only and i'm not quite there yet but close. There are lots of features to try!</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2155459,
              "author_name": "Jude Hunt",
              "author_url": "",
              "post_date": "2023-02-22T16:34:23.283000",
              "content": "<p>That's really interesting, I hadn't spotted the LB jump after the API was fixed. It's amazing that you're getting such a strong score just within the level groups - I'll get my head down and build some more features then! </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2155337,
      "author_name": "Shashwat Raman",
      "author_url": "",
      "post_date": "2023-02-22T15:23:55.290000",
      "content": "<p>Hello sir,<br>\nDo you first create all the features possible and then perform feature selection, or do you add the new features one by one and keep them if they improve the score?<br>\nThanks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2155341,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-02-22T15:28:36.943000",
          "content": "<p>I add dozens of new (related) features then compute CV score. If CV score increases I keep them all otherwise I discard them. Sometimes, when a new dozen fail (and I believe that it should help), I will give it another try by adding just half of them (that i think are best). I have not tried doing more careful feature selection yet. Right now i'm just exploring tons of new feature ideas.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2155362,
              "author_name": "Shashwat Raman",
              "author_url": "",
              "post_date": "2023-02-22T15:47:54.450000",
              "content": "<p>Thank you for the reply.<br>\nSo just to clarify, if the half dozen also fail, then I should discard those or try with the half of those again?<br>\nAnd sir, will you do any feature selection at the end?<br>\nThanks.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2142337,
      "author_name": "Shashwat Raman",
      "author_url": "",
      "post_date": "2023-02-13T13:12:02.887000",
      "content": "<p>Hello sir,<br>\nThank you for guiding us with such valuable information. <br>\nI have a question about hyperparameter optimization. Should I optimize the hyperparameters for the model after every feature that I create? <br>\nThe thing I want to ask is, at which steps of the modelling process should I do hyperparamater optimization?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2142468,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-02-13T14:52:32.093000",
          "content": "<p>Different Kagglers have different opinions. I <strong>do not</strong> optimize hyperparameters after adding each new feature. Furthermore, I do not use many hyperparameters so there isn't much to optimize anyway. For GBT models, I only use <code>max_depth</code>, <code>subsample</code>, and <code>colsample_bytree</code>. My current best submission (CV 704 LB 703) uses the hyperparameters in my public notebook.</p>\n<p>I optimze these 3 parameters in my first model which has dozens of initial features. Sometimes after a week goes by and i have hundreds more features, i will check these hyperparameters. I will also check how much different <code>learning_rate</code> affects CV and LB. But as you can see, I keep the hyperparameters simple and <strong>do not</strong> change them often. So far changing them (from my public notebook) has not improved my CV nor LB.</p>\n<p>I spend the majority of my time searching for new features. Engineering new features is the most important.</p>",
          "votes": 13,
          "replies": [
            {
              "id": 2142489,
              "author_name": "Shashwat Raman",
              "author_url": "",
              "post_date": "2023-02-13T15:10:02.910000",
              "content": "<p>Thank you so much, this will help me a lot. There's just one more thing. Do you optimize the hyperparameters manually or you use something like GridSearch/RandomSearch/BayesianOptimization?</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2142501,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-02-13T15:15:08.847000",
              "content": "<p>I do a manual grid search. Since I only adjust 3 parameters it is simple. I check different max depths of 3, 4, 5, 6 etc. Then my subsample is usually always 0.8. Then i check colsample_bytree by checking 0.2, 0.3, 0.4, 0.5, 0.6</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2142509,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-02-13T15:17:00.657000",
              "content": "<p>The <code>max_depth</code> is the most important to check (and will affect CV LB the most). In most competitions, we can just use <code>subsample = 0.8</code> and <code>colsample_bytree = 0.5</code> and then check max depths from 3, 4, 5, 6, 7, 8. Sometimes adjusting colsample_bytree helps. Changing subsample rarely helps.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2142537,
              "author_name": "Shashwat Raman",
              "author_url": "",
              "post_date": "2023-02-13T15:32:46.220000",
              "content": "<p>Thank you so muchh!!! 😁</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2141795,
      "author_name": "Ryan",
      "author_url": "",
      "post_date": "2023-02-13T05:11:30.633000",
      "content": "<p>That's a great strategy to boost your CV and LB score! Training on GPU and using CPU for inference can definitely speed up the process. And utilizing RAPID cuDF for feature engineering can make it even faster. I'm impressed with your effort to engineer over 1000 features, that's a lot of experimentation.</p>\n<p>Thanks for sharing! 🎉</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2140578,
      "author_name": "Reacher",
      "author_url": "",
      "post_date": "2023-02-11T21:16:56.723000",
      "content": "<p>Great post as usual , I just want to ask how much time did the 400 features inference took ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2140582,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-02-11T21:24:25.080000",
          "content": "<p>My current best submission takes 6 hours. (Note I have not done any CPU inference optimization. I just load question models, create features using 1 thread and infer using 1 thread).  There are 18 separate question models and each model uses a different subset of features. Each model uses approximately 250 features.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2140585,
              "author_name": "Reacher",
              "author_url": "",
              "post_date": "2023-02-11T21:29:12.890000",
              "content": "<p>Thanks for your reply !</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2981950,
      "author_name": "Oliynyk Oleksandr",
      "author_url": "",
      "post_date": "2024-09-07T09:59:28.667000",
      "content": "<p>The CPU is the base😉</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2282084,
      "author_name": "Anshul Yadav",
      "author_url": "",
      "post_date": "2023-05-31T11:02:53.340000",
      "content": "<p>Thanks for the guidance. I am getting <code>No module named 'jo_wilder.competition'</code> while trying to run Shashwat's inference notebook. I tried setting the environment option to <code>Pin to the original environment</code> but that is also not working. Am I the only one facing this problem? Could somebody help me with this?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2276314,
      "author_name": "Fabricio Braz",
      "author_url": "",
      "post_date": "2023-05-26T17:28:33.623000",
      "content": "<p>I encountered some errors during the feature engineering process, which were specifically raised during the ranking time. However, I managed to address those errors successfully. Although I was able to fix the issues, I encountered a new problem of exceeding the time limit during the ranking phase. Initially, my code took about 300 seconds to run. After refactoring the code, I managed to reduce the runtime to 120 seconds. Despite the improvement, I'm still facing the frustrating challenge of exceeding the time limit without receiving any feedback or specific information about the cause of the failure.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2246109,
      "author_name": "Hanmo Li",
      "author_url": "",
      "post_date": "2023-05-04T22:10:12.830000",
      "content": "<p>Amazing! It's really helpful!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2219021,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-12T07:52:50.353000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2226125,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-18T17:07:30.417000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2224576,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-17T12:42:24.750000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2224316,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-17T07:57:33.263000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2224326,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-04-17T08:06:47.277000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2224574,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-04-17T12:39:21.760000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2140613,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-11T22:40:28.547000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2140579,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-11T21:17:42.497000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2218143,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-11T13:14:23.013000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2216385,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-10T04:00:52.780000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2187292,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-18T15:44:23.133000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2146355,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-15T20:00:05.450000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3294817,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-09-26T20:13:08.593000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2140569": "I would like to point out that we do **not need** to feature engineer on CPU **nor** train on CPU in this competition. Only inference needs to be done on CPU. We should train in a one notebook and save model. Then infer with another notebook and load model.\n\n# Use CPU Kaggle Notebook for Inference\nDuring inference we must use CPU 8GB notebook, however there is plenty of memory because the Kaggle API uses a `for-loop`. In each iteration of the `for-loop` we are given two very small dataframes (`test.csv` and `sample_submission.csv`) that we process and infer. Each iteration we receive a new test.csv and sample_submission.csv. There are no memory issues because each test dataframe is a few hundred rows with 20 columns. Its size in memory is under 1MB. \n\nFYI, in each iteration of the `for-loop` we receive one chapter of one user (out of total of 3 chapters and 11k test users). A chapter is a few hundred rows of dataframe with one event per row (and 20 columns). When we receive  chapter one (i.e. `level_group = '0-4'`) then we use this data to predict questions `1-3` for the assoicated user. When we receive chapter two (i.e. `level_group = '5-12'`) then we use this data to predict questions `4-13`. When we receive chapter three (i.e. `level_group = '13-22'`) then we predict questions `14-18`.\n\n# Use GPU Kaggle Notebook for Train\nDuring training, we must load and process `train.csv` which has 13 million rows and 20 columns. It contains all 3 chapters for 11k train users. This file is a few GB and may cause memory problems when we feature engineer and/or train. Furthermore feature engineering and training takes time. Therefore I suggest using GPU Kaggle notebook during train. This notebook has 16GB GPU VRAM and 32GB CPU RAM for a total of 48GB of RAM! This is plenty of memory and speed to process `train.csv` and train our models.\n\n# XGBoost Starter Notebook - LB 0.680\nI published an XGBoost starter notebook [here][1]. Currently feature engineer, train, and infer are in one CPU notebook. I suggest creating two notebooks. Put the `train.csv` feature engineering and model training in the first GPU notebook. Then after training XGB Classifier model use \n\n    model.save_model(f'XGB_question_{t}.xgb')\n\nUpload these models to a Kaggle dataset. Then put the inference code in a second CPU notebook. And load model with\n\n    model = XGBClassifier()\n    model.load_model(f'XGB_question_{t}.xgb')\n\nUsing GPU to train XGB will give us **4x or more speed increase!** for model training and **10x or more speed increase!** for feature engineering\n\n# RAPIDS cuDF Feature Engineer\nIn the first notebook to utilize cuDF, we can change \n\n    import pandas as pd\n\ninto the following\n\n    import cudf as pd\n\nThen our feature engineering will be **10x or more faster!** There are two other changes we need to convert Pandas to RAPIDS cuDF. Use the two codes below\n\n    targets = pd.read_csv('train_labels.csv')\n    targets['session'] = pd.to_numeric( targets.session_id.str.split('_').list.get(0) )\n    targets['q'] = pd.to_numeric( targets.session_id.str.split('_q').list.get(1) )\n\nAnd because GroupKFold wants Pandas:\n\n    P1 = df.iloc[:,0].to_pandas()\n    P2 = df.index.to_pandas()\n    for i, (train_index, test_index) in enumerate(gkf.split(X=P1, groups=P2)):\n\n# How To Improve CV and LB\nTo secret to improving CV and LB is lots of fast experiments. Using GPU cuDF and GPU XGB, we can engineer a few new features and train 5x18=90 models in a few minutes. If the CV score increases, then we keep the new features, otherwise we discard and try again. \n\nSo far, using GPU, I have created and tested over 1000 features! My current best solution is XGB single model using my best 400 features. It achieves CV 0.703 and LB 0.702!\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Feb-2023/gpu.jpeg)\n\n# UPDATE - GPU Starter Notebook - LB 0.677\nShashwat has provided a RAPIDS cuDF GPU train/validate notebook [here][2] and CPU inference notebook [here][3]. Thanks @shashwatraman . His starter notebook is faster than my original CPU starter notebook and achieves `+0.001` CV boost and `+0.001` LB boost, just using GPU instead of CPU (without changing the features or model). Hooray!\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\n[2]: https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\n[3]: https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference",
    "2140625": "GPU is key, as usual 😃.",
    "2222841": "Thanks for all the info. I am learning quite a bit from your XGBoost notebook. @cdeotte ",
    "2187315": "I don't quite understand what kind of features can you add to the model to reach 400, I'm really confused about that part. Also thank you for you contribution!",
    "2170739": "Thanks for sharing.　I now have a better understanding of this competition.👍",
    "2163243": "Thanks Chris for all the info. I am learning quite a bit from your XGBoost notebook. I assume there is temporal information (subsequent sessions) for a player, so I am thinking to try a model that keeps track of this information (memory), for example a RNN. Any thoughts or recommendations for that?",
    "2162811": "Thanks Chris ! BTW if anyone isnt able to import cudf, you can easily install this `pip install cupy-cuda11x` in the notebook console. ",
    "2150685": "**UPDATE - GPU Starter Notebook - LB 0.677**\nShashwat has provided a RAPIDS cuDF GPU train/validate notebook [here][2] and CPU inference notebook [here][3]. Thanks @shashwatraman . His starter notebook is faster than my original CPU starter notebook and achieves `+0.001` CV boost and `+0.001` LB boost, just using GPU instead of CPU (without changing the features or model). Hooray!\n\n[2]: https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\n[3]: https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference",
    "2145873": "Nice, Amazing use of RAPIDS",
    "2145567": "Good job!!!",
    "2144188": "Amazing !!\n\n\n",
    "2143874": "Hmmm,  i think, what kaggle time limit for GPU too big for experiments(",
    "2143806": "Thanks for your sharing and I have a question.Will you try deep neural network methods like transformer, training on GPU and infer on CPU?I think the deep model may have better performance since we have complex sequence data.Maybe you think there is no way to use deep model with the computation constraint? Looking forward to your reply.",
    "2143600": "For the people familar with polars thats also an option.\nIn moment I create 430 features in under 1.5 minutes with polars. [this dataset](https://www.kaggle.com/datasets/zaemon1251hesty/polars01516) lets you install polars offline.",
    "2142155": "Amazing !!",
    "2141696": "that is amazing！",
    "2140847": "> Only inference needs to be done on CPU. \n\n@cdeotte my understanding, inference can be done on GPU for normal **Leaderboard Prizes**. Is it correct ?\nI think inference must be done on CPU only for **Efficiency Prizes** .",
    "2219917": "Hi, I have a question.\nFrom \" XGB single model using my best 400 features.\" so If I want to use 400 features out of 1000 features. Do I need to retrain my model?",
    "2193819": "Hi, i am a fan of you.I have a question.\nAfter the operation of Kfold,actually we get 18*5 models,will we get more well score in LB if we use the 90 models that conbine each model,for example,for each question,we can get the label that appear most frequently in 5 model output。 \nInstead，you retrain the model and keep one model for each question. why ?",
    "2184165": "Thanks a lot!\nFor those who are wondering what \n`targets['session'] = pd.to_numeric( targets.session_id.str.split('_').list.get(0) )`\ndoes (as I did), it is basically just the translation of\n`targets['session'] = targets.session_id.apply(lambda x: int(x.split('_')[0]) )` to cudf.\n\n\ntargets.session_id.str.split('_') gives a cudf series like\n0          [20090312431273200, q1]\n1          [20090312433251036, q1]\n\nand then .list.get(0) gives a cudf series like\n0         20090312431273200\n1         20090312433251036\n\nwith dtype = object\nand then pd.to_numeric is used to change the dtype to int",
    "2166817": "Hi @cdeotte,\nI'm facing a problem.\nAfter adding 80-90 features, a point comes where the score decreases with any new bunch of features added. \nIs this because the new features are just not helpful, there is too much noise in the previous features or there is some other reason? \nCan I do anything to fix this?\nThanks.",
    "2155379": "This is brilliant Chris, it's great to see how much fast experimentation you can do for feature engineering using a GPU - it's all the extra features that make the difference! Out of interest, how are you handling features from different level groups? \n\nI can imagine adding these would have quite a boost on the model rather than only including features where you group by session ID and level group, but saving all of the features for a previous level groups and concatenating these on seems excessive, inefficient, and might add noise to the dataset. On the other hand, adding the score a model predicted for a session ID's previous levels could help a lot, but that might miss some of the information the model could use to get there. How are you making features between different level groups, and how are you deciding which ones to use at inference?",
    "2155337": "Hello sir,\nDo you first create all the features possible and then perform feature selection, or do you add the new features one by one and keep them if they improve the score?\nThanks.",
    "2142337": "Hello sir,\nThank you for guiding us with such valuable information. \nI have a question about hyperparameter optimization. Should I optimize the hyperparameters for the model after every feature that I create? \nThe thing I want to ask is, at which steps of the modelling process should I do hyperparamater optimization?",
    "2141795": "That's a great strategy to boost your CV and LB score! Training on GPU and using CPU for inference can definitely speed up the process. And utilizing RAPID cuDF for feature engineering can make it even faster. I'm impressed with your effort to engineer over 1000 features, that's a lot of experimentation.\n\nThanks for sharing! 🎉",
    "2140578": "Great post as usual , I just want to ask how much time did the 400 features inference took ?",
    "2981950": "The CPU is the base😉",
    "2282084": "Thanks for the guidance. I am getting `No module named 'jo_wilder.competition'` while trying to run Shashwat's inference notebook. I tried setting the environment option to `Pin to the original environment` but that is also not working. Am I the only one facing this problem? Could somebody help me with this?",
    "2276314": "I encountered some errors during the feature engineering process, which were specifically raised during the ranking time. However, I managed to address those errors successfully. Although I was able to fix the issues, I encountered a new problem of exceeding the time limit during the ranking phase. Initially, my code took about 300 seconds to run. After refactoring the code, I managed to reduce the runtime to 120 seconds. Despite the improvement, I'm still facing the frustrating challenge of exceeding the time limit without receiving any feedback or specific information about the cause of the failure.",
    "2246109": "Amazing! It's really helpful!",
    "2219021": "Hi, Do you have a code snippet of using my best 400 features out of your 1000 features? How can I do that?",
    "2226125": "Thanks a lot for the information you shared. I never knew that feature engineering can get us such good results.  I am definitely going to try this way of training on GPU and trying inference on CPU. ",
    "2224576": "I have a question about the second notebook. There is a model created in the first notebook in the input tab, but how do I upload it to the input tab?（[here](https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference)）",
    "2224316": "Thank you for providing the notebook. When I ran your notebook, I encountered a memory error. Do you think there might be any settings that need to be adjusted outside of the notebook?（[here](https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train)）",
    "2140613": "",
    "2140579": "",
    "2218143": "Thanks a lot!",
    "2216385": "That's great! Thanks for sharing.",
    "2187292": "Thanks for sharing!",
    "2146355": "That's wow, thanks for sharing.",
    "3294817": "Thanks for all the useful information. "
  }
}