{
  "id": 388779,
  "title": "GPU Powered Feature Engineering and Training [Using RAPIDS cuDF]",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/388779",
  "author_name": "",
  "post_date": "2023-02-19T13:42:47.821104500Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Heyy,<br>\nI made a <strong>notebook series</strong> to implement Chris's idea to use <strong>GPU</strong> Kaggle notebook for feature engineering during training using <strong>RAPIDs cuDF</strong> and then CPU for inference from the discussion <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218\" target=\"_blank\">here</a>.</p>\n<h5>The series has two parts:</h5>\n<p>Part1 - Feature Engineering and XGBoost Training using GPU - <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\" target=\"_blank\">Link</a><br>\nPart2 - Making a Submission using CPU - <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference\" target=\"_blank\">Link</a></p>\n<p>Using GPU makes it really convenient to perform lightning fast experiments for feature engineering in this competition. In the CPU notebook, feature engineering takes about 1 min but utilizing the power of GPU takes the time down to like <strong>3 seconds!!</strong>!</p>\n<p>I decided to make these notebooks as I faced some problems in getting all this to work and in importing cuDF in the current Kaggle Environment.</p>\n<p>Hope it helps. Thanks you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for such an amazing idea.<br>\nThanks. 😊</p>",
  "messages": [
    {
      "id": "2150661",
      "postDate": "02/19/2023 13:42:47",
      "content": "<p>Heyy,<br>\nI made a <strong>notebook series</strong> to implement Chris's idea to use <strong>GPU</strong> Kaggle notebook for feature engineering during training using <strong>RAPIDs cuDF</strong> and then CPU for inference from the discussion <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218\" target=\"_blank\">here</a>.</p>\n<h5>The series has two parts:</h5>\n<p>Part1 - Feature Engineering and XGBoost Training using GPU - <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train\" target=\"_blank\">Link</a><br>\nPart2 - Making a Submission using CPU - <a href=\"https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference\" target=\"_blank\">Link</a></p>\n<p>Using GPU makes it really convenient to perform lightning fast experiments for feature engineering in this competition. In the CPU notebook, feature engineering takes about 1 min but utilizing the power of GPU takes the time down to like <strong>3 seconds!!</strong>!</p>\n<p>I decided to make these notebooks as I faced some problems in getting all this to work and in importing cuDF in the current Kaggle Environment.</p>\n<p>Hope it helps. Thanks you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for such an amazing idea.<br>\nThanks. 😊</p>",
      "rawMarkdown": "Heyy,\nI made a **notebook series** to implement Chris's idea to use **GPU** Kaggle notebook for feature engineering during training using **RAPIDs cuDF** and then CPU for inference from the discussion [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218).\n\n##### The series has two parts:\nPart1 - Feature Engineering and XGBoost Training using GPU - [Link](https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train)\nPart2 - Making a Submission using CPU - [Link](https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference)\n\nUsing GPU makes it really convenient to perform lightning fast experiments for feature engineering in this competition. In the CPU notebook, feature engineering takes about 1 min but utilizing the power of GPU takes the time down to like **3 seconds!!**!\n\nI decided to make these notebooks as I faced some problems in getting all this to work and in importing cuDF in the current Kaggle Environment.\n\nHope it helps. Thanks you @cdeotte for such an amazing idea.\nThanks. 😊",
      "votes": null
    },
    {
      "id": "2150674",
      "postDate": "02/19/2023 13:53:42",
      "content": "<p>Excellent job <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> It looks fantastic. It even gets <code>+0.001</code> CV and <code>+0.001</code> LB! This is helpful thanks! I link your two notebook from my discussion</p>\n<p>=============</p>\n<p>Note that your inference notebook reads from the disk 198,000 times. For each of 11k users and for each question, you read the XGB model from disk. Reading from disk is slow. It is best to read 18 times from disk <strong>ONCE</strong> before Kaggle's API and put the loaded models into a Python list like</p>\n<pre><code> QUESTION_MODELS = []\n for t in range(1,19):\n     clf = XGBClassifier()\n     clf.load_model(f'../input/gpu-xgb-baseline-using-rapids-cudf-train/XGB_question_{t}.xgb')\n     QUESTION_MODELS.append( clf )\n</code></pre>\n<p>Then during Kaggle's API we can do </p>\n<pre><code>clf = QUESTION_MODELS[t-1]\n</code></pre>",
      "rawMarkdown": "Excellent job @shashwatraman It looks fantastic. It even gets `+0.001` CV and `+0.001` LB! This is helpful thanks! I link your two notebook from my discussion\n\n=============\n\nNote that your inference notebook reads from the disk 198,000 times. For each of 11k users and for each question, you read the XGB model from disk. Reading from disk is slow. It is best to read 18 times from disk **ONCE** before Kaggle's API and put the loaded models into a Python list like\n\n     QUESTION_MODELS = []\n     for t in range(1,19):\n         clf = XGBClassifier()\n         clf.load_model(f'../input/gpu-xgb-baseline-using-rapids-cudf-train/XGB_question_{t}.xgb')\n         QUESTION_MODELS.append( clf )\n\nThen during Kaggle's API we can do \n\n    clf = QUESTION_MODELS[t-1]",
      "votes": null
    },
    {
      "id": "2150688",
      "postDate": "02/19/2023 14:03:41",
      "content": "<p>Thank you sir. You're too kind. </p>\n<p>I'll make these changes right away.<br>\nThanks.</p>",
      "rawMarkdown": "Thank you sir. You're too kind. \n\nI'll make these changes right away.\nThanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2150674,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/19/2023 13:53:42",
      "content": "<p>Excellent job <a href=\"https://www.kaggle.com/shashwatraman\" target=\"_blank\">@shashwatraman</a> It looks fantastic. It even gets <code>+0.001</code> CV and <code>+0.001</code> LB! This is helpful thanks! I link your two notebook from my discussion</p>\n<p>=============</p>\n<p>Note that your inference notebook reads from the disk 198,000 times. For each of 11k users and for each question, you read the XGB model from disk. Reading from disk is slow. It is best to read 18 times from disk <strong>ONCE</strong> before Kaggle's API and put the loaded models into a Python list like</p>\n<pre><code> QUESTION_MODELS = []\n for t in range(1,19):\n     clf = XGBClassifier()\n     clf.load_model(f'../input/gpu-xgb-baseline-using-rapids-cudf-train/XGB_question_{t}.xgb')\n     QUESTION_MODELS.append( clf )\n</code></pre>\n<p>Then during Kaggle's API we can do </p>\n<pre><code>clf = QUESTION_MODELS[t-1]\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2150688,
          "author_name": "shashwatraman",
          "author_url": "",
          "post_date": "02/19/2023 14:03:41",
          "content": "<p>Thank you sir. You're too kind. </p>\n<p>I'll make these changes right away.<br>\nThanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2150661": "Heyy,\nI made a **notebook series** to implement Chris's idea to use **GPU** Kaggle notebook for feature engineering during training using **RAPIDs cuDF** and then CPU for inference from the discussion [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218).\n\n##### The series has two parts:\nPart1 - Feature Engineering and XGBoost Training using GPU - [Link](https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-train)\nPart2 - Making a Submission using CPU - [Link](https://www.kaggle.com/code/shashwatraman/gpu-xgb-baseline-using-rapids-cudf-inference)\n\nUsing GPU makes it really convenient to perform lightning fast experiments for feature engineering in this competition. In the CPU notebook, feature engineering takes about 1 min but utilizing the power of GPU takes the time down to like **3 seconds!!**!\n\nI decided to make these notebooks as I faced some problems in getting all this to work and in importing cuDF in the current Kaggle Environment.\n\nHope it helps. Thanks you @cdeotte for such an amazing idea.\nThanks. 😊",
    "2150674": "Excellent job @shashwatraman It looks fantastic. It even gets `+0.001` CV and `+0.001` LB! This is helpful thanks! I link your two notebook from my discussion\n\n=============\n\nNote that your inference notebook reads from the disk 198,000 times. For each of 11k users and for each question, you read the XGB model from disk. Reading from disk is slow. It is best to read 18 times from disk **ONCE** before Kaggle's API and put the loaded models into a Python list like\n\n     QUESTION_MODELS = []\n     for t in range(1,19):\n         clf = XGBClassifier()\n         clf.load_model(f'../input/gpu-xgb-baseline-using-rapids-cudf-train/XGB_question_{t}.xgb')\n         QUESTION_MODELS.append( clf )\n\nThen during Kaggle's API we can do \n\n    clf = QUESTION_MODELS[t-1]",
    "2150688": "Thank you sir. You're too kind. \n\nI'll make these changes right away.\nThanks."
  },
  "source": "meta"
}