{
  "id": 399133,
  "title": "Achieve [CV 0.6914 | LB 0.694] with Only Four Features - Training Part Published",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/399133",
  "author_name": "",
  "post_date": "2023-04-02T15:41:44.577703900Z",
  "votes": 58,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>After spending some time refactoring the local pipeline into readable notebook, I finally publish the <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-tconv-with-4-features-training-part\" target=\"_blank\"><strong>training</strong> counterpart</a> of the inference notebook released about 3 days ago.</p>\n<p>In this kernel, I implement a complete pipeline, including data cleaning and processing, simple feature engineering, and model training with the specified CV scheme (i.e., <code>GroupKFold</code> with k=5). After the process is done, output objects (e.g., best model checkpoints) can be downloaded and used in the inference part.</p>\n<p>To be concrete, the experiment setup can be summarized as follows:</p>\n<ol>\n<li>Only 4 features are taken into consideration, including <code>elapsed_time</code>, <code>event_name</code>, <code>name</code>, and <code>room_fqid</code>.</li>\n<li>The model architecture is mainly based on 1D-Conv and categorical embeddings (as illustrated in following figure).<br>\n<a href=\"https://postimg.cc/vxDzT9RJ\" target=\"_blank\"><img src=\"https://i.postimg.cc/90YNxtLQ/2023-04-02-11-22-49.png\" alt=\"2023-04-02-11-22-49.png\"></a></li>\n<li><code>BCEWithLogitsLoss</code> is selected to be the loss criterion.</li>\n<li><code>Adam</code> and <code>CosineAnnealingWarmRestarts</code> are used.</li>\n<li><code>GroupKFold</code> with <code>k=5</code> is adopted to be the CV scheme, and all models are trained with 100 epochs without early stopping.</li>\n</ol>\n<p>Also, I think there's still a long way to go. In the future, I'm going to extend the experiments as follows:</p>\n<ol>\n<li>Add more features (<em>e.g.,</em> text information).</li>\n<li>Try different model architectures.<ul>\n<li>I've tried a simple modification of the model architecture, <a href=\"https://arxiv.org/abs/2010.12042\" target=\"_blank\">SAINT+</a>, but still can't obtain satisfying performance.</li></ul></li>\n<li>Add more training data (<em>i.e.,</em> released old test set).<ul>\n<li>As you can see in the training notebook, I use the old version of training set to train models.</li></ul></li>\n<li>Ensemble with tree-based models.</li>\n<li>Do some hyperparameter tuning near the end of the competition.</li>\n</ol>\n<p>Hope this quick summarization can help clarify the core setup of the experiment. If there's any mistake, please feel free to correct me. I appreciate you taking the time to look over the work and discuss with me.</p>",
  "messages": [
    {
      "id": "2206499",
      "postDate": "04/02/2023 15:41:44",
      "content": "<p>Hi everyone,</p>\n<p>After spending some time refactoring the local pipeline into readable notebook, I finally publish the <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-tconv-with-4-features-training-part\" target=\"_blank\"><strong>training</strong> counterpart</a> of the inference notebook released about 3 days ago.</p>\n<p>In this kernel, I implement a complete pipeline, including data cleaning and processing, simple feature engineering, and model training with the specified CV scheme (i.e., <code>GroupKFold</code> with k=5). After the process is done, output objects (e.g., best model checkpoints) can be downloaded and used in the inference part.</p>\n<p>To be concrete, the experiment setup can be summarized as follows:</p>\n<ol>\n<li>Only 4 features are taken into consideration, including <code>elapsed_time</code>, <code>event_name</code>, <code>name</code>, and <code>room_fqid</code>.</li>\n<li>The model architecture is mainly based on 1D-Conv and categorical embeddings (as illustrated in following figure).<br>\n<a href=\"https://postimg.cc/vxDzT9RJ\" target=\"_blank\"><img src=\"https://i.postimg.cc/90YNxtLQ/2023-04-02-11-22-49.png\" alt=\"2023-04-02-11-22-49.png\"></a></li>\n<li><code>BCEWithLogitsLoss</code> is selected to be the loss criterion.</li>\n<li><code>Adam</code> and <code>CosineAnnealingWarmRestarts</code> are used.</li>\n<li><code>GroupKFold</code> with <code>k=5</code> is adopted to be the CV scheme, and all models are trained with 100 epochs without early stopping.</li>\n</ol>\n<p>Also, I think there's still a long way to go. In the future, I'm going to extend the experiments as follows:</p>\n<ol>\n<li>Add more features (<em>e.g.,</em> text information).</li>\n<li>Try different model architectures.<ul>\n<li>I've tried a simple modification of the model architecture, <a href=\"https://arxiv.org/abs/2010.12042\" target=\"_blank\">SAINT+</a>, but still can't obtain satisfying performance.</li></ul></li>\n<li>Add more training data (<em>i.e.,</em> released old test set).<ul>\n<li>As you can see in the training notebook, I use the old version of training set to train models.</li></ul></li>\n<li>Ensemble with tree-based models.</li>\n<li>Do some hyperparameter tuning near the end of the competition.</li>\n</ol>\n<p>Hope this quick summarization can help clarify the core setup of the experiment. If there's any mistake, please feel free to correct me. I appreciate you taking the time to look over the work and discuss with me.</p>",
      "rawMarkdown": "Hi everyone,\n\nAfter spending some time refactoring the local pipeline into readable notebook, I finally publish the [**training** counterpart](https://www.kaggle.com/code/abaojiang/lb-0-694-tconv-with-4-features-training-part) of the inference notebook released about 3 days ago.\n\nIn this kernel, I implement a complete pipeline, including data cleaning and processing, simple feature engineering, and model training with the specified CV scheme (i.e., `GroupKFold` with k=5). After the process is done, output objects (e.g., best model checkpoints) can be downloaded and used in the inference part.\n\nTo be concrete, the experiment setup can be summarized as follows:\n1. Only 4 features are taken into consideration, including `elapsed_time`, `event_name`, `name`, and `room_fqid`.\n2. The model architecture is mainly based on 1D-Conv and categorical embeddings (as illustrated in following figure).\n[![2023-04-02-11-22-49.png](https://i.postimg.cc/90YNxtLQ/2023-04-02-11-22-49.png)](https://postimg.cc/vxDzT9RJ)\n3. `BCEWithLogitsLoss` is selected to be the loss criterion.\n4. `Adam` and `CosineAnnealingWarmRestarts` are used.\n5. `GroupKFold` with `k=5` is adopted to be the CV scheme, and all models are trained with 100 epochs without early stopping.\n\nAlso, I think there's still a long way to go. In the future, I'm going to extend the experiments as follows:\n1. Add more features (*e.g.,* text information).\n2. Try different model architectures.\n    * I've tried a simple modification of the model architecture, [SAINT+](https://arxiv.org/abs/2010.12042), but still can't obtain satisfying performance.\n3. Add more training data (*i.e.,* released old test set).\n    * As you can see in the training notebook, I use the old version of training set to train models.\n4. Ensemble with tree-based models.\n5. Do some hyperparameter tuning near the end of the competition.\n\nHope this quick summarization can help clarify the core setup of the experiment. If there's any mistake, please feel free to correct me. I appreciate you taking the time to look over the work and discuss with me.",
      "votes": null
    },
    {
      "id": "2206676",
      "postDate": "04/02/2023 18:25:01",
      "content": "<p>thank you so much for publishing this!</p>",
      "rawMarkdown": "thank you so much for publishing this!",
      "votes": null
    },
    {
      "id": "2206884",
      "postDate": "04/03/2023 01:25:19",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2208273",
      "postDate": "04/03/2023 23:55:49",
      "content": "<p>Thank you so much for sharing it!</p>",
      "rawMarkdown": "Thank you so much for sharing it!",
      "votes": null
    },
    {
      "id": "2208931",
      "postDate": "04/04/2023 11:25:12",
      "content": "<p>Thank you so much!! <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a> <br>\nThis is very helpful!</p>",
      "rawMarkdown": "Thank you so much!! @abaojiang \nThis is very helpful!",
      "votes": null
    },
    {
      "id": "2209209",
      "postDate": "04/04/2023 14:42:42",
      "content": "<p><strong>(Update, 0404):</strong> After I train models using the new training set (<em>i.e.,</em> old training set with old test set), CV and LB both boost a little. The performance report is shown as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Training Data</th>\n<th>CV (GroupKFold with k=5)</th>\n<th>Holdout (Released Old Test Set)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Old training set only</td>\n<td>0.6914</td>\n<td>0.6911</td>\n<td>0.694</td>\n</tr>\n<tr>\n<td>Old training set with old test set</td>\n<td>0.6931 <strong>(+0.0017)</strong></td>\n<td>X (Already act as training set)</td>\n<td>0.695</td>\n</tr>\n</tbody>\n</table>\n<p>The experimental result shows that adding more training data is beneficial, but not significant. Thanks!</p>",
      "rawMarkdown": "**(Update, 0404):** After I train models using the new training set (*i.e.,* old training set with old test set), CV and LB both boost a little. The performance report is shown as follows:\n\n|Training Data | CV (GroupKFold with k=5) | Holdout (Released Old Test Set) | LB    |\n|------------------------| ------------------------ | ------------------------------- | ----- |\n|Old training set only| 0.6914                   | 0.6911                          | 0.694 |\n|Old training set with old test set| 0.6931 **(+0.0017)**                   | X (Already act as training set)   | 0.695 |\n\nThe experimental result shows that adding more training data is beneficial, but not significant. Thanks!",
      "votes": null
    },
    {
      "id": "2209225",
      "postDate": "04/04/2023 14:56:22",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I've learned a lot from you!</p>",
      "rawMarkdown": "Thanks @cdeotte, I've learned a lot from you!",
      "votes": null
    },
    {
      "id": "2209228",
      "postDate": "04/04/2023 14:57:06",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a> 😊</p>",
      "rawMarkdown": "Thanks @hanaori 😊",
      "votes": null
    },
    {
      "id": "2209230",
      "postDate": "04/04/2023 14:57:49",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a>, glad it helps!</p>",
      "rawMarkdown": "Thanks @simonveitner, glad it helps!",
      "votes": null
    },
    {
      "id": "2221704",
      "postDate": "04/14/2023 13:37:35",
      "content": "<p>Thank you so much for sharing this!</p>",
      "rawMarkdown": "Thank you so much for sharing this!",
      "votes": null
    },
    {
      "id": "2221749",
      "postDate": "04/14/2023 14:37:24",
      "content": "<p>Excellent work done <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a> </p>",
      "rawMarkdown": "Excellent work done @abaojiang",
      "votes": null
    },
    {
      "id": "2226654",
      "postDate": "04/19/2023 05:33:55",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a> </p>",
      "rawMarkdown": "Thanks @abaojiang",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2206676,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "04/02/2023 18:25:01",
      "content": "<p>thank you so much for publishing this!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2209230,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "04/04/2023 14:57:49",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a>, glad it helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2206884,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/03/2023 01:25:19",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2209225,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "04/04/2023 14:56:22",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, I've learned a lot from you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2208273,
      "author_name": "hanaori",
      "author_url": "",
      "post_date": "04/03/2023 23:55:49",
      "content": "<p>Thank you so much for sharing it!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2209228,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "04/04/2023 14:57:06",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a> 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2208931,
      "author_name": "shashwatraman",
      "author_url": "",
      "post_date": "04/04/2023 11:25:12",
      "content": "<p>Thank you so much!! <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a> <br>\nThis is very helpful!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2209209,
      "author_name": "abaojiang",
      "author_url": "",
      "post_date": "04/04/2023 14:42:42",
      "content": "<p><strong>(Update, 0404):</strong> After I train models using the new training set (<em>i.e.,</em> old training set with old test set), CV and LB both boost a little. The performance report is shown as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Training Data</th>\n<th>CV (GroupKFold with k=5)</th>\n<th>Holdout (Released Old Test Set)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Old training set only</td>\n<td>0.6914</td>\n<td>0.6911</td>\n<td>0.694</td>\n</tr>\n<tr>\n<td>Old training set with old test set</td>\n<td>0.6931 <strong>(+0.0017)</strong></td>\n<td>X (Already act as training set)</td>\n<td>0.695</td>\n</tr>\n</tbody>\n</table>\n<p>The experimental result shows that adding more training data is beneficial, but not significant. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2221704,
      "author_name": "linchia6",
      "author_url": "",
      "post_date": "04/14/2023 13:37:35",
      "content": "<p>Thank you so much for sharing this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2221749,
      "author_name": "mohimaakter",
      "author_url": "",
      "post_date": "04/14/2023 14:37:24",
      "content": "<p>Excellent work done <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2226654,
      "author_name": "zainalinasir",
      "author_url": "",
      "post_date": "04/19/2023 05:33:55",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2206499": "Hi everyone,\n\nAfter spending some time refactoring the local pipeline into readable notebook, I finally publish the [**training** counterpart](https://www.kaggle.com/code/abaojiang/lb-0-694-tconv-with-4-features-training-part) of the inference notebook released about 3 days ago.\n\nIn this kernel, I implement a complete pipeline, including data cleaning and processing, simple feature engineering, and model training with the specified CV scheme (i.e., `GroupKFold` with k=5). After the process is done, output objects (e.g., best model checkpoints) can be downloaded and used in the inference part.\n\nTo be concrete, the experiment setup can be summarized as follows:\n1. Only 4 features are taken into consideration, including `elapsed_time`, `event_name`, `name`, and `room_fqid`.\n2. The model architecture is mainly based on 1D-Conv and categorical embeddings (as illustrated in following figure).\n[![2023-04-02-11-22-49.png](https://i.postimg.cc/90YNxtLQ/2023-04-02-11-22-49.png)](https://postimg.cc/vxDzT9RJ)\n3. `BCEWithLogitsLoss` is selected to be the loss criterion.\n4. `Adam` and `CosineAnnealingWarmRestarts` are used.\n5. `GroupKFold` with `k=5` is adopted to be the CV scheme, and all models are trained with 100 epochs without early stopping.\n\nAlso, I think there's still a long way to go. In the future, I'm going to extend the experiments as follows:\n1. Add more features (*e.g.,* text information).\n2. Try different model architectures.\n    * I've tried a simple modification of the model architecture, [SAINT+](https://arxiv.org/abs/2010.12042), but still can't obtain satisfying performance.\n3. Add more training data (*i.e.,* released old test set).\n    * As you can see in the training notebook, I use the old version of training set to train models.\n4. Ensemble with tree-based models.\n5. Do some hyperparameter tuning near the end of the competition.\n\nHope this quick summarization can help clarify the core setup of the experiment. If there's any mistake, please feel free to correct me. I appreciate you taking the time to look over the work and discuss with me.",
    "2206676": "thank you so much for publishing this!",
    "2206884": "Thanks for sharing!",
    "2208273": "Thank you so much for sharing it!",
    "2208931": "Thank you so much!! @abaojiang \nThis is very helpful!",
    "2209209": "**(Update, 0404):** After I train models using the new training set (*i.e.,* old training set with old test set), CV and LB both boost a little. The performance report is shown as follows:\n\n|Training Data | CV (GroupKFold with k=5) | Holdout (Released Old Test Set) | LB    |\n|------------------------| ------------------------ | ------------------------------- | ----- |\n|Old training set only| 0.6914                   | 0.6911                          | 0.694 |\n|Old training set with old test set| 0.6931 **(+0.0017)**                   | X (Already act as training set)   | 0.695 |\n\nThe experimental result shows that adding more training data is beneficial, but not significant. Thanks!",
    "2209225": "Thanks @cdeotte, I've learned a lot from you!",
    "2209228": "Thanks @hanaori 😊",
    "2209230": "Thanks @simonveitner, glad it helps!",
    "2221704": "Thank you so much for sharing this!",
    "2221749": "Excellent work done @abaojiang",
    "2226654": "Thanks @abaojiang"
  },
  "source": "meta"
}