{
  "id": 552677,
  "title": "Child Mind Institute PIU 3rd Place Solution ",
  "url": "/competitions/child-mind-institute-problematic-internet-use/writeups/jobayer-hossain-child-mind-institute-piu-3rd-place",
  "author_name": "",
  "post_date": "2025-01-01T12:52:00.237Z",
  "votes": 10,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, I’d like to thank the organizers for hosting this competition and everyone for making it such a thrilling experience. Despite the challenges posed by unpredictability, it provided a valuable opportunity to learn how to build robust solutions for small, noisy datasets.<br>\nMy approach was straightforward, and I’m excited to share it with you.</p>\n<p><strong>Cross-Validation</strong><br>\nOne of my key focuses, like many others, was to establish a stable and reliable CV framework. I avoided using any fixed random seed throughout the process. It took me 100 repetitions of 5-fold stratified KFold to achieve stable results, and I used 20 repetitions during Optuna hyperparameter tuning.<br>\nTo optimize the final QWK threshold, I used the OOF predictions from all these repetitions.</p>\n<p><strong>Model</strong><br>\nI stuck to LightGBM for the entire competition. I did start working on a CatBoost solution at one point but lost the energy to take it further or combine the two.</p>\n<p><strong>Feature Engineering</strong><br>\n<em>Actigraphy Data</em>:<br>\n-Calculated the standard deviation for X, Y, Z, and AngleZ, and the mean for Elmo.<br>\n-Derived features representing the five longest streaks of inactivity and activity using Elmo.<br>\n-Binned the \"light\" column into categories ranging from twilight to direct sunlight and took the value counts for each category.<br>\n<em>Instrument Data</em>:<br>\nI started with features from the public notebook, checking each one to see if it actually contributed to the model. After that, I added a handful of custom features based on my own experimentation.</p>\n<p><strong>Data Augmentation</strong><br>\n<em>NaN Augmentation</em>:<br>\nInitially, I imputed NaNs randomly in columns that already had missing values. Eventually, I just imputed NaN on all columns with NaNs for 20% of the data and combined this augmented data with the original dataset.<br>\n<em>Gaussian Noise and Imputation</em>:<br>\nI applied simple imputation and added Gaussian noise to 20% of the data. This augmented data was then merged with the original dataset.</p>\n<p><strong>Post-Processing</strong><br>\nI used the 'PCIAT-PCIAT_Total' column for training. To finalize predictions, I applied the optimized threshold to calculate the sii for each of the 100*5 models and took the mode to generate the final predictions.</p>\n<p><strong>Results</strong><br>\nInitially, my CV-LB correlation started to break down after achieving a leaderboard score of 0.46. At that point, I decided to focus entirely on CV and improve it further. I’m pleased with the results of this phase, which led to consistent private leaderboard scores. Below are the highlights from my last 5 submissions during this phase:</p>\n<p>LB Score     PB Score        Repeats<br>\n  0.445           0.477             1<br>\n  0.461           0.482          100<br>\n  0.461           0.479          100<br>\n  0.466           0.478          100 (best submission selected)<br>\n  0.458           0.480          100<br>\nAll of these submissions had nearly identical CV performance:<br>\nValidation QWK: 0.454 - 0.456<br>\nOptimized QWK: ~0.470</p>\n<p>After this phase, I switched strategies by fixing the random seed and focusing on achieving higher LB scores with minimal changes. While this led to a slight improvement in CV—validation QWK around 0.460 and optimized QWK around 0.471—the LB scores remained the same, and PB scores worsened, averaging around 0.470 on the private leaderboard.</p>\n<p>That’s all, thanks for reading!</p>\n<p><a href=\"https://www.kaggle.com/code/jobayerhossain/child-mind-piu-3rd-place-solution\" target=\"_blank\">https://www.kaggle.com/code/jobayerhossain/child-mind-piu-3rd-place-solution</a></p>",
  "messages": [
    {
      "id": "3077400",
      "postDate": "12/20/2024 23:07:28",
      "content": "<p>First of all, I’d like to thank the organizers for hosting this competition and everyone for making it such a thrilling experience. Despite the challenges posed by unpredictability, it provided a valuable opportunity to learn how to build robust solutions for small, noisy datasets.<br>\nMy approach was straightforward, and I’m excited to share it with you.</p>\n<p><strong>Cross-Validation</strong><br>\nOne of my key focuses, like many others, was to establish a stable and reliable CV framework. I avoided using any fixed random seed throughout the process. It took me 100 repetitions of 5-fold stratified KFold to achieve stable results, and I used 20 repetitions during Optuna hyperparameter tuning.<br>\nTo optimize the final QWK threshold, I used the OOF predictions from all these repetitions.</p>\n<p><strong>Model</strong><br>\nI stuck to LightGBM for the entire competition. I did start working on a CatBoost solution at one point but lost the energy to take it further or combine the two.</p>\n<p><strong>Feature Engineering</strong><br>\n<em>Actigraphy Data</em>:<br>\n-Calculated the standard deviation for X, Y, Z, and AngleZ, and the mean for Elmo.<br>\n-Derived features representing the five longest streaks of inactivity and activity using Elmo.<br>\n-Binned the \"light\" column into categories ranging from twilight to direct sunlight and took the value counts for each category.<br>\n<em>Instrument Data</em>:<br>\nI started with features from the public notebook, checking each one to see if it actually contributed to the model. After that, I added a handful of custom features based on my own experimentation.</p>\n<p><strong>Data Augmentation</strong><br>\n<em>NaN Augmentation</em>:<br>\nInitially, I imputed NaNs randomly in columns that already had missing values. Eventually, I just imputed NaN on all columns with NaNs for 20% of the data and combined this augmented data with the original dataset.<br>\n<em>Gaussian Noise and Imputation</em>:<br>\nI applied simple imputation and added Gaussian noise to 20% of the data. This augmented data was then merged with the original dataset.</p>\n<p><strong>Post-Processing</strong><br>\nI used the 'PCIAT-PCIAT_Total' column for training. To finalize predictions, I applied the optimized threshold to calculate the sii for each of the 100*5 models and took the mode to generate the final predictions.</p>\n<p><strong>Results</strong><br>\nInitially, my CV-LB correlation started to break down after achieving a leaderboard score of 0.46. At that point, I decided to focus entirely on CV and improve it further. I’m pleased with the results of this phase, which led to consistent private leaderboard scores. Below are the highlights from my last 5 submissions during this phase:</p>\n<p>LB Score     PB Score        Repeats<br>\n  0.445           0.477             1<br>\n  0.461           0.482          100<br>\n  0.461           0.479          100<br>\n  0.466           0.478          100 (best submission selected)<br>\n  0.458           0.480          100<br>\nAll of these submissions had nearly identical CV performance:<br>\nValidation QWK: 0.454 - 0.456<br>\nOptimized QWK: ~0.470</p>\n<p>After this phase, I switched strategies by fixing the random seed and focusing on achieving higher LB scores with minimal changes. While this led to a slight improvement in CV—validation QWK around 0.460 and optimized QWK around 0.471—the LB scores remained the same, and PB scores worsened, averaging around 0.470 on the private leaderboard.</p>\n<p>That’s all, thanks for reading!</p>\n<p><a href=\"https://www.kaggle.com/code/jobayerhossain/child-mind-piu-3rd-place-solution\" target=\"_blank\">https://www.kaggle.com/code/jobayerhossain/child-mind-piu-3rd-place-solution</a></p>",
      "rawMarkdown": "First of all, I’d like to thank the organizers for hosting this competition and everyone for making it such a thrilling experience. Despite the challenges posed by unpredictability, it provided a valuable opportunity to learn how to build robust solutions for small, noisy datasets.\nMy approach was straightforward, and I’m excited to share it with you.\n\n**Cross-Validation**\nOne of my key focuses, like many others, was to establish a stable and reliable CV framework. I avoided using any fixed random seed throughout the process. It took me 100 repetitions of 5-fold stratified KFold to achieve stable results, and I used 20 repetitions during Optuna hyperparameter tuning.\nTo optimize the final QWK threshold, I used the OOF predictions from all these repetitions.\n\n**Model**\nI stuck to LightGBM for the entire competition. I did start working on a CatBoost solution at one point but lost the energy to take it further or combine the two.\n\n**Feature Engineering**\n*Actigraphy Data*:\n-Calculated the standard deviation for X, Y, Z, and AngleZ, and the mean for Elmo.\n-Derived features representing the five longest streaks of inactivity and activity using Elmo.\n-Binned the \"light\" column into categories ranging from twilight to direct sunlight and took the value counts for each category.\n*Instrument Data*:\nI started with features from the public notebook, checking each one to see if it actually contributed to the model. After that, I added a handful of custom features based on my own experimentation.\n\n**Data Augmentation**\n*NaN Augmentation*:\nInitially, I imputed NaNs randomly in columns that already had missing values. Eventually, I just imputed NaN on all columns with NaNs for 20% of the data and combined this augmented data with the original dataset.\n*Gaussian Noise and Imputation*:\nI applied simple imputation and added Gaussian noise to 20% of the data. This augmented data was then merged with the original dataset.\n\n**Post-Processing**\nI used the 'PCIAT-PCIAT_Total' column for training. To finalize predictions, I applied the optimized threshold to calculate the sii for each of the 100*5 models and took the mode to generate the final predictions.\n\n**Results**\nInitially, my CV-LB correlation started to break down after achieving a leaderboard score of 0.46. At that point, I decided to focus entirely on CV and improve it further. I’m pleased with the results of this phase, which led to consistent private leaderboard scores. Below are the highlights from my last 5 submissions during this phase:\n\nLB Score     PB Score\t    Repeats\n  0.445\t       0.477\t         1\n  0.461\t       0.482\t      100\n  0.461\t       0.479\t      100\n  0.466\t       0.478\t      100 (best submission selected)\n  0.458\t       0.480\t      100\nAll of these submissions had nearly identical CV performance:\nValidation QWK: 0.454 - 0.456\nOptimized QWK: ~0.470\n\nAfter this phase, I switched strategies by fixing the random seed and focusing on achieving higher LB scores with minimal changes. While this led to a slight improvement in CV—validation QWK around 0.460 and optimized QWK around 0.471—the LB scores remained the same, and PB scores worsened, averaging around 0.470 on the private leaderboard.\n\nThat’s all, thanks for reading!\n\nhttps://www.kaggle.com/code/jobayerhossain/child-mind-piu-3rd-place-solution",
      "votes": null
    },
    {
      "id": "3077543",
      "postDate": "12/21/2024 05:22:07",
      "content": "<p>Congrats! Another simple original solution. Nice to see an example of Gaussian noise work with LGBM. If it is applied to a small subset of filled NaNs, it seems pretty reasonable. I usually experiment with including a Gaussian noise layer as a regularization to the NN architecture.</p>",
      "rawMarkdown": "Congrats! Another simple original solution. Nice to see an example of Gaussian noise work with LGBM. If it is applied to a small subset of filled NaNs, it seems pretty reasonable. I usually experiment with including a Gaussian noise layer as a regularization to the NN architecture.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3077543,
      "author_name": "yekenot",
      "author_url": "",
      "post_date": "12/21/2024 05:22:07",
      "content": "<p>Congrats! Another simple original solution. Nice to see an example of Gaussian noise work with LGBM. If it is applied to a small subset of filled NaNs, it seems pretty reasonable. I usually experiment with including a Gaussian noise layer as a regularization to the NN architecture.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3077400": "First of all, I’d like to thank the organizers for hosting this competition and everyone for making it such a thrilling experience. Despite the challenges posed by unpredictability, it provided a valuable opportunity to learn how to build robust solutions for small, noisy datasets.\nMy approach was straightforward, and I’m excited to share it with you.\n\n**Cross-Validation**\nOne of my key focuses, like many others, was to establish a stable and reliable CV framework. I avoided using any fixed random seed throughout the process. It took me 100 repetitions of 5-fold stratified KFold to achieve stable results, and I used 20 repetitions during Optuna hyperparameter tuning.\nTo optimize the final QWK threshold, I used the OOF predictions from all these repetitions.\n\n**Model**\nI stuck to LightGBM for the entire competition. I did start working on a CatBoost solution at one point but lost the energy to take it further or combine the two.\n\n**Feature Engineering**\n*Actigraphy Data*:\n-Calculated the standard deviation for X, Y, Z, and AngleZ, and the mean for Elmo.\n-Derived features representing the five longest streaks of inactivity and activity using Elmo.\n-Binned the \"light\" column into categories ranging from twilight to direct sunlight and took the value counts for each category.\n*Instrument Data*:\nI started with features from the public notebook, checking each one to see if it actually contributed to the model. After that, I added a handful of custom features based on my own experimentation.\n\n**Data Augmentation**\n*NaN Augmentation*:\nInitially, I imputed NaNs randomly in columns that already had missing values. Eventually, I just imputed NaN on all columns with NaNs for 20% of the data and combined this augmented data with the original dataset.\n*Gaussian Noise and Imputation*:\nI applied simple imputation and added Gaussian noise to 20% of the data. This augmented data was then merged with the original dataset.\n\n**Post-Processing**\nI used the 'PCIAT-PCIAT_Total' column for training. To finalize predictions, I applied the optimized threshold to calculate the sii for each of the 100*5 models and took the mode to generate the final predictions.\n\n**Results**\nInitially, my CV-LB correlation started to break down after achieving a leaderboard score of 0.46. At that point, I decided to focus entirely on CV and improve it further. I’m pleased with the results of this phase, which led to consistent private leaderboard scores. Below are the highlights from my last 5 submissions during this phase:\n\nLB Score     PB Score\t    Repeats\n  0.445\t       0.477\t         1\n  0.461\t       0.482\t      100\n  0.461\t       0.479\t      100\n  0.466\t       0.478\t      100 (best submission selected)\n  0.458\t       0.480\t      100\nAll of these submissions had nearly identical CV performance:\nValidation QWK: 0.454 - 0.456\nOptimized QWK: ~0.470\n\nAfter this phase, I switched strategies by fixing the random seed and focusing on achieving higher LB scores with minimal changes. While this led to a slight improvement in CV—validation QWK around 0.460 and optimized QWK around 0.471—the LB scores remained the same, and PB scores worsened, averaging around 0.470 on the private leaderboard.\n\nThat’s all, thanks for reading!\n\nhttps://www.kaggle.com/code/jobayerhossain/child-mind-piu-3rd-place-solution",
    "3077543": "Congrats! Another simple original solution. Nice to see an example of Gaussian noise work with LGBM. If it is applied to a small subset of filled NaNs, it seems pretty reasonable. I usually experiment with including a Gaussian noise layer as a regularization to the NN architecture."
  },
  "source": "meta"
}