{
  "id": 552492,
  "title": "Rank 106 approach - Simple feature sets and CV-LB stability focus",
  "url": "/competitions/child-mind-institute-problematic-internet-use/writeups/ravi-ramakrishnan-rank-106-approach-simple-feature",
  "author_name": "",
  "post_date": "2024-12-30T08:24:00.590Z",
  "votes": 30,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>Firstly thanks to Kaggle and CMI for the competition. I think this was one of the many churn-filled competitions this year, after Home Credit, ISIC, BirdClef and AES and I was in all of them (and shook down every time)! I participated in this challenge with the lessons learnt from all of them and tried to prevent a downward movement in the churn here and was successful at it!</p>\n<h1>Approach summary</h1>\n<ul>\n<li>I resorted to a very simple pipeline, comprising of single models that offered me at least a level of CV-LB relation and stability</li>\n<li>I used 3-5 feature sets and submitted a simple average blend of constituent models from the above step, choosing the best models that offered me a CV-LB stability </li>\n<li>I decided not to opt for any complex models like NNs, TabNet, AutoEncoder and resorted to simple boosted tree models only without early stopping</li>\n</ul>\n<h1>Training code</h1>\n<ul>\n<li>CMI2024|Final|Candidate1 - <a href=\"https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate1\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate1</a></li>\n<li>CMI2024|Final|Candidate2 - <a href=\"https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate2\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate2</a></li>\n</ul>\n<h1>Feature engineering</h1>\n<ul>\n<li>I tried a lot of features using the actigraphy data to no impact/ marginal impact. I restricted myself to public features at the end (descriptive statistics with a small addition/ removal of features in my final submissions)</li>\n<li>I cleaned up some noisy values in the dataset otherwise, based on domain knowledge (examples include blood-pressure values, BMI values, etc.)</li>\n<li>I did not impute any targets from the unknown sii section of the data and left this as-is (this was excluded from my model training)</li>\n<li>I resorted to 3 feature sets in one submission and 5 in another and landed up with the same private LB score across both of them</li>\n</ul>\n<h2>Target choices</h2>\n<p>I resorted to 2 target choices here</p>\n<ul>\n<li>Direct prediction of sii target</li>\n<li>Predicting sii using PCIAT-PCIAT1-20 and then adding up the predictions followed up with rule-based allocation as below-<br>\na. Predicted PCIAT-Total between 0 and 30 -&gt; 0<br>\nb. Predicted PCIAT-Total between 31 and 49 -&gt; 1<br>\nc. Predicted PCIAT-Total between 50 and 79 -&gt; 2<br>\nd. Predicted PCIAT-Total between 80 and 100 -&gt; 3</li>\n</ul>\n<p>Using 2 targets like this offered me better stability on the ensemble, between the CV and public LB scores</p>\n<h1>Model training</h1>\n<h2>Offline model</h2>\n<ul>\n<li>Cross-validation - 5 fold stratified by sii values </li>\n<li>A simple LightGBM with public notebook parameters worked well for me across the competition. I did not tune anything here as any attempt to tune my models led me to an undesirable CV-LB relation. A lot of my tuned models have also tanked on the private LB and I am happy I did not rely on any form of tuning here!</li>\n<li>I did a target tuning just like a lot of public kernels, but <strong>used the training data for tuning and not the OOF data</strong>. Tuning with a larger dataset across folds led me to a better CV-LB stability and relatively lesser impact on the public leaderboard on changing random seeds. </li>\n<li>I designed a simple class inheriting from a LightGBM regressor and tuned my thresholds as below-<br>\na. Train a regressor with the fold-level training data<br>\nb. Use the training data predictions and tune the thresholds using scipy.optimize.minimize like the ones in public kernels <br>\nc. Store the thresholds for later use</li>\n</ul>\n<h2>Full refit model</h2>\n<ul>\n<li>Once my offline CV scheme was ready, I refitted the model on the full training data (excluding the null target columns) and tuned the thresholds similar to the offline model process (using the full train set)</li>\n<li>I averaged the model across a large number of random states (typically 100+) and engendered stability with <strong>multistarts</strong></li>\n<li>I submitted this averaged model on the full training data to the leaderboard and varied the random states in the full refit (and ultimately the tuning process on full-fit) to ascertain my CV-LB stability. The chosen feature sets and submissions proved to be relatively least unstable and were included as candidates </li>\n</ul>\n<h2>Feature components and individual submission results</h2>\n<table>\n<thead>\n<tr>\n<th>Feature set</th>\n<th>CV</th>\n<th>Public LB score</th>\n<th>Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Set 1 - 163 features</td>\n<td>0.463429</td>\n<td>0.471</td>\n<td>PCIAT_*</td>\n</tr>\n<tr>\n<td>Set 2 - 163 features</td>\n<td>0.466009</td>\n<td>0.470</td>\n<td>PCIAT_*</td>\n</tr>\n<tr>\n<td>Set 3 - 163 features</td>\n<td>0.466455</td>\n<td>0.466</td>\n<td>PCIAT_*</td>\n</tr>\n<tr>\n<td>Set 4 - 163 features</td>\n<td>0.468296</td>\n<td>0.466</td>\n<td>sii</td>\n</tr>\n<tr>\n<td>Set 5 - 141 features</td>\n<td>0.463159</td>\n<td>0.470</td>\n<td>PCIAT_*</td>\n</tr>\n</tbody>\n</table>\n<p><br><strong>Overall result - Public/ Private LB -&gt; 0.471 / 0.454</strong> <br>\n<br></p>\n<table>\n<thead>\n<tr>\n<th>Feature set</th>\n<th>CV</th>\n<th>Public LB score</th>\n<th>Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Set 1 - 163 features</td>\n<td>0.466184</td>\n<td>0.472</td>\n<td>PCIAT_*</td>\n</tr>\n<tr>\n<td>Set 2 - 163 features</td>\n<td>0.468296</td>\n<td>0.466</td>\n<td>sii</td>\n</tr>\n<tr>\n<td>Set 3 - 141 features</td>\n<td>0.463159</td>\n<td>0.470</td>\n<td>PCIAT_*</td>\n</tr>\n</tbody>\n</table>\n<p><br><strong>Overall result - Public/ Private LB -&gt; 0.475 / 0.454</strong> </p>\n<h1>What I could have done better here</h1>\n<ul>\n<li>Better choice of final submissions </li>\n<li>Better choice of candidate feature sets</li>\n<li>One of my feature sets with a lower public LB score scored extremely well on the private set. I could have chosen it perhaps!</li>\n</ul>\n<h1>What I gained here</h1>\n<ul>\n<li>Training with threshold tuning on the training data rather than the OOF data was a good step for me and worked for me here</li>\n<li>My stability across the leaderboards was also a gain, after my falls in yester shakeup driven competitions</li>\n<li>I am happy I did not use any complex models and algorithms here and ended up with a relatively stable result. I think this itself feels like success given the churn in the competition</li>\n</ul>\n<h1>Concluding remarks</h1>\n<p>Good luck for all your competitions going ahead! All the best and happy learning!</p>\n<p>Best regards,<br>\nRavi Ramakrishnan</p>",
  "messages": [
    {
      "id": "3076459",
      "postDate": "12/20/2024 01:21:09",
      "content": "<p>Hello all,</p>\n<p>Firstly thanks to Kaggle and CMI for the competition. I think this was one of the many churn-filled competitions this year, after Home Credit, ISIC, BirdClef and AES and I was in all of them (and shook down every time)! I participated in this challenge with the lessons learnt from all of them and tried to prevent a downward movement in the churn here and was successful at it!</p>\n<h1>Approach summary</h1>\n<ul>\n<li>I resorted to a very simple pipeline, comprising of single models that offered me at least a level of CV-LB relation and stability</li>\n<li>I used 3-5 feature sets and submitted a simple average blend of constituent models from the above step, choosing the best models that offered me a CV-LB stability </li>\n<li>I decided not to opt for any complex models like NNs, TabNet, AutoEncoder and resorted to simple boosted tree models only without early stopping</li>\n</ul>\n<h1>Training code</h1>\n<ul>\n<li>CMI2024|Final|Candidate1 - <a href=\"https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate1\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate1</a></li>\n<li>CMI2024|Final|Candidate2 - <a href=\"https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate2\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate2</a></li>\n</ul>\n<h1>Feature engineering</h1>\n<ul>\n<li>I tried a lot of features using the actigraphy data to no impact/ marginal impact. I restricted myself to public features at the end (descriptive statistics with a small addition/ removal of features in my final submissions)</li>\n<li>I cleaned up some noisy values in the dataset otherwise, based on domain knowledge (examples include blood-pressure values, BMI values, etc.)</li>\n<li>I did not impute any targets from the unknown sii section of the data and left this as-is (this was excluded from my model training)</li>\n<li>I resorted to 3 feature sets in one submission and 5 in another and landed up with the same private LB score across both of them</li>\n</ul>\n<h2>Target choices</h2>\n<p>I resorted to 2 target choices here</p>\n<ul>\n<li>Direct prediction of sii target</li>\n<li>Predicting sii using PCIAT-PCIAT1-20 and then adding up the predictions followed up with rule-based allocation as below-<br>\na. Predicted PCIAT-Total between 0 and 30 -&gt; 0<br>\nb. Predicted PCIAT-Total between 31 and 49 -&gt; 1<br>\nc. Predicted PCIAT-Total between 50 and 79 -&gt; 2<br>\nd. Predicted PCIAT-Total between 80 and 100 -&gt; 3</li>\n</ul>\n<p>Using 2 targets like this offered me better stability on the ensemble, between the CV and public LB scores</p>\n<h1>Model training</h1>\n<h2>Offline model</h2>\n<ul>\n<li>Cross-validation - 5 fold stratified by sii values </li>\n<li>A simple LightGBM with public notebook parameters worked well for me across the competition. I did not tune anything here as any attempt to tune my models led me to an undesirable CV-LB relation. A lot of my tuned models have also tanked on the private LB and I am happy I did not rely on any form of tuning here!</li>\n<li>I did a target tuning just like a lot of public kernels, but <strong>used the training data for tuning and not the OOF data</strong>. Tuning with a larger dataset across folds led me to a better CV-LB stability and relatively lesser impact on the public leaderboard on changing random seeds. </li>\n<li>I designed a simple class inheriting from a LightGBM regressor and tuned my thresholds as below-<br>\na. Train a regressor with the fold-level training data<br>\nb. Use the training data predictions and tune the thresholds using scipy.optimize.minimize like the ones in public kernels <br>\nc. Store the thresholds for later use</li>\n</ul>\n<h2>Full refit model</h2>\n<ul>\n<li>Once my offline CV scheme was ready, I refitted the model on the full training data (excluding the null target columns) and tuned the thresholds similar to the offline model process (using the full train set)</li>\n<li>I averaged the model across a large number of random states (typically 100+) and engendered stability with <strong>multistarts</strong></li>\n<li>I submitted this averaged model on the full training data to the leaderboard and varied the random states in the full refit (and ultimately the tuning process on full-fit) to ascertain my CV-LB stability. The chosen feature sets and submissions proved to be relatively least unstable and were included as candidates </li>\n</ul>\n<h2>Feature components and individual submission results</h2>\n<table>\n<thead>\n<tr>\n<th>Feature set</th>\n<th>CV</th>\n<th>Public LB score</th>\n<th>Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Set 1 - 163 features</td>\n<td>0.463429</td>\n<td>0.471</td>\n<td>PCIAT_*</td>\n</tr>\n<tr>\n<td>Set 2 - 163 features</td>\n<td>0.466009</td>\n<td>0.470</td>\n<td>PCIAT_*</td>\n</tr>\n<tr>\n<td>Set 3 - 163 features</td>\n<td>0.466455</td>\n<td>0.466</td>\n<td>PCIAT_*</td>\n</tr>\n<tr>\n<td>Set 4 - 163 features</td>\n<td>0.468296</td>\n<td>0.466</td>\n<td>sii</td>\n</tr>\n<tr>\n<td>Set 5 - 141 features</td>\n<td>0.463159</td>\n<td>0.470</td>\n<td>PCIAT_*</td>\n</tr>\n</tbody>\n</table>\n<p><br><strong>Overall result - Public/ Private LB -&gt; 0.471 / 0.454</strong> <br>\n<br></p>\n<table>\n<thead>\n<tr>\n<th>Feature set</th>\n<th>CV</th>\n<th>Public LB score</th>\n<th>Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Set 1 - 163 features</td>\n<td>0.466184</td>\n<td>0.472</td>\n<td>PCIAT_*</td>\n</tr>\n<tr>\n<td>Set 2 - 163 features</td>\n<td>0.468296</td>\n<td>0.466</td>\n<td>sii</td>\n</tr>\n<tr>\n<td>Set 3 - 141 features</td>\n<td>0.463159</td>\n<td>0.470</td>\n<td>PCIAT_*</td>\n</tr>\n</tbody>\n</table>\n<p><br><strong>Overall result - Public/ Private LB -&gt; 0.475 / 0.454</strong> </p>\n<h1>What I could have done better here</h1>\n<ul>\n<li>Better choice of final submissions </li>\n<li>Better choice of candidate feature sets</li>\n<li>One of my feature sets with a lower public LB score scored extremely well on the private set. I could have chosen it perhaps!</li>\n</ul>\n<h1>What I gained here</h1>\n<ul>\n<li>Training with threshold tuning on the training data rather than the OOF data was a good step for me and worked for me here</li>\n<li>My stability across the leaderboards was also a gain, after my falls in yester shakeup driven competitions</li>\n<li>I am happy I did not use any complex models and algorithms here and ended up with a relatively stable result. I think this itself feels like success given the churn in the competition</li>\n</ul>\n<h1>Concluding remarks</h1>\n<p>Good luck for all your competitions going ahead! All the best and happy learning!</p>\n<p>Best regards,<br>\nRavi Ramakrishnan</p>",
      "rawMarkdown": "Hello all,\n\nFirstly thanks to Kaggle and CMI for the competition. I think this was one of the many churn-filled competitions this year, after Home Credit, ISIC, BirdClef and AES and I was in all of them (and shook down every time)! I participated in this challenge with the lessons learnt from all of them and tried to prevent a downward movement in the churn here and was successful at it!\n\n# Approach summary\n- I resorted to a very simple pipeline, comprising of single models that offered me at least a level of CV-LB relation and stability\n- I used 3-5 feature sets and submitted a simple average blend of constituent models from the above step, choosing the best models that offered me a CV-LB stability \n- I decided not to opt for any complex models like NNs, TabNet, AutoEncoder and resorted to simple boosted tree models only without early stopping\n\n# Training code\n- CMI2024|Final|Candidate1 - https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate1\n- CMI2024|Final|Candidate2 - https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate2\n\n# Feature engineering\n- I tried a lot of features using the actigraphy data to no impact/ marginal impact. I restricted myself to public features at the end (descriptive statistics with a small addition/ removal of features in my final submissions)\n- I cleaned up some noisy values in the dataset otherwise, based on domain knowledge (examples include blood-pressure values, BMI values, etc.)\n- I did not impute any targets from the unknown sii section of the data and left this as-is (this was excluded from my model training)\n- I resorted to 3 feature sets in one submission and 5 in another and landed up with the same private LB score across both of them\n\n## Target choices \nI resorted to 2 target choices here\n- Direct prediction of sii target\n- Predicting sii using PCIAT-PCIAT1-20 and then adding up the predictions followed up with rule-based allocation as below-\na. Predicted PCIAT-Total between 0 and 30 -> 0\nb. Predicted PCIAT-Total between 31 and 49 -> 1\nc. Predicted PCIAT-Total between 50 and 79 -> 2\nd. Predicted PCIAT-Total between 80 and 100 -> 3\n\nUsing 2 targets like this offered me better stability on the ensemble, between the CV and public LB scores\n\n# Model training  \n\n## Offline model\n- Cross-validation - 5 fold stratified by sii values \n- A simple LightGBM with public notebook parameters worked well for me across the competition. I did not tune anything here as any attempt to tune my models led me to an undesirable CV-LB relation. A lot of my tuned models have also tanked on the private LB and I am happy I did not rely on any form of tuning here!\n- I did a target tuning just like a lot of public kernels, but **used the training data for tuning and not the OOF data**. Tuning with a larger dataset across folds led me to a better CV-LB stability and relatively lesser impact on the public leaderboard on changing random seeds. \n- I designed a simple class inheriting from a LightGBM regressor and tuned my thresholds as below-\na. Train a regressor with the fold-level training data\nb. Use the training data predictions and tune the thresholds using scipy.optimize.minimize like the ones in public kernels \nc. Store the thresholds for later use\n\n## Full refit model\n- Once my offline CV scheme was ready, I refitted the model on the full training data (excluding the null target columns) and tuned the thresholds similar to the offline model process (using the full train set)\n- I averaged the model across a large number of random states (typically 100+) and engendered stability with **multistarts**\n- I submitted this averaged model on the full training data to the leaderboard and varied the random states in the full refit (and ultimately the tuning process on full-fit) to ascertain my CV-LB stability. The chosen feature sets and submissions proved to be relatively least unstable and were included as candidates \n\n## Feature components and individual submission results \n\n| Feature set | CV  | Public LB score | Target | \n| --- | --- | ----------- | ---- | \n| Set 1 - 163 features | 0.463429 | 0.471| PCIAT_*|\n| Set 2 - 163 features | 0.466009 | 0.470| PCIAT_*|\n| Set 3 - 163 features | 0.466455 | 0.466| PCIAT_*|\n| Set 4 - 163 features | 0.468296 | 0.466| sii |\n| Set 5 - 141 features | 0.463159 | 0.470 | PCIAT_*|\n\n<br>**Overall result - Public/ Private LB -> 0.471 / 0.454** \n<br>\n\n| Feature set | CV  | Public LB score | Target | \n| --- | --- | ----------- | ---- | \n| Set 1 - 163 features  | 0.466184 | 0.472| PCIAT_*|\n| Set 2 - 163 features | 0.468296 | 0.466| sii |\n| Set 3 - 141 features  | 0.463159 | 0.470 | PCIAT_*|\n\n<br>**Overall result - Public/ Private LB -> 0.475 / 0.454** \n\n# What I could have done better here\n- Better choice of final submissions \n- Better choice of candidate feature sets\n- One of my feature sets with a lower public LB score scored extremely well on the private set. I could have chosen it perhaps!\n\n# What I gained here\n- Training with threshold tuning on the training data rather than the OOF data was a good step for me and worked for me here\n- My stability across the leaderboards was also a gain, after my falls in yester shakeup driven competitions\n- I am happy I did not use any complex models and algorithms here and ended up with a relatively stable result. I think this itself feels like success given the churn in the competition\n\n# Concluding remarks\nGood luck for all your competitions going ahead! All the best and happy learning!\n\nBest regards,\nRavi Ramakrishnan",
      "votes": null
    },
    {
      "id": "3076463",
      "postDate": "12/20/2024 01:26:07",
      "content": "<p>Thanks for your sharing, I also tried to use PCIAT_Total to predict, but I failed. I was confused about the threshold selection of PCIAT_Total and sii, because I thought that the data of test and train might be quite different, so it was difficult to choose a suitable threshold</p>",
      "rawMarkdown": "Thanks for your sharing, I also tried to use PCIAT_Total to predict, but I failed. I was confused about the threshold selection of PCIAT_Total and sii, because I thought that the data of test and train might be quite different, so it was difficult to choose a suitable threshold",
      "votes": null
    },
    {
      "id": "3076464",
      "postDate": "12/20/2024 01:26:44",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>, Good Solution thanks for sharing</p>",
      "rawMarkdown": "Congrats @ravi20076, Good Solution thanks for sharing",
      "votes": null
    },
    {
      "id": "3076467",
      "postDate": "12/20/2024 01:30:49",
      "content": "<p><a href=\"https://www.kaggle.com/ruichardliu\" target=\"_blank\">@ruichardliu</a> I had this confusion and hence resorted to a rule based allocation, cutting the data into bins with cutoffs [0, 30, 50, 80] as explained above.<br>\nCongrats for your gold medal <a href=\"https://www.kaggle.com/ruichardliu\" target=\"_blank\">@ruichardliu</a> </p>",
      "rawMarkdown": "ruichardliu I had this confusion and hence resorted to a rule based allocation, cutting the data into bins with cutoffs [0, 30, 50, 80] as explained above.\nCongrats for your gold medal @ruichardliu",
      "votes": null
    },
    {
      "id": "3076643",
      "postDate": "12/20/2024 06:23:56",
      "content": "<p>Thanks Ravi, I have learnt a lot from you during this competition. Just a quick question, what did you change between these 3 sets? <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1198628%2F260000b9e1ce6423ae479f23476bca91%2FScreenshot%202024-12-20%20at%2007.19.23.png?generation=1734675809230160&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks Ravi, I have learnt a lot from you during this competition. Just a quick question, what did you change between these 3 sets? ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1198628%2F260000b9e1ce6423ae479f23476bca91%2FScreenshot%202024-12-20%20at%2007.19.23.png?generation=1734675809230160&alt=media)",
      "votes": null
    },
    {
      "id": "3076683",
      "postDate": "12/20/2024 07:25:50",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/diegoiglesias\" target=\"_blank\">@diegoiglesias</a> <br>\nI have shared my training code above, kindly peruse this to know more.<br>\nAll the best!</p>",
      "rawMarkdown": "Hello @diegoiglesias \nI have shared my training code above, kindly peruse this to know more.\nAll the best!",
      "votes": null
    },
    {
      "id": "3076711",
      "postDate": "12/20/2024 07:47:32",
      "content": "<p>Thanks Ravi, I have learnt a lot from you during this competition. !</p>",
      "rawMarkdown": "Thanks Ravi, I have learnt a lot from you during this competition. !",
      "votes": null
    },
    {
      "id": "3076933",
      "postDate": "12/20/2024 12:06:26",
      "content": "<p><a href=\"https://www.kaggle.com/mrsimple07\" target=\"_blank\">@mrsimple07</a> did you participate in the competition?</p>",
      "rawMarkdown": "mrsimple07 did you participate in the competition?",
      "votes": null
    },
    {
      "id": "3077239",
      "postDate": "12/20/2024 17:43:06",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> for this position! It was one of the most difficult competitions to stand in the LB</p>",
      "rawMarkdown": "Congrats @ravi20076 for this position! It was one of the most difficult competitions to stand in the LB",
      "votes": null
    },
    {
      "id": "3077276",
      "postDate": "12/20/2024 18:30:24",
      "content": "<p>Lol <br>\nGetting a medal here feels like success <a href=\"https://www.kaggle.com/octaviograu\" target=\"_blank\">@octaviograu</a> <br>\nAll the best my friend </p>",
      "rawMarkdown": "Lol \nGetting a medal here feels like success @octaviograu \nAll the best my friend",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3076463,
      "author_name": "ruichardliu",
      "author_url": "",
      "post_date": "12/20/2024 01:26:07",
      "content": "<p>Thanks for your sharing, I also tried to use PCIAT_Total to predict, but I failed. I was confused about the threshold selection of PCIAT_Total and sii, because I thought that the data of test and train might be quite different, so it was difficult to choose a suitable threshold</p>",
      "votes": null,
      "replies": [
        {
          "id": 3076467,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "12/20/2024 01:30:49",
          "content": "<p><a href=\"https://www.kaggle.com/ruichardliu\" target=\"_blank\">@ruichardliu</a> I had this confusion and hence resorted to a rule based allocation, cutting the data into bins with cutoffs [0, 30, 50, 80] as explained above.<br>\nCongrats for your gold medal <a href=\"https://www.kaggle.com/ruichardliu\" target=\"_blank\">@ruichardliu</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3076464,
      "author_name": "abdmental01",
      "author_url": "",
      "post_date": "12/20/2024 01:26:44",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>, Good Solution thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3076643,
      "author_name": "diegoiglesias",
      "author_url": "",
      "post_date": "12/20/2024 06:23:56",
      "content": "<p>Thanks Ravi, I have learnt a lot from you during this competition. Just a quick question, what did you change between these 3 sets? <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1198628%2F260000b9e1ce6423ae479f23476bca91%2FScreenshot%202024-12-20%20at%2007.19.23.png?generation=1734675809230160&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3076683,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "12/20/2024 07:25:50",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/diegoiglesias\" target=\"_blank\">@diegoiglesias</a> <br>\nI have shared my training code above, kindly peruse this to know more.<br>\nAll the best!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3076711,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "12/20/2024 07:47:32",
      "content": "<p>Thanks Ravi, I have learnt a lot from you during this competition. !</p>",
      "votes": null,
      "replies": [
        {
          "id": 3076933,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "12/20/2024 12:06:26",
          "content": "<p><a href=\"https://www.kaggle.com/mrsimple07\" target=\"_blank\">@mrsimple07</a> did you participate in the competition?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3077239,
      "author_name": "octaviograu",
      "author_url": "",
      "post_date": "12/20/2024 17:43:06",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> for this position! It was one of the most difficult competitions to stand in the LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 3077276,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "12/20/2024 18:30:24",
          "content": "<p>Lol <br>\nGetting a medal here feels like success <a href=\"https://www.kaggle.com/octaviograu\" target=\"_blank\">@octaviograu</a> <br>\nAll the best my friend </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3076459": "Hello all,\n\nFirstly thanks to Kaggle and CMI for the competition. I think this was one of the many churn-filled competitions this year, after Home Credit, ISIC, BirdClef and AES and I was in all of them (and shook down every time)! I participated in this challenge with the lessons learnt from all of them and tried to prevent a downward movement in the churn here and was successful at it!\n\n# Approach summary\n- I resorted to a very simple pipeline, comprising of single models that offered me at least a level of CV-LB relation and stability\n- I used 3-5 feature sets and submitted a simple average blend of constituent models from the above step, choosing the best models that offered me a CV-LB stability \n- I decided not to opt for any complex models like NNs, TabNet, AutoEncoder and resorted to simple boosted tree models only without early stopping\n\n# Training code\n- CMI2024|Final|Candidate1 - https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate1\n- CMI2024|Final|Candidate2 - https://www.kaggle.com/code/ravi20076/cmi2024-final-candidate2\n\n# Feature engineering\n- I tried a lot of features using the actigraphy data to no impact/ marginal impact. I restricted myself to public features at the end (descriptive statistics with a small addition/ removal of features in my final submissions)\n- I cleaned up some noisy values in the dataset otherwise, based on domain knowledge (examples include blood-pressure values, BMI values, etc.)\n- I did not impute any targets from the unknown sii section of the data and left this as-is (this was excluded from my model training)\n- I resorted to 3 feature sets in one submission and 5 in another and landed up with the same private LB score across both of them\n\n## Target choices \nI resorted to 2 target choices here\n- Direct prediction of sii target\n- Predicting sii using PCIAT-PCIAT1-20 and then adding up the predictions followed up with rule-based allocation as below-\na. Predicted PCIAT-Total between 0 and 30 -> 0\nb. Predicted PCIAT-Total between 31 and 49 -> 1\nc. Predicted PCIAT-Total between 50 and 79 -> 2\nd. Predicted PCIAT-Total between 80 and 100 -> 3\n\nUsing 2 targets like this offered me better stability on the ensemble, between the CV and public LB scores\n\n# Model training  \n\n## Offline model\n- Cross-validation - 5 fold stratified by sii values \n- A simple LightGBM with public notebook parameters worked well for me across the competition. I did not tune anything here as any attempt to tune my models led me to an undesirable CV-LB relation. A lot of my tuned models have also tanked on the private LB and I am happy I did not rely on any form of tuning here!\n- I did a target tuning just like a lot of public kernels, but **used the training data for tuning and not the OOF data**. Tuning with a larger dataset across folds led me to a better CV-LB stability and relatively lesser impact on the public leaderboard on changing random seeds. \n- I designed a simple class inheriting from a LightGBM regressor and tuned my thresholds as below-\na. Train a regressor with the fold-level training data\nb. Use the training data predictions and tune the thresholds using scipy.optimize.minimize like the ones in public kernels \nc. Store the thresholds for later use\n\n## Full refit model\n- Once my offline CV scheme was ready, I refitted the model on the full training data (excluding the null target columns) and tuned the thresholds similar to the offline model process (using the full train set)\n- I averaged the model across a large number of random states (typically 100+) and engendered stability with **multistarts**\n- I submitted this averaged model on the full training data to the leaderboard and varied the random states in the full refit (and ultimately the tuning process on full-fit) to ascertain my CV-LB stability. The chosen feature sets and submissions proved to be relatively least unstable and were included as candidates \n\n## Feature components and individual submission results \n\n| Feature set | CV  | Public LB score | Target | \n| --- | --- | ----------- | ---- | \n| Set 1 - 163 features | 0.463429 | 0.471| PCIAT_*|\n| Set 2 - 163 features | 0.466009 | 0.470| PCIAT_*|\n| Set 3 - 163 features | 0.466455 | 0.466| PCIAT_*|\n| Set 4 - 163 features | 0.468296 | 0.466| sii |\n| Set 5 - 141 features | 0.463159 | 0.470 | PCIAT_*|\n\n<br>**Overall result - Public/ Private LB -> 0.471 / 0.454** \n<br>\n\n| Feature set | CV  | Public LB score | Target | \n| --- | --- | ----------- | ---- | \n| Set 1 - 163 features  | 0.466184 | 0.472| PCIAT_*|\n| Set 2 - 163 features | 0.468296 | 0.466| sii |\n| Set 3 - 141 features  | 0.463159 | 0.470 | PCIAT_*|\n\n<br>**Overall result - Public/ Private LB -> 0.475 / 0.454** \n\n# What I could have done better here\n- Better choice of final submissions \n- Better choice of candidate feature sets\n- One of my feature sets with a lower public LB score scored extremely well on the private set. I could have chosen it perhaps!\n\n# What I gained here\n- Training with threshold tuning on the training data rather than the OOF data was a good step for me and worked for me here\n- My stability across the leaderboards was also a gain, after my falls in yester shakeup driven competitions\n- I am happy I did not use any complex models and algorithms here and ended up with a relatively stable result. I think this itself feels like success given the churn in the competition\n\n# Concluding remarks\nGood luck for all your competitions going ahead! All the best and happy learning!\n\nBest regards,\nRavi Ramakrishnan",
    "3076463": "Thanks for your sharing, I also tried to use PCIAT_Total to predict, but I failed. I was confused about the threshold selection of PCIAT_Total and sii, because I thought that the data of test and train might be quite different, so it was difficult to choose a suitable threshold",
    "3076464": "Congrats @ravi20076, Good Solution thanks for sharing",
    "3076467": "ruichardliu I had this confusion and hence resorted to a rule based allocation, cutting the data into bins with cutoffs [0, 30, 50, 80] as explained above.\nCongrats for your gold medal @ruichardliu",
    "3076643": "Thanks Ravi, I have learnt a lot from you during this competition. Just a quick question, what did you change between these 3 sets? ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1198628%2F260000b9e1ce6423ae479f23476bca91%2FScreenshot%202024-12-20%20at%2007.19.23.png?generation=1734675809230160&alt=media)",
    "3076683": "Hello @diegoiglesias \nI have shared my training code above, kindly peruse this to know more.\nAll the best!",
    "3076711": "Thanks Ravi, I have learnt a lot from you during this competition. !",
    "3076933": "mrsimple07 did you participate in the competition?",
    "3077239": "Congrats @ravi20076 for this position! It was one of the most difficult competitions to stand in the LB",
    "3077276": "Lol \nGetting a medal here feels like success @octaviograu \nAll the best my friend"
  },
  "source": "meta"
}