{
  "id": 553051,
  "title": "Top 100 Solution - Target Post-Processing Main Success Driver",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/553051",
  "author_name": "Meziane S",
  "post_date": "2024-12-23T11:32:13.147000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello, </p>\n<p>I am sharing the <a href=\"https://www.kaggle.com/code/mezianek/cmi-piu-lgbm-solutiontop100\" target=\"_blank\">notebook</a> that helped me making it to the top 💯 with my team 🚀 <a href=\"https://www.kaggle.com/pathofdata\" target=\"_blank\">@pathofdata</a> <a href=\"https://www.kaggle.com/alessiomiolla\" target=\"_blank\">@alessiomiolla</a>  🚀</p>\n<h2>Introduction</h2>\n<p>First, I'd like to thank Kaggle and the Child Mind Institute for providing such an interesting subject to work on. I hope the work done here will be useful for the CMI.</p>\n<p>You can make your own opinion checking the notebook, but I don't think what's been done is anything special. We joined the challenge late and haven't had much time to dedicate to it. Then, I'd like to focus on what I think has given us the little <strong>boost that made us enter the top 100 : 🔥target post-processing🔥.</strong></p>\n<h2>Notebook Summary</h2>\n<p>In case checking the notebook would be too tedious to go through, which I totally get haha, I try to summarize it below🫡</p>\n<ol>\n<li><strong>Parameterization of our work</strong> : to avoid getting lost in all the different tries we've done, we have a list of parameters at the very top of the notebook that helps us trying out different things (linear models vs. tree based models, target post-processing vs. None, etc)</li>\n<li><strong>Parquet Data Aggregation</strong>: we read and aggregate the parquet data 'at the same time'. The aggregation is very basic, we haven't treated this data as proper time series data, we have only produced statistics such as median, mean, max, etc.</li>\n<li>🔥<strong>Target Cleaning and Post-Processing</strong>: I will develop that point in the next section🔥</li>\n<li><strong>Feature Engineering</strong>: sadly, this part is not big enough. We have only taken care of 3 features - one related to BMI (we have used the lack of NAs in one to fill the other's NAs), Activity Summary Score (we put together in the same variable children and teens to reduce the # NAs), Sleep Disturbance Scale (same as for BMI).</li>\n<li><strong>Pseudo Feature Selection</strong>: again, no proper method here, we only tried to make our life easier at first (hoping we will have more time to get back there later). We dropped the features that are only in train and not in test (mostly the 20 questions), we dropped the 'seasons features' (thinking they will not have a huge predictive power). I also applied a filter on the features with too many NAs (my parameter called <code>ratio_na</code> ) - we decided to impute missing values for the features that have a 'reasonable' number of them.</li>\n<li><strong>Missing Imputation</strong> : here what's been done is around the variable <code>ratio_na</code>. Again, if one feature has a proportion of NAs above <code>ratio_na</code>, we just drop it; if not, we take care of it. We have noticed that imputing everything increases overfitting. In case we impute, we used the only two features that have a 100% presence - age and gender - to compute median per group for the numerical features and assign it to the NAs.</li>\n<li><strong>LGBM Regressor Selected on Manual HP Optimization</strong> : we built a manual loop to selected the best HP for our LGBM - <code>max_depth</code>, <code>learning_rate</code>, <code>n_estimators</code>, <code>feature_fraction</code>. Here again, it has to be understood as a first draft we wanted to polish afterwards (we haven't done so unfortunately)</li>\n<li><strong>Print of Metrics and Features Importance Chart</strong>: to get an idea on the overfitting, the most important features, the best HP, we print and plot a few metrics and charts at the end of the NB</li>\n</ol>\n<h2>🔥Target Cleaning and Post-Processing🔥</h2>\n<p>Thanks to the parametrizable logic of the notebook, I was able to identify the origin of the <strong>biggest performance boost we were able to achieve : the Target Cleaning and Post-Processing</strong>. </p>\n<h3>Target Cleaning</h3>\n<ol>\n<li>We <strong>don't consider IDs with a null target in the train dataset</strong>. Ideally, we would have liked trying some missing imputation technics but we haven't done so by lack of time mostly</li>\n<li>Among those, we drop the ones with at least 2 questions missing</li>\n<li>In the remaining ones, <strong>when there is exactly 1 question not answered, we drop those that have a score (<code>PCIAT-PCIAT_Total</code>) less than 5 points away from making them slide to another category (<code>sii</code>)</strong> . For example, if an ID has 19 questions answered out of 20, and has a score of 28 (then <code>ssi=0</code>), we'd drop that record because potentially  this record could be in <code>sii=1</code> if the question was answered (5 points max per question)</li>\n</ol>\n<p>Although we lost a considerable amount of data (~30% - mostly due to the drop of target being NA, let's be clear), it has proven helping the model. 🤔<strong>The intuition being : the model fits better on fewer but cleaner data</strong>🤔.</p>\n<h3>Target Post-Processing</h3>\n<p>It hasn't been said so far, although we believe it's implied, <strong>we have decided to deal with a regression problem</strong>. Although, we tried the classification approach at first, it has proven being unsuccessful for us. </p>\n<p>That being said, we naturally tried to predict <code>PCIAT-PCIAT_Total</code> and then produced the categories based on the definition of <code>sii</code> (0-30; 31-50; 51-80, 80+).  It was giving us <strong>better results than the classification approach</strong>, but we started to notice that the performance on the LB was not reflecting what we were seeing locally. Then we thought that maybe the distribution of the target was slightly different on the test dataset.</p>\n<p>To account for that possibility, <strong>instead of using the definition of <code>sii</code>'s classes we used its distribution in the train dataset and assigned each ID based on how we predicted <code>PCIAT-PCIAT_Total</code></strong>  🤯</p>\n<p>‼️<strong>Practically</strong>‼️ : </p>\n<ul>\n<li>we predict <code>PCIAT-PCIAT_Total</code></li>\n<li>we sort the predictions/IDs from smallest to highest</li>\n<li>the smallest values would go to <code>sii=0</code>. For example, if we noticed that 50% of the IDs in train where <code>sii=0</code> then we'd do the same for our predictions in the test set (check variable <code>target_distr</code> in the notebook)</li>\n<li>the highest values would go to <code>sii=3</code>. Similarly, if 10% of the train dataset had <code>sii=3</code>, then we'd reproduce that in our predictions.</li>\n<li>what's in between follows the same logic dictated by <code>target_distr</code> </li>\n</ul>\n<p><strong>Everything else being constant, this gave me a boost of ~0.04 in the leaderboard.</strong></p>\n<h2>What we did not have time to implement</h2>\n<p>To end this post - that is longer than I thought it will be 😅 - I'd like to share a list of ideas we didn't have time to implement and 'things' we have noticed along the way. The list has no particular order and I am probably missing a few points, but <strong>I am interested in your thoughts - please, feel free to share them in the comments section</strong>.</p>\n<ol>\n<li>Training the best (or any) model on more data (giving less data to test dataset) hasn't given a better performance on LB</li>\n<li>A bit of overfitting (never more than 0.05 gap between train and test though) helped me gain some points on public (and private) LB</li>\n<li>Parquet data gives a boost on tree based models but penalize the linear approaches (lasso included)</li>\n<li>Building ensemble models - even the same GBM with different random seeds was in the plan</li>\n<li>Use parquet data as proper time series data was in the plan</li>\n<li>Combine all features related to body built / weights / muscle - to have a macro 'body feature' with less missing values was in the plan</li>\n<li>Combine the SDS features to have less missing values was in the plan</li>\n<li>Combine the BIA features to have less missing values - PCA was in the plan here</li>\n</ol>\n<p>Thank you for reading my post, I hope you find it somehow useful. I am interested in your feedback and opinions regarding the target post-processing approach, whether you've used it or not!</p>",
  "messages": [
    {
      "id": 3079212,
      "postDate": "2024-12-23T11:32:13.147Z",
      "content": "<p>Hello, </p>\n<p>I am sharing the <a href=\"https://www.kaggle.com/code/mezianek/cmi-piu-lgbm-solutiontop100\" target=\"_blank\">notebook</a> that helped me making it to the top 💯 with my team 🚀 <a href=\"https://www.kaggle.com/pathofdata\" target=\"_blank\">@pathofdata</a> <a href=\"https://www.kaggle.com/alessiomiolla\" target=\"_blank\">@alessiomiolla</a>  🚀</p>\n<h2>Introduction</h2>\n<p>First, I'd like to thank Kaggle and the Child Mind Institute for providing such an interesting subject to work on. I hope the work done here will be useful for the CMI.</p>\n<p>You can make your own opinion checking the notebook, but I don't think what's been done is anything special. We joined the challenge late and haven't had much time to dedicate to it. Then, I'd like to focus on what I think has given us the little <strong>boost that made us enter the top 100 : 🔥target post-processing🔥.</strong></p>\n<h2>Notebook Summary</h2>\n<p>In case checking the notebook would be too tedious to go through, which I totally get haha, I try to summarize it below🫡</p>\n<ol>\n<li><strong>Parameterization of our work</strong> : to avoid getting lost in all the different tries we've done, we have a list of parameters at the very top of the notebook that helps us trying out different things (linear models vs. tree based models, target post-processing vs. None, etc)</li>\n<li><strong>Parquet Data Aggregation</strong>: we read and aggregate the parquet data 'at the same time'. The aggregation is very basic, we haven't treated this data as proper time series data, we have only produced statistics such as median, mean, max, etc.</li>\n<li>🔥<strong>Target Cleaning and Post-Processing</strong>: I will develop that point in the next section🔥</li>\n<li><strong>Feature Engineering</strong>: sadly, this part is not big enough. We have only taken care of 3 features - one related to BMI (we have used the lack of NAs in one to fill the other's NAs), Activity Summary Score (we put together in the same variable children and teens to reduce the # NAs), Sleep Disturbance Scale (same as for BMI).</li>\n<li><strong>Pseudo Feature Selection</strong>: again, no proper method here, we only tried to make our life easier at first (hoping we will have more time to get back there later). We dropped the features that are only in train and not in test (mostly the 20 questions), we dropped the 'seasons features' (thinking they will not have a huge predictive power). I also applied a filter on the features with too many NAs (my parameter called <code>ratio_na</code> ) - we decided to impute missing values for the features that have a 'reasonable' number of them.</li>\n<li><strong>Missing Imputation</strong> : here what's been done is around the variable <code>ratio_na</code>. Again, if one feature has a proportion of NAs above <code>ratio_na</code>, we just drop it; if not, we take care of it. We have noticed that imputing everything increases overfitting. In case we impute, we used the only two features that have a 100% presence - age and gender - to compute median per group for the numerical features and assign it to the NAs.</li>\n<li><strong>LGBM Regressor Selected on Manual HP Optimization</strong> : we built a manual loop to selected the best HP for our LGBM - <code>max_depth</code>, <code>learning_rate</code>, <code>n_estimators</code>, <code>feature_fraction</code>. Here again, it has to be understood as a first draft we wanted to polish afterwards (we haven't done so unfortunately)</li>\n<li><strong>Print of Metrics and Features Importance Chart</strong>: to get an idea on the overfitting, the most important features, the best HP, we print and plot a few metrics and charts at the end of the NB</li>\n</ol>\n<h2>🔥Target Cleaning and Post-Processing🔥</h2>\n<p>Thanks to the parametrizable logic of the notebook, I was able to identify the origin of the <strong>biggest performance boost we were able to achieve : the Target Cleaning and Post-Processing</strong>. </p>\n<h3>Target Cleaning</h3>\n<ol>\n<li>We <strong>don't consider IDs with a null target in the train dataset</strong>. Ideally, we would have liked trying some missing imputation technics but we haven't done so by lack of time mostly</li>\n<li>Among those, we drop the ones with at least 2 questions missing</li>\n<li>In the remaining ones, <strong>when there is exactly 1 question not answered, we drop those that have a score (<code>PCIAT-PCIAT_Total</code>) less than 5 points away from making them slide to another category (<code>sii</code>)</strong> . For example, if an ID has 19 questions answered out of 20, and has a score of 28 (then <code>ssi=0</code>), we'd drop that record because potentially  this record could be in <code>sii=1</code> if the question was answered (5 points max per question)</li>\n</ol>\n<p>Although we lost a considerable amount of data (~30% - mostly due to the drop of target being NA, let's be clear), it has proven helping the model. 🤔<strong>The intuition being : the model fits better on fewer but cleaner data</strong>🤔.</p>\n<h3>Target Post-Processing</h3>\n<p>It hasn't been said so far, although we believe it's implied, <strong>we have decided to deal with a regression problem</strong>. Although, we tried the classification approach at first, it has proven being unsuccessful for us. </p>\n<p>That being said, we naturally tried to predict <code>PCIAT-PCIAT_Total</code> and then produced the categories based on the definition of <code>sii</code> (0-30; 31-50; 51-80, 80+).  It was giving us <strong>better results than the classification approach</strong>, but we started to notice that the performance on the LB was not reflecting what we were seeing locally. Then we thought that maybe the distribution of the target was slightly different on the test dataset.</p>\n<p>To account for that possibility, <strong>instead of using the definition of <code>sii</code>'s classes we used its distribution in the train dataset and assigned each ID based on how we predicted <code>PCIAT-PCIAT_Total</code></strong>  🤯</p>\n<p>‼️<strong>Practically</strong>‼️ : </p>\n<ul>\n<li>we predict <code>PCIAT-PCIAT_Total</code></li>\n<li>we sort the predictions/IDs from smallest to highest</li>\n<li>the smallest values would go to <code>sii=0</code>. For example, if we noticed that 50% of the IDs in train where <code>sii=0</code> then we'd do the same for our predictions in the test set (check variable <code>target_distr</code> in the notebook)</li>\n<li>the highest values would go to <code>sii=3</code>. Similarly, if 10% of the train dataset had <code>sii=3</code>, then we'd reproduce that in our predictions.</li>\n<li>what's in between follows the same logic dictated by <code>target_distr</code> </li>\n</ul>\n<p><strong>Everything else being constant, this gave me a boost of ~0.04 in the leaderboard.</strong></p>\n<h2>What we did not have time to implement</h2>\n<p>To end this post - that is longer than I thought it will be 😅 - I'd like to share a list of ideas we didn't have time to implement and 'things' we have noticed along the way. The list has no particular order and I am probably missing a few points, but <strong>I am interested in your thoughts - please, feel free to share them in the comments section</strong>.</p>\n<ol>\n<li>Training the best (or any) model on more data (giving less data to test dataset) hasn't given a better performance on LB</li>\n<li>A bit of overfitting (never more than 0.05 gap between train and test though) helped me gain some points on public (and private) LB</li>\n<li>Parquet data gives a boost on tree based models but penalize the linear approaches (lasso included)</li>\n<li>Building ensemble models - even the same GBM with different random seeds was in the plan</li>\n<li>Use parquet data as proper time series data was in the plan</li>\n<li>Combine all features related to body built / weights / muscle - to have a macro 'body feature' with less missing values was in the plan</li>\n<li>Combine the SDS features to have less missing values was in the plan</li>\n<li>Combine the BIA features to have less missing values - PCA was in the plan here</li>\n</ol>\n<p>Thank you for reading my post, I hope you find it somehow useful. I am interested in your feedback and opinions regarding the target post-processing approach, whether you've used it or not!</p>",
      "rawMarkdown": "Hello, \n\nI am sharing the [notebook](https://www.kaggle.com/code/mezianek/cmi-piu-lgbm-solutiontop100) that helped me making it to the top 💯 with my team 🚀 @pathofdata @alessiomiolla  🚀\n\n\n## Introduction\n\nFirst, I'd like to thank Kaggle and the Child Mind Institute for providing such an interesting subject to work on. I hope the work done here will be useful for the CMI.\n\nYou can make your own opinion checking the notebook, but I don't think what's been done is anything special. We joined the challenge late and haven't had much time to dedicate to it. Then, I'd like to focus on what I think has given us the little **boost that made us enter the top 100 : 🔥target post-processing🔥.**\n\n\n## Notebook Summary\n\nIn case checking the notebook would be too tedious to go through, which I totally get haha, I try to summarize it below🫡\n\n1. **Parameterization of our work** : to avoid getting lost in all the different tries we've done, we have a list of parameters at the very top of the notebook that helps us trying out different things (linear models vs. tree based models, target post-processing vs. None, etc)\n2. **Parquet Data Aggregation**: we read and aggregate the parquet data 'at the same time'. The aggregation is very basic, we haven't treated this data as proper time series data, we have only produced statistics such as median, mean, max, etc.\n3. 🔥**Target Cleaning and Post-Processing**: I will develop that point in the next section🔥\n4. **Feature Engineering**: sadly, this part is not big enough. We have only taken care of 3 features - one related to BMI (we have used the lack of NAs in one to fill the other's NAs), Activity Summary Score (we put together in the same variable children and teens to reduce the # NAs), Sleep Disturbance Scale (same as for BMI).\n5. **Pseudo Feature Selection**: again, no proper method here, we only tried to make our life easier at first (hoping we will have more time to get back there later). We dropped the features that are only in train and not in test (mostly the 20 questions), we dropped the 'seasons features' (thinking they will not have a huge predictive power). I also applied a filter on the features with too many NAs (my parameter called `ratio_na` ) - we decided to impute missing values for the features that have a 'reasonable' number of them.\n6. **Missing Imputation** : here what's been done is around the variable `ratio_na`. Again, if one feature has a proportion of NAs above `ratio_na`, we just drop it; if not, we take care of it. We have noticed that imputing everything increases overfitting. In case we impute, we used the only two features that have a 100% presence - age and gender - to compute median per group for the numerical features and assign it to the NAs.\n7. **LGBM Regressor Selected on Manual HP Optimization** : we built a manual loop to selected the best HP for our LGBM - `max_depth`, `learning_rate`, `n_estimators`, `feature_fraction`. Here again, it has to be understood as a first draft we wanted to polish afterwards (we haven't done so unfortunately)\n8. **Print of Metrics and Features Importance Chart**: to get an idea on the overfitting, the most important features, the best HP, we print and plot a few metrics and charts at the end of the NB\n\n## 🔥Target Cleaning and Post-Processing🔥\n\nThanks to the parametrizable logic of the notebook, I was able to identify the origin of the **biggest performance boost we were able to achieve : the Target Cleaning and Post-Processing**. \n\n### Target Cleaning\n1. We **don't consider IDs with a null target in the train dataset**. Ideally, we would have liked trying some missing imputation technics but we haven't done so by lack of time mostly\n2. Among those, we drop the ones with at least 2 questions missing\n3. In the remaining ones, **when there is exactly 1 question not answered, we drop those that have a score (`PCIAT-PCIAT_Total`) less than 5 points away from making them slide to another category (`sii`)** . For example, if an ID has 19 questions answered out of 20, and has a score of 28 (then `ssi=0`), we'd drop that record because potentially  this record could be in `sii=1` if the question was answered (5 points max per question)\n\nAlthough we lost a considerable amount of data (~30% - mostly due to the drop of target being NA, let's be clear), it has proven helping the model. 🤔**The intuition being : the model fits better on fewer but cleaner data**🤔.\n\n\n### Target Post-Processing\n\nIt hasn't been said so far, although we believe it's implied, **we have decided to deal with a regression problem**. Although, we tried the classification approach at first, it has proven being unsuccessful for us. \n\nThat being said, we naturally tried to predict `PCIAT-PCIAT_Total` and then produced the categories based on the definition of `sii` (0-30; 31-50; 51-80, 80+).  It was giving us **better results than the classification approach**, but we started to notice that the performance on the LB was not reflecting what we were seeing locally. Then we thought that maybe the distribution of the target was slightly different on the test dataset.\n\nTo account for that possibility, **instead of using the definition of `sii`'s classes we used its distribution in the train dataset and assigned each ID based on how we predicted `PCIAT-PCIAT_Total`**  🤯\n\n‼️**Practically**‼️ : \n- we predict `PCIAT-PCIAT_Total`\n- we sort the predictions/IDs from smallest to highest\n- the smallest values would go to `sii=0`. For example, if we noticed that 50% of the IDs in train where `sii=0` then we'd do the same for our predictions in the test set (check variable `target_distr` in the notebook)\n- the highest values would go to `sii=3`. Similarly, if 10% of the train dataset had `sii=3`, then we'd reproduce that in our predictions.\n- what's in between follows the same logic dictated by `target_distr` \n\n**Everything else being constant, this gave me a boost of ~0.04 in the leaderboard.**\n\n\n## What we did not have time to implement\n\nTo end this post - that is longer than I thought it will be 😅 - I'd like to share a list of ideas we didn't have time to implement and 'things' we have noticed along the way. The list has no particular order and I am probably missing a few points, but **I am interested in your thoughts - please, feel free to share them in the comments section**.\n\n1. Training the best (or any) model on more data (giving less data to test dataset) hasn't given a better performance on LB\n2. A bit of overfitting (never more than 0.05 gap between train and test though) helped me gain some points on public (and private) LB\n3. Parquet data gives a boost on tree based models but penalize the linear approaches (lasso included)\n4. Building ensemble models - even the same GBM with different random seeds was in the plan\n5. Use parquet data as proper time series data was in the plan\n6. Combine all features related to body built / weights / muscle - to have a macro 'body feature' with less missing values was in the plan\n7. Combine the SDS features to have less missing values was in the plan\n8. Combine the BIA features to have less missing values - PCA was in the plan here\n\n\nThank you for reading my post, I hope you find it somehow useful. I am interested in your feedback and opinions regarding the target post-processing approach, whether you've used it or not!\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3079212": "Hello, \n\nI am sharing the [notebook](https://www.kaggle.com/code/mezianek/cmi-piu-lgbm-solutiontop100) that helped me making it to the top 💯 with my team 🚀 @pathofdata @alessiomiolla  🚀\n\n\n## Introduction\n\nFirst, I'd like to thank Kaggle and the Child Mind Institute for providing such an interesting subject to work on. I hope the work done here will be useful for the CMI.\n\nYou can make your own opinion checking the notebook, but I don't think what's been done is anything special. We joined the challenge late and haven't had much time to dedicate to it. Then, I'd like to focus on what I think has given us the little **boost that made us enter the top 100 : 🔥target post-processing🔥.**\n\n\n## Notebook Summary\n\nIn case checking the notebook would be too tedious to go through, which I totally get haha, I try to summarize it below🫡\n\n1. **Parameterization of our work** : to avoid getting lost in all the different tries we've done, we have a list of parameters at the very top of the notebook that helps us trying out different things (linear models vs. tree based models, target post-processing vs. None, etc)\n2. **Parquet Data Aggregation**: we read and aggregate the parquet data 'at the same time'. The aggregation is very basic, we haven't treated this data as proper time series data, we have only produced statistics such as median, mean, max, etc.\n3. 🔥**Target Cleaning and Post-Processing**: I will develop that point in the next section🔥\n4. **Feature Engineering**: sadly, this part is not big enough. We have only taken care of 3 features - one related to BMI (we have used the lack of NAs in one to fill the other's NAs), Activity Summary Score (we put together in the same variable children and teens to reduce the # NAs), Sleep Disturbance Scale (same as for BMI).\n5. **Pseudo Feature Selection**: again, no proper method here, we only tried to make our life easier at first (hoping we will have more time to get back there later). We dropped the features that are only in train and not in test (mostly the 20 questions), we dropped the 'seasons features' (thinking they will not have a huge predictive power). I also applied a filter on the features with too many NAs (my parameter called `ratio_na` ) - we decided to impute missing values for the features that have a 'reasonable' number of them.\n6. **Missing Imputation** : here what's been done is around the variable `ratio_na`. Again, if one feature has a proportion of NAs above `ratio_na`, we just drop it; if not, we take care of it. We have noticed that imputing everything increases overfitting. In case we impute, we used the only two features that have a 100% presence - age and gender - to compute median per group for the numerical features and assign it to the NAs.\n7. **LGBM Regressor Selected on Manual HP Optimization** : we built a manual loop to selected the best HP for our LGBM - `max_depth`, `learning_rate`, `n_estimators`, `feature_fraction`. Here again, it has to be understood as a first draft we wanted to polish afterwards (we haven't done so unfortunately)\n8. **Print of Metrics and Features Importance Chart**: to get an idea on the overfitting, the most important features, the best HP, we print and plot a few metrics and charts at the end of the NB\n\n## 🔥Target Cleaning and Post-Processing🔥\n\nThanks to the parametrizable logic of the notebook, I was able to identify the origin of the **biggest performance boost we were able to achieve : the Target Cleaning and Post-Processing**. \n\n### Target Cleaning\n1. We **don't consider IDs with a null target in the train dataset**. Ideally, we would have liked trying some missing imputation technics but we haven't done so by lack of time mostly\n2. Among those, we drop the ones with at least 2 questions missing\n3. In the remaining ones, **when there is exactly 1 question not answered, we drop those that have a score (`PCIAT-PCIAT_Total`) less than 5 points away from making them slide to another category (`sii`)** . For example, if an ID has 19 questions answered out of 20, and has a score of 28 (then `ssi=0`), we'd drop that record because potentially  this record could be in `sii=1` if the question was answered (5 points max per question)\n\nAlthough we lost a considerable amount of data (~30% - mostly due to the drop of target being NA, let's be clear), it has proven helping the model. 🤔**The intuition being : the model fits better on fewer but cleaner data**🤔.\n\n\n### Target Post-Processing\n\nIt hasn't been said so far, although we believe it's implied, **we have decided to deal with a regression problem**. Although, we tried the classification approach at first, it has proven being unsuccessful for us. \n\nThat being said, we naturally tried to predict `PCIAT-PCIAT_Total` and then produced the categories based on the definition of `sii` (0-30; 31-50; 51-80, 80+).  It was giving us **better results than the classification approach**, but we started to notice that the performance on the LB was not reflecting what we were seeing locally. Then we thought that maybe the distribution of the target was slightly different on the test dataset.\n\nTo account for that possibility, **instead of using the definition of `sii`'s classes we used its distribution in the train dataset and assigned each ID based on how we predicted `PCIAT-PCIAT_Total`**  🤯\n\n‼️**Practically**‼️ : \n- we predict `PCIAT-PCIAT_Total`\n- we sort the predictions/IDs from smallest to highest\n- the smallest values would go to `sii=0`. For example, if we noticed that 50% of the IDs in train where `sii=0` then we'd do the same for our predictions in the test set (check variable `target_distr` in the notebook)\n- the highest values would go to `sii=3`. Similarly, if 10% of the train dataset had `sii=3`, then we'd reproduce that in our predictions.\n- what's in between follows the same logic dictated by `target_distr` \n\n**Everything else being constant, this gave me a boost of ~0.04 in the leaderboard.**\n\n\n## What we did not have time to implement\n\nTo end this post - that is longer than I thought it will be 😅 - I'd like to share a list of ideas we didn't have time to implement and 'things' we have noticed along the way. The list has no particular order and I am probably missing a few points, but **I am interested in your thoughts - please, feel free to share them in the comments section**.\n\n1. Training the best (or any) model on more data (giving less data to test dataset) hasn't given a better performance on LB\n2. A bit of overfitting (never more than 0.05 gap between train and test though) helped me gain some points on public (and private) LB\n3. Parquet data gives a boost on tree based models but penalize the linear approaches (lasso included)\n4. Building ensemble models - even the same GBM with different random seeds was in the plan\n5. Use parquet data as proper time series data was in the plan\n6. Combine all features related to body built / weights / muscle - to have a macro 'body feature' with less missing values was in the plan\n7. Combine the SDS features to have less missing values was in the plan\n8. Combine the BIA features to have less missing values - PCA was in the plan here\n\n\nThank you for reading my post, I hope you find it somehow useful. I am interested in your feedback and opinions regarding the target post-processing approach, whether you've used it or not!\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n"
  }
}