{
  "id": 552638,
  "title": "First Place Write-Up: Or How I Won the Lottery",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552638",
  "author_name": "Lennart Haupts",
  "post_date": "2024-12-20T17:33:46.235000",
  "votes": 94,
  "comment_count": 50,
  "views": 0,
  "content": "<p><strong>First and foremost, I would like to thank the organizers and Kaggle for making this competition possible. Tackling real-world noisy data was both a challenging and rewarding experience.</strong></p>\n<p><strong>To be perfectly honest, luck played a major role in my success. I was especially lucky actually selecting the best possible notebook. Nevertheless, I’d like to present what I did. (<a href=\"https://www.kaggle.com/code/lennarthaupts/1st-place-cmi-model-v4-1-1-reduced?scriptVersionId=213769368\" target=\"_blank\">Link to the notebook</a>)</strong></p>\n<p>Interestingly, I dropped out of the competition two months ago only to re-enter recently. My focus during this return was on improving the robustness of the solution.</p>\n<h1>Final Model Overview</h1>\n<p>The final solution was a voted ensemble consisting of:<br>\n•    LGBMRegressor<br>\n•    Two XGBoost Regressors<br>\n•    CatBoostRegressor<br>\n•    ExtraTreesRegressor</p>\n<h1>Target, Cross-Validation, and Sample Weights</h1>\n<p>•    <strong>Target Variable:</strong> Instead of using the provided sii labels, I used the 'PCIAT-PCIAT_Total' score and converted the predictions to sii labels. <br>\n•    <strong>Distribution and Weighting:</strong> The target’s distribution, especially the excess zeros, led to two approaches: exploring sample weights and alternative objectives for regression like <strong>Tweedie</strong>. Equidistant bins were used to define the sample weights. Weighting did not improve the optimized scores directly but it brought the unoptimized scores closer to the optimized scores.<br>\n•    <strong>Cross-Validation:</strong> A 10-fold stratified KFold was used, stratified by the bins. Seeds were frequently changed to ensure stability. Submissions with different seeds were used to get further feedback from the public Leaderboard on the variance. The LB-score itself was mostly ignored to minimize overfitting to the leaderboard.</p>\n<h1>Data Cleaning, Feature Engineering, and Imputation</h1>\n<p><strong>Data Cleaning:</strong> </p>\n<ul>\n<li>Implausible values, such as body fat percentages over 60% or negative bone mineral content, were removed and replaced with NaN.</li>\n</ul>\n<p><strong>Feature Engineering:</strong> </p>\n<ul>\n<li>Various descriptive actigraph features were created, with separate masks for day and night.</li>\n<li>Dimensionality reduction of the actigraph data using PCA retained 15 components.</li>\n<li>Additional features included normalized values based on age group means and other features that seemed sensible like the difference between the daily energy expenditure and the basal metabolic rate.</li>\n<li>Quantile binning was applied to a good chunk of the features to deal with the noise. Which worked surprisingly well.</li>\n</ul>\n<p><strong>Imputation:</strong> </p>\n<ul>\n<li>Lasso was used for feature imputation, due to the moderately high dimensionality and noise.</li>\n<li>Features were imputed using Lasso. For each target column, a model was trained using features with fewer than 40% missing values, and missing values in these features were imputed based on the trained model. If no valid features (i.e., features with less than 40% missing values) were found for imputation, or if the number of valid samples was too small, the solution defaulted to mean imputation.</li>\n</ul>\n<h1>Parameter Tuning and Feature Selection</h1>\n<p>Early in the competition, it became apparent that typical parameter tuning with regular cross-validation setups resulted in unstable outcomes. To address this:<br>\n•    Repeated Stratified KFold was employed during parameter tuning. With 10 to 20 repeats. Which was computationally more expensive but yielded more robust results.<br>\n•    Feature selection was done manually based on feature importance, reducing the dataset to 39 features </p>",
  "messages": [
    {
      "id": 3077223,
      "postDate": "2024-12-20T17:33:46.237Z",
      "content": "<p><strong>First and foremost, I would like to thank the organizers and Kaggle for making this competition possible. Tackling real-world noisy data was both a challenging and rewarding experience.</strong></p>\n<p><strong>To be perfectly honest, luck played a major role in my success. I was especially lucky actually selecting the best possible notebook. Nevertheless, I’d like to present what I did. (<a href=\"https://www.kaggle.com/code/lennarthaupts/1st-place-cmi-model-v4-1-1-reduced?scriptVersionId=213769368\" target=\"_blank\">Link to the notebook</a>)</strong></p>\n<p>Interestingly, I dropped out of the competition two months ago only to re-enter recently. My focus during this return was on improving the robustness of the solution.</p>\n<h1>Final Model Overview</h1>\n<p>The final solution was a voted ensemble consisting of:<br>\n•    LGBMRegressor<br>\n•    Two XGBoost Regressors<br>\n•    CatBoostRegressor<br>\n•    ExtraTreesRegressor</p>\n<h1>Target, Cross-Validation, and Sample Weights</h1>\n<p>•    <strong>Target Variable:</strong> Instead of using the provided sii labels, I used the 'PCIAT-PCIAT_Total' score and converted the predictions to sii labels. <br>\n•    <strong>Distribution and Weighting:</strong> The target’s distribution, especially the excess zeros, led to two approaches: exploring sample weights and alternative objectives for regression like <strong>Tweedie</strong>. Equidistant bins were used to define the sample weights. Weighting did not improve the optimized scores directly but it brought the unoptimized scores closer to the optimized scores.<br>\n•    <strong>Cross-Validation:</strong> A 10-fold stratified KFold was used, stratified by the bins. Seeds were frequently changed to ensure stability. Submissions with different seeds were used to get further feedback from the public Leaderboard on the variance. The LB-score itself was mostly ignored to minimize overfitting to the leaderboard.</p>\n<h1>Data Cleaning, Feature Engineering, and Imputation</h1>\n<p><strong>Data Cleaning:</strong> </p>\n<ul>\n<li>Implausible values, such as body fat percentages over 60% or negative bone mineral content, were removed and replaced with NaN.</li>\n</ul>\n<p><strong>Feature Engineering:</strong> </p>\n<ul>\n<li>Various descriptive actigraph features were created, with separate masks for day and night.</li>\n<li>Dimensionality reduction of the actigraph data using PCA retained 15 components.</li>\n<li>Additional features included normalized values based on age group means and other features that seemed sensible like the difference between the daily energy expenditure and the basal metabolic rate.</li>\n<li>Quantile binning was applied to a good chunk of the features to deal with the noise. Which worked surprisingly well.</li>\n</ul>\n<p><strong>Imputation:</strong> </p>\n<ul>\n<li>Lasso was used for feature imputation, due to the moderately high dimensionality and noise.</li>\n<li>Features were imputed using Lasso. For each target column, a model was trained using features with fewer than 40% missing values, and missing values in these features were imputed based on the trained model. If no valid features (i.e., features with less than 40% missing values) were found for imputation, or if the number of valid samples was too small, the solution defaulted to mean imputation.</li>\n</ul>\n<h1>Parameter Tuning and Feature Selection</h1>\n<p>Early in the competition, it became apparent that typical parameter tuning with regular cross-validation setups resulted in unstable outcomes. To address this:<br>\n•    Repeated Stratified KFold was employed during parameter tuning. With 10 to 20 repeats. Which was computationally more expensive but yielded more robust results.<br>\n•    Feature selection was done manually based on feature importance, reducing the dataset to 39 features </p>",
      "rawMarkdown": "**First and foremost, I would like to thank the organizers and Kaggle for making this competition possible. Tackling real-world noisy data was both a challenging and rewarding experience.**\n\n**To be perfectly honest, luck played a major role in my success. I was especially lucky actually selecting the best possible notebook. Nevertheless, I’d like to present what I did. ([Link to the notebook](https://www.kaggle.com/code/lennarthaupts/1st-place-cmi-model-v4-1-1-reduced?scriptVersionId=213769368))**\n\nInterestingly, I dropped out of the competition two months ago only to re-enter recently. My focus during this return was on improving the robustness of the solution.\n\n# Final Model Overview\nThe final solution was a voted ensemble consisting of:\n•\tLGBMRegressor\n•\tTwo XGBoost Regressors\n•\tCatBoostRegressor\n•\tExtraTreesRegressor\n\n# Target, Cross-Validation, and Sample Weights\n•\t**Target Variable:** Instead of using the provided sii labels, I used the 'PCIAT-PCIAT_Total' score and converted the predictions to sii labels. \n•\t**Distribution and Weighting:** The target’s distribution, especially the excess zeros, led to two approaches: exploring sample weights and alternative objectives for regression like **Tweedie**. Equidistant bins were used to define the sample weights. Weighting did not improve the optimized scores directly but it brought the unoptimized scores closer to the optimized scores.\n•\t**Cross-Validation:** A 10-fold stratified KFold was used, stratified by the bins. Seeds were frequently changed to ensure stability. Submissions with different seeds were used to get further feedback from the public Leaderboard on the variance. The LB-score itself was mostly ignored to minimize overfitting to the leaderboard.\n# Data Cleaning, Feature Engineering, and Imputation\n**Data Cleaning:** \n- Implausible values, such as body fat percentages over 60% or negative bone mineral content, were removed and replaced with NaN.\n\n**Feature Engineering:** \n-\tVarious descriptive actigraph features were created, with separate masks for day and night.\n-\tDimensionality reduction of the actigraph data using PCA retained 15 components.\n-\tAdditional features included normalized values based on age group means and other features that seemed sensible like the difference between the daily energy expenditure and the basal metabolic rate.\n-      Quantile binning was applied to a good chunk of the features to deal with the noise. Which worked surprisingly well.\n\n**Imputation:** \n-\tLasso was used for feature imputation, due to the moderately high dimensionality and noise.\n-\tFeatures were imputed using Lasso. For each target column, a model was trained using features with fewer than 40% missing values, and missing values in these features were imputed based on the trained model. If no valid features (i.e., features with less than 40% missing values) were found for imputation, or if the number of valid samples was too small, the solution defaulted to mean imputation.\n# Parameter Tuning and Feature Selection\nEarly in the competition, it became apparent that typical parameter tuning with regular cross-validation setups resulted in unstable outcomes. To address this:\n•\tRepeated Stratified KFold was employed during parameter tuning. With 10 to 20 repeats. Which was computationally more expensive but yielded more robust results.\n•\tFeature selection was done manually based on feature importance, reducing the dataset to 39 features ",
      "votes": 94
    },
    {
      "id": 3077246,
      "postDate": "2024-12-20T17:51:34.420Z",
      "content": "<p>Simple but solid classical approach. Congrats! You overcome all bottlenecks by your methodology. This is inspiring.</p>",
      "rawMarkdown": "Simple but solid classical approach. Congrats! You overcome all bottlenecks by your methodology. This is inspiring.",
      "votes": 7,
      "replies": [
        {
          "id": 3077641,
          "postDate": "2024-12-21T08:21:31.750Z",
          "content": "<p>congrats <a href=\"https://www.kaggle.com/yekenot\" target=\"_blank\">@yekenot</a>… can u share ur approach.. u got 19th and I would love to know ur approach</p>",
          "rawMarkdown": "congrats @yekenot... can u share ur approach.. u got 19th and I would love to know ur approach"
        },
        {
          "id": 3078036,
          "postDate": "2024-12-21T17:49:21.127Z",
          "content": "<p>Thank you very much for your kind feedback. Very much appreciated.☺️</p>",
          "rawMarkdown": "Thank you very much for your kind feedback. Very much appreciated.☺️"
        }
      ]
    },
    {
      "id": 3097003,
      "postDate": "2025-01-14T23:34:28.510Z",
      "content": "<p>One strange phenomenon to be explained. I find the engineered features, like PU_norm, CU_norm etc., have lower Spearman's correlation score than the original features with the target, why using those normalized features while dropping the original features lead to better performance?</p>",
      "rawMarkdown": "One strange phenomenon to be explained. I find the engineered features, like PU_norm, CU_norm etc., have lower Spearman's correlation score than the original features with the target, why using those normalized features while dropping the original features lead to better performance?",
      "votes": 1,
      "replies": [
        {
          "id": 3100539,
          "postDate": "2025-01-19T14:06:16.173Z",
          "content": "<p>Hey, thank you for engaging with my solution. That's an insightful observation. PU and CU were strongly correlated with age, which also had a strong relationship with the outcome. By normalizing PU and CU based on the mean of the age groups, the age effect in these features was reduced. 😀</p>",
          "rawMarkdown": "Hey, thank you for engaging with my solution. That's an insightful observation. PU and CU were strongly correlated with age, which also had a strong relationship with the outcome. By normalizing PU and CU based on the mean of the age groups, the age effect in these features was reduced. 😀",
          "replies": [
            {
              "id": 3102993,
              "postDate": "2025-01-22T21:50:46.880Z",
              "content": "<p>Thank you very much for the explanation! The intuition totally makes sense. The confusion just comes from the lower correlation score of the normalized with respect to the target. So, during feature engineering, even if the normalized/derived features show lower correlation score with the target than original features, as long as we believe that the derived features are more meaningful, we should use those ones. Is this right?</p>",
              "rawMarkdown": "Thank you very much for the explanation! The intuition totally makes sense. The confusion just comes from the lower correlation score of the normalized with respect to the target. So, during feature engineering, even if the normalized/derived features show lower correlation score with the target than original features, as long as we believe that the derived features are more meaningful, we should use those ones. Is this right?"
            }
          ]
        }
      ]
    },
    {
      "id": 3083370,
      "postDate": "2024-12-29T11:43:11.533Z",
      "content": "<p>That's great, congrats</p>",
      "rawMarkdown": "That's great, congrats",
      "votes": 1
    },
    {
      "id": 3082809,
      "postDate": "2024-12-28T16:25:06.933Z",
      "content": "<p>very inspiring! </p>",
      "rawMarkdown": "very inspiring! ",
      "votes": 1
    },
    {
      "id": 3082575,
      "postDate": "2024-12-28T11:46:06.423Z",
      "content": "<p>Really nice feature engineering and congratulations!</p>",
      "rawMarkdown": "Really nice feature engineering and congratulations!",
      "votes": 1
    },
    {
      "id": 3081359,
      "postDate": "2024-12-26T16:29:14.670Z",
      "content": "<p>well done, great work</p>",
      "rawMarkdown": "well done, great work",
      "votes": 1
    },
    {
      "id": 3081184,
      "postDate": "2024-12-26T11:35:40.990Z",
      "content": "<p>I rlly admire u. As a person w/ deep ML knowledge, can u give me (a newbie) some advice on ML?</p>",
      "rawMarkdown": "I rlly admire u. As a person w/ deep ML knowledge, can u give me (a newbie) some advice on ML?",
      "votes": 1,
      "replies": [
        {
          "id": 3100544,
          "postDate": "2025-01-19T14:19:00.707Z",
          "content": "<p>Thank you so much, and apologies for my late response! 😊 I honestly wouldn’t consider myself as someone with particularly deep ML knowledge, but I really appreciate your kind words.</p>\n<p>If I could share four pieces of advice, they’d be:</p>\n<ul>\n<li>Do projects you enjoy. It’s the best way to stay motivated and ensure you keep learning.</li>\n<li>Focus on the fundamentals. Understanding the 'why' and 'how' behind the basics is crucial, and it’s easy to get distracted by the shiny appeal of fancy methods.</li>\n<li>Be critical of your validation approach. Only a solid validation strategy leads to trustworthy and usable results.</li>\n<li>Engage with others to exchange ideas. </li>\n</ul>\n<p><strong>Best of luck on your ML journey.</strong></p>",
          "rawMarkdown": "Thank you so much, and apologies for my late response! 😊 I honestly wouldn’t consider myself as someone with particularly deep ML knowledge, but I really appreciate your kind words.\n\nIf I could share four pieces of advice, they’d be:\n\n- Do projects you enjoy. It’s the best way to stay motivated and ensure you keep learning.\n- Focus on the fundamentals. Understanding the 'why' and 'how' behind the basics is crucial, and it’s easy to get distracted by the shiny appeal of fancy methods.\n- Be critical of your validation approach. Only a solid validation strategy leads to trustworthy and usable results.\n- Engage with others to exchange ideas. \n\n**Best of luck on your ML journey.**",
          "votes": 1,
          "replies": [
            {
              "id": 3118859,
              "postDate": "2025-02-08T14:56:55.463Z",
              "content": "<p>Thanks U for ur advice &lt;3</p>",
              "rawMarkdown": "Thanks U for ur advice <3"
            }
          ]
        }
      ]
    },
    {
      "id": 3081178,
      "postDate": "2024-12-26T11:23:17.253Z",
      "content": "<p>Wow, great work!</p>",
      "rawMarkdown": "Wow, great work!",
      "votes": 1
    },
    {
      "id": 3081075,
      "postDate": "2024-12-26T07:05:20.067Z",
      "content": "<p>Thanks for bringing this to my attention!</p>",
      "rawMarkdown": "Thanks for bringing this to my attention!\n",
      "votes": 1
    },
    {
      "id": 3080254,
      "postDate": "2024-12-24T23:30:51.783Z",
      "content": "<p>Wow, great work!</p>",
      "rawMarkdown": "Wow, great work!",
      "votes": 1
    },
    {
      "id": 3079982,
      "postDate": "2024-12-24T12:55:14.013Z",
      "content": "<p>Simple but solid classical approach. This is inspiring.</p>",
      "rawMarkdown": "Simple but solid classical approach. This is inspiring.\n\n\n",
      "votes": 1
    },
    {
      "id": 3079961,
      "postDate": "2024-12-24T12:13:18.220Z",
      "content": "<p>Congratulations, Lennart <a href=\"https://www.kaggle.com/lennarthaupts\" target=\"_blank\">@lennarthaupts</a> !</p>",
      "rawMarkdown": "Congratulations, Lennart @lennarthaupts !",
      "votes": 1
    },
    {
      "id": 3079935,
      "postDate": "2024-12-24T11:28:45.663Z",
      "content": "<p>great job, congra</p>",
      "rawMarkdown": "great job, congra",
      "votes": 1
    },
    {
      "id": 3079250,
      "postDate": "2024-12-23T12:43:19.597Z",
      "content": "<p>exellent👍</p>",
      "rawMarkdown": "exellent👍",
      "votes": 1
    },
    {
      "id": 3078926,
      "postDate": "2024-12-23T02:50:02.713Z",
      "content": "<p>Really nice feature engineering and congratulations!</p>",
      "rawMarkdown": "Really nice feature engineering and congratulations!",
      "votes": 1
    },
    {
      "id": 3078391,
      "postDate": "2024-12-22T08:31:08.667Z",
      "content": "<p>Congratulations！Solid work</p>",
      "rawMarkdown": "Congratulations！Solid work",
      "votes": 1
    },
    {
      "id": 3077913,
      "postDate": "2024-12-21T14:12:16.827Z",
      "content": "<p>Congratulations! Your solid work truly inspires me and highlights the attitude needed for success in Machine Learning. Thank you for sharing your insights!</p>",
      "rawMarkdown": "Congratulations! Your solid work truly inspires me and highlights the attitude needed for success in Machine Learning. Thank you for sharing your insights!\n",
      "votes": 1,
      "replies": [
        {
          "id": 3078038,
          "postDate": "2024-12-21T17:52:50.903Z",
          "content": "<p>Thanks a lot!</p>",
          "rawMarkdown": "Thanks a lot!",
          "votes": 1
        }
      ]
    },
    {
      "id": 3077639,
      "postDate": "2024-12-21T08:19:59.150Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/lennarthaupts\" target=\"_blank\">@lennarthaupts</a> 🎉🎉.. Your solution help me think beyond the box, especially the usage of ExtraTreesRegressor in the end… though I would like to know more about the details of dimensionality reduction using PCA and how did u come up with quantile binning and this method 'Features with fewer than 40% missing values were imputed with a Lasso model'.. I am just impressed with your methods, got to learn a lot and I am working my way upto Grandmaster…</p>\n<p>PS: would be really helpful if you can share your notebook </p>",
      "rawMarkdown": "Congratulations @lennarthaupts 🎉🎉.. Your solution help me think beyond the box, especially the usage of ExtraTreesRegressor in the end... though I would like to know more about the details of dimensionality reduction using PCA and how did u come up with quantile binning and this method 'Features with fewer than 40% missing values were imputed with a Lasso model'.. I am just impressed with your methods, got to learn a lot and I am working my way upto Grandmaster...\n\nPS: would be really helpful if you can share your notebook ",
      "votes": 1,
      "replies": [
        {
          "id": 3078058,
          "postDate": "2024-12-21T18:16:04.337Z",
          "content": "<p>Thank you very much! I’ve shared the notebook link above, but here it is again:</p>\n<p><a href=\"https://www.kaggle.com/code/lennarthaupts/1st-place-cmi-model-v4-1-1-reduced?scriptVersionId=213769368\" target=\"_blank\">Notebook Link</a></p>\n<p>I used quantile binning because the data was noisy. While binning loses some information by grouping values into intervals, it also reduces overfitting, the impact of outliers and measurement errors. I only imputed features with fewer than 40% missing values using Lasso to ensure there were enough samples for a robust model. The 40% threshold is a bit conservative, but it worked for reliability.</p>",
          "rawMarkdown": "Thank you very much! I’ve shared the notebook link above, but here it is again:\n\n[Notebook Link](https://www.kaggle.com/code/lennarthaupts/1st-place-cmi-model-v4-1-1-reduced?scriptVersionId=213769368)\n\nI used quantile binning because the data was noisy. While binning loses some information by grouping values into intervals, it also reduces overfitting, the impact of outliers and measurement errors. I only imputed features with fewer than 40% missing values using Lasso to ensure there were enough samples for a robust model. The 40% threshold is a bit conservative, but it worked for reliability."
        }
      ]
    },
    {
      "id": 3077611,
      "postDate": "2024-12-21T07:24:28.323Z",
      "content": "<p>Congratulations! And thank you for sharing the excellent strategies.</p>\n<p>Retaining unnecessary features can compromise the model's robustness. However, I hesitated to remove features because eliminating those with low feature importance did not necessarily lead to improvements in CV (Stratified 5 Folds *5). As a result, I ended up using about 70 features. Did you select 39 features by applying some criteria to determine their necessity based solely on feature importance values and your CV results? Looking at your notebook, it appears that the \"exclude\" list of unnecessary features was selected manually.</p>",
      "rawMarkdown": "Congratulations! And thank you for sharing the excellent strategies.\n\nRetaining unnecessary features can compromise the model's robustness. However, I hesitated to remove features because eliminating those with low feature importance did not necessarily lead to improvements in CV (Stratified 5 Folds *5). As a result, I ended up using about 70 features. Did you select 39 features by applying some criteria to determine their necessity based solely on feature importance values and your CV results? Looking at your notebook, it appears that the \"exclude\" list of unnecessary features was selected manually.",
      "votes": 1,
      "replies": [
        {
          "id": 3078042,
          "postDate": "2024-12-21T17:56:53.613Z",
          "content": "<p>Thank you for your kind words!</p>\n<p>Yes, I approached this manually and iteratively without really hard criteria. I focused on features that showed low importance across three models, removing one to three at a time and observing the impact on CV. Features were excluded if the CV score improved or only dropped marginally (typically in the third decimal). It was essentially a manual, eyeballed version of recursive feature elimination.</p>",
          "rawMarkdown": "Thank you for your kind words!\n\nYes, I approached this manually and iteratively without really hard criteria. I focused on features that showed low importance across three models, removing one to three at a time and observing the impact on CV. Features were excluded if the CV score improved or only dropped marginally (typically in the third decimal). It was essentially a manual, eyeballed version of recursive feature elimination."
        }
      ]
    },
    {
      "id": 3077540,
      "postDate": "2024-12-21T05:17:38.427Z",
      "content": "<p>Congratulations！Solid work</p>",
      "rawMarkdown": "Congratulations！Solid work",
      "votes": 1,
      "replies": [
        {
          "id": 3078039,
          "postDate": "2024-12-21T17:53:27.763Z",
          "content": "<p>Thank you.</p>",
          "rawMarkdown": "Thank you."
        }
      ]
    },
    {
      "id": 3077513,
      "postDate": "2024-12-21T04:36:30.750Z",
      "content": "<p>Congratulations. Really efficient strategy.</p>",
      "rawMarkdown": "Congratulations. Really efficient strategy.",
      "votes": 1,
      "replies": [
        {
          "id": 3078040,
          "postDate": "2024-12-21T17:53:53.390Z",
          "content": "<p>Many thanks!</p>",
          "rawMarkdown": "Many thanks!"
        }
      ]
    },
    {
      "id": 3077357,
      "postDate": "2024-12-20T20:30:45.657Z",
      "content": "<p>Congrats, Very Efficient ⭐</p>",
      "rawMarkdown": "Congrats, Very Efficient ⭐",
      "votes": 1,
      "replies": [
        {
          "id": 3078046,
          "postDate": "2024-12-21T17:59:28.590Z",
          "content": "<p>Thanks so much☺️</p>",
          "rawMarkdown": "Thanks so much☺️"
        }
      ]
    },
    {
      "id": 3077349,
      "postDate": "2024-12-20T20:13:42.753Z",
      "content": "<p>Congratulations! Very different approach compared to \"best\" public notebooks, who didn't drop any features at all. The robustness and stability of the model seems to have been an important aspect of the competition that many of us missed. <br>\nWhen adding new features, what is it that made you decide \"I should combine X and Y feature in Z way\"? Is it only feature importance or something else? </p>",
      "rawMarkdown": "Congratulations! Very different approach compared to \"best\" public notebooks, who didn't drop any features at all. The robustness and stability of the model seems to have been an important aspect of the competition that many of us missed. \nWhen adding new features, what is it that made you decide \"I should combine X and Y feature in Z way\"? Is it only feature importance or something else? ",
      "votes": 1,
      "replies": [
        {
          "id": 3078049,
          "postDate": "2024-12-21T18:01:03.823Z",
          "content": "<p>Thank you for the kind comment.</p>\n<p>When constructing new features, I usually focused on finding a domain reason; what the new feature might capture. Some features were also created to reduce dimensionality. For instance, I calculated mean, min, and max for FGC_zones and then dropped the original columns.</p>",
          "rawMarkdown": "Thank you for the kind comment.\n\nWhen constructing new features, I usually focused on finding a domain reason; what the new feature might capture. Some features were also created to reduce dimensionality. For instance, I calculated mean, min, and max for FGC_zones and then dropped the original columns."
        }
      ]
    },
    {
      "id": 3077295,
      "postDate": "2024-12-20T18:56:04.313Z",
      "content": "<p>Many congratulations and well thought out feature engineering approaches given the unreliable nature of this data. </p>",
      "rawMarkdown": "Many congratulations and well thought out feature engineering approaches given the unreliable nature of this data. ",
      "votes": 1,
      "replies": [
        {
          "id": 3078041,
          "postDate": "2024-12-21T17:54:45.587Z",
          "content": "<p>I appreciate your kind words.</p>",
          "rawMarkdown": "I appreciate your kind words."
        }
      ]
    },
    {
      "id": 3077244,
      "postDate": "2024-12-20T17:51:00.257Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/lennarthaupts\" target=\"_blank\">@lennarthaupts</a> !</p>",
      "rawMarkdown": "Congratulations @lennarthaupts !",
      "votes": 2,
      "replies": [
        {
          "id": 3078037,
          "postDate": "2024-12-21T17:52:06.333Z",
          "content": "<p>Much appreciated!☺️</p>",
          "rawMarkdown": "Much appreciated!☺️",
          "votes": 1
        }
      ]
    },
    {
      "id": 3077751,
      "postDate": "2024-12-21T11:07:00.617Z",
      "content": "<p>Thank you for sharing your approach to the problem. Great work !</p>",
      "rawMarkdown": "Thank you for sharing your approach to the problem. Great work !"
    },
    {
      "id": 3151004,
      "postDate": "2025-03-16T07:46:17.503Z",
      "content": "<p>Congratulations, but how to get the best set of ensemble models.</p>",
      "rawMarkdown": "Congratulations, but how to get the best set of ensemble models."
    },
    {
      "id": 3089547,
      "postDate": "2025-01-06T07:37:58.560Z",
      "content": "<p>nice work!</p>\n<p>Could I ask, </p>\n<ol>\n<li>why 2 XGboost models? Since the rest of the model types you only went with one implementation.</li>\n<li>the final model submission is trained on the whole dataset right? not an ensemble of CV models using the mean prediction</li>\n</ol>",
      "rawMarkdown": "nice work!\n\nCould I ask, \n\n1. why 2 XGboost models? Since the rest of the model types you only went with one implementation.\n2. the final model submission is trained on the whole dataset right? not an ensemble of CV models using the mean prediction"
    },
    {
      "id": 3079828,
      "postDate": "2024-12-24T07:50:19.917Z",
      "content": "<p>congrats!!😀</p>",
      "rawMarkdown": "congrats!!😀",
      "votes": 1
    },
    {
      "id": 3079786,
      "postDate": "2024-12-24T06:07:56.777Z",
      "content": "<p>Congratulations! Thanks for sharing!</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 3079538,
      "postDate": "2024-12-23T19:09:15.247Z",
      "content": "<p>Thank you for your sharing</p>",
      "rawMarkdown": "Thank you for your sharing",
      "votes": 1
    },
    {
      "id": 3083444,
      "postDate": "2024-12-29T13:51:55.923Z",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations"
    },
    {
      "id": 3083400,
      "postDate": "2024-12-29T12:47:18.917Z",
      "content": "<p>✨🎉Awesome!</p>",
      "rawMarkdown": "✨🎉Awesome!"
    },
    {
      "id": 3083229,
      "postDate": "2024-12-29T07:11:25.900Z",
      "content": "<p>congratulation….</p>",
      "rawMarkdown": "congratulation...."
    },
    {
      "id": 3081307,
      "postDate": "2024-12-26T15:18:52.897Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3077246,
      "author_name": "Vladimir Demidov",
      "author_url": "",
      "post_date": "2024-12-20T17:51:34.420000",
      "content": "<p>Simple but solid classical approach. Congrats! You overcome all bottlenecks by your methodology. This is inspiring.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 3077641,
          "author_name": "Abhisek Das",
          "author_url": "",
          "post_date": "2024-12-21T08:21:31.750000",
          "content": "<p>congrats <a href=\"https://www.kaggle.com/yekenot\" target=\"_blank\">@yekenot</a>… can u share ur approach.. u got 19th and I would love to know ur approach</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3078036,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2024-12-21T17:49:21.127000",
          "content": "<p>Thank you very much for your kind feedback. Very much appreciated.☺️</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3097003,
      "author_name": "Miles",
      "author_url": "",
      "post_date": "2025-01-14T23:34:28.510000",
      "content": "<p>One strange phenomenon to be explained. I find the engineered features, like PU_norm, CU_norm etc., have lower Spearman's correlation score than the original features with the target, why using those normalized features while dropping the original features lead to better performance?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3100539,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2025-01-19T14:06:16.173000",
          "content": "<p>Hey, thank you for engaging with my solution. That's an insightful observation. PU and CU were strongly correlated with age, which also had a strong relationship with the outcome. By normalizing PU and CU based on the mean of the age groups, the age effect in these features was reduced. 😀</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3102993,
              "author_name": "Miles",
              "author_url": "",
              "post_date": "2025-01-22T21:50:46.880000",
              "content": "<p>Thank you very much for the explanation! The intuition totally makes sense. The confusion just comes from the lower correlation score of the normalized with respect to the target. So, during feature engineering, even if the normalized/derived features show lower correlation score with the target than original features, as long as we believe that the derived features are more meaningful, we should use those ones. Is this right?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3083370,
      "author_name": "Syed Rabeet",
      "author_url": "",
      "post_date": "2024-12-29T11:43:11.533000",
      "content": "<p>That's great, congrats</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3082809,
      "author_name": "Bikram Chatterjee",
      "author_url": "",
      "post_date": "2024-12-28T16:25:06.933000",
      "content": "<p>very inspiring! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3082575,
      "author_name": "Harsh Darji",
      "author_url": "",
      "post_date": "2024-12-28T11:46:06.423000",
      "content": "<p>Really nice feature engineering and congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3081359,
      "author_name": "Mohan Nellore",
      "author_url": "",
      "post_date": "2024-12-26T16:29:14.670000",
      "content": "<p>well done, great work</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3081184,
      "author_name": "Dang An",
      "author_url": "",
      "post_date": "2024-12-26T11:35:40.990000",
      "content": "<p>I rlly admire u. As a person w/ deep ML knowledge, can u give me (a newbie) some advice on ML?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3100544,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2025-01-19T14:19:00.707000",
          "content": "<p>Thank you so much, and apologies for my late response! 😊 I honestly wouldn’t consider myself as someone with particularly deep ML knowledge, but I really appreciate your kind words.</p>\n<p>If I could share four pieces of advice, they’d be:</p>\n<ul>\n<li>Do projects you enjoy. It’s the best way to stay motivated and ensure you keep learning.</li>\n<li>Focus on the fundamentals. Understanding the 'why' and 'how' behind the basics is crucial, and it’s easy to get distracted by the shiny appeal of fancy methods.</li>\n<li>Be critical of your validation approach. Only a solid validation strategy leads to trustworthy and usable results.</li>\n<li>Engage with others to exchange ideas. </li>\n</ul>\n<p><strong>Best of luck on your ML journey.</strong></p>",
          "votes": 1,
          "replies": [
            {
              "id": 3118859,
              "author_name": "Dang An",
              "author_url": "",
              "post_date": "2025-02-08T14:56:55.463000",
              "content": "<p>Thanks U for ur advice &lt;3</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3081178,
      "author_name": "Mai Phong",
      "author_url": "",
      "post_date": "2024-12-26T11:23:17.253000",
      "content": "<p>Wow, great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3081075,
      "author_name": "CodeCavalier",
      "author_url": "",
      "post_date": "2024-12-26T07:05:20.067000",
      "content": "<p>Thanks for bringing this to my attention!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3080254,
      "author_name": "Juliet Kreisel",
      "author_url": "",
      "post_date": "2024-12-24T23:30:51.783000",
      "content": "<p>Wow, great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3079982,
      "author_name": "Eya Jlassi",
      "author_url": "",
      "post_date": "2024-12-24T12:55:14.013000",
      "content": "<p>Simple but solid classical approach. This is inspiring.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3079961,
      "author_name": "Alexander Kapturov",
      "author_url": "",
      "post_date": "2024-12-24T12:13:18.220000",
      "content": "<p>Congratulations, Lennart <a href=\"https://www.kaggle.com/lennarthaupts\" target=\"_blank\">@lennarthaupts</a> !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3079935,
      "author_name": "Zhixing Zhang",
      "author_url": "",
      "post_date": "2024-12-24T11:28:45.663000",
      "content": "<p>great job, congra</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3079250,
      "author_name": "Azadeh Razmi",
      "author_url": "",
      "post_date": "2024-12-23T12:43:19.597000",
      "content": "<p>exellent👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3078926,
      "author_name": "Yinuo Huang",
      "author_url": "",
      "post_date": "2024-12-23T02:50:02.713000",
      "content": "<p>Really nice feature engineering and congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3078391,
      "author_name": "Nadav Cherry",
      "author_url": "",
      "post_date": "2024-12-22T08:31:08.667000",
      "content": "<p>Congratulations！Solid work</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3077913,
      "author_name": "Eunki_H",
      "author_url": "",
      "post_date": "2024-12-21T14:12:16.827000",
      "content": "<p>Congratulations! Your solid work truly inspires me and highlights the attitude needed for success in Machine Learning. Thank you for sharing your insights!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3078038,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2024-12-21T17:52:50.903000",
          "content": "<p>Thanks a lot!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3077639,
      "author_name": "Abhisek Das",
      "author_url": "",
      "post_date": "2024-12-21T08:19:59.150000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/lennarthaupts\" target=\"_blank\">@lennarthaupts</a> 🎉🎉.. Your solution help me think beyond the box, especially the usage of ExtraTreesRegressor in the end… though I would like to know more about the details of dimensionality reduction using PCA and how did u come up with quantile binning and this method 'Features with fewer than 40% missing values were imputed with a Lasso model'.. I am just impressed with your methods, got to learn a lot and I am working my way upto Grandmaster…</p>\n<p>PS: would be really helpful if you can share your notebook </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3078058,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2024-12-21T18:16:04.337000",
          "content": "<p>Thank you very much! I’ve shared the notebook link above, but here it is again:</p>\n<p><a href=\"https://www.kaggle.com/code/lennarthaupts/1st-place-cmi-model-v4-1-1-reduced?scriptVersionId=213769368\" target=\"_blank\">Notebook Link</a></p>\n<p>I used quantile binning because the data was noisy. While binning loses some information by grouping values into intervals, it also reduces overfitting, the impact of outliers and measurement errors. I only imputed features with fewer than 40% missing values using Lasso to ensure there were enough samples for a robust model. The 40% threshold is a bit conservative, but it worked for reliability.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3077611,
      "author_name": "taroiwai",
      "author_url": "",
      "post_date": "2024-12-21T07:24:28.323000",
      "content": "<p>Congratulations! And thank you for sharing the excellent strategies.</p>\n<p>Retaining unnecessary features can compromise the model's robustness. However, I hesitated to remove features because eliminating those with low feature importance did not necessarily lead to improvements in CV (Stratified 5 Folds *5). As a result, I ended up using about 70 features. Did you select 39 features by applying some criteria to determine their necessity based solely on feature importance values and your CV results? Looking at your notebook, it appears that the \"exclude\" list of unnecessary features was selected manually.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3078042,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2024-12-21T17:56:53.613000",
          "content": "<p>Thank you for your kind words!</p>\n<p>Yes, I approached this manually and iteratively without really hard criteria. I focused on features that showed low importance across three models, removing one to three at a time and observing the impact on CV. Features were excluded if the CV score improved or only dropped marginally (typically in the third decimal). It was essentially a manual, eyeballed version of recursive feature elimination.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3077540,
      "author_name": "muteman",
      "author_url": "",
      "post_date": "2024-12-21T05:17:38.427000",
      "content": "<p>Congratulations！Solid work</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3078039,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2024-12-21T17:53:27.763000",
          "content": "<p>Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3077513,
      "author_name": "ArthurT800",
      "author_url": "",
      "post_date": "2024-12-21T04:36:30.750000",
      "content": "<p>Congratulations. Really efficient strategy.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3078040,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2024-12-21T17:53:53.390000",
          "content": "<p>Many thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3077357,
      "author_name": "Mahmoud Elshahed",
      "author_url": "",
      "post_date": "2024-12-20T20:30:45.657000",
      "content": "<p>Congrats, Very Efficient ⭐</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3078046,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2024-12-21T17:59:28.590000",
          "content": "<p>Thanks so much☺️</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3077349,
      "author_name": "vandl",
      "author_url": "",
      "post_date": "2024-12-20T20:13:42.753000",
      "content": "<p>Congratulations! Very different approach compared to \"best\" public notebooks, who didn't drop any features at all. The robustness and stability of the model seems to have been an important aspect of the competition that many of us missed. <br>\nWhen adding new features, what is it that made you decide \"I should combine X and Y feature in Z way\"? Is it only feature importance or something else? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3078049,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2024-12-21T18:01:03.823000",
          "content": "<p>Thank you for the kind comment.</p>\n<p>When constructing new features, I usually focused on finding a domain reason; what the new feature might capture. Some features were also created to reduce dimensionality. For instance, I calculated mean, min, and max for FGC_zones and then dropped the original columns.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3077295,
      "author_name": "Deepak Saldanha",
      "author_url": "",
      "post_date": "2024-12-20T18:56:04.313000",
      "content": "<p>Many congratulations and well thought out feature engineering approaches given the unreliable nature of this data. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3078041,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2024-12-21T17:54:45.587000",
          "content": "<p>I appreciate your kind words.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3077244,
      "author_name": "aldparis",
      "author_url": "",
      "post_date": "2024-12-20T17:51:00.257000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/lennarthaupts\" target=\"_blank\">@lennarthaupts</a> !</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3078037,
          "author_name": "Lennart Haupts",
          "author_url": "",
          "post_date": "2024-12-21T17:52:06.333000",
          "content": "<p>Much appreciated!☺️</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3077751,
      "author_name": "EvelynSuryadi",
      "author_url": "",
      "post_date": "2024-12-21T11:07:00.617000",
      "content": "<p>Thank you for sharing your approach to the problem. Great work !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3151004,
      "author_name": "Saurabh Tiwari",
      "author_url": "",
      "post_date": "2025-03-16T07:46:17.503000",
      "content": "<p>Congratulations, but how to get the best set of ensemble models.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3089547,
      "author_name": "just another tuesday",
      "author_url": "",
      "post_date": "2025-01-06T07:37:58.560000",
      "content": "<p>nice work!</p>\n<p>Could I ask, </p>\n<ol>\n<li>why 2 XGboost models? Since the rest of the model types you only went with one implementation.</li>\n<li>the final model submission is trained on the whole dataset right? not an ensemble of CV models using the mean prediction</li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3079828,
      "author_name": "quietseong",
      "author_url": "",
      "post_date": "2024-12-24T07:50:19.917000",
      "content": "<p>congrats!!😀</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3079786,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-24T06:07:56.777000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3079538,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-23T19:09:15.247000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3083444,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-29T13:51:55.923000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3083400,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-29T12:47:18.917000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3083229,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-29T07:11:25.900000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3081307,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-26T15:18:52.897000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3077223": "**First and foremost, I would like to thank the organizers and Kaggle for making this competition possible. Tackling real-world noisy data was both a challenging and rewarding experience.**\n\n**To be perfectly honest, luck played a major role in my success. I was especially lucky actually selecting the best possible notebook. Nevertheless, I’d like to present what I did. ([Link to the notebook](https://www.kaggle.com/code/lennarthaupts/1st-place-cmi-model-v4-1-1-reduced?scriptVersionId=213769368))**\n\nInterestingly, I dropped out of the competition two months ago only to re-enter recently. My focus during this return was on improving the robustness of the solution.\n\n# Final Model Overview\nThe final solution was a voted ensemble consisting of:\n•\tLGBMRegressor\n•\tTwo XGBoost Regressors\n•\tCatBoostRegressor\n•\tExtraTreesRegressor\n\n# Target, Cross-Validation, and Sample Weights\n•\t**Target Variable:** Instead of using the provided sii labels, I used the 'PCIAT-PCIAT_Total' score and converted the predictions to sii labels. \n•\t**Distribution and Weighting:** The target’s distribution, especially the excess zeros, led to two approaches: exploring sample weights and alternative objectives for regression like **Tweedie**. Equidistant bins were used to define the sample weights. Weighting did not improve the optimized scores directly but it brought the unoptimized scores closer to the optimized scores.\n•\t**Cross-Validation:** A 10-fold stratified KFold was used, stratified by the bins. Seeds were frequently changed to ensure stability. Submissions with different seeds were used to get further feedback from the public Leaderboard on the variance. The LB-score itself was mostly ignored to minimize overfitting to the leaderboard.\n# Data Cleaning, Feature Engineering, and Imputation\n**Data Cleaning:** \n- Implausible values, such as body fat percentages over 60% or negative bone mineral content, were removed and replaced with NaN.\n\n**Feature Engineering:** \n-\tVarious descriptive actigraph features were created, with separate masks for day and night.\n-\tDimensionality reduction of the actigraph data using PCA retained 15 components.\n-\tAdditional features included normalized values based on age group means and other features that seemed sensible like the difference between the daily energy expenditure and the basal metabolic rate.\n-      Quantile binning was applied to a good chunk of the features to deal with the noise. Which worked surprisingly well.\n\n**Imputation:** \n-\tLasso was used for feature imputation, due to the moderately high dimensionality and noise.\n-\tFeatures were imputed using Lasso. For each target column, a model was trained using features with fewer than 40% missing values, and missing values in these features were imputed based on the trained model. If no valid features (i.e., features with less than 40% missing values) were found for imputation, or if the number of valid samples was too small, the solution defaulted to mean imputation.\n# Parameter Tuning and Feature Selection\nEarly in the competition, it became apparent that typical parameter tuning with regular cross-validation setups resulted in unstable outcomes. To address this:\n•\tRepeated Stratified KFold was employed during parameter tuning. With 10 to 20 repeats. Which was computationally more expensive but yielded more robust results.\n•\tFeature selection was done manually based on feature importance, reducing the dataset to 39 features ",
    "3077246": "Simple but solid classical approach. Congrats! You overcome all bottlenecks by your methodology. This is inspiring.",
    "3097003": "One strange phenomenon to be explained. I find the engineered features, like PU_norm, CU_norm etc., have lower Spearman's correlation score than the original features with the target, why using those normalized features while dropping the original features lead to better performance?",
    "3083370": "That's great, congrats",
    "3082809": "very inspiring! ",
    "3082575": "Really nice feature engineering and congratulations!",
    "3081359": "well done, great work",
    "3081184": "I rlly admire u. As a person w/ deep ML knowledge, can u give me (a newbie) some advice on ML?",
    "3081178": "Wow, great work!",
    "3081075": "Thanks for bringing this to my attention!\n",
    "3080254": "Wow, great work!",
    "3079982": "Simple but solid classical approach. This is inspiring.\n\n\n",
    "3079961": "Congratulations, Lennart @lennarthaupts !",
    "3079935": "great job, congra",
    "3079250": "exellent👍",
    "3078926": "Really nice feature engineering and congratulations!",
    "3078391": "Congratulations！Solid work",
    "3077913": "Congratulations! Your solid work truly inspires me and highlights the attitude needed for success in Machine Learning. Thank you for sharing your insights!\n",
    "3077639": "Congratulations @lennarthaupts 🎉🎉.. Your solution help me think beyond the box, especially the usage of ExtraTreesRegressor in the end... though I would like to know more about the details of dimensionality reduction using PCA and how did u come up with quantile binning and this method 'Features with fewer than 40% missing values were imputed with a Lasso model'.. I am just impressed with your methods, got to learn a lot and I am working my way upto Grandmaster...\n\nPS: would be really helpful if you can share your notebook ",
    "3077611": "Congratulations! And thank you for sharing the excellent strategies.\n\nRetaining unnecessary features can compromise the model's robustness. However, I hesitated to remove features because eliminating those with low feature importance did not necessarily lead to improvements in CV (Stratified 5 Folds *5). As a result, I ended up using about 70 features. Did you select 39 features by applying some criteria to determine their necessity based solely on feature importance values and your CV results? Looking at your notebook, it appears that the \"exclude\" list of unnecessary features was selected manually.",
    "3077540": "Congratulations！Solid work",
    "3077513": "Congratulations. Really efficient strategy.",
    "3077357": "Congrats, Very Efficient ⭐",
    "3077349": "Congratulations! Very different approach compared to \"best\" public notebooks, who didn't drop any features at all. The robustness and stability of the model seems to have been an important aspect of the competition that many of us missed. \nWhen adding new features, what is it that made you decide \"I should combine X and Y feature in Z way\"? Is it only feature importance or something else? ",
    "3077295": "Many congratulations and well thought out feature engineering approaches given the unreliable nature of this data. ",
    "3077244": "Congratulations @lennarthaupts !",
    "3077751": "Thank you for sharing your approach to the problem. Great work !",
    "3151004": "Congratulations, but how to get the best set of ensemble models.",
    "3089547": "nice work!\n\nCould I ask, \n\n1. why 2 XGboost models? Since the rest of the model types you only went with one implementation.\n2. the final model submission is trained on the whole dataset right? not an ensemble of CV models using the mean prediction",
    "3079828": "congrats!!😀",
    "3079786": "Congratulations! Thanks for sharing!",
    "3079538": "Thank you for your sharing",
    "3083444": "congratulations",
    "3083400": "✨🎉Awesome!",
    "3083229": "congratulation....",
    "3081307": "Thanks for sharing!"
  }
}