{
  "id": 552940,
  "title": "10th Place Solution: Hierarchical Bayesian Model with 5 features",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552940",
  "author_name": "",
  "post_date": "2024-12-22T15:22:11.555457300Z",
  "votes": 8,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First of all, I would like to express my gratitude to the competition organizers and participants.<br>\nEnsuring accuracy with a noisy dataset was highly challenging and offered many valuable lessons.</p>\n<p>The solution presented below uses only five features and employs a relatively simple hierarchical Bayesian model (CV: 0.441, Public: 0.447, Private: 0.473).</p>\n<p>At the beginning of the competition, like many others, I considered GBDT-based solutions. However, I realized that using a similar approach as other participants would likely reduce the chance of a shake-up. Taking into account the limited number of features contributing to predictions, I shifted to a simpler modeling strategy, aiming for a shake-up.</p>\n<p>I attribute the success of this strategy purely to luck, but I hope this approach can serve as a reference for others.</p>\n<h1>Notebook</h1>\n<p><a href=\"https://www.kaggle.com/code/junpeimorioka/10th-place-5-features-hierarchical-bayes\" target=\"_blank\">https://www.kaggle.com/code/junpeimorioka/10th-place-5-features-hierarchical-bayes</a></p>\n<h1>Methodology</h1>\n<h2>Target Variable</h2>\n<p>To build a regression model, I selected <strong>PCIAT-PCIAT_Total</strong> as the target variable, as it is easier to handle as a quantitative variable than sii.</p>\n<h2>Features</h2>\n<p>Based on EDA, the following five features were used:</p>\n<ul>\n<li>Age_Category (new)</li>\n<li>Basic_Demos-Sex (raw)</li>\n<li>SDS-SDS_Total_T (raw)</li>\n<li>PreInt_EduHx-computerinternet_hoursday (raw)</li>\n<li>Binary_CGAS_Score (new)</li>\n</ul>\n<p>The actigraph data allowed us to extract features correlated with the target variable; however, due to the high proportion of missing values, it was excluded from this analysis.</p>\n<h2>Missing Value Handling</h2>\n<p>Missing values were imputed using the median of the training data.</p>\n<h2>Model</h2>\n<p>A hierarchical Bayesian model was only used. Random effects were modeled based on age categories, excluding gender. Additionally, since the target variable is truncated at a lower bound of 0, this was reflected in the model.</p>\n<h2>Predictions</h2>\n<p>The predictions for PCIAT-PCIAT_Total were obtained as the mean of the posterior distribution samples.<br>\nTo derive predictions for SII, thresholds were applied. These thresholds were determined using random grid search to maximize QWK on the training data.</p>\n<h2>Cross-Validation and Prediction of Test Data</h2>\n<p>Seven-fold stratified cross-validation was used. Predictions are made on the test data for each model. The final prediction is obtained by averaging and rounding the results of these predictions.</p>\n<p>Thank you for reading until the end!</p>",
  "messages": [
    {
      "id": "3078626",
      "postDate": "12/22/2024 15:22:11",
      "content": "<p>First of all, I would like to express my gratitude to the competition organizers and participants.<br>\nEnsuring accuracy with a noisy dataset was highly challenging and offered many valuable lessons.</p>\n<p>The solution presented below uses only five features and employs a relatively simple hierarchical Bayesian model (CV: 0.441, Public: 0.447, Private: 0.473).</p>\n<p>At the beginning of the competition, like many others, I considered GBDT-based solutions. However, I realized that using a similar approach as other participants would likely reduce the chance of a shake-up. Taking into account the limited number of features contributing to predictions, I shifted to a simpler modeling strategy, aiming for a shake-up.</p>\n<p>I attribute the success of this strategy purely to luck, but I hope this approach can serve as a reference for others.</p>\n<h1>Notebook</h1>\n<p><a href=\"https://www.kaggle.com/code/junpeimorioka/10th-place-5-features-hierarchical-bayes\" target=\"_blank\">https://www.kaggle.com/code/junpeimorioka/10th-place-5-features-hierarchical-bayes</a></p>\n<h1>Methodology</h1>\n<h2>Target Variable</h2>\n<p>To build a regression model, I selected <strong>PCIAT-PCIAT_Total</strong> as the target variable, as it is easier to handle as a quantitative variable than sii.</p>\n<h2>Features</h2>\n<p>Based on EDA, the following five features were used:</p>\n<ul>\n<li>Age_Category (new)</li>\n<li>Basic_Demos-Sex (raw)</li>\n<li>SDS-SDS_Total_T (raw)</li>\n<li>PreInt_EduHx-computerinternet_hoursday (raw)</li>\n<li>Binary_CGAS_Score (new)</li>\n</ul>\n<p>The actigraph data allowed us to extract features correlated with the target variable; however, due to the high proportion of missing values, it was excluded from this analysis.</p>\n<h2>Missing Value Handling</h2>\n<p>Missing values were imputed using the median of the training data.</p>\n<h2>Model</h2>\n<p>A hierarchical Bayesian model was only used. Random effects were modeled based on age categories, excluding gender. Additionally, since the target variable is truncated at a lower bound of 0, this was reflected in the model.</p>\n<h2>Predictions</h2>\n<p>The predictions for PCIAT-PCIAT_Total were obtained as the mean of the posterior distribution samples.<br>\nTo derive predictions for SII, thresholds were applied. These thresholds were determined using random grid search to maximize QWK on the training data.</p>\n<h2>Cross-Validation and Prediction of Test Data</h2>\n<p>Seven-fold stratified cross-validation was used. Predictions are made on the test data for each model. The final prediction is obtained by averaging and rounding the results of these predictions.</p>\n<p>Thank you for reading until the end!</p>",
      "rawMarkdown": "First of all, I would like to express my gratitude to the competition organizers and participants.\nEnsuring accuracy with a noisy dataset was highly challenging and offered many valuable lessons.\n\nThe solution presented below uses only five features and employs a relatively simple hierarchical Bayesian model (CV: 0.441, Public: 0.447, Private: 0.473).\n\nAt the beginning of the competition, like many others, I considered GBDT-based solutions. However, I realized that using a similar approach as other participants would likely reduce the chance of a shake-up. Taking into account the limited number of features contributing to predictions, I shifted to a simpler modeling strategy, aiming for a shake-up.\n\nI attribute the success of this strategy purely to luck, but I hope this approach can serve as a reference for others.\n\n# Notebook\n[https://www.kaggle.com/code/junpeimorioka/10th-place-5-features-hierarchical-bayes](https://www.kaggle.com/code/junpeimorioka/10th-place-5-features-hierarchical-bayes)\n\n# Methodology\n\n## Target Variable\nTo build a regression model, I selected **PCIAT-PCIAT_Total** as the target variable, as it is easier to handle as a quantitative variable than sii.\n\n## Features\nBased on EDA, the following five features were used:\n- Age_Category (new)\n- Basic_Demos-Sex (raw)\n- SDS-SDS_Total_T (raw)\n- PreInt_EduHx-computerinternet_hoursday (raw)\n- Binary_CGAS_Score (new)\n\nThe actigraph data allowed us to extract features correlated with the target variable; however, due to the high proportion of missing values, it was excluded from this analysis.\n\n## Missing Value Handling\nMissing values were imputed using the median of the training data.\n\n## Model\nA hierarchical Bayesian model was only used. Random effects were modeled based on age categories, excluding gender. Additionally, since the target variable is truncated at a lower bound of 0, this was reflected in the model.\n\n## Predictions\nThe predictions for PCIAT-PCIAT_Total were obtained as the mean of the posterior distribution samples.\nTo derive predictions for SII, thresholds were applied. These thresholds were determined using random grid search to maximize QWK on the training data.\n\n## Cross-Validation and Prediction of Test Data\nSeven-fold stratified cross-validation was used. Predictions are made on the test data for each model. The final prediction is obtained by averaging and rounding the results of these predictions.\n\nThank you for reading until the end!",
      "votes": null
    },
    {
      "id": "3079432",
      "postDate": "12/23/2024 16:14:41",
      "content": "<p>It does not load the notebook :-)</p>",
      "rawMarkdown": "It does not load the notebook :-)",
      "votes": null
    },
    {
      "id": "3079567",
      "postDate": "12/23/2024 20:19:48",
      "content": "<p>Thank you for pointing out the issue. I have corrected the link settings.</p>",
      "rawMarkdown": "Thank you for pointing out the issue. I have corrected the link settings.",
      "votes": null
    },
    {
      "id": "3132056",
      "postDate": "02/23/2025 16:30:57",
      "content": "<p>thank you for your work</p>",
      "rawMarkdown": "thank you for your work",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3079432,
      "author_name": "faibioss",
      "author_url": "",
      "post_date": "12/23/2024 16:14:41",
      "content": "<p>It does not load the notebook :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3079567,
          "author_name": "junpeimorioka",
          "author_url": "",
          "post_date": "12/23/2024 20:19:48",
          "content": "<p>Thank you for pointing out the issue. I have corrected the link settings.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3132056,
      "author_name": "hiyato",
      "author_url": "",
      "post_date": "02/23/2025 16:30:57",
      "content": "<p>thank you for your work</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3078626": "First of all, I would like to express my gratitude to the competition organizers and participants.\nEnsuring accuracy with a noisy dataset was highly challenging and offered many valuable lessons.\n\nThe solution presented below uses only five features and employs a relatively simple hierarchical Bayesian model (CV: 0.441, Public: 0.447, Private: 0.473).\n\nAt the beginning of the competition, like many others, I considered GBDT-based solutions. However, I realized that using a similar approach as other participants would likely reduce the chance of a shake-up. Taking into account the limited number of features contributing to predictions, I shifted to a simpler modeling strategy, aiming for a shake-up.\n\nI attribute the success of this strategy purely to luck, but I hope this approach can serve as a reference for others.\n\n# Notebook\n[https://www.kaggle.com/code/junpeimorioka/10th-place-5-features-hierarchical-bayes](https://www.kaggle.com/code/junpeimorioka/10th-place-5-features-hierarchical-bayes)\n\n# Methodology\n\n## Target Variable\nTo build a regression model, I selected **PCIAT-PCIAT_Total** as the target variable, as it is easier to handle as a quantitative variable than sii.\n\n## Features\nBased on EDA, the following five features were used:\n- Age_Category (new)\n- Basic_Demos-Sex (raw)\n- SDS-SDS_Total_T (raw)\n- PreInt_EduHx-computerinternet_hoursday (raw)\n- Binary_CGAS_Score (new)\n\nThe actigraph data allowed us to extract features correlated with the target variable; however, due to the high proportion of missing values, it was excluded from this analysis.\n\n## Missing Value Handling\nMissing values were imputed using the median of the training data.\n\n## Model\nA hierarchical Bayesian model was only used. Random effects were modeled based on age categories, excluding gender. Additionally, since the target variable is truncated at a lower bound of 0, this was reflected in the model.\n\n## Predictions\nThe predictions for PCIAT-PCIAT_Total were obtained as the mean of the posterior distribution samples.\nTo derive predictions for SII, thresholds were applied. These thresholds were determined using random grid search to maximize QWK on the training data.\n\n## Cross-Validation and Prediction of Test Data\nSeven-fold stratified cross-validation was used. Predictions are made on the test data for each model. The final prediction is obtained by averaging and rounding the results of these predictions.\n\nThank you for reading until the end!",
    "3079432": "It does not load the notebook :-)",
    "3079567": "Thank you for pointing out the issue. I have corrected the link settings.",
    "3132056": "thank you for your work"
  },
  "source": "meta"
}