{
  "id": 552656,
  "title": "5th Place Solution",
  "url": "/competitions/child-mind-institute-problematic-internet-use/writeups/peyman-armaghan-5th-place-solution",
  "author_name": "",
  "post_date": "2024-12-20T22:21:00.030Z",
  "votes": 14,
  "comment_count": 4,
  "views": 0,
  "content": "<p>this is my approach for this competition:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19925008%2F03cb6b2ac8013f2081da24a8a9581bcf%2FUntitled.png?generation=1734719836792270&amp;alt=media\" alt=\"\"></p>\n<p>1- <strong>For Time Series:</strong><br>\nTo make some useful features from the time series, I used clustering. I picked 6 parquet files (2 most variant files from each class), combined them, and fit 15 clusters using the KMeans algorithm. then i used this fitted model to extract cluster from other parquets<br>\nI assumed each cluster could represent a specific movement in the parquet data. Then, I averaged each cluster's values over the time duration and used that average as a feature for each user. So, in the end, I had 15 features per user, which represented the average of some kind of activity during the time the user wore the watch.<br>\nI was late in the competition and didn’t have enough time to clean the data or extract more useful features from the time series. I believe there was a lot of room to extract much better features from it.</p>\n<p>2- <strong>Choosing the Right Target:</strong><br>\nAfter some initial submissions, I found that using threshold optimization led to overfitting and inconsistent results. So, I ignored optimization and used PCIAT-PCIAT_Total as the target labels, and applied fixed thresholds [31, 50, 80]. Since the exact values between these thresholds were less important than their relative position to the thresholds for the final prediction, I binned the PCIAT-PCIAT_Total into 10 bins and adjusted the thresholds accordingly.</p>\n<p>3 - <strong>For Unlabeled Data:</strong><br>\nI used pseudo-labeling. I trained three GBDT models (CatBoost, LGBM, XGB), a Lasso regression model, and a neural network (256-128-64 architecture). Then I ensembled these models to predict labels for the unlabeled data. These new labels were then used to train my final models.</p>\n<p>4-<strong>For Final Prediction:</strong><br>\nFor the final prediction, I used the same four models I had used in pseudo-labeling, trained on combination of original and pseudo labels. I trained the models on the entire dataset and validated only on the original labels. This approach gave me the following results for my winning submission (I didn’t use time series for this prediction):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19925008%2F9638ad309fe3ee9fe931c7ec9dbbc7c0%2Ftble.jpg?generation=1734721244023159&amp;alt=media\" alt=\"\"></p>\n<p>Then, I trained GBDT models on the whole dataset (tabular data + time series features). And used all these models for final ensemble </p>\n<p>Combining these  models gave me a CV of 4.6 and a private LB of 4.81, which put me in second place. However, since QWK is a noisy metric, I tried adding some other models (logistic regression and decision tree regression) with lower CV scores. This improved my final CV and public LB slightly, but it dropped my private score from 4.81 to 4.77. I couldn’t blame for this because at the end the CV and LB and also the consistency between these two are the only metrics to choose final submission. </p>\n<p>special Thanks to  organizers and Kaggle and all participants who shared their knowledge. Best of luck in the next competition, and Happy New Year!</p>",
  "messages": [
    {
      "id": "3077317",
      "postDate": "12/20/2024 19:31:27",
      "content": "<p>this is my approach for this competition:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19925008%2F03cb6b2ac8013f2081da24a8a9581bcf%2FUntitled.png?generation=1734719836792270&amp;alt=media\" alt=\"\"></p>\n<p>1- <strong>For Time Series:</strong><br>\nTo make some useful features from the time series, I used clustering. I picked 6 parquet files (2 most variant files from each class), combined them, and fit 15 clusters using the KMeans algorithm. then i used this fitted model to extract cluster from other parquets<br>\nI assumed each cluster could represent a specific movement in the parquet data. Then, I averaged each cluster's values over the time duration and used that average as a feature for each user. So, in the end, I had 15 features per user, which represented the average of some kind of activity during the time the user wore the watch.<br>\nI was late in the competition and didn’t have enough time to clean the data or extract more useful features from the time series. I believe there was a lot of room to extract much better features from it.</p>\n<p>2- <strong>Choosing the Right Target:</strong><br>\nAfter some initial submissions, I found that using threshold optimization led to overfitting and inconsistent results. So, I ignored optimization and used PCIAT-PCIAT_Total as the target labels, and applied fixed thresholds [31, 50, 80]. Since the exact values between these thresholds were less important than their relative position to the thresholds for the final prediction, I binned the PCIAT-PCIAT_Total into 10 bins and adjusted the thresholds accordingly.</p>\n<p>3 - <strong>For Unlabeled Data:</strong><br>\nI used pseudo-labeling. I trained three GBDT models (CatBoost, LGBM, XGB), a Lasso regression model, and a neural network (256-128-64 architecture). Then I ensembled these models to predict labels for the unlabeled data. These new labels were then used to train my final models.</p>\n<p>4-<strong>For Final Prediction:</strong><br>\nFor the final prediction, I used the same four models I had used in pseudo-labeling, trained on combination of original and pseudo labels. I trained the models on the entire dataset and validated only on the original labels. This approach gave me the following results for my winning submission (I didn’t use time series for this prediction):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19925008%2F9638ad309fe3ee9fe931c7ec9dbbc7c0%2Ftble.jpg?generation=1734721244023159&amp;alt=media\" alt=\"\"></p>\n<p>Then, I trained GBDT models on the whole dataset (tabular data + time series features). And used all these models for final ensemble </p>\n<p>Combining these  models gave me a CV of 4.6 and a private LB of 4.81, which put me in second place. However, since QWK is a noisy metric, I tried adding some other models (logistic regression and decision tree regression) with lower CV scores. This improved my final CV and public LB slightly, but it dropped my private score from 4.81 to 4.77. I couldn’t blame for this because at the end the CV and LB and also the consistency between these two are the only metrics to choose final submission. </p>\n<p>special Thanks to  organizers and Kaggle and all participants who shared their knowledge. Best of luck in the next competition, and Happy New Year!</p>",
      "rawMarkdown": "this is my approach for this competition:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19925008%2F03cb6b2ac8013f2081da24a8a9581bcf%2FUntitled.png?generation=1734719836792270&alt=media)\n\n\n1- **For Time Series:**\nTo make some useful features from the time series, I used clustering. I picked 6 parquet files (2 most variant files from each class), combined them, and fit 15 clusters using the KMeans algorithm. then i used this fitted model to extract cluster from other parquets\nI assumed each cluster could represent a specific movement in the parquet data. Then, I averaged each cluster's values over the time duration and used that average as a feature for each user. So, in the end, I had 15 features per user, which represented the average of some kind of activity during the time the user wore the watch.\nI was late in the competition and didn’t have enough time to clean the data or extract more useful features from the time series. I believe there was a lot of room to extract much better features from it.\n\n2- **Choosing the Right Target:**\nAfter some initial submissions, I found that using threshold optimization led to overfitting and inconsistent results. So, I ignored optimization and used PCIAT-PCIAT_Total as the target labels, and applied fixed thresholds [31, 50, 80]. Since the exact values between these thresholds were less important than their relative position to the thresholds for the final prediction, I binned the PCIAT-PCIAT_Total into 10 bins and adjusted the thresholds accordingly.\n\n3 - **For Unlabeled Data:**\nI used pseudo-labeling. I trained three GBDT models (CatBoost, LGBM, XGB), a Lasso regression model, and a neural network (256-128-64 architecture). Then I ensembled these models to predict labels for the unlabeled data. These new labels were then used to train my final models.\n\n4-**For Final Prediction:**\nFor the final prediction, I used the same four models I had used in pseudo-labeling, trained on combination of original and pseudo labels. I trained the models on the entire dataset and validated only on the original labels. This approach gave me the following results for my winning submission (I didn’t use time series for this prediction):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19925008%2F9638ad309fe3ee9fe931c7ec9dbbc7c0%2Ftble.jpg?generation=1734721244023159&alt=media)\n\nThen, I trained GBDT models on the whole dataset (tabular data + time series features). And used all these models for final ensemble \n\nCombining these  models gave me a CV of 4.6 and a private LB of 4.81, which put me in second place. However, since QWK is a noisy metric, I tried adding some other models (logistic regression and decision tree regression) with lower CV scores. This improved my final CV and public LB slightly, but it dropped my private score from 4.81 to 4.77. I couldn’t blame for this because at the end the CV and LB and also the consistency between these two are the only metrics to choose final submission. \n\nspecial Thanks to  organizers and Kaggle and all participants who shared their knowledge. Best of luck in the next competition, and Happy New Year!",
      "votes": null
    },
    {
      "id": "3077419",
      "postDate": "12/21/2024 00:05:05",
      "content": "<p>Thanks for sharing. It's fab. And congratulations\"<br>\nMay I ask about point 3 (Unlabelled data)? Do you only fill features or do you fill PCIAT_total as well?<br>\nI spent quite a lot of time filling the unlabelled data but I only used LGBM. I somehow dropped mine since it didn't improve CV/LB.<br>\nYours sound robust.</p>",
      "rawMarkdown": "Thanks for sharing. It's fab. And congratulations\"\nMay I ask about point 3 (Unlabelled data)? Do you only fill features or do you fill PCIAT_total as well?\nI spent quite a lot of time filling the unlabelled data but I only used LGBM. I somehow dropped mine since it didn't improve CV/LB.\nYours sound robust.",
      "votes": null
    },
    {
      "id": "3078056",
      "postDate": "12/21/2024 18:11:53",
      "content": "<p>thank you tom<br>\nI fill the NaN values in the features using KNN imputation for the entire dataset then I train some base models based on the available PCIAT_total  and ensemble them using hill climbing. finally, I use the ensembled model to predict the missing PCIAT_total . throughout the process, I use rmse as the evaluation metric.</p>",
      "rawMarkdown": "thank you tom\nI fill the NaN values in the features using KNN imputation for the entire dataset then I train some base models based on the available PCIAT_total  and ensemble them using hill climbing. finally, I use the ensembled model to predict the missing PCIAT_total . throughout the process, I use rmse as the evaluation metric.",
      "votes": null
    },
    {
      "id": "3078448",
      "postDate": "12/22/2024 10:21:37",
      "content": "<p>Congrats, Peyman! great execution and well-deserved success!</p>",
      "rawMarkdown": "Congrats, Peyman! great execution and well-deserved success!",
      "votes": null
    },
    {
      "id": "3078464",
      "postDate": "12/22/2024 11:03:36",
      "content": "<p>Great work and congratulations on the 5th place finish! 🎉  </p>\n<p>Your approach is really insightful, especially the use of clustering for time series feature extraction. It's impressive how you managed to simplify complex movements into representative features, even with limited time to refine.  </p>\n<p>I also appreciate the clarity in your reasoning for choosing the PCIAT-PCIAT_Total as the target and the fixed thresholds. It’s a practical way to address overfitting issues, and your binning strategy makes a lot of sense for stabilizing predictions.  </p>\n<p>Pseudo-labeling combined with an ensemble of diverse models like GBDTs, Lasso, and a neural network was a smart way to leverage the unlabeled data effectively. The added robustness from this step clearly paid off!  </p>\n<p>The trade-off between CV and LB scores is always tricky, especially with a noisy metric like QWK, but your experimentation with additional models to balance this was a great learning point.  </p>\n<p>Thanks for sharing such a detailed breakdown, and best of luck in future competitions. Your strategy will definitely inspire others in tackling similar challenges! 👏</p>",
      "rawMarkdown": "Great work and congratulations on the 5th place finish! 🎉  \n\nYour approach is really insightful, especially the use of clustering for time series feature extraction. It's impressive how you managed to simplify complex movements into representative features, even with limited time to refine.  \n\nI also appreciate the clarity in your reasoning for choosing the PCIAT-PCIAT_Total as the target and the fixed thresholds. It’s a practical way to address overfitting issues, and your binning strategy makes a lot of sense for stabilizing predictions.  \n\nPseudo-labeling combined with an ensemble of diverse models like GBDTs, Lasso, and a neural network was a smart way to leverage the unlabeled data effectively. The added robustness from this step clearly paid off!  \n\nThe trade-off between CV and LB scores is always tricky, especially with a noisy metric like QWK, but your experimentation with additional models to balance this was a great learning point.  \n\nThanks for sharing such a detailed breakdown, and best of luck in future competitions. Your strategy will definitely inspire others in tackling similar challenges! 👏",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3077419,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "12/21/2024 00:05:05",
      "content": "<p>Thanks for sharing. It's fab. And congratulations\"<br>\nMay I ask about point 3 (Unlabelled data)? Do you only fill features or do you fill PCIAT_total as well?<br>\nI spent quite a lot of time filling the unlabelled data but I only used LGBM. I somehow dropped mine since it didn't improve CV/LB.<br>\nYours sound robust.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3078056,
          "author_name": "peymanarmaghan",
          "author_url": "",
          "post_date": "12/21/2024 18:11:53",
          "content": "<p>thank you tom<br>\nI fill the NaN values in the features using KNN imputation for the entire dataset then I train some base models based on the available PCIAT_total  and ensemble them using hill climbing. finally, I use the ensembled model to predict the missing PCIAT_total . throughout the process, I use rmse as the evaluation metric.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3078448,
      "author_name": "hamedabedi",
      "author_url": "",
      "post_date": "12/22/2024 10:21:37",
      "content": "<p>Congrats, Peyman! great execution and well-deserved success!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3078464,
      "author_name": "sudhansu1234",
      "author_url": "",
      "post_date": "12/22/2024 11:03:36",
      "content": "<p>Great work and congratulations on the 5th place finish! 🎉  </p>\n<p>Your approach is really insightful, especially the use of clustering for time series feature extraction. It's impressive how you managed to simplify complex movements into representative features, even with limited time to refine.  </p>\n<p>I also appreciate the clarity in your reasoning for choosing the PCIAT-PCIAT_Total as the target and the fixed thresholds. It’s a practical way to address overfitting issues, and your binning strategy makes a lot of sense for stabilizing predictions.  </p>\n<p>Pseudo-labeling combined with an ensemble of diverse models like GBDTs, Lasso, and a neural network was a smart way to leverage the unlabeled data effectively. The added robustness from this step clearly paid off!  </p>\n<p>The trade-off between CV and LB scores is always tricky, especially with a noisy metric like QWK, but your experimentation with additional models to balance this was a great learning point.  </p>\n<p>Thanks for sharing such a detailed breakdown, and best of luck in future competitions. Your strategy will definitely inspire others in tackling similar challenges! 👏</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3077317": "this is my approach for this competition:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19925008%2F03cb6b2ac8013f2081da24a8a9581bcf%2FUntitled.png?generation=1734719836792270&alt=media)\n\n\n1- **For Time Series:**\nTo make some useful features from the time series, I used clustering. I picked 6 parquet files (2 most variant files from each class), combined them, and fit 15 clusters using the KMeans algorithm. then i used this fitted model to extract cluster from other parquets\nI assumed each cluster could represent a specific movement in the parquet data. Then, I averaged each cluster's values over the time duration and used that average as a feature for each user. So, in the end, I had 15 features per user, which represented the average of some kind of activity during the time the user wore the watch.\nI was late in the competition and didn’t have enough time to clean the data or extract more useful features from the time series. I believe there was a lot of room to extract much better features from it.\n\n2- **Choosing the Right Target:**\nAfter some initial submissions, I found that using threshold optimization led to overfitting and inconsistent results. So, I ignored optimization and used PCIAT-PCIAT_Total as the target labels, and applied fixed thresholds [31, 50, 80]. Since the exact values between these thresholds were less important than their relative position to the thresholds for the final prediction, I binned the PCIAT-PCIAT_Total into 10 bins and adjusted the thresholds accordingly.\n\n3 - **For Unlabeled Data:**\nI used pseudo-labeling. I trained three GBDT models (CatBoost, LGBM, XGB), a Lasso regression model, and a neural network (256-128-64 architecture). Then I ensembled these models to predict labels for the unlabeled data. These new labels were then used to train my final models.\n\n4-**For Final Prediction:**\nFor the final prediction, I used the same four models I had used in pseudo-labeling, trained on combination of original and pseudo labels. I trained the models on the entire dataset and validated only on the original labels. This approach gave me the following results for my winning submission (I didn’t use time series for this prediction):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19925008%2F9638ad309fe3ee9fe931c7ec9dbbc7c0%2Ftble.jpg?generation=1734721244023159&alt=media)\n\nThen, I trained GBDT models on the whole dataset (tabular data + time series features). And used all these models for final ensemble \n\nCombining these  models gave me a CV of 4.6 and a private LB of 4.81, which put me in second place. However, since QWK is a noisy metric, I tried adding some other models (logistic regression and decision tree regression) with lower CV scores. This improved my final CV and public LB slightly, but it dropped my private score from 4.81 to 4.77. I couldn’t blame for this because at the end the CV and LB and also the consistency between these two are the only metrics to choose final submission. \n\nspecial Thanks to  organizers and Kaggle and all participants who shared their knowledge. Best of luck in the next competition, and Happy New Year!",
    "3077419": "Thanks for sharing. It's fab. And congratulations\"\nMay I ask about point 3 (Unlabelled data)? Do you only fill features or do you fill PCIAT_total as well?\nI spent quite a lot of time filling the unlabelled data but I only used LGBM. I somehow dropped mine since it didn't improve CV/LB.\nYours sound robust.",
    "3078056": "thank you tom\nI fill the NaN values in the features using KNN imputation for the entire dataset then I train some base models based on the available PCIAT_total  and ensemble them using hill climbing. finally, I use the ensembled model to predict the missing PCIAT_total . throughout the process, I use rmse as the evaluation metric.",
    "3078448": "Congrats, Peyman! great execution and well-deserved success!",
    "3078464": "Great work and congratulations on the 5th place finish! 🎉  \n\nYour approach is really insightful, especially the use of clustering for time series feature extraction. It's impressive how you managed to simplify complex movements into representative features, even with limited time to refine.  \n\nI also appreciate the clarity in your reasoning for choosing the PCIAT-PCIAT_Total as the target and the fixed thresholds. It’s a practical way to address overfitting issues, and your binning strategy makes a lot of sense for stabilizing predictions.  \n\nPseudo-labeling combined with an ensemble of diverse models like GBDTs, Lasso, and a neural network was a smart way to leverage the unlabeled data effectively. The added robustness from this step clearly paid off!  \n\nThe trade-off between CV and LB scores is always tricky, especially with a noisy metric like QWK, but your experimentation with additional models to balance this was a great learning point.  \n\nThanks for sharing such a detailed breakdown, and best of luck in future competitions. Your strategy will definitely inspire others in tackling similar challenges! 👏"
  },
  "source": "meta"
}