{
  "id": 552513,
  "title": "19th Place Solution",
  "url": "/competitions/child-mind-institute-problematic-internet-use/writeups/vladimir-demidov-19th-place-solution",
  "author_name": "",
  "post_date": "2024-12-20T04:11:18.900Z",
  "votes": 25,
  "comment_count": 4,
  "views": 0,
  "content": "<p><strong>Thoughts</strong></p>\n<p>I have heavy feelings after it is finalized. I mean, obviously I'm not smart enough to compete in serious competitions. Solving a difficult task on clean data, with the huge amount of famous skilled folks in the top LB, is pretty challenging even for strong professionals. The only chance to get gold for me is highly shakeable games like this one. I realized it after a few of my first competitions in the past. And by the background's violet color in my profile picture, I am symbolizing that, honestly, I have no hope for reaching Master tier.</p>\n<p>I would be happy to surrender and not try again, but here is the quintessence of forming addiction. A person will leave the game if he can't get close to the desired result for a long time, he will also quit the game if he gets what he wants and be satisfied with it. But if the environment leads the game to the fact that a person <strong>almost</strong> achieves what he wants, a person will try again and again in this vicious cycle, comforted by the anticipation of success. This is how it works, and I would really like to reduce the time spent on Kaggle, but unfortunately it seems to me that I will become even more addicted to it.</p>\n<p><strong>Baseline</strong></p>\n<p>I joined lately but have watched this competition since the very beginning. On the one side, I was lucky enough to choose the right strategy; on the other side, I didn't bring my effort to victory. I described my plan for solution development in one discussion <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551850#3073586\" target=\"_blank\">here</a>; you can find my post. The basis of solution is my public <a href=\"https://www.kaggle.com/code/yekenot/cmi-piu-deeptables-nn-cv-0-482\" target=\"_blank\">notebook</a>, Version 60, CV 0.482 | Public 0.448 | Private 0.459, around 60th place on Private LB, so it is a high silver medal just from fork and submit. I've not planned something grand, just tried to explore data by the model's response.</p>\n<p>According to the plan, I've worked with main data only, reduced by missing target, with a tuned threshold ratio (more than 30% missing) for dropping high-NaN columns. So it has 21 features in total after dropping: 6 categorical and 15 numerical. I've tried to fill remaining numerical NaNs with <code>SimpleImputer(strategy='median')</code>, linear interpolation, and <code>KNNImputer</code>, but the best results are achieved by <code>IterativeImputer(max_iter=19)</code> - this was something new to me. <code>MinMaxScaler()</code> was used for scaling. Categorical NaNs were filled automatically by the model's internal option <code>(SimpleImputer(strategy='constant'))</code>. I will not describe all the parameters of the model; you can explore it in the notebook. It's tuned DeepTables NN, pretty simple as you can see, with moderate dropout rates and an LR scheduler, nothing fancy here. But it worked very well, as in my other experiments, and produced an even better score after applying the global maximum search for the threshold optimization procedure. This may be considered a good single model but not the best, I think.</p>\n<p><strong>Ensemble</strong></p>\n<p>For the 19th-place ensemble, I just forked this NN's stuff and added a couple of GBDT models, untuned CatBoost and LightGBM(only adjusted <code>\"reg_alpha\": 3.1</code>), with early stopping 100 rounds. Same folds, all data preparation the same as in the NN's pipeline. Global maximum optimization for the thresholds was included. As a result I received 3 models with close CV performance. The CV of GBDT is a little higher as expected, and this could be a better choice to fit an untuned CatBoost as a single model instead of NN, although I haven't checked this yet. For the ensemble, I applied the <code>majority_vote</code> function from other notebooks to the rounded predictions of these 3 models to get Public LB 0.458 | Private LB 0.469 | &gt;2000 places jump due to shake and 19th-place terrible silver medal in two steps from the gold. That's all I can say. Thanks for your attention.</p>\n<p>Sincerely yours,</p>\n<p>Forever Expert</p>",
  "messages": [
    {
      "id": "3076546",
      "postDate": "12/20/2024 03:49:09",
      "content": "<p><strong>Thoughts</strong></p>\n<p>I have heavy feelings after it is finalized. I mean, obviously I'm not smart enough to compete in serious competitions. Solving a difficult task on clean data, with the huge amount of famous skilled folks in the top LB, is pretty challenging even for strong professionals. The only chance to get gold for me is highly shakeable games like this one. I realized it after a few of my first competitions in the past. And by the background's violet color in my profile picture, I am symbolizing that, honestly, I have no hope for reaching Master tier.</p>\n<p>I would be happy to surrender and not try again, but here is the quintessence of forming addiction. A person will leave the game if he can't get close to the desired result for a long time, he will also quit the game if he gets what he wants and be satisfied with it. But if the environment leads the game to the fact that a person <strong>almost</strong> achieves what he wants, a person will try again and again in this vicious cycle, comforted by the anticipation of success. This is how it works, and I would really like to reduce the time spent on Kaggle, but unfortunately it seems to me that I will become even more addicted to it.</p>\n<p><strong>Baseline</strong></p>\n<p>I joined lately but have watched this competition since the very beginning. On the one side, I was lucky enough to choose the right strategy; on the other side, I didn't bring my effort to victory. I described my plan for solution development in one discussion <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551850#3073586\" target=\"_blank\">here</a>; you can find my post. The basis of solution is my public <a href=\"https://www.kaggle.com/code/yekenot/cmi-piu-deeptables-nn-cv-0-482\" target=\"_blank\">notebook</a>, Version 60, CV 0.482 | Public 0.448 | Private 0.459, around 60th place on Private LB, so it is a high silver medal just from fork and submit. I've not planned something grand, just tried to explore data by the model's response.</p>\n<p>According to the plan, I've worked with main data only, reduced by missing target, with a tuned threshold ratio (more than 30% missing) for dropping high-NaN columns. So it has 21 features in total after dropping: 6 categorical and 15 numerical. I've tried to fill remaining numerical NaNs with <code>SimpleImputer(strategy='median')</code>, linear interpolation, and <code>KNNImputer</code>, but the best results are achieved by <code>IterativeImputer(max_iter=19)</code> - this was something new to me. <code>MinMaxScaler()</code> was used for scaling. Categorical NaNs were filled automatically by the model's internal option <code>(SimpleImputer(strategy='constant'))</code>. I will not describe all the parameters of the model; you can explore it in the notebook. It's tuned DeepTables NN, pretty simple as you can see, with moderate dropout rates and an LR scheduler, nothing fancy here. But it worked very well, as in my other experiments, and produced an even better score after applying the global maximum search for the threshold optimization procedure. This may be considered a good single model but not the best, I think.</p>\n<p><strong>Ensemble</strong></p>\n<p>For the 19th-place ensemble, I just forked this NN's stuff and added a couple of GBDT models, untuned CatBoost and LightGBM(only adjusted <code>\"reg_alpha\": 3.1</code>), with early stopping 100 rounds. Same folds, all data preparation the same as in the NN's pipeline. Global maximum optimization for the thresholds was included. As a result I received 3 models with close CV performance. The CV of GBDT is a little higher as expected, and this could be a better choice to fit an untuned CatBoost as a single model instead of NN, although I haven't checked this yet. For the ensemble, I applied the <code>majority_vote</code> function from other notebooks to the rounded predictions of these 3 models to get Public LB 0.458 | Private LB 0.469 | &gt;2000 places jump due to shake and 19th-place terrible silver medal in two steps from the gold. That's all I can say. Thanks for your attention.</p>\n<p>Sincerely yours,</p>\n<p>Forever Expert</p>",
      "rawMarkdown": "**Thoughts**\n\nI have heavy feelings after it is finalized. I mean, obviously I'm not smart enough to compete in serious competitions. Solving a difficult task on clean data, with the huge amount of famous skilled folks in the top LB, is pretty challenging even for strong professionals. The only chance to get gold for me is highly shakeable games like this one. I realized it after a few of my first competitions in the past. And by the background's violet color in my profile picture, I am symbolizing that, honestly, I have no hope for reaching Master tier.\n\nI would be happy to surrender and not try again, but here is the quintessence of forming addiction. A person will leave the game if he can't get close to the desired result for a long time, he will also quit the game if he gets what he wants and be satisfied with it. But if the environment leads the game to the fact that a person **almost** achieves what he wants, a person will try again and again in this vicious cycle, comforted by the anticipation of success. This is how it works, and I would really like to reduce the time spent on Kaggle, but unfortunately it seems to me that I will become even more addicted to it.\n\n**Baseline**\n\nI joined lately but have watched this competition since the very beginning. On the one side, I was lucky enough to choose the right strategy; on the other side, I didn't bring my effort to victory. I described my plan for solution development in one discussion [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551850#3073586); you can find my post. The basis of solution is my public [notebook](https://www.kaggle.com/code/yekenot/cmi-piu-deeptables-nn-cv-0-482), Version 60, CV 0.482 | Public 0.448 | Private 0.459, around 60th place on Private LB, so it is a high silver medal just from fork and submit. I've not planned something grand, just tried to explore data by the model's response.\n\nAccording to the plan, I've worked with main data only, reduced by missing target, with a tuned threshold ratio (more than 30% missing) for dropping high-NaN columns. So it has 21 features in total after dropping: 6 categorical and 15 numerical. I've tried to fill remaining numerical NaNs with `SimpleImputer(strategy='median')`, linear interpolation, and `KNNImputer`, but the best results are achieved by `IterativeImputer(max_iter=19)` - this was something new to me. `MinMaxScaler()` was used for scaling. Categorical NaNs were filled automatically by the model's internal option `(SimpleImputer(strategy='constant'))`. I will not describe all the parameters of the model; you can explore it in the notebook. It's tuned DeepTables NN, pretty simple as you can see, with moderate dropout rates and an LR scheduler, nothing fancy here. But it worked very well, as in my other experiments, and produced an even better score after applying the global maximum search for the threshold optimization procedure. This may be considered a good single model but not the best, I think.\n\n**Ensemble**\n\nFor the 19th-place ensemble, I just forked this NN's stuff and added a couple of GBDT models, untuned CatBoost and LightGBM(only adjusted `\"reg_alpha\": 3.1`), with early stopping 100 rounds. Same folds, all data preparation the same as in the NN's pipeline. Global maximum optimization for the thresholds was included. As a result I received 3 models with close CV performance. The CV of GBDT is a little higher as expected, and this could be a better choice to fit an untuned CatBoost as a single model instead of NN, although I haven't checked this yet. For the ensemble, I applied the `majority_vote` function from other notebooks to the rounded predictions of these 3 models to get Public LB 0.458 | Private LB 0.469 | >2000 places jump due to shake and 19th-place terrible silver medal in two steps from the gold. That's all I can say. Thanks for your attention.\n\nSincerely yours,\n\nForever Expert",
      "votes": null
    },
    {
      "id": "3076685",
      "postDate": "12/20/2024 07:27:11",
      "content": "<p>Thank you for your sharing and you are very modest. You do not rely on the luck but to rely on your correct solution. You did not use the time-series data and autoencoder like others, avoiding severe data leakage and getting high score on CV. From my perspective, your high rank is predictable.</p>",
      "rawMarkdown": "Thank you for your sharing and you are very modest. You do not rely on the luck but to rely on your correct solution. You did not use the time-series data and autoencoder like others, avoiding severe data leakage and getting high score on CV. From my perspective, your high rank is predictable.",
      "votes": null
    },
    {
      "id": "3076701",
      "postDate": "12/20/2024 07:41:10",
      "content": "<p>This is perhaps the correct way to deal with such competitions <a href=\"https://www.kaggle.com/yekenot\" target=\"_blank\">@yekenot</a> <br>\nLearnt a lot from this approach!</p>",
      "rawMarkdown": "This is perhaps the correct way to deal with such competitions @yekenot \nLearnt a lot from this approach!",
      "votes": null
    },
    {
      "id": "3076767",
      "postDate": "12/20/2024 08:42:48",
      "content": "<p>Congratulations on your great placement and thanks for sharing your solution and your thoughts. All of your Deeptable models I know about perform really well and I profited much from them. I think your strategy for this competition was right, adding untuned models with a simple ensembling technique and spending not too much time on it. I wish you all the best for the next competition. Gold is far away for me but I still earn a lot of programming skills and data science practice. I hope the game is still working for you, as I would like to see you again here sometimes.</p>",
      "rawMarkdown": "Congratulations on your great placement and thanks for sharing your solution and your thoughts. All of your Deeptable models I know about perform really well and I profited much from them. I think your strategy for this competition was right, adding untuned models with a simple ensembling technique and spending not too much time on it. I wish you all the best for the next competition. Gold is far away for me but I still earn a lot of programming skills and data science practice. I hope the game is still working for you, as I would like to see you again here sometimes.",
      "votes": null
    },
    {
      "id": "3076783",
      "postDate": "12/20/2024 09:03:14",
      "content": "<p>Thanks for the kind words. I appreciate your support very much. It would be easy to switch out on other things, say, in the middle of the 2010s, but now it won't let it go because ML is everywhere and all attention is on it. I can't imagine a more suitable place on the internet than Kaggle for mind training.</p>",
      "rawMarkdown": "Thanks for the kind words. I appreciate your support very much. It would be easy to switch out on other things, say, in the middle of the 2010s, but now it won't let it go because ML is everywhere and all attention is on it. I can't imagine a more suitable place on the internet than Kaggle for mind training.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3076685,
      "author_name": "qufangcq",
      "author_url": "",
      "post_date": "12/20/2024 07:27:11",
      "content": "<p>Thank you for your sharing and you are very modest. You do not rely on the luck but to rely on your correct solution. You did not use the time-series data and autoencoder like others, avoiding severe data leakage and getting high score on CV. From my perspective, your high rank is predictable.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3076701,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "12/20/2024 07:41:10",
      "content": "<p>This is perhaps the correct way to deal with such competitions <a href=\"https://www.kaggle.com/yekenot\" target=\"_blank\">@yekenot</a> <br>\nLearnt a lot from this approach!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3076767,
      "author_name": "martinapreusse",
      "author_url": "",
      "post_date": "12/20/2024 08:42:48",
      "content": "<p>Congratulations on your great placement and thanks for sharing your solution and your thoughts. All of your Deeptable models I know about perform really well and I profited much from them. I think your strategy for this competition was right, adding untuned models with a simple ensembling technique and spending not too much time on it. I wish you all the best for the next competition. Gold is far away for me but I still earn a lot of programming skills and data science practice. I hope the game is still working for you, as I would like to see you again here sometimes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3076783,
          "author_name": "yekenot",
          "author_url": "",
          "post_date": "12/20/2024 09:03:14",
          "content": "<p>Thanks for the kind words. I appreciate your support very much. It would be easy to switch out on other things, say, in the middle of the 2010s, but now it won't let it go because ML is everywhere and all attention is on it. I can't imagine a more suitable place on the internet than Kaggle for mind training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3076546": "**Thoughts**\n\nI have heavy feelings after it is finalized. I mean, obviously I'm not smart enough to compete in serious competitions. Solving a difficult task on clean data, with the huge amount of famous skilled folks in the top LB, is pretty challenging even for strong professionals. The only chance to get gold for me is highly shakeable games like this one. I realized it after a few of my first competitions in the past. And by the background's violet color in my profile picture, I am symbolizing that, honestly, I have no hope for reaching Master tier.\n\nI would be happy to surrender and not try again, but here is the quintessence of forming addiction. A person will leave the game if he can't get close to the desired result for a long time, he will also quit the game if he gets what he wants and be satisfied with it. But if the environment leads the game to the fact that a person **almost** achieves what he wants, a person will try again and again in this vicious cycle, comforted by the anticipation of success. This is how it works, and I would really like to reduce the time spent on Kaggle, but unfortunately it seems to me that I will become even more addicted to it.\n\n**Baseline**\n\nI joined lately but have watched this competition since the very beginning. On the one side, I was lucky enough to choose the right strategy; on the other side, I didn't bring my effort to victory. I described my plan for solution development in one discussion [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551850#3073586); you can find my post. The basis of solution is my public [notebook](https://www.kaggle.com/code/yekenot/cmi-piu-deeptables-nn-cv-0-482), Version 60, CV 0.482 | Public 0.448 | Private 0.459, around 60th place on Private LB, so it is a high silver medal just from fork and submit. I've not planned something grand, just tried to explore data by the model's response.\n\nAccording to the plan, I've worked with main data only, reduced by missing target, with a tuned threshold ratio (more than 30% missing) for dropping high-NaN columns. So it has 21 features in total after dropping: 6 categorical and 15 numerical. I've tried to fill remaining numerical NaNs with `SimpleImputer(strategy='median')`, linear interpolation, and `KNNImputer`, but the best results are achieved by `IterativeImputer(max_iter=19)` - this was something new to me. `MinMaxScaler()` was used for scaling. Categorical NaNs were filled automatically by the model's internal option `(SimpleImputer(strategy='constant'))`. I will not describe all the parameters of the model; you can explore it in the notebook. It's tuned DeepTables NN, pretty simple as you can see, with moderate dropout rates and an LR scheduler, nothing fancy here. But it worked very well, as in my other experiments, and produced an even better score after applying the global maximum search for the threshold optimization procedure. This may be considered a good single model but not the best, I think.\n\n**Ensemble**\n\nFor the 19th-place ensemble, I just forked this NN's stuff and added a couple of GBDT models, untuned CatBoost and LightGBM(only adjusted `\"reg_alpha\": 3.1`), with early stopping 100 rounds. Same folds, all data preparation the same as in the NN's pipeline. Global maximum optimization for the thresholds was included. As a result I received 3 models with close CV performance. The CV of GBDT is a little higher as expected, and this could be a better choice to fit an untuned CatBoost as a single model instead of NN, although I haven't checked this yet. For the ensemble, I applied the `majority_vote` function from other notebooks to the rounded predictions of these 3 models to get Public LB 0.458 | Private LB 0.469 | >2000 places jump due to shake and 19th-place terrible silver medal in two steps from the gold. That's all I can say. Thanks for your attention.\n\nSincerely yours,\n\nForever Expert",
    "3076685": "Thank you for your sharing and you are very modest. You do not rely on the luck but to rely on your correct solution. You did not use the time-series data and autoencoder like others, avoiding severe data leakage and getting high score on CV. From my perspective, your high rank is predictable.",
    "3076701": "This is perhaps the correct way to deal with such competitions @yekenot \nLearnt a lot from this approach!",
    "3076767": "Congratulations on your great placement and thanks for sharing your solution and your thoughts. All of your Deeptable models I know about perform really well and I profited much from them. I think your strategy for this competition was right, adding untuned models with a simple ensembling technique and spending not too much time on it. I wish you all the best for the next competition. Gold is far away for me but I still earn a lot of programming skills and data science practice. I hope the game is still working for you, as I would like to see you again here sometimes.",
    "3076783": "Thanks for the kind words. I appreciate your support very much. It would be easy to switch out on other things, say, in the middle of the 2010s, but now it won't let it go because ML is everywhere and all attention is on it. I can't imagine a more suitable place on the internet than Kaggle for mind training."
  },
  "source": "meta"
}