{
  "id": 361765,
  "title": "💡 Final Week Plan | 📗Things learned on the 3rd week",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/361765",
  "author_name": "",
  "post_date": "2022-10-23T14:51:25.966729Z",
  "votes": 13,
  "comment_count": 2,
  "views": 0,
  "content": "<p>On the final week we still have things to learn, there are amazing notebooks that shows how to make things easier.</p>\n<h2>Missing values</h2>\n<ul>\n<li>XGBoost can handle missing values without imputation</li>\n<li>Some imputation methods that I tried makes no big difference (I tried using a place on the field that is unlikely to add up to the probabilities of the team and binary feature to say if it is missing or not)</li>\n</ul>\n<h2>Save memory</h2>\n<ul>\n<li>Dask allows to create dataframe that speed up the process compare to pandas, which is a really great tool for our datasets.</li>\n<li>use gc.collect() to allow clearing up the memory on each step, so it's something that we can take advantage between models. </li>\n</ul>\n<h2>Models</h2>\n<ul>\n<li>There are some AutoML that might work but not sure if they are able to handle big datasets</li>\n<li>From this <a href=\"https://www.kaggle.com/code/donatoriccio/how-to-load-21m-rows-in-1-minute-using-2-lines/notebook\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/donatoriccio\" target=\"_blank\">@donatoriccio</a> I was able to get a better score without using any feature engineering or data manipulation</li>\n</ul>\n<h2>Next week plan</h2>\n<ul>\n<li>I have two notebooks, one for my own experiment and another one to test things from other notebooks</li>\n<li>Try things learned on my own notebook </li>\n<li>Experiment with things of public notebooks</li>\n<li>My rule will be no copy notebooks, I'll be coding along and see if I am able to understand, if not I will try to search until I learn something, my final submission will be on my own notebook from the things learned.</li>\n</ul>\n<p>The notebooks I will like to explore are:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/sergiosaharovskiy/tps-oct-2022-viz-players-positions-animated/notebook\" target=\"_blank\">His visualizations are amazing</a> by <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> </li>\n<li><a href=\"https://www.kaggle.com/code/paddykb/tps-2022-10-fastai\" target=\"_blank\">FastAI model</a> by <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a></li>\n<li><a href=\"https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta/notebook\" target=\"_blank\">Amazing notebook with a lot of things to learn</a> by <a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">@alexryzhkov</a> (This is an improvement of the <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> FastAI model, I want to explore how improvements are made and try to understand their process). </li>\n</ul>",
  "messages": [
    {
      "id": "2000814",
      "postDate": "10/23/2022 14:51:25",
      "content": "<p>On the final week we still have things to learn, there are amazing notebooks that shows how to make things easier.</p>\n<h2>Missing values</h2>\n<ul>\n<li>XGBoost can handle missing values without imputation</li>\n<li>Some imputation methods that I tried makes no big difference (I tried using a place on the field that is unlikely to add up to the probabilities of the team and binary feature to say if it is missing or not)</li>\n</ul>\n<h2>Save memory</h2>\n<ul>\n<li>Dask allows to create dataframe that speed up the process compare to pandas, which is a really great tool for our datasets.</li>\n<li>use gc.collect() to allow clearing up the memory on each step, so it's something that we can take advantage between models. </li>\n</ul>\n<h2>Models</h2>\n<ul>\n<li>There are some AutoML that might work but not sure if they are able to handle big datasets</li>\n<li>From this <a href=\"https://www.kaggle.com/code/donatoriccio/how-to-load-21m-rows-in-1-minute-using-2-lines/notebook\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/donatoriccio\" target=\"_blank\">@donatoriccio</a> I was able to get a better score without using any feature engineering or data manipulation</li>\n</ul>\n<h2>Next week plan</h2>\n<ul>\n<li>I have two notebooks, one for my own experiment and another one to test things from other notebooks</li>\n<li>Try things learned on my own notebook </li>\n<li>Experiment with things of public notebooks</li>\n<li>My rule will be no copy notebooks, I'll be coding along and see if I am able to understand, if not I will try to search until I learn something, my final submission will be on my own notebook from the things learned.</li>\n</ul>\n<p>The notebooks I will like to explore are:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/sergiosaharovskiy/tps-oct-2022-viz-players-positions-animated/notebook\" target=\"_blank\">His visualizations are amazing</a> by <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> </li>\n<li><a href=\"https://www.kaggle.com/code/paddykb/tps-2022-10-fastai\" target=\"_blank\">FastAI model</a> by <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a></li>\n<li><a href=\"https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta/notebook\" target=\"_blank\">Amazing notebook with a lot of things to learn</a> by <a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">@alexryzhkov</a> (This is an improvement of the <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> FastAI model, I want to explore how improvements are made and try to understand their process). </li>\n</ul>",
      "rawMarkdown": "On the final week we still have things to learn, there are amazing notebooks that shows how to make things easier.\n\n## Missing values\n\n- XGBoost can handle missing values without imputation\n- Some imputation methods that I tried makes no big difference (I tried using a place on the field that is unlikely to add up to the probabilities of the team and binary feature to say if it is missing or not)\n\n## Save memory\n\n- Dask allows to create dataframe that speed up the process compare to pandas, which is a really great tool for our datasets.\n- use gc.collect() to allow clearing up the memory on each step, so it's something that we can take advantage between models. \n\n## Models\n\n- There are some AutoML that might work but not sure if they are able to handle big datasets\n- From this [notebook](https://www.kaggle.com/code/donatoriccio/how-to-load-21m-rows-in-1-minute-using-2-lines/notebook) by @donatoriccio I was able to get a better score without using any feature engineering or data manipulation\n\n## Next week plan\n\n- I have two notebooks, one for my own experiment and another one to test things from other notebooks\n- Try things learned on my own notebook \n- Experiment with things of public notebooks\n- My rule will be no copy notebooks, I'll be coding along and see if I am able to understand, if not I will try to search until I learn something, my final submission will be on my own notebook from the things learned.\n\nThe notebooks I will like to explore are:\n- [His visualizations are amazing](https://www.kaggle.com/code/sergiosaharovskiy/tps-oct-2022-viz-players-positions-animated/notebook) by @sergiosaharovskiy \n- [FastAI model](https://www.kaggle.com/code/paddykb/tps-2022-10-fastai) by @paddykb\n- [Amazing notebook with a lot of things to learn](https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta/notebook) by @alexryzhkov (This is an improvement of the @paddykb FastAI model, I want to explore how improvements are made and try to understand their process).",
      "votes": null
    },
    {
      "id": "2001161",
      "postDate": "10/23/2022 19:58:42",
      "content": "<p><a href=\"https://www.kaggle.com/pastorsoto\" target=\"_blank\">@pastorsoto</a> thanks for mentioning me 😎</p>\n<p>I will add my 5 cents. First of all, to figure out what are the improvements in my kernel, take a look on this posts (all of these tricks are used):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/360084\" target=\"_blank\">Why Mish activation works here?</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359502\" target=\"_blank\">Test time data augmentation (TTA) in the nutshell</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359714\" target=\"_blank\">What validation should be used to decrease the metric gap and prevent overfitting?</a></li>\n</ul>\n<p>Secondly, if you are thinking about the AutoML solution - I'm the head of the <a href=\"https://github.com/sb-ai-lab/LightAutoML\" target=\"_blank\">LightAutoML</a> project, the opensource AutoML library which can help you to create good models in the fast automatic way. <a href=\"https://www.kaggle.com/mukaseevru\" target=\"_blank\">@mukaseevru</a> has created the <a href=\"https://www.kaggle.com/code/mukaseevru/tps-oct-22-lama-lightautoml-fe-sampling\" target=\"_blank\">example kernel</a> how to use it on this competition. At the early stages of the competition I have used it on my PC and if you give it a bigger timeout and change f1-score metric to the logloss in <code>Task</code> object you can receive <strong>0.19714</strong> LB (leaderboard) score on full data and <strong>0.19788</strong> LB score on the 20% sample. After that I have created the augmented dataset shuffling players and teams to receive better score - I have created the 163 million rows train dataset with 92 features and 25 million rows test dataset. LightAutoML model was built on the such a huge amount of data in ~6*2 hours (using 32 CPU cores) for the both targets and the both predictions take ~18*2 minutes. This makes me to receive <strong>0.19604</strong> on LB, so LightAutoML can handle this amount of data without any problem.</p>\n<p>Finally, I'm also going to create the topic with some important notes about the competition metric, the LogLoss - so stay tuned and follow me not to miss it 🙃</p>\n<p>Hope this helps and good luck in your work!</p>\n<p>Alex</p>",
      "rawMarkdown": "pastorsoto thanks for mentioning me 😎\n\nI will add my 5 cents. First of all, to figure out what are the improvements in my kernel, take a look on this posts (all of these tricks are used):\n- [Why Mish activation works here?](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/360084)\n- [Test time data augmentation (TTA) in the nutshell](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359502)\n- [What validation should be used to decrease the metric gap and prevent overfitting?](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359714)\n\nSecondly, if you are thinking about the AutoML solution - I'm the head of the [LightAutoML](https://github.com/sb-ai-lab/LightAutoML) project, the opensource AutoML library which can help you to create good models in the fast automatic way. @mukaseevru has created the [example kernel](https://www.kaggle.com/code/mukaseevru/tps-oct-22-lama-lightautoml-fe-sampling) how to use it on this competition. At the early stages of the competition I have used it on my PC and if you give it a bigger timeout and change f1-score metric to the logloss in `Task` object you can receive **0.19714** LB (leaderboard) score on full data and **0.19788** LB score on the 20% sample. After that I have created the augmented dataset shuffling players and teams to receive better score - I have created the 163 million rows train dataset with 92 features and 25 million rows test dataset. LightAutoML model was built on the such a huge amount of data in ~6\\*2 hours (using 32 CPU cores) for the both targets and the both predictions take ~18\\*2 minutes. This makes me to receive **0.19604** on LB, so LightAutoML can handle this amount of data without any problem.\n\nFinally, I'm also going to create the topic with some important notes about the competition metric, the LogLoss - so stay tuned and follow me not to miss it 🙃\n\nHope this helps and good luck in your work!\n\nAlex",
      "votes": null
    },
    {
      "id": "2001270",
      "postDate": "10/23/2022 21:36:34",
      "content": "<p>Hi. That's amazing insight!!! Thank you so much. I will test your AutoML project, I think can give me a great benchmark, I will study your answer, it's really helpful for my learning journey. </p>",
      "rawMarkdown": "Hi. That's amazing insight!!! Thank you so much. I will test your AutoML project, I think can give me a great benchmark, I will study your answer, it's really helpful for my learning journey.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2001161,
      "author_name": "alexryzhkov",
      "author_url": "",
      "post_date": "10/23/2022 19:58:42",
      "content": "<p><a href=\"https://www.kaggle.com/pastorsoto\" target=\"_blank\">@pastorsoto</a> thanks for mentioning me 😎</p>\n<p>I will add my 5 cents. First of all, to figure out what are the improvements in my kernel, take a look on this posts (all of these tricks are used):</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/360084\" target=\"_blank\">Why Mish activation works here?</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359502\" target=\"_blank\">Test time data augmentation (TTA) in the nutshell</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359714\" target=\"_blank\">What validation should be used to decrease the metric gap and prevent overfitting?</a></li>\n</ul>\n<p>Secondly, if you are thinking about the AutoML solution - I'm the head of the <a href=\"https://github.com/sb-ai-lab/LightAutoML\" target=\"_blank\">LightAutoML</a> project, the opensource AutoML library which can help you to create good models in the fast automatic way. <a href=\"https://www.kaggle.com/mukaseevru\" target=\"_blank\">@mukaseevru</a> has created the <a href=\"https://www.kaggle.com/code/mukaseevru/tps-oct-22-lama-lightautoml-fe-sampling\" target=\"_blank\">example kernel</a> how to use it on this competition. At the early stages of the competition I have used it on my PC and if you give it a bigger timeout and change f1-score metric to the logloss in <code>Task</code> object you can receive <strong>0.19714</strong> LB (leaderboard) score on full data and <strong>0.19788</strong> LB score on the 20% sample. After that I have created the augmented dataset shuffling players and teams to receive better score - I have created the 163 million rows train dataset with 92 features and 25 million rows test dataset. LightAutoML model was built on the such a huge amount of data in ~6*2 hours (using 32 CPU cores) for the both targets and the both predictions take ~18*2 minutes. This makes me to receive <strong>0.19604</strong> on LB, so LightAutoML can handle this amount of data without any problem.</p>\n<p>Finally, I'm also going to create the topic with some important notes about the competition metric, the LogLoss - so stay tuned and follow me not to miss it 🙃</p>\n<p>Hope this helps and good luck in your work!</p>\n<p>Alex</p>",
      "votes": null,
      "replies": [
        {
          "id": 2001270,
          "author_name": "pastorsoto",
          "author_url": "",
          "post_date": "10/23/2022 21:36:34",
          "content": "<p>Hi. That's amazing insight!!! Thank you so much. I will test your AutoML project, I think can give me a great benchmark, I will study your answer, it's really helpful for my learning journey. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2000814": "On the final week we still have things to learn, there are amazing notebooks that shows how to make things easier.\n\n## Missing values\n\n- XGBoost can handle missing values without imputation\n- Some imputation methods that I tried makes no big difference (I tried using a place on the field that is unlikely to add up to the probabilities of the team and binary feature to say if it is missing or not)\n\n## Save memory\n\n- Dask allows to create dataframe that speed up the process compare to pandas, which is a really great tool for our datasets.\n- use gc.collect() to allow clearing up the memory on each step, so it's something that we can take advantage between models. \n\n## Models\n\n- There are some AutoML that might work but not sure if they are able to handle big datasets\n- From this [notebook](https://www.kaggle.com/code/donatoriccio/how-to-load-21m-rows-in-1-minute-using-2-lines/notebook) by @donatoriccio I was able to get a better score without using any feature engineering or data manipulation\n\n## Next week plan\n\n- I have two notebooks, one for my own experiment and another one to test things from other notebooks\n- Try things learned on my own notebook \n- Experiment with things of public notebooks\n- My rule will be no copy notebooks, I'll be coding along and see if I am able to understand, if not I will try to search until I learn something, my final submission will be on my own notebook from the things learned.\n\nThe notebooks I will like to explore are:\n- [His visualizations are amazing](https://www.kaggle.com/code/sergiosaharovskiy/tps-oct-2022-viz-players-positions-animated/notebook) by @sergiosaharovskiy \n- [FastAI model](https://www.kaggle.com/code/paddykb/tps-2022-10-fastai) by @paddykb\n- [Amazing notebook with a lot of things to learn](https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta/notebook) by @alexryzhkov (This is an improvement of the @paddykb FastAI model, I want to explore how improvements are made and try to understand their process).",
    "2001161": "pastorsoto thanks for mentioning me 😎\n\nI will add my 5 cents. First of all, to figure out what are the improvements in my kernel, take a look on this posts (all of these tricks are used):\n- [Why Mish activation works here?](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/360084)\n- [Test time data augmentation (TTA) in the nutshell](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359502)\n- [What validation should be used to decrease the metric gap and prevent overfitting?](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/359714)\n\nSecondly, if you are thinking about the AutoML solution - I'm the head of the [LightAutoML](https://github.com/sb-ai-lab/LightAutoML) project, the opensource AutoML library which can help you to create good models in the fast automatic way. @mukaseevru has created the [example kernel](https://www.kaggle.com/code/mukaseevru/tps-oct-22-lama-lightautoml-fe-sampling) how to use it on this competition. At the early stages of the competition I have used it on my PC and if you give it a bigger timeout and change f1-score metric to the logloss in `Task` object you can receive **0.19714** LB (leaderboard) score on full data and **0.19788** LB score on the 20% sample. After that I have created the augmented dataset shuffling players and teams to receive better score - I have created the 163 million rows train dataset with 92 features and 25 million rows test dataset. LightAutoML model was built on the such a huge amount of data in ~6\\*2 hours (using 32 CPU cores) for the both targets and the both predictions take ~18\\*2 minutes. This makes me to receive **0.19604** on LB, so LightAutoML can handle this amount of data without any problem.\n\nFinally, I'm also going to create the topic with some important notes about the competition metric, the LogLoss - so stay tuned and follow me not to miss it 🙃\n\nHope this helps and good luck in your work!\n\nAlex",
    "2001270": "Hi. That's amazing insight!!! Thank you so much. I will test your AutoML project, I think can give me a great benchmark, I will study your answer, it's really helpful for my learning journey."
  },
  "source": "meta"
}