{
  "id": 351021,
  "title": "Kaggle wisdom from Grandmaster senkin13",
  "url": "/competitions/open-problems-multimodal/discussion/351021",
  "author_name": "Alexander Chervov",
  "post_date": "2022-09-08T07:51:21.054000",
  "votes": 34,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Twitter recommended me Kaggle Grandmaster <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>  twitter - it is very cool to read - let me share.<br>\nEspecially Kaggle novices MUST read.</p>\n<p><a href=\"https://twitter.com/senkin13\" target=\"_blank\">https://twitter.com/senkin13</a></p>\n<p>\"Sometimes kaggle is exhausted, frustrated.But if you take kaggle as a hobby, an part of Life, you will feel better.\"</p>\n<p>\"Most of kagglers may have these mistakes: ensemble too early,hyper-parameter tuning too early,team up too early,celebrate too early. 2) give up too early,imitate kernel notebook too early.😅\"</p>\n<p>\"When I am at top rank of LB,if it is a single model I hope competitors thought it is ensemble, if it is ensemble, I hope thought it is single model.😂\"</p>\n<p>\"When I decide to join a kaggle competition, I will estimate skill/luck ratio,if &lt;0.5 give up,if &gt;0.9 all in,else see if it is interesting then decide effort level.For example H&amp;M,Riiid,Molecular,Avito,Talkingdata are very good competitions.\"<br>\nHow do you estimate luck factor?<br>\n\"I check CV LB correlation with small improvement first, make sure train and test have similar distribution .then check if single model has big improvement space that need skills,creativity,hard work,not depend on lucky seed and ensemble.\"</p>\n<p>\"If you want to win a kaggle competition, remember [The devil is in the details], the winner solution seems not too much different with other solutions, they just want to keep a little secret 😅\"</p>\n<p>\"H&amp;M Competition is not only high usage of CPU&amp;Memory, but also hard disk,I can not create more data to train&amp;inference because full of my 4 TB disk. Now I can release my disk😀\"</p>\n<p>\"I remember lightgbm was always better than xgboost and catboost before.But recent tabular data competitions I joined like (kaggle days championship 2&amp;3, riiid, nishika, atmacup), xgboost or catboost were much better than lightgbm,so don’t trust experience,trust experiment.\"</p>\n<p>\"I write a blog to introduce using data science competitions skills to improve business value.\"<br>\n<a href=\"https://twitter.com/senkin13/status/1552819057090387969?t=tptk1nnnRaM7GGIgKV4s3A&amp;s=19\" target=\"_blank\">https://twitter.com/senkin13/status/1552819057090387969?t=tptk1nnnRaM7GGIgKV4s3A&amp;s=19</a></p>\n<p>PS <br>\nHe is top1 on LB at the moment. </p>\n<p>PS PS<br>\nIf someone knows other cool twitters from Kagglers - please share</p>",
  "messages": [
    {
      "id": 1930799,
      "postDate": "2022-09-08T07:51:21.053Z",
      "content": "<p>Twitter recommended me Kaggle Grandmaster <a href=\"https://www.kaggle.com/senkin13\" target=\"_blank\">@senkin13</a>  twitter - it is very cool to read - let me share.<br>\nEspecially Kaggle novices MUST read.</p>\n<p><a href=\"https://twitter.com/senkin13\" target=\"_blank\">https://twitter.com/senkin13</a></p>\n<p>\"Sometimes kaggle is exhausted, frustrated.But if you take kaggle as a hobby, an part of Life, you will feel better.\"</p>\n<p>\"Most of kagglers may have these mistakes: ensemble too early,hyper-parameter tuning too early,team up too early,celebrate too early. 2) give up too early,imitate kernel notebook too early.😅\"</p>\n<p>\"When I am at top rank of LB,if it is a single model I hope competitors thought it is ensemble, if it is ensemble, I hope thought it is single model.😂\"</p>\n<p>\"When I decide to join a kaggle competition, I will estimate skill/luck ratio,if &lt;0.5 give up,if &gt;0.9 all in,else see if it is interesting then decide effort level.For example H&amp;M,Riiid,Molecular,Avito,Talkingdata are very good competitions.\"<br>\nHow do you estimate luck factor?<br>\n\"I check CV LB correlation with small improvement first, make sure train and test have similar distribution .then check if single model has big improvement space that need skills,creativity,hard work,not depend on lucky seed and ensemble.\"</p>\n<p>\"If you want to win a kaggle competition, remember [The devil is in the details], the winner solution seems not too much different with other solutions, they just want to keep a little secret 😅\"</p>\n<p>\"H&amp;M Competition is not only high usage of CPU&amp;Memory, but also hard disk,I can not create more data to train&amp;inference because full of my 4 TB disk. Now I can release my disk😀\"</p>\n<p>\"I remember lightgbm was always better than xgboost and catboost before.But recent tabular data competitions I joined like (kaggle days championship 2&amp;3, riiid, nishika, atmacup), xgboost or catboost were much better than lightgbm,so don’t trust experience,trust experiment.\"</p>\n<p>\"I write a blog to introduce using data science competitions skills to improve business value.\"<br>\n<a href=\"https://twitter.com/senkin13/status/1552819057090387969?t=tptk1nnnRaM7GGIgKV4s3A&amp;s=19\" target=\"_blank\">https://twitter.com/senkin13/status/1552819057090387969?t=tptk1nnnRaM7GGIgKV4s3A&amp;s=19</a></p>\n<p>PS <br>\nHe is top1 on LB at the moment. </p>\n<p>PS PS<br>\nIf someone knows other cool twitters from Kagglers - please share</p>",
      "rawMarkdown": "Twitter recommended me Kaggle Grandmaster @senkin13  twitter - it is very cool to read - let me share.\nEspecially Kaggle novices MUST read.\n\nhttps://twitter.com/senkin13\n\n\"Sometimes kaggle is exhausted, frustrated.But if you take kaggle as a hobby, an part of Life, you will feel better.\"\n\n\"Most of kagglers may have these mistakes: ensemble too early,hyper-parameter tuning too early,team up too early,celebrate too early. 2) give up too early,imitate kernel notebook too early.😅\"\n\n\"When I am at top rank of LB,if it is a single model I hope competitors thought it is ensemble, if it is ensemble, I hope thought it is single model.😂\"\n\n\"When I decide to join a kaggle competition, I will estimate skill/luck ratio,if <0.5 give up,if >0.9 all in,else see if it is interesting then decide effort level.For example H&M,Riiid,Molecular,Avito,Talkingdata are very good competitions.\"\nHow do you estimate luck factor?\n\"I check CV LB correlation with small improvement first, make sure train and test have similar distribution .then check if single model has big improvement space that need skills,creativity,hard work,not depend on lucky seed and ensemble.\"\n\n\n\"If you want to win a kaggle competition, remember [The devil is in the details], the winner solution seems not too much different with other solutions, they just want to keep a little secret 😅\"\n\n\"H&M Competition is not only high usage of CPU&Memory, but also hard disk,I can not create more data to train&inference because full of my 4 TB disk. Now I can release my disk😀\"\n\n\"I remember lightgbm was always better than xgboost and catboost before.But recent tabular data competitions I joined like (kaggle days championship 2&3, riiid, nishika, atmacup), xgboost or catboost were much better than lightgbm,so don’t trust experience,trust experiment.\"\n\n\"I write a blog to introduce using data science competitions skills to improve business value.\"\nhttps://twitter.com/senkin13/status/1552819057090387969?t=tptk1nnnRaM7GGIgKV4s3A&s=19\n\n PS \nHe is top1 on LB at the moment. \n\nPS PS\nIf someone knows other cool twitters from Kagglers - please share\n\n",
      "votes": 34
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1930799": "Twitter recommended me Kaggle Grandmaster @senkin13  twitter - it is very cool to read - let me share.\nEspecially Kaggle novices MUST read.\n\nhttps://twitter.com/senkin13\n\n\"Sometimes kaggle is exhausted, frustrated.But if you take kaggle as a hobby, an part of Life, you will feel better.\"\n\n\"Most of kagglers may have these mistakes: ensemble too early,hyper-parameter tuning too early,team up too early,celebrate too early. 2) give up too early,imitate kernel notebook too early.😅\"\n\n\"When I am at top rank of LB,if it is a single model I hope competitors thought it is ensemble, if it is ensemble, I hope thought it is single model.😂\"\n\n\"When I decide to join a kaggle competition, I will estimate skill/luck ratio,if <0.5 give up,if >0.9 all in,else see if it is interesting then decide effort level.For example H&M,Riiid,Molecular,Avito,Talkingdata are very good competitions.\"\nHow do you estimate luck factor?\n\"I check CV LB correlation with small improvement first, make sure train and test have similar distribution .then check if single model has big improvement space that need skills,creativity,hard work,not depend on lucky seed and ensemble.\"\n\n\n\"If you want to win a kaggle competition, remember [The devil is in the details], the winner solution seems not too much different with other solutions, they just want to keep a little secret 😅\"\n\n\"H&M Competition is not only high usage of CPU&Memory, but also hard disk,I can not create more data to train&inference because full of my 4 TB disk. Now I can release my disk😀\"\n\n\"I remember lightgbm was always better than xgboost and catboost before.But recent tabular data competitions I joined like (kaggle days championship 2&3, riiid, nishika, atmacup), xgboost or catboost were much better than lightgbm,so don’t trust experience,trust experiment.\"\n\n\"I write a blog to introduce using data science competitions skills to improve business value.\"\nhttps://twitter.com/senkin13/status/1552819057090387969?t=tptk1nnnRaM7GGIgKV4s3A&s=19\n\n PS \nHe is top1 on LB at the moment. \n\nPS PS\nIf someone knows other cool twitters from Kagglers - please share\n\n"
  }
}