{
  "id": 361815,
  "title": "LB shakeup prediction",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/361815",
  "author_name": "",
  "post_date": "2022-10-23T22:32:21.607619900Z",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>It is 8 days before this competition ends and it's a great time to see what we get for now and what we can receive on the private LB. There are 2 important points to mention:</p>\n<ul>\n<li>The dataset for this competition is huge so our models can get all the necessary insights from the data to receive good scores. This usually goes to almost no shakeup at the end</li>\n<li>The public/private split is 24%/76% so the private data is more than 3 times bigger than public. This can lead for the great shakeup if the majority of people use the LB scores to check their models instead of proper validation techniques.</li>\n</ul>\n<p>My prediction is that the first point will beat the second one (almost no shakeup).</p>\n<p>But what do you think about it?</p>",
  "messages": [
    {
      "id": "2001291",
      "postDate": "10/23/2022 22:32:21",
      "content": "<p>It is 8 days before this competition ends and it's a great time to see what we get for now and what we can receive on the private LB. There are 2 important points to mention:</p>\n<ul>\n<li>The dataset for this competition is huge so our models can get all the necessary insights from the data to receive good scores. This usually goes to almost no shakeup at the end</li>\n<li>The public/private split is 24%/76% so the private data is more than 3 times bigger than public. This can lead for the great shakeup if the majority of people use the LB scores to check their models instead of proper validation techniques.</li>\n</ul>\n<p>My prediction is that the first point will beat the second one (almost no shakeup).</p>\n<p>But what do you think about it?</p>",
      "rawMarkdown": "It is 8 days before this competition ends and it's a great time to see what we get for now and what we can receive on the private LB. There are 2 important points to mention:\n- The dataset for this competition is huge so our models can get all the necessary insights from the data to receive good scores. This usually goes to almost no shakeup at the end\n- The public/private split is 24%/76% so the private data is more than 3 times bigger than public. This can lead for the great shakeup if the majority of people use the LB scores to check their models instead of proper validation techniques.\n\nMy prediction is that the first point will beat the second one (almost no shakeup).\n\nBut what do you think about it?",
      "votes": null
    },
    {
      "id": "2001314",
      "postDate": "10/23/2022 23:33:53",
      "content": "<p><a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">@alexryzhkov</a> I like the topic and I am glad you posted it.  I have been thinking about this too and have an additional consideration to add to the conversation.  </p>\n<p>Some people believe it is nearly impossible to overfit with a NN.  What do you think?</p>\n<p><a href=\"https://www.kaggle.com/TomasVdB\" target=\"_blank\">@TomasVdB</a>, <a href=\"https://www.kaggle.com/fabianbong\" target=\"_blank\">@fabianbong</a> <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> <a href=\"https://www.kaggle.com/samuelcortinhas\" target=\"_blank\">@samuelcortinhas</a>, <a href=\"https://www.kaggle.com/hiro5299834\" target=\"_blank\">@hiro5299834</a>, &amp; <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> - What are your thoughts?</p>",
      "rawMarkdown": "alexryzhkov I like the topic and I am glad you posted it.  I have been thinking about this too and have an additional consideration to add to the conversation.  \n\nSome people believe it is nearly impossible to overfit with a NN.  What do you think?\n\n@TomasVdB, @fabianbong @sergiosaharovskiy @samuelcortinhas, @hiro5299834, & @paddykb - What are your thoughts?",
      "votes": null
    },
    {
      "id": "2001592",
      "postDate": "10/24/2022 06:52:36",
      "content": "<p>I wouldn't go as far as to say overfitting isn't possible, but it definitely takes a long time before it seems to happen on this dataset. I think I've gotten lucky in this competition by hitting on an LR, model architecture and epochs combination that suddenly improved my score quite a bit. I actually use less features in my later models than in my earlier ones which seemed to make the model more robust. </p>",
      "rawMarkdown": "I wouldn't go as far as to say overfitting isn't possible, but it definitely takes a long time before it seems to happen on this dataset. I think I've gotten lucky in this competition by hitting on an LR, model architecture and epochs combination that suddenly improved my score quite a bit. I actually use less features in my later models than in my earlier ones which seemed to make the model more robust.",
      "votes": null
    },
    {
      "id": "2001839",
      "postDate": "10/24/2022 10:35:32",
      "content": "<p>My predictions of shakeups have always been bad so I try not to predict it anymore haha</p>",
      "rawMarkdown": "My predictions of shakeups have always been bad so I try not to predict it anymore haha",
      "votes": null
    },
    {
      "id": "2002640",
      "postDate": "10/25/2022 01:18:03",
      "content": "<p><a href=\"https://www.kaggle.com/tomasvdb\" target=\"_blank\">@tomasvdb</a> I started to look into a feature importance approach for NN.  I am finding LIME, SHAP, &amp; permutation approaches.  However, I have not settled on one to try.  Any suggestions?</p>",
      "rawMarkdown": "tomasvdb I started to look into a feature importance approach for NN.  I am finding LIME, SHAP, & permutation approaches.  However, I have not settled on one to try.  Any suggestions?",
      "votes": null
    },
    {
      "id": "2002642",
      "postDate": "10/25/2022 01:19:24",
      "content": "<p>Also, what I have noticed is using a larger layer number vs more layers has been working for me.  It also takes longer to run, too.  😜</p>",
      "rawMarkdown": "Also, what I have noticed is using a larger layer number vs more layers has been working for me.  It also takes longer to run, too.  😜",
      "votes": null
    },
    {
      "id": "2003817",
      "postDate": "10/25/2022 19:45:28",
      "content": "<p>Just saw this! Personally this is my first competition that I have been part of from beginning to end so your guess is as good as mine :) </p>",
      "rawMarkdown": "Just saw this! Personally this is my first competition that I have been part of from beginning to end so your guess is as good as mine :)",
      "votes": null
    },
    {
      "id": "2003970",
      "postDate": "10/25/2022 23:36:33",
      "content": "<p>Way to go!  </p>",
      "rawMarkdown": "Way to go!",
      "votes": null
    },
    {
      "id": "2006007",
      "postDate": "10/27/2022 11:51:07",
      "content": "<p>I don't trust the 4th decimal place, so that's a lot of potential churn in the top 40 :/</p>",
      "rawMarkdown": "I don't trust the 4th decimal place, so that's a lot of potential churn in the top 40 :/",
      "votes": null
    },
    {
      "id": "2006037",
      "postDate": "10/27/2022 12:08:25",
      "content": "<p>Yep, that sounds like a true story, especially for neural nets :))</p>",
      "rawMarkdown": "Yep, that sounds like a true story, especially for neural nets :))",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2001314,
      "author_name": "michaeljf",
      "author_url": "",
      "post_date": "10/23/2022 23:33:53",
      "content": "<p><a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">@alexryzhkov</a> I like the topic and I am glad you posted it.  I have been thinking about this too and have an additional consideration to add to the conversation.  </p>\n<p>Some people believe it is nearly impossible to overfit with a NN.  What do you think?</p>\n<p><a href=\"https://www.kaggle.com/TomasVdB\" target=\"_blank\">@TomasVdB</a>, <a href=\"https://www.kaggle.com/fabianbong\" target=\"_blank\">@fabianbong</a> <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> <a href=\"https://www.kaggle.com/samuelcortinhas\" target=\"_blank\">@samuelcortinhas</a>, <a href=\"https://www.kaggle.com/hiro5299834\" target=\"_blank\">@hiro5299834</a>, &amp; <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> - What are your thoughts?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2001592,
          "author_name": "tomasvdb",
          "author_url": "",
          "post_date": "10/24/2022 06:52:36",
          "content": "<p>I wouldn't go as far as to say overfitting isn't possible, but it definitely takes a long time before it seems to happen on this dataset. I think I've gotten lucky in this competition by hitting on an LR, model architecture and epochs combination that suddenly improved my score quite a bit. I actually use less features in my later models than in my earlier ones which seemed to make the model more robust. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2001839,
          "author_name": "samuelcortinhas",
          "author_url": "",
          "post_date": "10/24/2022 10:35:32",
          "content": "<p>My predictions of shakeups have always been bad so I try not to predict it anymore haha</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2002640,
          "author_name": "michaeljf",
          "author_url": "",
          "post_date": "10/25/2022 01:18:03",
          "content": "<p><a href=\"https://www.kaggle.com/tomasvdb\" target=\"_blank\">@tomasvdb</a> I started to look into a feature importance approach for NN.  I am finding LIME, SHAP, &amp; permutation approaches.  However, I have not settled on one to try.  Any suggestions?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2002642,
          "author_name": "michaeljf",
          "author_url": "",
          "post_date": "10/25/2022 01:19:24",
          "content": "<p>Also, what I have noticed is using a larger layer number vs more layers has been working for me.  It also takes longer to run, too.  😜</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2003817,
          "author_name": "fabianbong",
          "author_url": "",
          "post_date": "10/25/2022 19:45:28",
          "content": "<p>Just saw this! Personally this is my first competition that I have been part of from beginning to end so your guess is as good as mine :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2003970,
          "author_name": "michaeljf",
          "author_url": "",
          "post_date": "10/25/2022 23:36:33",
          "content": "<p>Way to go!  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2006007,
      "author_name": "paddykb",
      "author_url": "",
      "post_date": "10/27/2022 11:51:07",
      "content": "<p>I don't trust the 4th decimal place, so that's a lot of potential churn in the top 40 :/</p>",
      "votes": null,
      "replies": [
        {
          "id": 2006037,
          "author_name": "alexryzhkov",
          "author_url": "",
          "post_date": "10/27/2022 12:08:25",
          "content": "<p>Yep, that sounds like a true story, especially for neural nets :))</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2001291": "It is 8 days before this competition ends and it's a great time to see what we get for now and what we can receive on the private LB. There are 2 important points to mention:\n- The dataset for this competition is huge so our models can get all the necessary insights from the data to receive good scores. This usually goes to almost no shakeup at the end\n- The public/private split is 24%/76% so the private data is more than 3 times bigger than public. This can lead for the great shakeup if the majority of people use the LB scores to check their models instead of proper validation techniques.\n\nMy prediction is that the first point will beat the second one (almost no shakeup).\n\nBut what do you think about it?",
    "2001314": "alexryzhkov I like the topic and I am glad you posted it.  I have been thinking about this too and have an additional consideration to add to the conversation.  \n\nSome people believe it is nearly impossible to overfit with a NN.  What do you think?\n\n@TomasVdB, @fabianbong @sergiosaharovskiy @samuelcortinhas, @hiro5299834, & @paddykb - What are your thoughts?",
    "2001592": "I wouldn't go as far as to say overfitting isn't possible, but it definitely takes a long time before it seems to happen on this dataset. I think I've gotten lucky in this competition by hitting on an LR, model architecture and epochs combination that suddenly improved my score quite a bit. I actually use less features in my later models than in my earlier ones which seemed to make the model more robust.",
    "2001839": "My predictions of shakeups have always been bad so I try not to predict it anymore haha",
    "2002640": "tomasvdb I started to look into a feature importance approach for NN.  I am finding LIME, SHAP, & permutation approaches.  However, I have not settled on one to try.  Any suggestions?",
    "2002642": "Also, what I have noticed is using a larger layer number vs more layers has been working for me.  It also takes longer to run, too.  😜",
    "2003817": "Just saw this! Personally this is my first competition that I have been part of from beginning to end so your guess is as good as mine :)",
    "2003970": "Way to go!",
    "2006007": "I don't trust the 4th decimal place, so that's a lot of potential churn in the top 40 :/",
    "2006037": "Yep, that sounds like a true story, especially for neural nets :))"
  },
  "source": "meta"
}