{
  "id": 547433,
  "title": "Will the same result as ISIC2024 be repeated again?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/547433",
  "author_name": "",
  "post_date": "2024-11-21T14:13:52.333681200Z",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>The private LB of ISIC2024 showed a huge shake-up and changed the rank. Rethinking about the public LB of ISIC2024, the score around 0.185 contains nearly 400 teams, which is similar to this competition. So do you think the huge shake-up will repeat again?</p>",
  "messages": [
    {
      "id": "3051668",
      "postDate": "11/21/2024 14:13:52",
      "content": "<p>The private LB of ISIC2024 showed a huge shake-up and changed the rank. Rethinking about the public LB of ISIC2024, the score around 0.185 contains nearly 400 teams, which is similar to this competition. So do you think the huge shake-up will repeat again?</p>",
      "rawMarkdown": "The private LB of ISIC2024 showed a huge shake-up and changed the rank. Rethinking about the public LB of ISIC2024, the score around 0.185 contains nearly 400 teams, which is similar to this competition. So do you think the huge shake-up will repeat again?",
      "votes": null
    },
    {
      "id": "3051744",
      "postDate": "11/21/2024 15:41:43",
      "content": "<p>Mostly yes <a href=\"https://www.kaggle.com/qufangcq\" target=\"_blank\">@qufangcq</a> <br>\nThis is most likely to happen here, perhaps more than ISIC </p>",
      "rawMarkdown": "Mostly yes @qufangcq \nThis is most likely to happen here, perhaps more than ISIC",
      "votes": null
    },
    {
      "id": "3051753",
      "postDate": "11/21/2024 15:52:05",
      "content": "<p>The question for me is: How do you choose your final model? All my models seems very sensitive to random seed/fold choices. Even if the model is overall solid ~0.46 it can be 0.39 at its worst fold and 0.51 at its best.</p>",
      "rawMarkdown": "The question for me is: How do you choose your final model? All my models seems very sensitive to random seed/fold choices. Even if the model is overall solid ~0.46 it can be 0.39 at its worst fold and 0.51 at its best.",
      "votes": null
    },
    {
      "id": "3051789",
      "postDate": "11/21/2024 16:21:14",
      "content": "<p>Thank you for your insight. Do you think the missing values and some unreliable data are the main reasons?</p>",
      "rawMarkdown": "Thank you for your insight. Do you think the missing values and some unreliable data are the main reasons?",
      "votes": null
    },
    {
      "id": "3051794",
      "postDate": "11/21/2024 16:23:07",
      "content": "<p>Thank you. It seems that luck is also what we need. Good luck.</p>",
      "rawMarkdown": "Thank you. It seems that luck is also what we need. Good luck.",
      "votes": null
    },
    {
      "id": "3051937",
      "postDate": "11/21/2024 19:26:50",
      "content": "<p>That is one of the reasons <a href=\"https://www.kaggle.com/qufangcq\" target=\"_blank\">@qufangcq</a> </p>",
      "rawMarkdown": "That is one of the reasons @qufangcq",
      "votes": null
    },
    {
      "id": "3054705",
      "postDate": "11/25/2024 02:57:42",
      "content": "<p>Indeed, I guess this is a luck competition.🤔</p>",
      "rawMarkdown": "Indeed, I guess this is a luck competition.🤔",
      "votes": null
    },
    {
      "id": "3054950",
      "postDate": "11/25/2024 11:13:57",
      "content": "<p>Yes, Chances are very high for this to happen. Many people who are high on the LB are just copy pasting a solution which has incorrectly applied KNNImputer and Encoding. Under this scenario, it feels that chances are high for a big shakeup. </p>",
      "rawMarkdown": "Yes, Chances are very high for this to happen. Many people who are high on the LB are just copy pasting a solution which has incorrectly applied KNNImputer and Encoding. Under this scenario, it feels that chances are high for a big shakeup.",
      "votes": null
    },
    {
      "id": "3054956",
      "postDate": "11/25/2024 11:21:01",
      "content": "<p>Yes, I noticed that. Especially the usage of auto-encoder seems to be incorrect.</p>",
      "rawMarkdown": "Yes, I noticed that. Especially the usage of auto-encoder seems to be incorrect.",
      "votes": null
    },
    {
      "id": "3054966",
      "postDate": "11/25/2024 11:32:42",
      "content": "<p>Yeah! and KNNImputer is applied on the target variable sii also. KNNImputer can't be applied to target variables as it leads to data leakage.</p>",
      "rawMarkdown": "Yeah! and KNNImputer is applied on the target variable sii also. KNNImputer can't be applied to target variables as it leads to data leakage.",
      "votes": null
    },
    {
      "id": "3055294",
      "postDate": "11/25/2024 17:09:19",
      "content": "<p>I think it's highly likely that this will occur.</p>",
      "rawMarkdown": "I think it's highly likely that this will occur.",
      "votes": null
    },
    {
      "id": "3055463",
      "postDate": "11/25/2024 19:21:04",
      "content": "<p>the two submissions I'm planning to select are:</p>\n<ul>\n<li>model with highest mean QWK across CV folds</li>\n<li>model with lowest coefficient of variation of QWK across CV folds that is near the highest mean</li>\n</ul>\n<p>the hypothesis being that the lowest coefficient of variation should be the most 'stable' and robust against whatever the private test set has to offer..</p>\n<p>would be curious to know if there are better strategies :)</p>",
      "rawMarkdown": "the two submissions I'm planning to select are:\n\n- model with highest mean QWK across CV folds\n- model with lowest coefficient of variation of QWK across CV folds that is near the highest mean\n\nthe hypothesis being that the lowest coefficient of variation should be the most 'stable' and robust against whatever the private test set has to offer..\n\nwould be curious to know if there are better strategies :)",
      "votes": null
    },
    {
      "id": "3055482",
      "postDate": "11/25/2024 19:52:08",
      "content": "<p>My strategy is basically the same, but I also take public leaderboard into the equation. Since it is 38% of the total database, it is also 38% of the final score, right?</p>",
      "rawMarkdown": "My strategy is basically the same, but I also take public leaderboard into the equation. Since it is 38% of the total database, it is also 38% of the final score, right?",
      "votes": null
    },
    {
      "id": "3056787",
      "postDate": "11/27/2024 11:39:41",
      "content": "<p>I haven't read through all the details of the high scoring notebook. At first glance an autoencoder for the actigraphy time series seems reasonable to me?<br>\nIs this what you meant by encoding or something else?<br>\nThanks for the insight anyway!</p>",
      "rawMarkdown": "I haven't read through all the details of the high scoring notebook. At first glance an autoencoder for the actigraphy time series seems reasonable to me?\nIs this what you meant by encoding or something else?\nThanks for the insight anyway!",
      "votes": null
    },
    {
      "id": "3059070",
      "postDate": "11/30/2024 09:38:06",
      "content": "<p>The use of an autoencoder in general isn't unreasonable, but it's implemented incorrectly making learned patterns meaningless for the test-data. </p>",
      "rawMarkdown": "The use of an autoencoder in general isn't unreasonable, but it's implemented incorrectly making learned patterns meaningless for the test-data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3051744,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "11/21/2024 15:41:43",
      "content": "<p>Mostly yes <a href=\"https://www.kaggle.com/qufangcq\" target=\"_blank\">@qufangcq</a> <br>\nThis is most likely to happen here, perhaps more than ISIC </p>",
      "votes": null,
      "replies": [
        {
          "id": 3051789,
          "author_name": "qufangcq",
          "author_url": "",
          "post_date": "11/21/2024 16:21:14",
          "content": "<p>Thank you for your insight. Do you think the missing values and some unreliable data are the main reasons?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3051937,
              "author_name": "ravi20076",
              "author_url": "",
              "post_date": "11/21/2024 19:26:50",
              "content": "<p>That is one of the reasons <a href=\"https://www.kaggle.com/qufangcq\" target=\"_blank\">@qufangcq</a> </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3051753,
      "author_name": "mariusheuser",
      "author_url": "",
      "post_date": "11/21/2024 15:52:05",
      "content": "<p>The question for me is: How do you choose your final model? All my models seems very sensitive to random seed/fold choices. Even if the model is overall solid ~0.46 it can be 0.39 at its worst fold and 0.51 at its best.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3051794,
          "author_name": "qufangcq",
          "author_url": "",
          "post_date": "11/21/2024 16:23:07",
          "content": "<p>Thank you. It seems that luck is also what we need. Good luck.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3054705,
              "author_name": "fredrickunderwood",
              "author_url": "",
              "post_date": "11/25/2024 02:57:42",
              "content": "<p>Indeed, I guess this is a luck competition.🤔</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3055463,
          "author_name": "josephmarturano",
          "author_url": "",
          "post_date": "11/25/2024 19:21:04",
          "content": "<p>the two submissions I'm planning to select are:</p>\n<ul>\n<li>model with highest mean QWK across CV folds</li>\n<li>model with lowest coefficient of variation of QWK across CV folds that is near the highest mean</li>\n</ul>\n<p>the hypothesis being that the lowest coefficient of variation should be the most 'stable' and robust against whatever the private test set has to offer..</p>\n<p>would be curious to know if there are better strategies :)</p>",
          "votes": null,
          "replies": [
            {
              "id": 3055482,
              "author_name": "mariusheuser",
              "author_url": "",
              "post_date": "11/25/2024 19:52:08",
              "content": "<p>My strategy is basically the same, but I also take public leaderboard into the equation. Since it is 38% of the total database, it is also 38% of the final score, right?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3054950,
      "author_name": "taimour",
      "author_url": "",
      "post_date": "11/25/2024 11:13:57",
      "content": "<p>Yes, Chances are very high for this to happen. Many people who are high on the LB are just copy pasting a solution which has incorrectly applied KNNImputer and Encoding. Under this scenario, it feels that chances are high for a big shakeup. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3054956,
          "author_name": "qufangcq",
          "author_url": "",
          "post_date": "11/25/2024 11:21:01",
          "content": "<p>Yes, I noticed that. Especially the usage of auto-encoder seems to be incorrect.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3054966,
              "author_name": "taimour",
              "author_url": "",
              "post_date": "11/25/2024 11:32:42",
              "content": "<p>Yeah! and KNNImputer is applied on the target variable sii also. KNNImputer can't be applied to target variables as it leads to data leakage.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3056787,
          "author_name": "kontoudi",
          "author_url": "",
          "post_date": "11/27/2024 11:39:41",
          "content": "<p>I haven't read through all the details of the high scoring notebook. At first glance an autoencoder for the actigraphy time series seems reasonable to me?<br>\nIs this what you meant by encoding or something else?<br>\nThanks for the insight anyway!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3059070,
              "author_name": "lennarthaupts",
              "author_url": "",
              "post_date": "11/30/2024 09:38:06",
              "content": "<p>The use of an autoencoder in general isn't unreasonable, but it's implemented incorrectly making learned patterns meaningless for the test-data. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3055294,
      "author_name": "humayrakhanomrime",
      "author_url": "",
      "post_date": "11/25/2024 17:09:19",
      "content": "<p>I think it's highly likely that this will occur.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3051668": "The private LB of ISIC2024 showed a huge shake-up and changed the rank. Rethinking about the public LB of ISIC2024, the score around 0.185 contains nearly 400 teams, which is similar to this competition. So do you think the huge shake-up will repeat again?",
    "3051744": "Mostly yes @qufangcq \nThis is most likely to happen here, perhaps more than ISIC",
    "3051753": "The question for me is: How do you choose your final model? All my models seems very sensitive to random seed/fold choices. Even if the model is overall solid ~0.46 it can be 0.39 at its worst fold and 0.51 at its best.",
    "3051789": "Thank you for your insight. Do you think the missing values and some unreliable data are the main reasons?",
    "3051794": "Thank you. It seems that luck is also what we need. Good luck.",
    "3051937": "That is one of the reasons @qufangcq",
    "3054705": "Indeed, I guess this is a luck competition.🤔",
    "3054950": "Yes, Chances are very high for this to happen. Many people who are high on the LB are just copy pasting a solution which has incorrectly applied KNNImputer and Encoding. Under this scenario, it feels that chances are high for a big shakeup.",
    "3054956": "Yes, I noticed that. Especially the usage of auto-encoder seems to be incorrect.",
    "3054966": "Yeah! and KNNImputer is applied on the target variable sii also. KNNImputer can't be applied to target variables as it leads to data leakage.",
    "3055294": "I think it's highly likely that this will occur.",
    "3055463": "the two submissions I'm planning to select are:\n\n- model with highest mean QWK across CV folds\n- model with lowest coefficient of variation of QWK across CV folds that is near the highest mean\n\nthe hypothesis being that the lowest coefficient of variation should be the most 'stable' and robust against whatever the private test set has to offer..\n\nwould be curious to know if there are better strategies :)",
    "3055482": "My strategy is basically the same, but I also take public leaderboard into the equation. Since it is 38% of the total database, it is also 38% of the final score, right?",
    "3056787": "I haven't read through all the details of the high scoring notebook. At first glance an autoencoder for the actigraphy time series seems reasonable to me?\nIs this what you meant by encoding or something else?\nThanks for the insight anyway!",
    "3059070": "The use of an autoencoder in general isn't unreasonable, but it's implemented incorrectly making learned patterns meaningless for the test-data."
  },
  "source": "meta"
}