{
  "id": 536441,
  "title": "Discussion: The Impact of Parameters and Random Seeds on Competition Performance ﻿| UPDATED| FINAL SUBMISSION discussion",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/536441",
  "author_name": "",
  "post_date": "2024-09-27T15:14:16.396710Z",
  "votes": 23,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In this competition, it became clear that the results were <strong>highly sensitive</strong> to <strong>parameter tuning</strong> and the choice of <strong>random seeds</strong>. Even small changes in these factors caused noticeable differences in <strong>leaderboard scores</strong>, showing that the model was significantly affected by them.<br>\n﻿<br>\nThis means that the model's <strong>robustness</strong> was somewhat compromised, as it relied heavily on precise <strong>parameter adjustments</strong> for optimal performance. While this highlights the need for thorough <strong>hyperparameter tuning</strong>, it also raises concerns about how well the model might <strong>generalize</strong> to new, unseen data where such fine-tuning isn’t possible.<br>\n﻿<br>\nThe impact of <strong>random seed initialization</strong> was especially striking. Running the same notebook with different seeds resulted in <strong>leaderboard scores</strong> ranging from <strong>0.452 to 0.471</strong>, even with all other settings kept constant. This suggests that the model's performance was not <strong>stable</strong> and could be <strong>overfitting</strong> to specific patterns determined by certain seeds. Interestingly, the score variations caused by changing seeds (<strong>±0.02</strong>) were similar to the improvements I achieved through <strong>feature engineering</strong>, making it difficult to determine which changes were truly effective.<br>\n﻿<br>\nThese observations suggest that the competition results may be influenced more by <strong>luck</strong> than by actual <strong>modeling skills</strong>. A more robust <strong>evaluation framework</strong> could help address this issue, such as by averaging results across multiple runs with different seeds or by focusing on how well models <strong>generalize</strong> to diverse testing scenarios.<br>\n﻿<br>\nWe welcome everyone's thoughts and insights on this matter, as a broader discussion could help improve the <strong>fairness</strong> and <strong>reliability</strong> of future competitions. And if you agree, an <strong>upvote</strong> would be greatly appreciated!🤩</p>\n<hr>\n<p>As the final submission date is approaching, does anyone have any thoughts on the final submission?<br>\n Should we aim to integrate as many methods as possible？focus on achieving the highest LB score, or prioritize the highest CV performance? </p>\n<hr>\n<p>Share ur opinions below!!</p>",
  "messages": [
    {
      "id": "3000352",
      "postDate": "09/27/2024 15:14:16",
      "content": "<p>In this competition, it became clear that the results were <strong>highly sensitive</strong> to <strong>parameter tuning</strong> and the choice of <strong>random seeds</strong>. Even small changes in these factors caused noticeable differences in <strong>leaderboard scores</strong>, showing that the model was significantly affected by them.<br>\n﻿<br>\nThis means that the model's <strong>robustness</strong> was somewhat compromised, as it relied heavily on precise <strong>parameter adjustments</strong> for optimal performance. While this highlights the need for thorough <strong>hyperparameter tuning</strong>, it also raises concerns about how well the model might <strong>generalize</strong> to new, unseen data where such fine-tuning isn’t possible.<br>\n﻿<br>\nThe impact of <strong>random seed initialization</strong> was especially striking. Running the same notebook with different seeds resulted in <strong>leaderboard scores</strong> ranging from <strong>0.452 to 0.471</strong>, even with all other settings kept constant. This suggests that the model's performance was not <strong>stable</strong> and could be <strong>overfitting</strong> to specific patterns determined by certain seeds. Interestingly, the score variations caused by changing seeds (<strong>±0.02</strong>) were similar to the improvements I achieved through <strong>feature engineering</strong>, making it difficult to determine which changes were truly effective.<br>\n﻿<br>\nThese observations suggest that the competition results may be influenced more by <strong>luck</strong> than by actual <strong>modeling skills</strong>. A more robust <strong>evaluation framework</strong> could help address this issue, such as by averaging results across multiple runs with different seeds or by focusing on how well models <strong>generalize</strong> to diverse testing scenarios.<br>\n﻿<br>\nWe welcome everyone's thoughts and insights on this matter, as a broader discussion could help improve the <strong>fairness</strong> and <strong>reliability</strong> of future competitions. And if you agree, an <strong>upvote</strong> would be greatly appreciated!🤩</p>\n<hr>\n<p>As the final submission date is approaching, does anyone have any thoughts on the final submission?<br>\n Should we aim to integrate as many methods as possible？focus on achieving the highest LB score, or prioritize the highest CV performance? </p>\n<hr>\n<p>Share ur opinions below!!</p>",
      "rawMarkdown": "In this competition, it became clear that the results were **highly sensitive** to **parameter tuning** and the choice of **random seeds**. Even small changes in these factors caused noticeable differences in **leaderboard scores**, showing that the model was significantly affected by them.\n﻿\nThis means that the model's **robustness** was somewhat compromised, as it relied heavily on precise **parameter adjustments** for optimal performance. While this highlights the need for thorough **hyperparameter tuning**, it also raises concerns about how well the model might **generalize** to new, unseen data where such fine-tuning isn’t possible.\n﻿\nThe impact of **random seed initialization** was especially striking. Running the same notebook with different seeds resulted in **leaderboard scores** ranging from **0.452 to 0.471**, even with all other settings kept constant. This suggests that the model's performance was not **stable** and could be **overfitting** to specific patterns determined by certain seeds. Interestingly, the score variations caused by changing seeds (**±0.02**) were similar to the improvements I achieved through **feature engineering**, making it difficult to determine which changes were truly effective.\n﻿\nThese observations suggest that the competition results may be influenced more by **luck** than by actual **modeling skills**. A more robust **evaluation framework** could help address this issue, such as by averaging results across multiple runs with different seeds or by focusing on how well models **generalize** to diverse testing scenarios.\n﻿\nWe welcome everyone's thoughts and insights on this matter, as a broader discussion could help improve the **fairness** and **reliability** of future competitions. And if you agree, an **upvote** would be greatly appreciated!🤩\n\n________________________________________________________________________________________________________________________________\n\nAs the final submission date is approaching, does anyone have any thoughts on the final submission?\n Should we aim to integrate as many methods as possible？focus on achieving the highest LB score, or prioritize the highest CV performance? \n_________________\nShare ur opinions below!!",
      "votes": null
    },
    {
      "id": "3000916",
      "postDate": "09/28/2024 08:47:28",
      "content": "<p>I think so too. Changing the seed jitter even reached 0.04 (0.417~0.453). The jitter issue in this competition is even more severe than in um-mcts. The noise and missing values inherent in the data are a significant problem. By the way, how did you handle these missing values?</p>",
      "rawMarkdown": "I think so too. Changing the seed jitter even reached 0.04 (0.417~0.453). The jitter issue in this competition is even more severe than in um-mcts. The noise and missing values inherent in the data are a significant problem. By the way, how did you handle these missing values?",
      "votes": null
    },
    {
      "id": "3001570",
      "postDate": "09/29/2024 02:15:05",
      "content": "<p>That’s how small data works! It's probably an econometric method problem. </p>",
      "rawMarkdown": "That’s how small data works! It's probably an econometric method problem.",
      "votes": null
    },
    {
      "id": "3032762",
      "postDate": "10/31/2024 10:47:12",
      "content": "<p>It will be important not to give up to use an approach after getting low validation scores a few times.</p>",
      "rawMarkdown": "It will be important not to give up to use an approach after getting low validation scores a few times.",
      "votes": null
    },
    {
      "id": "3038252",
      "postDate": "11/06/2024 18:56:56",
      "content": "<p><a href=\"https://www.kaggle.com/wayne127\" target=\"_blank\">@wayne127</a> yes  i experiment the same issues with the submitting process , that lead to a huge shaking into the final , but even though the skillest competitors will be the winners such we must they know how to handle those issues effecientelly</p>",
      "rawMarkdown": "wayne127 yes  i experiment the same issues with the submitting process , that lead to a huge shaking into the final , but even though the skillest competitors will be the winners such we must they know how to handle those issues effecientelly",
      "votes": null
    },
    {
      "id": "3056606",
      "postDate": "11/27/2024 07:01:19",
      "content": "<p>I noticed that there are hardly any Kaggle GMs among the top ranks on the LB for this competition, which might suggest that this competition isn't as successful as expected 😭</p>",
      "rawMarkdown": "I noticed that there are hardly any Kaggle GMs among the top ranks on the LB for this competition, which might suggest that this competition isn't as successful as expected 😭",
      "votes": null
    },
    {
      "id": "3056609",
      "postDate": "11/27/2024 07:03:05",
      "content": "<p>Yes, I believe the final leaderboard shakeup will bring a pleasant surprise to those heroes who persevere until the end.</p>",
      "rawMarkdown": "Yes, I believe the final leaderboard shakeup will bring a pleasant surprise to those heroes who persevere until the end.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3000916,
      "author_name": "wscmisx",
      "author_url": "",
      "post_date": "09/28/2024 08:47:28",
      "content": "<p>I think so too. Changing the seed jitter even reached 0.04 (0.417~0.453). The jitter issue in this competition is even more severe than in um-mcts. The noise and missing values inherent in the data are a significant problem. By the way, how did you handle these missing values?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3001570,
      "author_name": "thiagobgil",
      "author_url": "",
      "post_date": "09/29/2024 02:15:05",
      "content": "<p>That’s how small data works! It's probably an econometric method problem. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3032762,
      "author_name": "ykawakita",
      "author_url": "",
      "post_date": "10/31/2024 10:47:12",
      "content": "<p>It will be important not to give up to use an approach after getting low validation scores a few times.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3056609,
          "author_name": "wayne127",
          "author_url": "",
          "post_date": "11/27/2024 07:03:05",
          "content": "<p>Yes, I believe the final leaderboard shakeup will bring a pleasant surprise to those heroes who persevere until the end.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3038252,
      "author_name": "saidkoussi",
      "author_url": "",
      "post_date": "11/06/2024 18:56:56",
      "content": "<p><a href=\"https://www.kaggle.com/wayne127\" target=\"_blank\">@wayne127</a> yes  i experiment the same issues with the submitting process , that lead to a huge shaking into the final , but even though the skillest competitors will be the winners such we must they know how to handle those issues effecientelly</p>",
      "votes": null,
      "replies": [
        {
          "id": 3056606,
          "author_name": "wayne127",
          "author_url": "",
          "post_date": "11/27/2024 07:01:19",
          "content": "<p>I noticed that there are hardly any Kaggle GMs among the top ranks on the LB for this competition, which might suggest that this competition isn't as successful as expected 😭</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3000352": "In this competition, it became clear that the results were **highly sensitive** to **parameter tuning** and the choice of **random seeds**. Even small changes in these factors caused noticeable differences in **leaderboard scores**, showing that the model was significantly affected by them.\n﻿\nThis means that the model's **robustness** was somewhat compromised, as it relied heavily on precise **parameter adjustments** for optimal performance. While this highlights the need for thorough **hyperparameter tuning**, it also raises concerns about how well the model might **generalize** to new, unseen data where such fine-tuning isn’t possible.\n﻿\nThe impact of **random seed initialization** was especially striking. Running the same notebook with different seeds resulted in **leaderboard scores** ranging from **0.452 to 0.471**, even with all other settings kept constant. This suggests that the model's performance was not **stable** and could be **overfitting** to specific patterns determined by certain seeds. Interestingly, the score variations caused by changing seeds (**±0.02**) were similar to the improvements I achieved through **feature engineering**, making it difficult to determine which changes were truly effective.\n﻿\nThese observations suggest that the competition results may be influenced more by **luck** than by actual **modeling skills**. A more robust **evaluation framework** could help address this issue, such as by averaging results across multiple runs with different seeds or by focusing on how well models **generalize** to diverse testing scenarios.\n﻿\nWe welcome everyone's thoughts and insights on this matter, as a broader discussion could help improve the **fairness** and **reliability** of future competitions. And if you agree, an **upvote** would be greatly appreciated!🤩\n\n________________________________________________________________________________________________________________________________\n\nAs the final submission date is approaching, does anyone have any thoughts on the final submission?\n Should we aim to integrate as many methods as possible？focus on achieving the highest LB score, or prioritize the highest CV performance? \n_________________\nShare ur opinions below!!",
    "3000916": "I think so too. Changing the seed jitter even reached 0.04 (0.417~0.453). The jitter issue in this competition is even more severe than in um-mcts. The noise and missing values inherent in the data are a significant problem. By the way, how did you handle these missing values?",
    "3001570": "That’s how small data works! It's probably an econometric method problem.",
    "3032762": "It will be important not to give up to use an approach after getting low validation scores a few times.",
    "3038252": "wayne127 yes  i experiment the same issues with the submitting process , that lead to a huge shaking into the final , but even though the skillest competitors will be the winners such we must they know how to handle those issues effecientelly",
    "3056606": "I noticed that there are hardly any Kaggle GMs among the top ranks on the LB for this competition, which might suggest that this competition isn't as successful as expected 😭",
    "3056609": "Yes, I believe the final leaderboard shakeup will bring a pleasant surprise to those heroes who persevere until the end."
  },
  "source": "meta"
}