{
  "id": 549795,
  "title": "The impact of random seeds can be much bigger than the minor changes of approaches.",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/549795",
  "author_name": "",
  "post_date": "2024-12-03T23:37:28.074453800Z",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I tried to find out how random seeds affect on LB score, using Yu Yang Chang's notebook: <a href=\"https://www.kaggle.com/code/cchangyyy/0-494-notebook\" target=\"_blank\">https://www.kaggle.com/code/cchangyyy/0-494-notebook</a></p>\n<p>I changed two kinds of random seeds, which are originally, SEED=42(In [2]) and seed_everything(2024) (In[3]). And, as the input to first model I tried two changes of approachs: adding CGAS-Season and adding all season categories.</p>\n<p>The result is shown below.</p>\n<table>\n<thead>\n<tr>\n<th>seeds</th>\n<th>without changes</th>\n<th>use CGAS-Season</th>\n<th>use all season categories</th>\n<th><em>standard deviation</em></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>42, 2024</td>\n<td>0.494</td>\n<td>0.484</td>\n<td>0.487</td>\n<td>0.0042</td>\n</tr>\n<tr>\n<td>43, 2025</td>\n<td>0.454</td>\n<td>0.445</td>\n<td>0.446</td>\n<td>0.0040</td>\n</tr>\n<tr>\n<td>44, 2026</td>\n<td>0.431</td>\n<td>0.432</td>\n<td>0.432</td>\n<td>0.00047</td>\n</tr>\n<tr>\n<td>45, 2027</td>\n<td>0.470</td>\n<td>0.474</td>\n<td>0.471</td>\n<td>0.0017</td>\n</tr>\n<tr>\n<td><em>standard deviation</em></td>\n<td>0.023</td>\n<td>0.021</td>\n<td>0.021</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>It means that changes in random seeds can cause bigger effects than minor changes of approaches.</p>\n<p>Probably, LB score is useless to see the minor changes' impact.</p>",
  "messages": [
    {
      "id": "3062814",
      "postDate": "12/03/2024 23:37:28",
      "content": "<p>I tried to find out how random seeds affect on LB score, using Yu Yang Chang's notebook: <a href=\"https://www.kaggle.com/code/cchangyyy/0-494-notebook\" target=\"_blank\">https://www.kaggle.com/code/cchangyyy/0-494-notebook</a></p>\n<p>I changed two kinds of random seeds, which are originally, SEED=42(In [2]) and seed_everything(2024) (In[3]). And, as the input to first model I tried two changes of approachs: adding CGAS-Season and adding all season categories.</p>\n<p>The result is shown below.</p>\n<table>\n<thead>\n<tr>\n<th>seeds</th>\n<th>without changes</th>\n<th>use CGAS-Season</th>\n<th>use all season categories</th>\n<th><em>standard deviation</em></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>42, 2024</td>\n<td>0.494</td>\n<td>0.484</td>\n<td>0.487</td>\n<td>0.0042</td>\n</tr>\n<tr>\n<td>43, 2025</td>\n<td>0.454</td>\n<td>0.445</td>\n<td>0.446</td>\n<td>0.0040</td>\n</tr>\n<tr>\n<td>44, 2026</td>\n<td>0.431</td>\n<td>0.432</td>\n<td>0.432</td>\n<td>0.00047</td>\n</tr>\n<tr>\n<td>45, 2027</td>\n<td>0.470</td>\n<td>0.474</td>\n<td>0.471</td>\n<td>0.0017</td>\n</tr>\n<tr>\n<td><em>standard deviation</em></td>\n<td>0.023</td>\n<td>0.021</td>\n<td>0.021</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>It means that changes in random seeds can cause bigger effects than minor changes of approaches.</p>\n<p>Probably, LB score is useless to see the minor changes' impact.</p>",
      "rawMarkdown": "I tried to find out how random seeds affect on LB score, using Yu Yang Chang's notebook: https://www.kaggle.com/code/cchangyyy/0-494-notebook\n\nI changed two kinds of random seeds, which are originally, SEED=42(In [2]) and seed_everything(2024) (In[3]). And, as the input to first model I tried two changes of approachs: adding CGAS-Season and adding all season categories.\n\nThe result is shown below.\n\n| seeds | without changes | use CGAS-Season | use all season categories | *standard deviation* |\n| --- | --- | --- |\n| 42, 2024 | 0.494 | 0.484 | 0.487 | 0.0042 |\n| 43, 2025 | 0.454 | 0.445 | 0.446 | 0.0040 |\n| 44, 2026 | 0.431 | 0.432 | 0.432 | 0.00047 |\n| 45, 2027 | 0.470 | 0.474 | 0.471 | 0.0017 |\n| *standard deviation* | 0.023 | 0.021 | 0.021 |  |\n\nIt means that changes in random seeds can cause bigger effects than minor changes of approaches.\n\nProbably, LB score is useless to see the minor changes' impact.",
      "votes": null
    },
    {
      "id": "3062951",
      "postDate": "12/04/2024 03:43:20",
      "content": "<p>Similar result here! It seems like the big difference between different seeds shows that the LB version is useless. I tried submitting different seed versions and it started giving me values similar to yours.</p>",
      "rawMarkdown": "Similar result here! It seems like the big difference between different seeds shows that the LB version is useless. I tried submitting different seed versions and it started giving me values similar to yours.",
      "votes": null
    },
    {
      "id": "3063199",
      "postDate": "12/04/2024 08:45:21",
      "content": "<p>Yes, it will be difficult to predict our final ranking.</p>",
      "rawMarkdown": "Yes, it will be difficult to predict our final ranking.",
      "votes": null
    },
    {
      "id": "3064786",
      "postDate": "12/06/2024 02:17:14",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": null
    },
    {
      "id": "3064943",
      "postDate": "12/06/2024 06:48:49",
      "content": "<p>I think there must be some impossible sii values so the models are unstable. Parents will react differently. For example two students with all other fields the same, but one has sii=0 and the other has sii=3</p>",
      "rawMarkdown": "I think there must be some impossible sii values so the models are unstable. Parents will react differently. For example two students with all other fields the same, but one has sii=0 and the other has sii=3",
      "votes": null
    },
    {
      "id": "3064983",
      "postDate": "12/06/2024 07:34:42",
      "content": "<p>That's true. Probably, some data about the parents will be useful.</p>",
      "rawMarkdown": "That's true. Probably, some data about the parents will be useful.",
      "votes": null
    },
    {
      "id": "3065211",
      "postDate": "12/06/2024 14:37:57",
      "content": "<p>such randomness </p>",
      "rawMarkdown": "such randomness",
      "votes": null
    },
    {
      "id": "3065987",
      "postDate": "12/07/2024 14:14:13",
      "content": "<p>I also obtained similar results. And when I improve the performance of one individual model in my ensemble model, it always performs worse than the original results, which means that my LB score is somehow just luck. </p>",
      "rawMarkdown": "I also obtained similar results. And when I improve the performance of one individual model in my ensemble model, it always performs worse than the original results, which means that my LB score is somehow just luck.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3062951,
      "author_name": "alperenduru",
      "author_url": "",
      "post_date": "12/04/2024 03:43:20",
      "content": "<p>Similar result here! It seems like the big difference between different seeds shows that the LB version is useless. I tried submitting different seed versions and it started giving me values similar to yours.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3063199,
          "author_name": "ykawakita",
          "author_url": "",
          "post_date": "12/04/2024 08:45:21",
          "content": "<p>Yes, it will be difficult to predict our final ranking.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3064786,
      "author_name": "shanyun",
      "author_url": "",
      "post_date": "12/06/2024 02:17:14",
      "content": "<p>Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3064943,
      "author_name": "lawrencechernin",
      "author_url": "",
      "post_date": "12/06/2024 06:48:49",
      "content": "<p>I think there must be some impossible sii values so the models are unstable. Parents will react differently. For example two students with all other fields the same, but one has sii=0 and the other has sii=3</p>",
      "votes": null,
      "replies": [
        {
          "id": 3064983,
          "author_name": "ykawakita",
          "author_url": "",
          "post_date": "12/06/2024 07:34:42",
          "content": "<p>That's true. Probably, some data about the parents will be useful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3065211,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "12/06/2024 14:37:57",
      "content": "<p>such randomness </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3065987,
      "author_name": "desolade",
      "author_url": "",
      "post_date": "12/07/2024 14:14:13",
      "content": "<p>I also obtained similar results. And when I improve the performance of one individual model in my ensemble model, it always performs worse than the original results, which means that my LB score is somehow just luck. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3062814": "I tried to find out how random seeds affect on LB score, using Yu Yang Chang's notebook: https://www.kaggle.com/code/cchangyyy/0-494-notebook\n\nI changed two kinds of random seeds, which are originally, SEED=42(In [2]) and seed_everything(2024) (In[3]). And, as the input to first model I tried two changes of approachs: adding CGAS-Season and adding all season categories.\n\nThe result is shown below.\n\n| seeds | without changes | use CGAS-Season | use all season categories | *standard deviation* |\n| --- | --- | --- |\n| 42, 2024 | 0.494 | 0.484 | 0.487 | 0.0042 |\n| 43, 2025 | 0.454 | 0.445 | 0.446 | 0.0040 |\n| 44, 2026 | 0.431 | 0.432 | 0.432 | 0.00047 |\n| 45, 2027 | 0.470 | 0.474 | 0.471 | 0.0017 |\n| *standard deviation* | 0.023 | 0.021 | 0.021 |  |\n\nIt means that changes in random seeds can cause bigger effects than minor changes of approaches.\n\nProbably, LB score is useless to see the minor changes' impact.",
    "3062951": "Similar result here! It seems like the big difference between different seeds shows that the LB version is useless. I tried submitting different seed versions and it started giving me values similar to yours.",
    "3063199": "Yes, it will be difficult to predict our final ranking.",
    "3064786": "Thank you for sharing!",
    "3064943": "I think there must be some impossible sii values so the models are unstable. Parents will react differently. For example two students with all other fields the same, but one has sii=0 and the other has sii=3",
    "3064983": "That's true. Probably, some data about the parents will be useful.",
    "3065211": "such randomness",
    "3065987": "I also obtained similar results. And when I improve the performance of one individual model in my ensemble model, it always performs worse than the original results, which means that my LB score is somehow just luck."
  },
  "source": "meta"
}