{
  "id": 545096,
  "title": "CV - LB relationship",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/545096",
  "author_name": "",
  "post_date": "2024-11-08T10:34:26.559368700Z",
  "votes": 13,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Does anyone have a good CV-LB relationship?</p>",
  "messages": [
    {
      "id": "3039743",
      "postDate": "11/08/2024 10:34:26",
      "content": "<p>Does anyone have a good CV-LB relationship?</p>",
      "rawMarkdown": "Does anyone have a good CV-LB relationship?",
      "votes": null
    },
    {
      "id": "3039793",
      "postDate": "11/08/2024 11:41:33",
      "content": "<p>Hey, this is pretty much my first competition, so I might be wrong about how it works, but this is my point of view:</p>\n<p>The LB is completely irrelevant. We make 20 predictions(test data) for it, while the data of these 20 predictions is already part of the training data set (eventhough some information is missing). Just because we can make good prediction on such a small data set, doesn't mean it translates to any other small data set well.<br>\nSo it is all about CV. Having a consistent CV score (mean without alot of variance) is the only important thing.<br>\nOf course you still need some luck to get the right seed / parameters for the ~52 data on the privat LB. Im still confused why it is such a small sample size. I hope I dont misunderstand how it works :)</p>\n<p>My relationship between CV and LB is good, but low. So my CV score matches my LB score. Dont know if that is any good, but most of the LB is probably overfitted.</p>",
      "rawMarkdown": "Hey, this is pretty much my first competition, so I might be wrong about how it works, but this is my point of view:\n\nThe LB is completely irrelevant. We make 20 predictions(test data) for it, while the data of these 20 predictions is already part of the training data set (eventhough some information is missing). Just because we can make good prediction on such a small data set, doesn't mean it translates to any other small data set well.\nSo it is all about CV. Having a consistent CV score (mean without alot of variance) is the only important thing.\nOf course you still need some luck to get the right seed / parameters for the ~52 data on the privat LB. Im still confused why it is such a small sample size. I hope I dont misunderstand how it works :)\n\nMy relationship between CV and LB is good, but low. So my CV score matches my LB score. Dont know if that is any good, but most of the LB is probably overfitted.",
      "votes": null
    },
    {
      "id": "3039809",
      "postDate": "11/08/2024 12:10:40",
      "content": "<p>The LB evaluation is on 1444 out of 3800. </p>",
      "rawMarkdown": "The LB evaluation is on 1444 out of 3800.",
      "votes": null
    },
    {
      "id": "3039831",
      "postDate": "11/08/2024 12:28:50",
      "content": "<p>Ah, so the test.csv is fake and just a placeholder and gets replaced by the real one after submitting? I was confused how it works. But makes sense now. Thank you.<br>\nThen take my post with a grain of salt. Eventhough I still believe in those points. LB is just another fold basically :)</p>",
      "rawMarkdown": "Ah, so the test.csv is fake and just a placeholder and gets replaced by the real one after submitting? I was confused how it works. But makes sense now. Thank you.\nThen take my post with a grain of salt. Eventhough I still believe in those points. LB is just another fold basically :)",
      "votes": null
    },
    {
      "id": "3039906",
      "postDate": "11/08/2024 14:00:51",
      "content": "<p>Short answer <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> <strong>NO</strong></p>",
      "rawMarkdown": "Short answer @abdmental01 **NO**",
      "votes": null
    },
    {
      "id": "3040567",
      "postDate": "11/09/2024 09:34:04",
      "content": "<p>actually CV and LB are not related for me</p>",
      "rawMarkdown": "actually CV and LB are not related for me",
      "votes": null
    },
    {
      "id": "3041459",
      "postDate": "11/10/2024 11:45:17",
      "content": "<p>As per my experience, there is not any strong relationship between CV and LB.</p>",
      "rawMarkdown": "As per my experience, there is not any strong relationship between CV and LB.",
      "votes": null
    },
    {
      "id": "3044626",
      "postDate": "11/13/2024 15:58:16",
      "content": "<p>Maybe not the top LB score, but pretty balanced model in my case:</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.468</td>\n<td>0.447</td>\n</tr>\n<tr>\n<td>0.465</td>\n<td>0.447</td>\n</tr>\n<tr>\n<td>0.470</td>\n<td>0.441</td>\n</tr>\n<tr>\n<td>0.470</td>\n<td>0.460</td>\n</tr>\n<tr>\n<td>0.468</td>\n<td>0.452</td>\n</tr>\n<tr>\n<td>0.475</td>\n<td>0.456</td>\n</tr>\n<tr>\n<td>0.468</td>\n<td>0.445</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Maybe not the top LB score, but pretty balanced model in my case:\n| CV |  LB|\n| --- | --- |\n| 0.468 | 0.447 |\n| 0.465 | 0.447 |\n| 0.470 | 0.441 |\n| 0.470 | 0.460 |\n| 0.468 | 0.452 |\n| 0.475 | 0.456 |\n| 0.468 | 0.445 |",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3039793,
      "author_name": "mariusheuser",
      "author_url": "",
      "post_date": "11/08/2024 11:41:33",
      "content": "<p>Hey, this is pretty much my first competition, so I might be wrong about how it works, but this is my point of view:</p>\n<p>The LB is completely irrelevant. We make 20 predictions(test data) for it, while the data of these 20 predictions is already part of the training data set (eventhough some information is missing). Just because we can make good prediction on such a small data set, doesn't mean it translates to any other small data set well.<br>\nSo it is all about CV. Having a consistent CV score (mean without alot of variance) is the only important thing.<br>\nOf course you still need some luck to get the right seed / parameters for the ~52 data on the privat LB. Im still confused why it is such a small sample size. I hope I dont misunderstand how it works :)</p>\n<p>My relationship between CV and LB is good, but low. So my CV score matches my LB score. Dont know if that is any good, but most of the LB is probably overfitted.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3039809,
          "author_name": "bsmelbs",
          "author_url": "",
          "post_date": "11/08/2024 12:10:40",
          "content": "<p>The LB evaluation is on 1444 out of 3800. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3039831,
              "author_name": "mariusheuser",
              "author_url": "",
              "post_date": "11/08/2024 12:28:50",
              "content": "<p>Ah, so the test.csv is fake and just a placeholder and gets replaced by the real one after submitting? I was confused how it works. But makes sense now. Thank you.<br>\nThen take my post with a grain of salt. Eventhough I still believe in those points. LB is just another fold basically :)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3039906,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "11/08/2024 14:00:51",
      "content": "<p>Short answer <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> <strong>NO</strong></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3040567,
      "author_name": "desolade",
      "author_url": "",
      "post_date": "11/09/2024 09:34:04",
      "content": "<p>actually CV and LB are not related for me</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3041459,
      "author_name": "taimour",
      "author_url": "",
      "post_date": "11/10/2024 11:45:17",
      "content": "<p>As per my experience, there is not any strong relationship between CV and LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3044626,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "11/13/2024 15:58:16",
      "content": "<p>Maybe not the top LB score, but pretty balanced model in my case:</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.468</td>\n<td>0.447</td>\n</tr>\n<tr>\n<td>0.465</td>\n<td>0.447</td>\n</tr>\n<tr>\n<td>0.470</td>\n<td>0.441</td>\n</tr>\n<tr>\n<td>0.470</td>\n<td>0.460</td>\n</tr>\n<tr>\n<td>0.468</td>\n<td>0.452</td>\n</tr>\n<tr>\n<td>0.475</td>\n<td>0.456</td>\n</tr>\n<tr>\n<td>0.468</td>\n<td>0.445</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3039743": "Does anyone have a good CV-LB relationship?",
    "3039793": "Hey, this is pretty much my first competition, so I might be wrong about how it works, but this is my point of view:\n\nThe LB is completely irrelevant. We make 20 predictions(test data) for it, while the data of these 20 predictions is already part of the training data set (eventhough some information is missing). Just because we can make good prediction on such a small data set, doesn't mean it translates to any other small data set well.\nSo it is all about CV. Having a consistent CV score (mean without alot of variance) is the only important thing.\nOf course you still need some luck to get the right seed / parameters for the ~52 data on the privat LB. Im still confused why it is such a small sample size. I hope I dont misunderstand how it works :)\n\nMy relationship between CV and LB is good, but low. So my CV score matches my LB score. Dont know if that is any good, but most of the LB is probably overfitted.",
    "3039809": "The LB evaluation is on 1444 out of 3800.",
    "3039831": "Ah, so the test.csv is fake and just a placeholder and gets replaced by the real one after submitting? I was confused how it works. But makes sense now. Thank you.\nThen take my post with a grain of salt. Eventhough I still believe in those points. LB is just another fold basically :)",
    "3039906": "Short answer @abdmental01 **NO**",
    "3040567": "actually CV and LB are not related for me",
    "3041459": "As per my experience, there is not any strong relationship between CV and LB.",
    "3044626": "Maybe not the top LB score, but pretty balanced model in my case:\n| CV |  LB|\n| --- | --- |\n| 0.468 | 0.447 |\n| 0.465 | 0.447 |\n| 0.470 | 0.441 |\n| 0.470 | 0.460 |\n| 0.468 | 0.452 |\n| 0.475 | 0.456 |\n| 0.468 | 0.445 |"
  },
  "source": "meta"
}