{
  "id": 518461,
  "title": "My validation score and Public LB is quite different. What is the reason of this?",
  "url": "/competitions/uspto-explainable-ai/discussion/518461",
  "author_name": "",
  "post_date": "2024-07-06T15:49:16.773959800Z",
  "votes": 10,
  "comment_count": 8,
  "views": 0,
  "content": "<p>These days, I am struggling with getting the reliable correlation of validation and LB score. When I achieved 0.76 in validation set, I only got 0.43 in the submission. </p>\n<p>The details are as follows:</p>\n<ol>\n<li>Extracting about 200,000 rows of patent_metadata(2,500 rows of nearest_neighbors.csv and 75,000 rows of random patents)</li>\n<li>Creating query by data made in procedure 1. (cpc_codes only)</li>\n<li>Creating Whoosh index by another subset of patent_metadata(same setting in procedure 1, but using another 75,000 rows of random patents)</li>\n<li>Calculating score by whoosh index created in procedure 3.</li>\n</ol>\n<p>Is there anyone who faced the same issue? How to tackle this problem? I have no idea where this gab of scores come from…</p>",
  "messages": [
    {
      "id": "2908872",
      "postDate": "07/06/2024 15:49:16",
      "content": "<p>These days, I am struggling with getting the reliable correlation of validation and LB score. When I achieved 0.76 in validation set, I only got 0.43 in the submission. </p>\n<p>The details are as follows:</p>\n<ol>\n<li>Extracting about 200,000 rows of patent_metadata(2,500 rows of nearest_neighbors.csv and 75,000 rows of random patents)</li>\n<li>Creating query by data made in procedure 1. (cpc_codes only)</li>\n<li>Creating Whoosh index by another subset of patent_metadata(same setting in procedure 1, but using another 75,000 rows of random patents)</li>\n<li>Calculating score by whoosh index created in procedure 3.</li>\n</ol>\n<p>Is there anyone who faced the same issue? How to tackle this problem? I have no idea where this gab of scores come from…</p>",
      "rawMarkdown": "These days, I am struggling with getting the reliable correlation of validation and LB score. When I achieved 0.76 in validation set, I only got 0.43 in the submission. \n\nThe details are as follows:\n1. Extracting about 200,000 rows of patent_metadata(2,500 rows of nearest_neighbors.csv and 75,000 rows of random patents)\n2. Creating query by data made in procedure 1. (cpc_codes only)\n3. Creating Whoosh index by another subset of patent_metadata(same setting in procedure 1, but using another 75,000 rows of random patents)\n4. Calculating score by whoosh index created in procedure 3.\n\nIs there anyone who faced the same issue? How to tackle this problem? I have no idea where this gab of scores come from...",
      "votes": null
    },
    {
      "id": "2908893",
      "postDate": "07/06/2024 16:03:47",
      "content": "<p>I still have a certain gap between val score and lb at this point and have not been able to determine the cause. I expect that the sampling method is being done in a way that makes the problem more difficult than random sampling.</p>",
      "rawMarkdown": "I still have a certain gap between val score and lb at this point and have not been able to determine the cause. I expect that the sampling method is being done in a way that makes the problem more difficult than random sampling.",
      "votes": null
    },
    {
      "id": "2909459",
      "postDate": "07/07/2024 02:33:56",
      "content": "<p>Thank you for the explanation! Though I have difficulties with the gap of validation and public scores, a certain correlation of these scores can be seen(ex. val: 0.76, public: 0.43 and val 0.81, public: 0.47). So, I will try to get more reliable validation method and reach higher score.</p>",
      "rawMarkdown": "Thank you for the explanation! Though I have difficulties with the gap of validation and public scores, a certain correlation of these scores can be seen(ex. val: 0.76, public: 0.43 and val 0.81, public: 0.47). So, I will try to get more reliable validation method and reach higher score.",
      "votes": null
    },
    {
      "id": "2912872",
      "postDate": "07/09/2024 05:45:10",
      "content": "<p>I seem to be in a very similar situation to you. My CV:0.86, LB:0.52. <br>\nI also don't know if this is due to a difference in the distribution of the test data or if there is a mistake in the way I create the validation index…</p>",
      "rawMarkdown": "I seem to be in a very similar situation to you. My CV:0.86, LB:0.52. \nI also don't know if this is due to a difference in the distribution of the test data or if there is a mistake in the way I create the validation index...",
      "votes": null
    },
    {
      "id": "2913334",
      "postDate": "07/09/2024 12:33:56",
      "content": "<p>Why do we need the local validation? I mean the private and public LB use the same test_index, so the public LB is the best indicator for the final result. Did I miss something?</p>",
      "rawMarkdown": "Why do we need the local validation? I mean the private and public LB use the same test_index, so the public LB is the best indicator for the final result. Did I miss something?",
      "votes": null
    },
    {
      "id": "2913855",
      "postDate": "07/09/2024 16:45:39",
      "content": "<p>For me there's two reasons:</p>\n<ol>\n<li>You can only submit 5 times per day, local validation makes it possible to test many more ideas per day and only submit those that seem like improvements in your local setting.</li>\n<li>Submissions can take a while, local validation can be much quicker, thereby shortening the time it takes to verify whether a certain change is a (likely) improvement or not.</li>\n</ol>\n<p>Nonetheless, the value of local validation depends a lot on how much you trust your own implementation of it. Like others in this thread, mine isn't that great either at the moment.</p>",
      "rawMarkdown": "For me there's two reasons:\n1. You can only submit 5 times per day, local validation makes it possible to test many more ideas per day and only submit those that seem like improvements in your local setting.\n2. Submissions can take a while, local validation can be much quicker, thereby shortening the time it takes to verify whether a certain change is a (likely) improvement or not.\n\nNonetheless, the value of local validation depends a lot on how much you trust your own implementation of it. Like others in this thread, mine isn't that great either at the moment.",
      "votes": null
    },
    {
      "id": "2914946",
      "postDate": "07/10/2024 08:56:21",
      "content": "<p>So the diif between local validation and public LB doesn't matter, since local validation is just used to test ideas, and we should trust LB at the end. Right?</p>",
      "rawMarkdown": "So the diif between local validation and public LB doesn't matter, since local validation is just used to test ideas, and we should trust LB at the end. Right?",
      "votes": null
    },
    {
      "id": "2916636",
      "postDate": "07/11/2024 06:32:18",
      "content": "<p>are you fixing the same 2500 rows of nearest_neighbors.csv? I would suggest creating multiple Whoosh indexes for validation and resampling 2500 rows for each index</p>",
      "rawMarkdown": "are you fixing the same 2500 rows of nearest_neighbors.csv? I would suggest creating multiple Whoosh indexes for validation and resampling 2500 rows for each index",
      "votes": null
    },
    {
      "id": "2916638",
      "postDate": "07/11/2024 06:38:36",
      "content": "<p>I think public LB is a gauge just like scores from 1 validation index, so it is better to estimate with multiple indexes and the public LB also. The distribution of terms in the full test set could be quite different, it seems to me that public LB is somewhat easier and does not contain very difficult patents</p>",
      "rawMarkdown": "I think public LB is a gauge just like scores from 1 validation index, so it is better to estimate with multiple indexes and the public LB also. The distribution of terms in the full test set could be quite different, it seems to me that public LB is somewhat easier and does not contain very difficult patents",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2908893,
      "author_name": "ryotayoshinobu",
      "author_url": "",
      "post_date": "07/06/2024 16:03:47",
      "content": "<p>I still have a certain gap between val score and lb at this point and have not been able to determine the cause. I expect that the sampling method is being done in a way that makes the problem more difficult than random sampling.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2909459,
          "author_name": "shigeria",
          "author_url": "",
          "post_date": "07/07/2024 02:33:56",
          "content": "<p>Thank you for the explanation! Though I have difficulties with the gap of validation and public scores, a certain correlation of these scores can be seen(ex. val: 0.76, public: 0.43 and val 0.81, public: 0.47). So, I will try to get more reliable validation method and reach higher score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2912872,
      "author_name": "ebinan92",
      "author_url": "",
      "post_date": "07/09/2024 05:45:10",
      "content": "<p>I seem to be in a very similar situation to you. My CV:0.86, LB:0.52. <br>\nI also don't know if this is due to a difference in the distribution of the test data or if there is a mistake in the way I create the validation index…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2913334,
      "author_name": "huanligong",
      "author_url": "",
      "post_date": "07/09/2024 12:33:56",
      "content": "<p>Why do we need the local validation? I mean the private and public LB use the same test_index, so the public LB is the best indicator for the final result. Did I miss something?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2913855,
          "author_name": "jmerle",
          "author_url": "",
          "post_date": "07/09/2024 16:45:39",
          "content": "<p>For me there's two reasons:</p>\n<ol>\n<li>You can only submit 5 times per day, local validation makes it possible to test many more ideas per day and only submit those that seem like improvements in your local setting.</li>\n<li>Submissions can take a while, local validation can be much quicker, thereby shortening the time it takes to verify whether a certain change is a (likely) improvement or not.</li>\n</ol>\n<p>Nonetheless, the value of local validation depends a lot on how much you trust your own implementation of it. Like others in this thread, mine isn't that great either at the moment.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2914946,
              "author_name": "huanligong",
              "author_url": "",
              "post_date": "07/10/2024 08:56:21",
              "content": "<p>So the diif between local validation and public LB doesn't matter, since local validation is just used to test ideas, and we should trust LB at the end. Right?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2916638,
                  "author_name": "dilliontan",
                  "author_url": "",
                  "post_date": "07/11/2024 06:38:36",
                  "content": "<p>I think public LB is a gauge just like scores from 1 validation index, so it is better to estimate with multiple indexes and the public LB also. The distribution of terms in the full test set could be quite different, it seems to me that public LB is somewhat easier and does not contain very difficult patents</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2916636,
      "author_name": "dilliontan",
      "author_url": "",
      "post_date": "07/11/2024 06:32:18",
      "content": "<p>are you fixing the same 2500 rows of nearest_neighbors.csv? I would suggest creating multiple Whoosh indexes for validation and resampling 2500 rows for each index</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2908872": "These days, I am struggling with getting the reliable correlation of validation and LB score. When I achieved 0.76 in validation set, I only got 0.43 in the submission. \n\nThe details are as follows:\n1. Extracting about 200,000 rows of patent_metadata(2,500 rows of nearest_neighbors.csv and 75,000 rows of random patents)\n2. Creating query by data made in procedure 1. (cpc_codes only)\n3. Creating Whoosh index by another subset of patent_metadata(same setting in procedure 1, but using another 75,000 rows of random patents)\n4. Calculating score by whoosh index created in procedure 3.\n\nIs there anyone who faced the same issue? How to tackle this problem? I have no idea where this gab of scores come from...",
    "2908893": "I still have a certain gap between val score and lb at this point and have not been able to determine the cause. I expect that the sampling method is being done in a way that makes the problem more difficult than random sampling.",
    "2909459": "Thank you for the explanation! Though I have difficulties with the gap of validation and public scores, a certain correlation of these scores can be seen(ex. val: 0.76, public: 0.43 and val 0.81, public: 0.47). So, I will try to get more reliable validation method and reach higher score.",
    "2912872": "I seem to be in a very similar situation to you. My CV:0.86, LB:0.52. \nI also don't know if this is due to a difference in the distribution of the test data or if there is a mistake in the way I create the validation index...",
    "2913334": "Why do we need the local validation? I mean the private and public LB use the same test_index, so the public LB is the best indicator for the final result. Did I miss something?",
    "2913855": "For me there's two reasons:\n1. You can only submit 5 times per day, local validation makes it possible to test many more ideas per day and only submit those that seem like improvements in your local setting.\n2. Submissions can take a while, local validation can be much quicker, thereby shortening the time it takes to verify whether a certain change is a (likely) improvement or not.\n\nNonetheless, the value of local validation depends a lot on how much you trust your own implementation of it. Like others in this thread, mine isn't that great either at the moment.",
    "2914946": "So the diif between local validation and public LB doesn't matter, since local validation is just used to test ideas, and we should trust LB at the end. Right?",
    "2916636": "are you fixing the same 2500 rows of nearest_neighbors.csv? I would suggest creating multiple Whoosh indexes for validation and resampling 2500 rows for each index",
    "2916638": "I think public LB is a gauge just like scores from 1 validation index, so it is better to estimate with multiple indexes and the public LB also. The distribution of terms in the full test set could be quite different, it seems to me that public LB is somewhat easier and does not contain very difficult patents"
  },
  "source": "meta"
}