{
  "id": 126253,
  "title": "Let's discuss about private leaderboard shakeup",
  "url": "/competitions/pku-autonomous-driving/discussion/126253",
  "author_name": "",
  "post_date": "2020-01-16T13:36:38.844244Z",
  "votes": 5,
  "comment_count": 16,
  "views": 0,
  "content": "<p>91% data waiting to test our models in private lb,i have observed few times that good local cv leading to poor public lb score and not so much impressive cv can get us good lb score,so what's your thought? do you expect large shakeup? how you are going to select your best 2 submissions? 1 with high local cv and 1 with high public lb score? </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F786b0fb00bcf5d3956bd979109feeb97%2FshakeUP.jpg?generation=1579181565934147&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "720506",
      "postDate": "01/16/2020 13:36:38",
      "content": "<p>91% data waiting to test our models in private lb,i have observed few times that good local cv leading to poor public lb score and not so much impressive cv can get us good lb score,so what's your thought? do you expect large shakeup? how you are going to select your best 2 submissions? 1 with high local cv and 1 with high public lb score? </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F786b0fb00bcf5d3956bd979109feeb97%2FshakeUP.jpg?generation=1579181565934147&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "91% data waiting to test our models in private lb,i have observed few times that good local cv leading to poor public lb score and not so much impressive cv can get us good lb score,so what's your thought? do you expect large shakeup? how you are going to select your best 2 submissions? 1 with high local cv and 1 with high public lb score? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F786b0fb00bcf5d3956bd979109feeb97%2FshakeUP.jpg?generation=1579181565934147&amp;alt=media)",
      "votes": null
    },
    {
      "id": "720560",
      "postDate": "01/16/2020 14:26:10",
      "content": "<p>I had local CV of .137 but only LB of .032 so I'm probably overfitting. If you have high CV and LB score then I don't think there'll be much shakeup, in fact your score would probably increase.</p>",
      "rawMarkdown": "I had local CV of .137 but only LB of .032 so I'm probably overfitting. If you have high CV and LB score then I don't think there'll be much shakeup, in fact your score would probably increase.",
      "votes": null
    },
    {
      "id": "720612",
      "postDate": "01/16/2020 15:21:47",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>  i think you are right</p>",
      "rawMarkdown": "greatgamedota  i think you are right",
      "votes": null
    },
    {
      "id": "720689",
      "postDate": "01/16/2020 16:42:24",
      "content": "<p>I think that high CV is just because LB metric is not clear, so we are probably using different calculation methods for CV.</p>",
      "rawMarkdown": "I think that high CV is just because LB metric is not clear, so we are probably using different calculation methods for CV.",
      "votes": null
    },
    {
      "id": "720705",
      "postDate": "01/16/2020 16:53:16",
      "content": "<p>to me : high cv depends on threshold </p>",
      "rawMarkdown": "to me : high cv depends on threshold",
      "votes": null
    },
    {
      "id": "720793",
      "postDate": "01/16/2020 18:37:36",
      "content": "<p>I haven't been tuning the threshold (0.5 for sigmoid = 0 in logit val) and my approach so far has been both increasing in CV and LB, while the values themselves are very off. I do not think there would be much shakeup unless many people have been tuning threshold based on LB</p>",
      "rawMarkdown": "I haven't been tuning the threshold (0.5 for sigmoid = 0 in logit val) and my approach so far has been both increasing in CV and LB, while the values themselves are very off. I do not think there would be much shakeup unless many people have been tuning threshold based on LB",
      "votes": null
    },
    {
      "id": "720796",
      "postDate": "01/16/2020 18:42:30",
      "content": "<p>What's your local cv score for logits&gt;0? \nAnd how much data you are using for   local cv calculation?</p>",
      "rawMarkdown": "What's your local cv score for logits&gt;0? \nAnd how much data you are using for   local cv calculation?",
      "votes": null
    },
    {
      "id": "720805",
      "postDate": "01/16/2020 18:53:37",
      "content": "<p>~0.16 cv and 0.79 LB, within same approach and it is my highest score in both by far. I'm using 0.1 split!</p>\n\n<p>To add more details, I'm using the evaluation kernel but without random scoring (hence the high value), though in the least for me significantly higher cv meant higher LB so far</p>",
      "rawMarkdown": "~0.16 cv and 0.79 LB, within same approach and it is my highest score in both by far. I'm using 0.1 split!\n\nTo add more details, I'm using the evaluation kernel but without random scoring (hence the high value), though in the least for me significantly higher cv meant higher LB so far",
      "votes": null
    },
    {
      "id": "721096",
      "postDate": "01/17/2020 03:26:02",
      "content": "<p>Today I made a submission that had cv 0.108 (tito's script) while resulted in lb 0.090. I think I'd better check my predictions and verify whether errors exist or not.</p>",
      "rawMarkdown": "Today I made a submission that had cv 0.108 (tito's script) while resulted in lb 0.090. I think I'd better check my predictions and verify whether errors exist or not.",
      "votes": null
    },
    {
      "id": "721125",
      "postDate": "01/17/2020 04:31:31",
      "content": "<p>One strange thing is that I have tested thresholds from 0.3 to 0.535 (sigmoid), the CV dropped\n while lb didn't change.</p>",
      "rawMarkdown": "One strange thing is that I have tested thresholds from 0.3 to 0.535 (sigmoid), the CV dropped\n while lb didn't change.",
      "votes": null
    },
    {
      "id": "721126",
      "postDate": "01/17/2020 04:33:43",
      "content": "<p>are you using logits&gt;0 for your best score?</p>",
      "rawMarkdown": "are you using logits&gt;0 for your best score?",
      "votes": null
    },
    {
      "id": "721155",
      "postDate": "01/17/2020 05:30:45",
      "content": "<p>I'm using sigmoid&gt;0.3, and it is logits&gt;-0.85. </p>",
      "rawMarkdown": "I'm using sigmoid&gt;0.3, and it is logits&gt;-0.85.",
      "votes": null
    },
    {
      "id": "721260",
      "postDate": "01/17/2020 08:30:38",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> Hey I have also encountered this problem and finally, I've find out that it's due to the mistake in my implementation. Your CV is satisfying and I don't think overfitting could lead to such a large gap. Please check your code carefully and do some visualizations.</p>",
      "rawMarkdown": "greatgamedota Hey I have also encountered this problem and finally, I've find out that it's due to the mistake in my implementation. Your CV is satisfying and I don't think overfitting could lead to such a large gap. Please check your code carefully and do some visualizations.",
      "votes": null
    },
    {
      "id": "721698",
      "postDate": "01/17/2020 16:09:22",
      "content": "<p>So you are not using the original version of mAP script from public kernel? <a href=\"/syoya1997\">@syoya1997</a> </p>",
      "rawMarkdown": "So you are not using the original version of mAP script from public kernel? @syoya1997",
      "votes": null
    },
    {
      "id": "721895",
      "postDate": "01/17/2020 21:01:26",
      "content": "<p>I had a feeling the test images are not truly randomly selected from the same pool by looking at the images. They are consistently harder to recognize. Not sure whether that has implications on the local/lb difference.</p>",
      "rawMarkdown": "I had a feeling the test images are not truly randomly selected from the same pool by looking at the images. They are consistently harder to recognize. Not sure whether that has implications on the local/lb difference.",
      "votes": null
    },
    {
      "id": "724940",
      "postDate": "01/21/2020 16:26:37",
      "content": "<p>I've got CV0.080 and LB0.050.</p>",
      "rawMarkdown": "I've got CV0.080 and LB0.050.",
      "votes": null
    },
    {
      "id": "724978",
      "postDate": "01/21/2020 17:00:23",
      "content": "<p>I also experienced localCV(I use racall*precision) and public LB inconsistency.\nHowever, it seems that public lb itself is somewhat stable for heatmap threthold ( +-0.002 lb for 0.1 threhold change),\nand stable for my models trained with different seeds ( around +-0.002 lb jittering for my models).</p>\n\n<p>Anyway, I hope I stay in the same position (this is my first competition!)</p>",
      "rawMarkdown": "I also experienced localCV(I use racall*precision) and public LB inconsistency.\nHowever, it seems that public lb itself is somewhat stable for heatmap threthold ( +-0.002 lb for 0.1 threhold change),\nand stable for my models trained with different seeds ( around +-0.002 lb jittering for my models).\n\nAnyway, I hope I stay in the same position (this is my first competition!)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 720560,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "01/16/2020 14:26:10",
      "content": "<p>I had local CV of .137 but only LB of .032 so I'm probably overfitting. If you have high CV and LB score then I don't think there'll be much shakeup, in fact your score would probably increase.</p>",
      "votes": null,
      "replies": [
        {
          "id": 720612,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/16/2020 15:21:47",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>  i think you are right</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 720689,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "01/16/2020 16:42:24",
          "content": "<p>I think that high CV is just because LB metric is not clear, so we are probably using different calculation methods for CV.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 720705,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/16/2020 16:53:16",
          "content": "<p>to me : high cv depends on threshold </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 721260,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "01/17/2020 08:30:38",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> Hey I have also encountered this problem and finally, I've find out that it's due to the mistake in my implementation. Your CV is satisfying and I don't think overfitting could lead to such a large gap. Please check your code carefully and do some visualizations.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 721698,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "01/17/2020 16:09:22",
          "content": "<p>So you are not using the original version of mAP script from public kernel? <a href=\"/syoya1997\">@syoya1997</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 720793,
      "author_name": "joonl04",
      "author_url": "",
      "post_date": "01/16/2020 18:37:36",
      "content": "<p>I haven't been tuning the threshold (0.5 for sigmoid = 0 in logit val) and my approach so far has been both increasing in CV and LB, while the values themselves are very off. I do not think there would be much shakeup unless many people have been tuning threshold based on LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 720796,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/16/2020 18:42:30",
          "content": "<p>What's your local cv score for logits&gt;0? \nAnd how much data you are using for   local cv calculation?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 720805,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "01/16/2020 18:53:37",
          "content": "<p>~0.16 cv and 0.79 LB, within same approach and it is my highest score in both by far. I'm using 0.1 split!</p>\n\n<p>To add more details, I'm using the evaluation kernel but without random scoring (hence the high value), though in the least for me significantly higher cv meant higher LB so far</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 721096,
      "author_name": "richardwong1994",
      "author_url": "",
      "post_date": "01/17/2020 03:26:02",
      "content": "<p>Today I made a submission that had cv 0.108 (tito's script) while resulted in lb 0.090. I think I'd better check my predictions and verify whether errors exist or not.</p>",
      "votes": null,
      "replies": [
        {
          "id": 721125,
          "author_name": "richardwong1994",
          "author_url": "",
          "post_date": "01/17/2020 04:31:31",
          "content": "<p>One strange thing is that I have tested thresholds from 0.3 to 0.535 (sigmoid), the CV dropped\n while lb didn't change.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 721126,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/17/2020 04:33:43",
          "content": "<p>are you using logits&gt;0 for your best score?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 721155,
          "author_name": "richardwong1994",
          "author_url": "",
          "post_date": "01/17/2020 05:30:45",
          "content": "<p>I'm using sigmoid&gt;0.3, and it is logits&gt;-0.85. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 721895,
      "author_name": "hhaoyan",
      "author_url": "",
      "post_date": "01/17/2020 21:01:26",
      "content": "<p>I had a feeling the test images are not truly randomly selected from the same pool by looking at the images. They are consistently harder to recognize. Not sure whether that has implications on the local/lb difference.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 724940,
      "author_name": "nxrprime",
      "author_url": "",
      "post_date": "01/21/2020 16:26:37",
      "content": "<p>I've got CV0.080 and LB0.050.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 724978,
      "author_name": "lisosia",
      "author_url": "",
      "post_date": "01/21/2020 17:00:23",
      "content": "<p>I also experienced localCV(I use racall*precision) and public LB inconsistency.\nHowever, it seems that public lb itself is somewhat stable for heatmap threthold ( +-0.002 lb for 0.1 threhold change),\nand stable for my models trained with different seeds ( around +-0.002 lb jittering for my models).</p>\n\n<p>Anyway, I hope I stay in the same position (this is my first competition!)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "720506": "91% data waiting to test our models in private lb,i have observed few times that good local cv leading to poor public lb score and not so much impressive cv can get us good lb score,so what's your thought? do you expect large shakeup? how you are going to select your best 2 submissions? 1 with high local cv and 1 with high public lb score? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F786b0fb00bcf5d3956bd979109feeb97%2FshakeUP.jpg?generation=1579181565934147&amp;alt=media)",
    "720560": "I had local CV of .137 but only LB of .032 so I'm probably overfitting. If you have high CV and LB score then I don't think there'll be much shakeup, in fact your score would probably increase.",
    "720612": "greatgamedota  i think you are right",
    "720689": "I think that high CV is just because LB metric is not clear, so we are probably using different calculation methods for CV.",
    "720705": "to me : high cv depends on threshold",
    "720793": "I haven't been tuning the threshold (0.5 for sigmoid = 0 in logit val) and my approach so far has been both increasing in CV and LB, while the values themselves are very off. I do not think there would be much shakeup unless many people have been tuning threshold based on LB",
    "720796": "What's your local cv score for logits&gt;0? \nAnd how much data you are using for   local cv calculation?",
    "720805": "~0.16 cv and 0.79 LB, within same approach and it is my highest score in both by far. I'm using 0.1 split!\n\nTo add more details, I'm using the evaluation kernel but without random scoring (hence the high value), though in the least for me significantly higher cv meant higher LB so far",
    "721096": "Today I made a submission that had cv 0.108 (tito's script) while resulted in lb 0.090. I think I'd better check my predictions and verify whether errors exist or not.",
    "721125": "One strange thing is that I have tested thresholds from 0.3 to 0.535 (sigmoid), the CV dropped\n while lb didn't change.",
    "721126": "are you using logits&gt;0 for your best score?",
    "721155": "I'm using sigmoid&gt;0.3, and it is logits&gt;-0.85.",
    "721260": "greatgamedota Hey I have also encountered this problem and finally, I've find out that it's due to the mistake in my implementation. Your CV is satisfying and I don't think overfitting could lead to such a large gap. Please check your code carefully and do some visualizations.",
    "721698": "So you are not using the original version of mAP script from public kernel? @syoya1997",
    "721895": "I had a feeling the test images are not truly randomly selected from the same pool by looking at the images. They are consistently harder to recognize. Not sure whether that has implications on the local/lb difference.",
    "724940": "I've got CV0.080 and LB0.050.",
    "724978": "I also experienced localCV(I use racall*precision) and public LB inconsistency.\nHowever, it seems that public lb itself is somewhat stable for heatmap threthold ( +-0.002 lb for 0.1 threhold change),\nand stable for my models trained with different seeds ( around +-0.002 lb jittering for my models).\n\nAnyway, I hope I stay in the same position (this is my first competition!)"
  },
  "source": "meta"
}