{
  "id": 106842,
  "title": "Tuning regression threshold?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/106842",
  "author_name": "",
  "post_date": "2019-08-31T06:32:40.770562100Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>So I know that optimizing qwk does not work that well since the validation data derived from training dataset that we are provided does not generalize to the test dataset. However, lately I've experienced noticeable score boost with some tuning of threshold based on public LB score. Is this just overfitting to the public LB? Or is there some way I could validate this approach to help overall score boost but not necessarily overfit to the public LB? Is anyone else doing this?</p>",
  "messages": [
    {
      "id": "614158",
      "postDate": "08/31/2019 06:32:40",
      "content": "<p>So I know that optimizing qwk does not work that well since the validation data derived from training dataset that we are provided does not generalize to the test dataset. However, lately I've experienced noticeable score boost with some tuning of threshold based on public LB score. Is this just overfitting to the public LB? Or is there some way I could validate this approach to help overall score boost but not necessarily overfit to the public LB? Is anyone else doing this?</p>",
      "rawMarkdown": "So I know that optimizing qwk does not work that well since the validation data derived from training dataset that we are provided does not generalize to the test dataset. However, lately I've experienced noticeable score boost with some tuning of threshold based on public LB score. Is this just overfitting to the public LB? Or is there some way I could validate this approach to help overall score boost but not necessarily overfit to the public LB? Is anyone else doing this?",
      "votes": null
    },
    {
      "id": "614163",
      "postDate": "08/31/2019 06:40:20",
      "content": "<p>Don't trust LB. Public test set is crazy, god knows what private test is gonna be like.</p>",
      "rawMarkdown": "Don't trust LB. Public test set is crazy, god knows what private test is gonna be like.",
      "votes": null
    },
    {
      "id": "614172",
      "postDate": "08/31/2019 06:56:01",
      "content": "<p>If you don't mind sharing, have you been using LB at all to tune your models so far? It seems quite impossible for me to reach anything above 0.8 purely based on CV, and given there is a kernel indicating there may not as big of a shake up, I'm quite tempted to use LB as tuning signal :/</p>",
      "rawMarkdown": "If you don't mind sharing, have you been using LB at all to tune your models so far? It seems quite impossible for me to reach anything above 0.8 purely based on CV, and given there is a kernel indicating there may not as big of a shake up, I'm quite tempted to use LB as tuning signal :/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 614163,
      "author_name": "rishabhiitbhu",
      "author_url": "",
      "post_date": "08/31/2019 06:40:20",
      "content": "<p>Don't trust LB. Public test set is crazy, god knows what private test is gonna be like.</p>",
      "votes": null,
      "replies": [
        {
          "id": 614172,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "08/31/2019 06:56:01",
          "content": "<p>If you don't mind sharing, have you been using LB at all to tune your models so far? It seems quite impossible for me to reach anything above 0.8 purely based on CV, and given there is a kernel indicating there may not as big of a shake up, I'm quite tempted to use LB as tuning signal :/</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "614158": "So I know that optimizing qwk does not work that well since the validation data derived from training dataset that we are provided does not generalize to the test dataset. However, lately I've experienced noticeable score boost with some tuning of threshold based on public LB score. Is this just overfitting to the public LB? Or is there some way I could validate this approach to help overall score boost but not necessarily overfit to the public LB? Is anyone else doing this?",
    "614163": "Don't trust LB. Public test set is crazy, god knows what private test is gonna be like.",
    "614172": "If you don't mind sharing, have you been using LB at all to tune your models so far? It seems quite impossible for me to reach anything above 0.8 purely based on CV, and given there is a kernel indicating there may not as big of a shake up, I'm quite tempted to use LB as tuning signal :/"
  },
  "source": "meta"
}