{
  "id": 106935,
  "title": "QWK optimization vs. bayesian statistics",
  "url": "/competitions/aptos2019-blindness-detection/discussion/106935",
  "author_name": "",
  "post_date": "2019-08-31T22:56:22.565470400Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>as many people are using a regression model in that competition I was just wondering how you optimize your thresholds for class association. I saw the common way of using a validation set, which is a subset of the training data with the same data and label distribution, to optimize the thresholds against the QWK metric. But this approach is levering the domain gap problem even more or ? Not only the input data from test/train changes (different crop etc.) but also because of less appearance of one class the thresholds are set in a worse way. </p>\n\n<p>For example here are the training data distribution for 0.91 QWK:\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943054%2F68bee3c28cada4deaa9aff7b591a1a86%2FTTA_traind_dist.PNG?generation=1567291483166280&amp;alt=media\" alt=\"\"></p>\n\n<p>You can see that there are lots of 0s, but not so many 2,3s \nAlso the optimizer gives me a 2.68 as the threshold between class 2 and 3.</p>\n\n<p>Now that same model predicts this distribution for the test-set:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943054%2F1e90cef874dd43cdd46b1d6f5a60d832%2FTTADist.PNG?generation=1567291552263940&amp;alt=media\" alt=\"\"></p>\n\n<p>As you can see there will be a huge problem because the separation between class 2 and 3 is deciding mostly about the model accuracy. </p>\n\n<p>So my question now: Related to the bayesian statistics, the guess of thresholds from a sample distribution of a validation set adds information to the system which is not actually wrong because we dont know the target distribution of right labels. As we want to minimize the information influence for minimal domain gap (from train to test), it makes sense to set the thresholds to 0.5,1.5,2.5.....so equidistant between the classes. This configuration would ensure maximal entropy.\nDoes that make sense ? Or am I on a wrong way here.\nI would love to have a discussion about that 💥 </p>",
  "messages": [
    {
      "id": "614695",
      "postDate": "08/31/2019 22:56:22",
      "content": "<p>Hello,</p>\n\n<p>as many people are using a regression model in that competition I was just wondering how you optimize your thresholds for class association. I saw the common way of using a validation set, which is a subset of the training data with the same data and label distribution, to optimize the thresholds against the QWK metric. But this approach is levering the domain gap problem even more or ? Not only the input data from test/train changes (different crop etc.) but also because of less appearance of one class the thresholds are set in a worse way. </p>\n\n<p>For example here are the training data distribution for 0.91 QWK:\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943054%2F68bee3c28cada4deaa9aff7b591a1a86%2FTTA_traind_dist.PNG?generation=1567291483166280&amp;alt=media\" alt=\"\"></p>\n\n<p>You can see that there are lots of 0s, but not so many 2,3s \nAlso the optimizer gives me a 2.68 as the threshold between class 2 and 3.</p>\n\n<p>Now that same model predicts this distribution for the test-set:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943054%2F1e90cef874dd43cdd46b1d6f5a60d832%2FTTADist.PNG?generation=1567291552263940&amp;alt=media\" alt=\"\"></p>\n\n<p>As you can see there will be a huge problem because the separation between class 2 and 3 is deciding mostly about the model accuracy. </p>\n\n<p>So my question now: Related to the bayesian statistics, the guess of thresholds from a sample distribution of a validation set adds information to the system which is not actually wrong because we dont know the target distribution of right labels. As we want to minimize the information influence for minimal domain gap (from train to test), it makes sense to set the thresholds to 0.5,1.5,2.5.....so equidistant between the classes. This configuration would ensure maximal entropy.\nDoes that make sense ? Or am I on a wrong way here.\nI would love to have a discussion about that 💥 </p>",
      "rawMarkdown": "Hello,\n\nas many people are using a regression model in that competition I was just wondering how you optimize your thresholds for class association. I saw the common way of using a validation set, which is a subset of the training data with the same data and label distribution, to optimize the thresholds against the QWK metric. But this approach is levering the domain gap problem even more or ? Not only the input data from test/train changes (different crop etc.) but also because of less appearance of one class the thresholds are set in a worse way. \n\nFor example here are the training data distribution for 0.91 QWK:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943054%2F68bee3c28cada4deaa9aff7b591a1a86%2FTTA_traind_dist.PNG?generation=1567291483166280&amp;alt=media)\n\nYou can see that there are lots of 0s, but not so many 2,3s \nAlso the optimizer gives me a 2.68 as the threshold between class 2 and 3.\n\nNow that same model predicts this distribution for the test-set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943054%2F1e90cef874dd43cdd46b1d6f5a60d832%2FTTADist.PNG?generation=1567291552263940&amp;alt=media)\n\nAs you can see there will be a huge problem because the separation between class 2 and 3 is deciding mostly about the model accuracy. \n\nSo my question now: Related to the bayesian statistics, the guess of thresholds from a sample distribution of a validation set adds information to the system which is not actually wrong because we dont know the target distribution of right labels. As we want to minimize the information influence for minimal domain gap (from train to test), it makes sense to set the thresholds to 0.5,1.5,2.5.....so equidistant between the classes. This configuration would ensure maximal entropy.\nDoes that make sense ? Or am I on a wrong way here.\nI would love to have a discussion about that 💥",
      "votes": null
    },
    {
      "id": "614774",
      "postDate": "09/01/2019 03:38:16",
      "content": "<p>Hope not to overfit the data, our team use only thresholds 0.5, 1.5, ..., 3.5 .</p>",
      "rawMarkdown": "Hope not to overfit the data, our team use only thresholds 0.5, 1.5, ..., 3.5 .",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 614774,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "09/01/2019 03:38:16",
      "content": "<p>Hope not to overfit the data, our team use only thresholds 0.5, 1.5, ..., 3.5 .</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "614695": "Hello,\n\nas many people are using a regression model in that competition I was just wondering how you optimize your thresholds for class association. I saw the common way of using a validation set, which is a subset of the training data with the same data and label distribution, to optimize the thresholds against the QWK metric. But this approach is levering the domain gap problem even more or ? Not only the input data from test/train changes (different crop etc.) but also because of less appearance of one class the thresholds are set in a worse way. \n\nFor example here are the training data distribution for 0.91 QWK:\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943054%2F68bee3c28cada4deaa9aff7b591a1a86%2FTTA_traind_dist.PNG?generation=1567291483166280&amp;alt=media)\n\nYou can see that there are lots of 0s, but not so many 2,3s \nAlso the optimizer gives me a 2.68 as the threshold between class 2 and 3.\n\nNow that same model predicts this distribution for the test-set:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943054%2F1e90cef874dd43cdd46b1d6f5a60d832%2FTTADist.PNG?generation=1567291552263940&amp;alt=media)\n\nAs you can see there will be a huge problem because the separation between class 2 and 3 is deciding mostly about the model accuracy. \n\nSo my question now: Related to the bayesian statistics, the guess of thresholds from a sample distribution of a validation set adds information to the system which is not actually wrong because we dont know the target distribution of right labels. As we want to minimize the information influence for minimal domain gap (from train to test), it makes sense to set the thresholds to 0.5,1.5,2.5.....so equidistant between the classes. This configuration would ensure maximal entropy.\nDoes that make sense ? Or am I on a wrong way here.\nI would love to have a discussion about that 💥",
    "614774": "Hope not to overfit the data, our team use only thresholds 0.5, 1.5, ..., 3.5 ."
  },
  "source": "meta"
}