{
  "id": 98042,
  "title": "Data Augmentation or Adversarial Validation or Both or ... ?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/98042",
  "author_name": "",
  "post_date": "2019-06-30T22:47:22.073887300Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F838971%2F593dd1666b71d6131f9f6d8f28b8dff5%2FCapture2.PNG?generation=1561956311367886&amp;alt=media\" alt=\"\">\nAs others have mentioned, the test set and training set appear to have different frequencies of the 5 classes. What are the approaches people normally take for dealing with this in a NN-centric competition? I am aware of data augmentation techniques and adversarial validation. Is there anything else? Perhaps rank-pruning?</p>",
  "messages": [
    {
      "id": "565402",
      "postDate": "06/30/2019 22:47:22",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F838971%2F593dd1666b71d6131f9f6d8f28b8dff5%2FCapture2.PNG?generation=1561956311367886&amp;alt=media\" alt=\"\">\nAs others have mentioned, the test set and training set appear to have different frequencies of the 5 classes. What are the approaches people normally take for dealing with this in a NN-centric competition? I am aware of data augmentation techniques and adversarial validation. Is there anything else? Perhaps rank-pruning?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F838971%2F593dd1666b71d6131f9f6d8f28b8dff5%2FCapture2.PNG?generation=1561956311367886&amp;alt=media)\nAs others have mentioned, the test set and training set appear to have different frequencies of the 5 classes. What are the approaches people normally take for dealing with this in a NN-centric competition? I am aware of data augmentation techniques and adversarial validation. Is there anything else? Perhaps rank-pruning?",
      "votes": null
    },
    {
      "id": "565409",
      "postDate": "06/30/2019 22:57:07",
      "content": "<p>Some rank-pruning references:</p>\n\n<ul>\n<li><p><a href=\"https://github.com/cgnorthcutt/cleanlab/\">https://github.com/cgnorthcutt/cleanlab/</a></p></li>\n<li><p><a href=\"https://arxiv.org/abs/1705.01936\">https://arxiv.org/abs/1705.01936</a></p></li>\n</ul>",
      "rawMarkdown": "Some rank-pruning references:\n\n* https://github.com/cgnorthcutt/cleanlab/\n\n* https://arxiv.org/abs/1705.01936",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 565409,
      "author_name": "puremath86",
      "author_url": "",
      "post_date": "06/30/2019 22:57:07",
      "content": "<p>Some rank-pruning references:</p>\n\n<ul>\n<li><p><a href=\"https://github.com/cgnorthcutt/cleanlab/\">https://github.com/cgnorthcutt/cleanlab/</a></p></li>\n<li><p><a href=\"https://arxiv.org/abs/1705.01936\">https://arxiv.org/abs/1705.01936</a></p></li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "565402": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F838971%2F593dd1666b71d6131f9f6d8f28b8dff5%2FCapture2.PNG?generation=1561956311367886&amp;alt=media)\nAs others have mentioned, the test set and training set appear to have different frequencies of the 5 classes. What are the approaches people normally take for dealing with this in a NN-centric competition? I am aware of data augmentation techniques and adversarial validation. Is there anything else? Perhaps rank-pruning?",
    "565409": "Some rank-pruning references:\n\n* https://github.com/cgnorthcutt/cleanlab/\n\n* https://arxiv.org/abs/1705.01936"
  },
  "source": "meta"
}