{
  "id": 97808,
  "title": "Is there even enough training data?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/97808",
  "author_name": "",
  "post_date": "2019-06-29T00:42:40.521770600Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p><img src=\"https://i.imgur.com/WfLIyz7.png\" alt=\"Imgur\"></p>\n\n<p>My understanding is that there should be a minimum 1000 images per class. Am I wrong?</p>",
  "messages": [
    {
      "id": "564083",
      "postDate": "06/29/2019 00:42:40",
      "content": "<p><img src=\"https://i.imgur.com/WfLIyz7.png\" alt=\"Imgur\"></p>\n\n<p>My understanding is that there should be a minimum 1000 images per class. Am I wrong?</p>",
      "rawMarkdown": "![Imgur](https://i.imgur.com/WfLIyz7.png)\n\nMy understanding is that there should be a minimum 1000 images per class. Am I wrong?",
      "votes": null
    },
    {
      "id": "564203",
      "postDate": "06/29/2019 05:42:48",
      "content": "<p>From our data and it is stated in papers also predicting for rare classes (1, 3, 4) is going to be especially difficult. Probably using data augmentation can help partly with this problem.</p>",
      "rawMarkdown": "From our data and it is stated in papers also predicting for rare classes (1, 3, 4) is going to be especially difficult. Probably using data augmentation can help partly with this problem.",
      "votes": null
    },
    {
      "id": "564680",
      "postDate": "06/29/2019 19:42:43",
      "content": "<p>Sorry, you're wrong. If you can distinguish these classes using your own neural network, an artificial NN also can. There are systems for one-shot learning, for example.</p>\n\n<p>Having 1000 samples per class would be nice but it's not mandatory.</p>",
      "rawMarkdown": "Sorry, you're wrong. If you can distinguish these classes using your own neural network, an artificial NN also can. There are systems for one-shot learning, for example.\n\nHaving 1000 samples per class would be nice but it's not mandatory.",
      "votes": null
    },
    {
      "id": "564744",
      "postDate": "06/29/2019 22:59:16",
      "content": "<p>I see what you're saying. My other issue is that the full test(~13,000) is 3.5x train, but that's another story. </p>",
      "rawMarkdown": "I see what you're saying. My other issue is that the full test(~13,000) is 3.5x train, but that's another story.",
      "votes": null
    },
    {
      "id": "565400",
      "postDate": "06/30/2019 22:37:42",
      "content": "<p>You may be thinking about the learning problem the wrong way. If you don't think you have enough representative samples from the data and augmented data --then go find more data. There was a similar competition four years ago: <a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/data\">https://www.kaggle.com/c/diabetic-retinopathy-detection/data</a></p>",
      "rawMarkdown": "You may be thinking about the learning problem the wrong way. If you don't think you have enough representative samples from the data and augmented data --then go find more data. There was a similar competition four years ago: https://www.kaggle.com/c/diabetic-retinopathy-detection/data",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 564203,
      "author_name": "demonplus",
      "author_url": "",
      "post_date": "06/29/2019 05:42:48",
      "content": "<p>From our data and it is stated in papers also predicting for rare classes (1, 3, 4) is going to be especially difficult. Probably using data augmentation can help partly with this problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 564680,
      "author_name": "artyomp",
      "author_url": "",
      "post_date": "06/29/2019 19:42:43",
      "content": "<p>Sorry, you're wrong. If you can distinguish these classes using your own neural network, an artificial NN also can. There are systems for one-shot learning, for example.</p>\n\n<p>Having 1000 samples per class would be nice but it's not mandatory.</p>",
      "votes": null,
      "replies": [
        {
          "id": 564744,
          "author_name": "teeyee314",
          "author_url": "",
          "post_date": "06/29/2019 22:59:16",
          "content": "<p>I see what you're saying. My other issue is that the full test(~13,000) is 3.5x train, but that's another story. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 565400,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "06/30/2019 22:37:42",
          "content": "<p>You may be thinking about the learning problem the wrong way. If you don't think you have enough representative samples from the data and augmented data --then go find more data. There was a similar competition four years ago: <a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/data\">https://www.kaggle.com/c/diabetic-retinopathy-detection/data</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "564083": "![Imgur](https://i.imgur.com/WfLIyz7.png)\n\nMy understanding is that there should be a minimum 1000 images per class. Am I wrong?",
    "564203": "From our data and it is stated in papers also predicting for rare classes (1, 3, 4) is going to be especially difficult. Probably using data augmentation can help partly with this problem.",
    "564680": "Sorry, you're wrong. If you can distinguish these classes using your own neural network, an artificial NN also can. There are systems for one-shot learning, for example.\n\nHaving 1000 samples per class would be nice but it's not mandatory.",
    "564744": "I see what you're saying. My other issue is that the full test(~13,000) is 3.5x train, but that's another story.",
    "565400": "You may be thinking about the learning problem the wrong way. If you don't think you have enough representative samples from the data and augmented data --then go find more data. There was a similar competition four years ago: https://www.kaggle.com/c/diabetic-retinopathy-detection/data"
  },
  "source": "meta"
}