{
  "id": 20353,
  "title": "optimum input size",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20353",
  "author_name": "",
  "post_date": "2016-04-23T02:43:23.843Z",
  "votes": 3,
  "comment_count": 1,
  "views": 567,
  "content": "<p>If a person is to do the classification, I believe the classification error is independent of the image size as long as it is large enough (Width&gt;96), but it seems that the neural network does depend on the size very much. Width=96 image already contains all the information to classify each image, so why couldn't the CNN work as well for width=96 as for width say ~128 or larger ? The con-nets need to see more trivial details in order to make a correct classification. Is it because some intrinsic faults of the designing of the Cnets?  Any one has a good insight?</p>",
  "messages": [
    {
      "id": "116304",
      "postDate": "04/23/2016 02:43:23",
      "content": "<p>If a person is to do the classification, I believe the classification error is independent of the image size as long as it is large enough (Width&gt;96), but it seems that the neural network does depend on the size very much. Width=96 image already contains all the information to classify each image, so why couldn't the CNN work as well for width=96 as for width say ~128 or larger ? The con-nets need to see more trivial details in order to make a correct classification. Is it because some intrinsic faults of the designing of the Cnets?  Any one has a good insight?</p>",
      "rawMarkdown": "If a person is to do the classification, I believe the classification error is independent of the image size as long as it is large enough (Width>96), but it seems that the neural network does depend on the size very much. Width=96 image already contains all the information to classify each image, so why couldn't the CNN work as well for width=96 as for width say ~128 or larger ? The con-nets need to see more trivial details in order to make a correct classification. Is it because some intrinsic faults of the designing of the Cnets?  Any one has a good insight?",
      "votes": null
    },
    {
      "id": "116460",
      "postDate": "04/24/2016 14:36:46",
      "content": "<p>But sometimes you do not need to see more trivial details to make a correct classification, like in this case, the problem is more like a pose estimation problem, where the relative positions of your left, right (hand, shoulder, elbow)  matters, and details like other objects in the picture won't help and might even confuse the network</p>",
      "rawMarkdown": "But sometimes you do not need to see more trivial details to make a correct classification, like in this case, the problem is more like a pose estimation problem, where the relative positions of your left, right (hand, shoulder, elbow)  matters, and details like other objects in the picture won't help and might even confuse the network",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 116460,
      "author_name": "usixuz",
      "author_url": "",
      "post_date": "04/24/2016 14:36:46",
      "content": "<p>But sometimes you do not need to see more trivial details to make a correct classification, like in this case, the problem is more like a pose estimation problem, where the relative positions of your left, right (hand, shoulder, elbow)  matters, and details like other objects in the picture won't help and might even confuse the network</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "116304": "If a person is to do the classification, I believe the classification error is independent of the image size as long as it is large enough (Width>96), but it seems that the neural network does depend on the size very much. Width=96 image already contains all the information to classify each image, so why couldn't the CNN work as well for width=96 as for width say ~128 or larger ? The con-nets need to see more trivial details in order to make a correct classification. Is it because some intrinsic faults of the designing of the Cnets?  Any one has a good insight?",
    "116460": "But sometimes you do not need to see more trivial details to make a correct classification, like in this case, the problem is more like a pose estimation problem, where the relative positions of your left, right (hand, shoulder, elbow)  matters, and details like other objects in the picture won't help and might even confuse the network"
  },
  "source": "meta"
}