{
  "id": 23099,
  "title": "Preprocessed Dataset Available",
  "url": "/competitions/painter-by-numbers/discussion/23099",
  "author_name": "Flynn, Michael",
  "post_date": "2016-08-25T22:45:47.843000",
  "votes": 15,
  "comment_count": 8,
  "views": 1183,
  "content": "<p>For those of you who are having trouble downloading files and are going to be resizing the images to a smaller size (e.g. for a CNN), I've made a preprocessed dataset available <a href=\"https://github.com/zo7/painter-by-numbers/releases/tag/data-v1.0\">here</a>. (Roughly 2.2GB for both the train and test sets)</p>\n\n<p>This dataset resizes all of the images so that their smallest side is 256 pixels long, giving you enough wiggle room to do data augmentation and speed up reading/processing. Since I standardized all of the formats this should solve any issues people have been experiencing by not being able to read some images.</p>\n\n<p>Hope this helps!</p>",
  "messages": [
    {
      "id": 132480,
      "postDate": "2016-08-25T22:45:47.843Z",
      "content": "<p>For those of you who are having trouble downloading files and are going to be resizing the images to a smaller size (e.g. for a CNN), I've made a preprocessed dataset available <a href=\"https://github.com/zo7/painter-by-numbers/releases/tag/data-v1.0\">here</a>. (Roughly 2.2GB for both the train and test sets)</p>\n\n<p>This dataset resizes all of the images so that their smallest side is 256 pixels long, giving you enough wiggle room to do data augmentation and speed up reading/processing. Since I standardized all of the formats this should solve any issues people have been experiencing by not being able to read some images.</p>\n\n<p>Hope this helps!</p>",
      "rawMarkdown": "For those of you who are having trouble downloading files and are going to be resizing the images to a smaller size (e.g. for a CNN), I've made a preprocessed dataset available [here](https://github.com/zo7/painter-by-numbers/releases/tag/data-v1.0). (Roughly 2.2GB for both the train and test sets)\r\n\r\nThis dataset resizes all of the images so that their smallest side is 256 pixels long, giving you enough wiggle room to do data augmentation and speed up reading/processing. Since I standardized all of the formats this should solve any issues people have been experiencing by not being able to read some images.\r\n\r\nHope this helps!",
      "votes": 15
    },
    {
      "id": 164321,
      "postDate": "2017-02-28T18:32:18.877Z",
      "content": "<p>Hey, are you still hosting this dataset through your github account. It looks like the download link is broken.\nUpdate: Seems to have been an issue with the Amazon S3 bucket which hosted the data on GitHub's end. Works now</p>",
      "rawMarkdown": "Hey, are you still hosting this dataset through your github account. It looks like the download link is broken.\nUpdate: Seems to have been an issue with the Amazon S3 bucket which hosted the data on GitHub's end. Works now"
    },
    {
      "id": 132481,
      "postDate": "2016-08-25T23:11:16.817Z",
      "content": "<p>Thanks for putting the resized data set together!</p>",
      "rawMarkdown": "Thanks for putting the resized data set together!"
    },
    {
      "id": 1101575,
      "postDate": "2020-12-04T03:23:44.277Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1105558,
          "postDate": "2020-12-08T02:10:45.037Z",
          "content": "<p>It depends on how you're feeding the images to your model. Many CNNs use an input size of 224 or 256 pixels, so if you're going to end up resizing them anyway this would be good enough. If you wanted to try a more robust model or data augmentation strategy that requires more resolution (e.g. analyzing patches at different resolutions or rotating images) then this may not work.</p>",
          "rawMarkdown": "It depends on how you're feeding the images to your model. Many CNNs use an input size of 224 or 256 pixels, so if you're going to end up resizing them anyway this would be good enough. If you wanted to try a more robust model or data augmentation strategy that requires more resolution (e.g. analyzing patches at different resolutions or rotating images) then this may not work."
        }
      ]
    },
    {
      "id": 142257,
      "postDate": "2016-11-01T02:38:26.410Z",
      "content": "<p>Thanks a lot for this!! :D</p>",
      "rawMarkdown": "Thanks a lot for this!! :D"
    },
    {
      "id": 141164,
      "postDate": "2016-10-25T14:11:15.520Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!"
    },
    {
      "id": 134694,
      "postDate": "2016-09-07T23:17:25.793Z",
      "content": "<p>Thank you Michael!</p>",
      "rawMarkdown": "Thank you Michael!"
    },
    {
      "id": 133504,
      "postDate": "2016-09-02T15:49:06.470Z",
      "content": "<p>Thanks. This really helps.</p>",
      "rawMarkdown": "Thanks. This really helps."
    }
  ],
  "comments": [
    {
      "id": 164321,
      "author_name": "Siddharth Dinesh",
      "author_url": "",
      "post_date": "2017-02-28T18:32:18.877000",
      "content": "<p>Hey, are you still hosting this dataset through your github account. It looks like the download link is broken.\nUpdate: Seems to have been an issue with the Amazon S3 bucket which hosted the data on GitHub's end. Works now</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 132481,
      "author_name": "small yellow duck",
      "author_url": "",
      "post_date": "2016-08-25T23:11:16.817000",
      "content": "<p>Thanks for putting the resized data set together!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1101575,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-04T03:23:44.277000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1105558,
          "author_name": "Flynn, Michael",
          "author_url": "",
          "post_date": "2020-12-08T02:10:45.037000",
          "content": "<p>It depends on how you're feeding the images to your model. Many CNNs use an input size of 224 or 256 pixels, so if you're going to end up resizing them anyway this would be good enough. If you wanted to try a more robust model or data augmentation strategy that requires more resolution (e.g. analyzing patches at different resolutions or rotating images) then this may not work.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 142257,
      "author_name": "Apoorv",
      "author_url": "",
      "post_date": "2016-11-01T02:38:26.410000",
      "content": "<p>Thanks a lot for this!! :D</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 141164,
      "author_name": "Vadim Nazarov",
      "author_url": "",
      "post_date": "2016-10-25T14:11:15.520000",
      "content": "<p>Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 134694,
      "author_name": "ckdelta",
      "author_url": "",
      "post_date": "2016-09-07T23:17:25.793000",
      "content": "<p>Thank you Michael!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 133504,
      "author_name": "Sonali Chawla",
      "author_url": "",
      "post_date": "2016-09-02T15:49:06.470000",
      "content": "<p>Thanks. This really helps.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "132480": "For those of you who are having trouble downloading files and are going to be resizing the images to a smaller size (e.g. for a CNN), I've made a preprocessed dataset available [here](https://github.com/zo7/painter-by-numbers/releases/tag/data-v1.0). (Roughly 2.2GB for both the train and test sets)\r\n\r\nThis dataset resizes all of the images so that their smallest side is 256 pixels long, giving you enough wiggle room to do data augmentation and speed up reading/processing. Since I standardized all of the formats this should solve any issues people have been experiencing by not being able to read some images.\r\n\r\nHope this helps!",
    "164321": "Hey, are you still hosting this dataset through your github account. It looks like the download link is broken.\nUpdate: Seems to have been an issue with the Amazon S3 bucket which hosted the data on GitHub's end. Works now",
    "132481": "Thanks for putting the resized data set together!",
    "1101575": "",
    "142257": "Thanks a lot for this!! :D",
    "141164": "Thanks!",
    "134694": "Thank you Michael!",
    "133504": "Thanks. This really helps."
  }
}