{
  "id": 122467,
  "title": "Train image dataset (256px)",
  "url": "/competitions/bengaliai-cv19/discussion/122467",
  "author_name": "",
  "post_date": "2019-12-20T11:08:13.987674400Z",
  "votes": 23,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I saved and uploaded the training images to a dataset. You can download or use it as an external source. <a href=\"https://www.kaggle.com/dataset/a318f9ccd11aea9ede828487914dbbcb76776b72aeb4ef85b51709cfbbe004d3\">You can find it here</a>.</p>\n\n<p>The dataset contains all of the images from the training set converted to 256x256 png format.</p>",
  "messages": [
    {
      "id": "699370",
      "postDate": "12/20/2019 11:08:13",
      "content": "<p>I saved and uploaded the training images to a dataset. You can download or use it as an external source. <a href=\"https://www.kaggle.com/dataset/a318f9ccd11aea9ede828487914dbbcb76776b72aeb4ef85b51709cfbbe004d3\">You can find it here</a>.</p>\n\n<p>The dataset contains all of the images from the training set converted to 256x256 png format.</p>",
      "rawMarkdown": "I saved and uploaded the training images to a dataset. You can download or use it as an external source. [You can find it here](https://www.kaggle.com/dataset/a318f9ccd11aea9ede828487914dbbcb76776b72aeb4ef85b51709cfbbe004d3).\n\nThe dataset contains all of the images from the training set converted to 256x256 png format.",
      "votes": null
    },
    {
      "id": "699412",
      "postDate": "12/20/2019 12:17:46",
      "content": "<p>What process did you use to do the conversion?</p>",
      "rawMarkdown": "What process did you use to do the conversion?",
      "votes": null
    },
    {
      "id": "699474",
      "postDate": "12/20/2019 13:53:38",
      "content": "<p>Nothing extra, just padding with numpy</p>",
      "rawMarkdown": "Nothing extra, just padding with numpy",
      "votes": null
    },
    {
      "id": "699501",
      "postDate": "12/20/2019 14:16:02",
      "content": "<p>Kindly check link.Link not working <a href=\"/pestipeti\">@pestipeti</a> </p>",
      "rawMarkdown": "Kindly check link.Link not working @pestipeti",
      "votes": null
    },
    {
      "id": "699506",
      "postDate": "12/20/2019 14:20:48",
      "content": "<p>Thanks <a href=\"/harunshimanto\">@harunshimanto</a> I've fixed it.</p>",
      "rawMarkdown": "Thanks @harunshimanto I've fixed it.",
      "votes": null
    },
    {
      "id": "699543",
      "postDate": "12/20/2019 15:08:14",
      "content": "<p>That makes sense. I have not been able to add your dataset to my kernel though. :(</p>",
      "rawMarkdown": "That makes sense. I have not been able to add your dataset to my kernel though. :(",
      "votes": null
    },
    {
      "id": "699544",
      "postDate": "12/20/2019 15:11:25",
      "content": "<p>This is the first time I use Kaggle Dataset, I enabled link sharing, but maybe it was not enough. I've just made it public. It should work now.</p>",
      "rawMarkdown": "This is the first time I use Kaggle Dataset, I enabled link sharing, but maybe it was not enough. I've just made it public. It should work now.",
      "votes": null
    },
    {
      "id": "699575",
      "postDate": "12/20/2019 15:49:18",
      "content": "<p>Thank you. It is working now.</p>",
      "rawMarkdown": "Thank you. It is working now.",
      "votes": null
    },
    {
      "id": "699587",
      "postDate": "12/20/2019 16:04:50",
      "content": "<p>I want to point out that it is much much faster to load data from this format. </p>\n\n<p>Loading just 1/4th of the original training dataset in parquet format required 13 minutes 40 seconds. That means it would take about 54 minutes 40 seconds run one single epoch.</p>\n\n<p>On the other hand, loading the full training dataset from PNG images required 4 minutes 50 seconds. </p>\n\n<p>I guess this illustrates the importance of using proper data format. </p>",
      "rawMarkdown": "I want to point out that it is much much faster to load data from this format. \n\nLoading just 1/4th of the original training dataset in parquet format required 13 minutes 40 seconds. That means it would take about 54 minutes 40 seconds run one single epoch.\n\nOn the other hand, loading the full training dataset from PNG images required 4 minutes 50 seconds. \n\nI guess this illustrates the importance of using proper data format.",
      "votes": null
    },
    {
      "id": "699611",
      "postDate": "12/20/2019 16:27:37",
      "content": "<p>Thanks for this info, I did not check it.</p>",
      "rawMarkdown": "Thanks for this info, I did not check it.",
      "votes": null
    },
    {
      "id": "700749",
      "postDate": "12/22/2019 14:48:43",
      "content": "<p><a href=\"/pestipeti\">@pestipeti</a>  Can you please share the code that you used for creating the dataset? Thanks.</p>",
      "rawMarkdown": "pestipeti  Can you please share the code that you used for creating the dataset? Thanks.",
      "votes": null
    },
    {
      "id": "700754",
      "postDate": "12/22/2019 14:54:59",
      "content": "<p>Besides the data loading/image saving, I only used simple padding. You can find the code (<code>make_square</code>) in the inference notebook.</p>",
      "rawMarkdown": "Besides the data loading/image saving, I only used simple padding. You can find the code (`make_square`) in the inference notebook.",
      "votes": null
    },
    {
      "id": "700761",
      "postDate": "12/22/2019 15:07:44",
      "content": "<p>Thanks :)</p>",
      "rawMarkdown": "Thanks :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 699412,
      "author_name": "ibraheemmoosa",
      "author_url": "",
      "post_date": "12/20/2019 12:17:46",
      "content": "<p>What process did you use to do the conversion?</p>",
      "votes": null,
      "replies": [
        {
          "id": 699474,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "12/20/2019 13:53:38",
          "content": "<p>Nothing extra, just padding with numpy</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699543,
          "author_name": "ibraheemmoosa",
          "author_url": "",
          "post_date": "12/20/2019 15:08:14",
          "content": "<p>That makes sense. I have not been able to add your dataset to my kernel though. :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699544,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "12/20/2019 15:11:25",
          "content": "<p>This is the first time I use Kaggle Dataset, I enabled link sharing, but maybe it was not enough. I've just made it public. It should work now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699575,
          "author_name": "ibraheemmoosa",
          "author_url": "",
          "post_date": "12/20/2019 15:49:18",
          "content": "<p>Thank you. It is working now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 699501,
      "author_name": "harunshimanto",
      "author_url": "",
      "post_date": "12/20/2019 14:16:02",
      "content": "<p>Kindly check link.Link not working <a href=\"/pestipeti\">@pestipeti</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 699506,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "12/20/2019 14:20:48",
          "content": "<p>Thanks <a href=\"/harunshimanto\">@harunshimanto</a> I've fixed it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 699587,
      "author_name": "ibraheemmoosa",
      "author_url": "",
      "post_date": "12/20/2019 16:04:50",
      "content": "<p>I want to point out that it is much much faster to load data from this format. </p>\n\n<p>Loading just 1/4th of the original training dataset in parquet format required 13 minutes 40 seconds. That means it would take about 54 minutes 40 seconds run one single epoch.</p>\n\n<p>On the other hand, loading the full training dataset from PNG images required 4 minutes 50 seconds. </p>\n\n<p>I guess this illustrates the importance of using proper data format. </p>",
      "votes": null,
      "replies": [
        {
          "id": 699611,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "12/20/2019 16:27:37",
          "content": "<p>Thanks for this info, I did not check it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 700749,
      "author_name": "tahsin",
      "author_url": "",
      "post_date": "12/22/2019 14:48:43",
      "content": "<p><a href=\"/pestipeti\">@pestipeti</a>  Can you please share the code that you used for creating the dataset? Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 700754,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "12/22/2019 14:54:59",
          "content": "<p>Besides the data loading/image saving, I only used simple padding. You can find the code (<code>make_square</code>) in the inference notebook.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 700761,
          "author_name": "tahsin",
          "author_url": "",
          "post_date": "12/22/2019 15:07:44",
          "content": "<p>Thanks :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "699370": "I saved and uploaded the training images to a dataset. You can download or use it as an external source. [You can find it here](https://www.kaggle.com/dataset/a318f9ccd11aea9ede828487914dbbcb76776b72aeb4ef85b51709cfbbe004d3).\n\nThe dataset contains all of the images from the training set converted to 256x256 png format.",
    "699412": "What process did you use to do the conversion?",
    "699474": "Nothing extra, just padding with numpy",
    "699501": "Kindly check link.Link not working @pestipeti",
    "699506": "Thanks @harunshimanto I've fixed it.",
    "699543": "That makes sense. I have not been able to add your dataset to my kernel though. :(",
    "699544": "This is the first time I use Kaggle Dataset, I enabled link sharing, but maybe it was not enough. I've just made it public. It should work now.",
    "699575": "Thank you. It is working now.",
    "699587": "I want to point out that it is much much faster to load data from this format. \n\nLoading just 1/4th of the original training dataset in parquet format required 13 minutes 40 seconds. That means it would take about 54 minutes 40 seconds run one single epoch.\n\nOn the other hand, loading the full training dataset from PNG images required 4 minutes 50 seconds. \n\nI guess this illustrates the importance of using proper data format.",
    "699611": "Thanks for this info, I did not check it.",
    "700749": "pestipeti  Can you please share the code that you used for creating the dataset? Thanks.",
    "700754": "Besides the data loading/image saving, I only used simple padding. You can find the code (`make_square`) in the inference notebook.",
    "700761": "Thanks :)"
  },
  "source": "meta"
}