{
  "id": 123248,
  "title": "How long did your model train.",
  "url": "/competitions/bengaliai-cv19/discussion/123248",
  "author_name": "",
  "post_date": "2019-12-26T05:35:00.858284300Z",
  "votes": 10,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I was training xception and it took soooo long. It was out of 9 hours limit so I have to train on colab.(PS I trained three separate models for different targets).</p>",
  "messages": [
    {
      "id": "703403",
      "postDate": "12/26/2019 05:35:00",
      "content": "<p>I was training xception and it took soooo long. It was out of 9 hours limit so I have to train on colab.(PS I trained three separate models for different targets).</p>",
      "rawMarkdown": "I was training xception and it took soooo long. It was out of 9 hours limit so I have to train on colab.(PS I trained three separate models for different targets).",
      "votes": null
    },
    {
      "id": "703936",
      "postDate": "12/26/2019 20:46:57",
      "content": "<p>You can speed up your training time by optimizing your codes that was critical to me. From 1 hour to 20min per epoch by small changing in customized Dataset.</p>\n\n<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/122993\">https://www.kaggle.com/c/bengaliai-cv19/discussion/122993</a> This was helpful to me.</p>\n\n<p>And I am using Numba when preprocessing the images, This seems a bit helpful...?</p>",
      "rawMarkdown": "You can speed up your training time by optimizing your codes that was critical to me. From 1 hour to 20min per epoch by small changing in customized Dataset.\n\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/122993 This was helpful to me.\n\nAnd I am using Numba when preprocessing the images, This seems a bit helpful...?",
      "votes": null
    },
    {
      "id": "704015",
      "postDate": "12/26/2019 23:46:05",
      "content": "<p>I preprocessed all the images in another kernel so it doesn't have anything to do with parquets. What is probably my bottleneck is my data generator. \nI preprocesssed the images are zipped and the zip file is too big to unzip(because the hdd is only 4gb) and I cannot load them all at once(because they will report an OOM error) so I made a data generator that load from zipfile. But loading directly from zipfile cost a lot of time. In colab, I can unzip them all at once and load them at once(because they offer more hdd and more RAM). </p>",
      "rawMarkdown": "I preprocessed all the images in another kernel so it doesn't have anything to do with parquets. What is probably my bottleneck is my data generator. \nI preprocesssed the images are zipped and the zip file is too big to unzip(because the hdd is only 4gb) and I cannot load them all at once(because they will report an OOM error) so I made a data generator that load from zipfile. But loading directly from zipfile cost a lot of time. In colab, I can unzip them all at once and load them at once(because they offer more hdd and more RAM).",
      "votes": null
    },
    {
      "id": "704017",
      "postDate": "12/26/2019 23:56:51",
      "content": "<p>I don't fully understand, but it seems like you converted tabular data to something like png jpg format with zipped, and you unzip the data one by one whenever you feed these into data generator(is it keras?). Why don't you use tabular form instead? I think that's redundant process. People are more familiar with dealing with the image format files as usual image classification competition providing, so that it seems like they tend to convert the data to image file format, there are more memory efficient ways like feather format that I am using // no need to convert it image file and again convert it to 0~255 int coding…</p>",
      "rawMarkdown": "I don't fully understand, but it seems like you converted tabular data to something like png jpg format with zipped, and you unzip the data one by one whenever you feed these into data generator(is it keras?). Why don't you use tabular form instead? I think that's redundant process. People are more familiar with dealing with the image format files as usual image classification competition providing, so that it seems like they tend to convert the data to image file format, there are more memory efficient ways like feather format that I am using // no need to convert it image file and again convert it to 0~255 int coding…",
      "votes": null
    },
    {
      "id": "704019",
      "postDate": "12/27/2019 00:02:16",
      "content": "<p>Yes it is keras. I converted it into image format because I thought it is more convient to deal with. Thank you for your advice. I will give it a shot.</p>",
      "rawMarkdown": "Yes it is keras. I converted it into image format because I thought it is more convient to deal with. Thank you for your advice. I will give it a shot.",
      "votes": null
    },
    {
      "id": "704192",
      "postDate": "12/27/2019 06:36:38",
      "content": "<p>You can preprocess the data then save it to something like .feather format , that's what I did. Trains for about ~3-4 hours at most. Also as Peter said in the discussions , in the data generator , using numpy arrays instead of Pandas DF as your data type saves WAY too much time (about 20x) . Do that with inference too if you haven't done that , because with DenseNet , my inference took &gt; 2 hours so it timed out and submission wasn't accepted. </p>",
      "rawMarkdown": "You can preprocess the data then save it to something like .feather format , that's what I did. Trains for about ~3-4 hours at most. Also as Peter said in the discussions , in the data generator , using numpy arrays instead of Pandas DF as your data type saves WAY too much time (about 20x) . Do that with inference too if you haven't done that , because with DenseNet , my inference took &gt; 2 hours so it timed out and submission wasn't accepted.",
      "votes": null
    },
    {
      "id": "706145",
      "postDate": "12/30/2019 01:12:01",
      "content": "<p>Thanks. Won't the test data be swopped  by other test data? Then the .feather data wouldn't be the same. </p>",
      "rawMarkdown": "Thanks. Won't the test data be swopped  by other test data? Then the .feather data wouldn't be the same.",
      "votes": null
    },
    {
      "id": "706244",
      "postDate": "12/30/2019 06:08:25",
      "content": "<p>Yes , the test data will be replaced with the complete parquet files containing all the data. In the inference we won't be having any feather data , feather data is just used for faster loading in the training kernel. In the inference , we just use the cropping function , and pass it to the dataloader , not save it as .feather files. Basically , in the inference , we are not using feather data at all.</p>",
      "rawMarkdown": "Yes , the test data will be replaced with the complete parquet files containing all the data. In the inference we won't be having any feather data , feather data is just used for faster loading in the training kernel. In the inference , we just use the cropping function , and pass it to the dataloader , not save it as .feather files. Basically , in the inference , we are not using feather data at all.",
      "votes": null
    },
    {
      "id": "706901",
      "postDate": "12/31/2019 01:04:02",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "706995",
      "postDate": "12/31/2019 06:18:23",
      "content": "<p>Yesterday I found out that , if you use a normalization transform , it significantly reduces the training times. Compared to ~5 hours training without normalization , the models with normalization take about 1.5-2 hours only!</p>",
      "rawMarkdown": "Yesterday I found out that , if you use a normalization transform , it significantly reduces the training times. Compared to ~5 hours training without normalization , the models with normalization take about 1.5-2 hours only!",
      "votes": null
    },
    {
      "id": "707326",
      "postDate": "12/31/2019 16:35:28",
      "content": "<p>Do you mean <code>sklearn.preprocessing.Normalizer</code> or the pytorch one?</p>",
      "rawMarkdown": "Do you mean `sklearn.preprocessing.Normalizer` or the pytorch one?",
      "votes": null
    },
    {
      "id": "707342",
      "postDate": "12/31/2019 17:09:50",
      "content": "<p>I used torchvision.transforms.Normalize in the transforms parameter for the dataloader.\nYou could use the sklearn one too I think.</p>",
      "rawMarkdown": "I used torchvision.transforms.Normalize in the transforms parameter for the dataloader.\nYou could use the sklearn one too I think.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 703936,
      "author_name": "hanjoonchoe",
      "author_url": "",
      "post_date": "12/26/2019 20:46:57",
      "content": "<p>You can speed up your training time by optimizing your codes that was critical to me. From 1 hour to 20min per epoch by small changing in customized Dataset.</p>\n\n<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/122993\">https://www.kaggle.com/c/bengaliai-cv19/discussion/122993</a> This was helpful to me.</p>\n\n<p>And I am using Numba when preprocessing the images, This seems a bit helpful...?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 704015,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "12/26/2019 23:46:05",
      "content": "<p>I preprocessed all the images in another kernel so it doesn't have anything to do with parquets. What is probably my bottleneck is my data generator. \nI preprocesssed the images are zipped and the zip file is too big to unzip(because the hdd is only 4gb) and I cannot load them all at once(because they will report an OOM error) so I made a data generator that load from zipfile. But loading directly from zipfile cost a lot of time. In colab, I can unzip them all at once and load them at once(because they offer more hdd and more RAM). </p>",
      "votes": null,
      "replies": [
        {
          "id": 704017,
          "author_name": "hanjoonchoe",
          "author_url": "",
          "post_date": "12/26/2019 23:56:51",
          "content": "<p>I don't fully understand, but it seems like you converted tabular data to something like png jpg format with zipped, and you unzip the data one by one whenever you feed these into data generator(is it keras?). Why don't you use tabular form instead? I think that's redundant process. People are more familiar with dealing with the image format files as usual image classification competition providing, so that it seems like they tend to convert the data to image file format, there are more memory efficient ways like feather format that I am using // no need to convert it image file and again convert it to 0~255 int coding…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 704019,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "12/27/2019 00:02:16",
          "content": "<p>Yes it is keras. I converted it into image format because I thought it is more convient to deal with. Thank you for your advice. I will give it a shot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 704192,
      "author_name": "p4rallax",
      "author_url": "",
      "post_date": "12/27/2019 06:36:38",
      "content": "<p>You can preprocess the data then save it to something like .feather format , that's what I did. Trains for about ~3-4 hours at most. Also as Peter said in the discussions , in the data generator , using numpy arrays instead of Pandas DF as your data type saves WAY too much time (about 20x) . Do that with inference too if you haven't done that , because with DenseNet , my inference took &gt; 2 hours so it timed out and submission wasn't accepted. </p>",
      "votes": null,
      "replies": [
        {
          "id": 706145,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "12/30/2019 01:12:01",
          "content": "<p>Thanks. Won't the test data be swopped  by other test data? Then the .feather data wouldn't be the same. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 706244,
          "author_name": "p4rallax",
          "author_url": "",
          "post_date": "12/30/2019 06:08:25",
          "content": "<p>Yes , the test data will be replaced with the complete parquet files containing all the data. In the inference we won't be having any feather data , feather data is just used for faster loading in the training kernel. In the inference , we just use the cropping function , and pass it to the dataloader , not save it as .feather files. Basically , in the inference , we are not using feather data at all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 706901,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "12/31/2019 01:04:02",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 706995,
          "author_name": "p4rallax",
          "author_url": "",
          "post_date": "12/31/2019 06:18:23",
          "content": "<p>Yesterday I found out that , if you use a normalization transform , it significantly reduces the training times. Compared to ~5 hours training without normalization , the models with normalization take about 1.5-2 hours only!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707326,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "12/31/2019 16:35:28",
          "content": "<p>Do you mean <code>sklearn.preprocessing.Normalizer</code> or the pytorch one?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707342,
          "author_name": "p4rallax",
          "author_url": "",
          "post_date": "12/31/2019 17:09:50",
          "content": "<p>I used torchvision.transforms.Normalize in the transforms parameter for the dataloader.\nYou could use the sklearn one too I think.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "703403": "I was training xception and it took soooo long. It was out of 9 hours limit so I have to train on colab.(PS I trained three separate models for different targets).",
    "703936": "You can speed up your training time by optimizing your codes that was critical to me. From 1 hour to 20min per epoch by small changing in customized Dataset.\n\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/122993 This was helpful to me.\n\nAnd I am using Numba when preprocessing the images, This seems a bit helpful...?",
    "704015": "I preprocessed all the images in another kernel so it doesn't have anything to do with parquets. What is probably my bottleneck is my data generator. \nI preprocesssed the images are zipped and the zip file is too big to unzip(because the hdd is only 4gb) and I cannot load them all at once(because they will report an OOM error) so I made a data generator that load from zipfile. But loading directly from zipfile cost a lot of time. In colab, I can unzip them all at once and load them at once(because they offer more hdd and more RAM).",
    "704017": "I don't fully understand, but it seems like you converted tabular data to something like png jpg format with zipped, and you unzip the data one by one whenever you feed these into data generator(is it keras?). Why don't you use tabular form instead? I think that's redundant process. People are more familiar with dealing with the image format files as usual image classification competition providing, so that it seems like they tend to convert the data to image file format, there are more memory efficient ways like feather format that I am using // no need to convert it image file and again convert it to 0~255 int coding…",
    "704019": "Yes it is keras. I converted it into image format because I thought it is more convient to deal with. Thank you for your advice. I will give it a shot.",
    "704192": "You can preprocess the data then save it to something like .feather format , that's what I did. Trains for about ~3-4 hours at most. Also as Peter said in the discussions , in the data generator , using numpy arrays instead of Pandas DF as your data type saves WAY too much time (about 20x) . Do that with inference too if you haven't done that , because with DenseNet , my inference took &gt; 2 hours so it timed out and submission wasn't accepted.",
    "706145": "Thanks. Won't the test data be swopped  by other test data? Then the .feather data wouldn't be the same.",
    "706244": "Yes , the test data will be replaced with the complete parquet files containing all the data. In the inference we won't be having any feather data , feather data is just used for faster loading in the training kernel. In the inference , we just use the cropping function , and pass it to the dataloader , not save it as .feather files. Basically , in the inference , we are not using feather data at all.",
    "706901": "Thanks!",
    "706995": "Yesterday I found out that , if you use a normalization transform , it significantly reduces the training times. Compared to ~5 hours training without normalization , the models with normalization take about 1.5-2 hours only!",
    "707326": "Do you mean `sklearn.preprocessing.Normalizer` or the pytorch one?",
    "707342": "I used torchvision.transforms.Normalize in the transforms parameter for the dataloader.\nYou could use the sklearn one too I think."
  },
  "source": "meta"
}