{
  "id": 39800,
  "title": "how long your model trained? ",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/39800",
  "author_name": "Five Years",
  "post_date": "2017-09-21T07:58:46.655000",
  "votes": 1,
  "comment_count": 39,
  "views": 0,
  "content": "<p>I trained all image and 2 epoch cost about 1.5 day</p>",
  "messages": [
    {
      "id": 223135,
      "postDate": "2017-09-21T07:58:46.657Z",
      "content": "<p>I trained all image and 2 epoch cost about 1.5 day</p>",
      "rawMarkdown": "I trained all image and 2 epoch cost about 1.5 day",
      "votes": 1
    },
    {
      "id": 224967,
      "postDate": "2017-09-28T01:52:50.160Z",
      "content": "<p>waiting for your result</p>",
      "rawMarkdown": "waiting for your result",
      "votes": -1
    },
    {
      "id": 226344,
      "postDate": "2017-10-02T00:08:21.650Z",
      "content": "<p>My solution is based on  cxflow framework with Xception model and 128 image size, batch_size=300. Using two 1080ti gpus to train , it is  over 10 hours per epoch.  Very slow.  : (</p>",
      "rawMarkdown": "My solution is based on  cxflow framework with Xception model and 128 image size, batch_size=300. Using two 1080ti gpus to train , it is  over 10 hours per epoch.  Very slow.  : ("
    },
    {
      "id": 225946,
      "postDate": "2017-09-30T13:58:12.287Z",
      "content": "<p>I've used SE-ResNet-50 with a 5270-neuron classification layer on Keras + TensorFlow on a single 1080 Ti, batch size 256, training/validation split is 80%-20%, input image size 160x160. It takes ~6 hours to do a single epoch (with 8 workers). </p>\n\n<p>This uses my BSON generator for Keras that I posted in Kernels, but with additional data augmentation, train.bson stored on SSD drive. </p>\n\n<p>Currently I'm creating predictions for the test set using 10 crops per image and this takes about 18 hours... I should try to optimize this a bit. ;-)</p>",
      "rawMarkdown": "I've used SE-ResNet-50 with a 5270-neuron classification layer on Keras + TensorFlow on a single 1080 Ti, batch size 256, training/validation split is 80%-20%, input image size 160x160. It takes ~6 hours to do a single epoch (with 8 workers). \n\nThis uses my BSON generator for Keras that I posted in Kernels, but with additional data augmentation, train.bson stored on SSD drive. \n\nCurrently I'm creating predictions for the test set using 10 crops per image and this takes about 18 hours... I should try to optimize this a bit. ;-)",
      "replies": [
        {
          "id": 226029,
          "postDate": "2017-09-30T19:19:50.877Z",
          "content": "<p>Thanks for sharing, your BSON generator is really helpful! Do you train from scratch or fine-tuning on SE-ResNet-50? I tried to fine-tuning on Xception, it took me 6 hours on a single epoch. (1080, 32RAM, I7, batch size 32) Is that normal?</p>",
          "rawMarkdown": "Thanks for sharing, your BSON generator is really helpful! Do you train from scratch or fine-tuning on SE-ResNet-50? I tried to fine-tuning on Xception, it took me 6 hours on a single epoch. (1080, 32RAM, I7, batch size 32) Is that normal?"
        },
        {
          "id": 226030,
          "postDate": "2017-09-30T19:25:40.793Z",
          "content": "<p>But from scratch, it will take me more than 30 hours (now it takes me 17 hours on half of the training set)</p>",
          "rawMarkdown": "But from scratch, it will take me more than 30 hours (now it takes me 17 hours on half of the training set)"
        },
        {
          "id": 226144,
          "postDate": "2017-10-01T06:52:21.337Z",
          "content": "<p>By the way, how did you do with 8 workers?</p>",
          "rawMarkdown": "By the way, how did you do with 8 workers?"
        },
        {
          "id": 226147,
          "postDate": "2017-10-01T07:07:57.920Z",
          "content": "<p>By the way , how did you do with 8 workers?</p>",
          "rawMarkdown": "By the way , how did you do with 8 workers?"
        },
        {
          "id": 226183,
          "postDate": "2017-10-01T10:00:33.613Z",
          "content": "<p>I did not train from scratch, fine-tuning only (on the last block of the network). It might help to eventually fine-tune <em>all</em> the layers but with a tiny learning rate, but I'd save that for until the network already gets a good score.</p>",
          "rawMarkdown": "I did not train from scratch, fine-tuning only (on the last block of the network). It might help to eventually fine-tune *all* the layers but with a tiny learning rate, but I'd save that for until the network already gets a good score."
        },
        {
          "id": 226201,
          "postDate": "2017-10-01T12:31:59.527Z",
          "content": "<p>Hi Analog, I use your BSONIterator class to fine-tune ResNet18 with the same database(the same machine, the same batch size). It takes about 33 hours for one epoch. I profile the program and find the lock code scope  takes about 70%  of the total time. </p>\n\n<p>Is it normal? Is there any way to speed up the training process?</p>\n\n<p>And as a green hand, I don't how to train it on 8 workers as you,  could you share more details?</p>\n\n<p>Best regard. </p>\n\n<hr>\n\n<pre><code>        with self.lock:\n            image_row = self.images_df.iloc[j]\n            product_id = image_row[\"product_id\"]\n            offset_row = self.offsets_df.loc[product_id]\n            # Read this product's data from the BSON file.\n            self.file.seek(offset_row[\"offset\"])\n            item_data = self.file.read(offset_row[\"length\"])\n</code></pre>",
          "rawMarkdown": "Hi Analog, I use your BSONIterator class to fine-tune ResNet18 with the same database(the same machine, the same batch size). It takes about 33 hours for one epoch. I profile the program and find the lock code scope  takes about 70%  of the total time. \n\nIs it normal? Is there any way to speed up the training process?\n\nAnd as a green hand, I don't how to train it on 8 workers as you,  could you share more details?\n\nBest regard. \n\n\n----------\n\n\n            with self.lock:\n                image_row = self.images_df.iloc[j]\n                product_id = image_row[\"product_id\"]\n                offset_row = self.offsets_df.loc[product_id]\n                # Read this product's data from the BSON file.\n                self.file.seek(offset_row[\"offset\"])\n                item_data = self.file.read(offset_row[\"length\"])\n",
          "votes": 1
        },
        {
          "id": 226238,
          "postDate": "2017-10-01T14:06:35.890Z",
          "content": "<p>It makes sense that it spends most of the time there, since disk I/O is slow (especially since I think you mentioned you're not using an SSD). If it takes 33 hours for one epoch, then your computer spends most of its time generating the batches, and your GPU is mostly waiting (not computing anything).</p>\n\n<p>To use 8 workers, just add the parameter <code>workers=8</code> to the <code>model.fit_generator()</code> call.</p>",
          "rawMarkdown": "It makes sense that it spends most of the time there, since disk I/O is slow (especially since I think you mentioned you're not using an SSD). If it takes 33 hours for one epoch, then your computer spends most of its time generating the batches, and your GPU is mostly waiting (not computing anything).\n\nTo use 8 workers, just add the parameter `workers=8` to the `model.fit_generator()` call."
        },
        {
          "id": 226246,
          "postDate": "2017-10-01T14:42:37.473Z",
          "content": "<p>hi @zhangsongwei which program did you use to profile the code, thanks</p>",
          "rawMarkdown": "hi @zhangsongwei which program did you use to profile the code, thanks"
        },
        {
          "id": 226351,
          "postDate": "2017-10-02T01:01:18.450Z",
          "content": "<p>Just compute the time.</p>",
          "rawMarkdown": "Just compute the time."
        }
      ]
    },
    {
      "id": 225860,
      "postDate": "2017-09-30T06:44:13.263Z",
      "content": "<p>Using the random access kernel, I can train 1 epoch in 10-11hours with the normal resolutions and train-time augmentations (random rotate, flip, and transpose). batch size is 96 for training and 32 for validation. 98-2% split match the lb +- 10^-3</p>",
      "rawMarkdown": "Using the random access kernel, I can train 1 epoch in 10-11hours with the normal resolutions and train-time augmentations (random rotate, flip, and transpose). batch size is 96 for training and 32 for validation. 98-2% split match the lb +- 10^-3",
      "replies": [
        {
          "id": 225862,
          "postDate": "2017-09-30T06:47:38.307Z",
          "content": "<p>Oh, thank you, and what's your hardware?</p>",
          "rawMarkdown": "Oh, thank you, and what's your hardware?"
        },
        {
          "id": 225867,
          "postDate": "2017-09-30T07:34:20.207Z",
          "content": "<p>1080 (unfortunately not Ti), 64 GB of ram, i7, and ~3TB of storage</p>",
          "rawMarkdown": "1080 (unfortunately not Ti), 64 GB of ram, i7, and ~3TB of storage"
        },
        {
          "id": 225868,
          "postDate": "2017-09-30T07:38:37.580Z",
          "content": "<p>16GB ram is so sad</p>",
          "rawMarkdown": "16GB ram is so sad",
          "votes": 1
        },
        {
          "id": 225879,
          "postDate": "2017-09-30T08:38:08.203Z",
          "content": "<p>16GB ram is what I have, too. 32GB would be perfect, but 16GB is just enough!</p>",
          "rawMarkdown": "16GB ram is what I have, too. 32GB would be perfect, but 16GB is just enough!",
          "votes": -1
        },
        {
          "id": 226007,
          "postDate": "2017-09-30T17:05:41.883Z",
          "content": "<p>Yeah 16 is more than enough for this competition if you're using random access</p>",
          "rawMarkdown": "Yeah 16 is more than enough for this competition if you're using random access"
        },
        {
          "id": 226114,
          "postDate": "2017-10-01T04:22:57.740Z",
          "content": "<p>Yes, i just use this method</p>",
          "rawMarkdown": "Yes, i just use this method"
        },
        {
          "id": 230990,
          "postDate": "2017-10-13T11:17:41.147Z",
          "content": "<p>SpruceMoose how many images You use for train and validation set?</p>\n\n<p>You train all layers?</p>\n\n<p>Thanks</p>",
          "rawMarkdown": "SpruceMoose how many images You use for train and validation set?\n\nYou train all layers?\n\nThanks"
        }
      ]
    },
    {
      "id": 225769,
      "postDate": "2017-09-30T01:27:26.097Z",
      "content": "<p>My net is ResNet18, and a epoch took about one day.</p>\n\n<p>Mine is 1080TI. So sad!</p>",
      "rawMarkdown": "My net is ResNet18, and a epoch took about one day.\n\nMine is 1080TI. So sad!",
      "replies": [
        {
          "id": 225865,
          "postDate": "2017-09-30T07:27:07.013Z",
          "content": "<p>It's ok, I use googlenet, the  net is faster.</p>",
          "rawMarkdown": "It's ok, I use googlenet, the  net is faster."
        },
        {
          "id": 225880,
          "postDate": "2017-09-30T08:38:39.060Z",
          "content": "<p>Something must be wrong with your setup. 1 Epoch takes 2 hours for me on my 1080TI with resnet-18</p>",
          "rawMarkdown": "Something must be wrong with your setup. 1 Epoch takes 2 hours for me on my 1080TI with resnet-18"
        },
        {
          "id": 226087,
          "postDate": "2017-10-01T00:40:53.557Z",
          "content": "<p>@Tim Joseph</p>\n\n<p>I used Human's BSON generator for Keras that  posted in his Kernels(<a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson</a>), but without additional data augmentation. And I put my data into ResNet-18 using pytorch. My batch size is 256.</p>\n\n<p>I tried to find out what happened, but didn't succeed. Could you please show more details about you setup?</p>\n\n<p>Thank you!</p>",
          "rawMarkdown": "@Tim Joseph\n\nI used Human's BSON generator for Keras that  posted in his Kernels(https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson), but without additional data augmentation. And I put my data into ResNet-18 using pytorch. My batch size is 256.\n\nI tried to find out what happened, but didn't succeed. Could you please show more details about you setup?\n\nThank you!"
        },
        {
          "id": 226089,
          "postDate": "2017-10-01T00:55:45.253Z",
          "content": "<p>@zhangsongwei does that mean you're using decode without decode_images? (the combining part) If so, I recommend combining product images into one and then resizing. It speeds things up and doesn't seem to hurt acc.</p>",
          "rawMarkdown": "@zhangsongwei does that mean you're using decode without decode_images? (the combining part) If so, I recommend combining product images into one and then resizing. It speeds things up and doesn't seem to hurt acc."
        },
        {
          "id": 226095,
          "postDate": "2017-10-01T02:26:02.630Z",
          "content": "<p>Hi , @SpruceMoose, thanks for your advice. I think you are right , I am using decode  after checking my code. But I am confused about what's the meaning of decode_images, could you please show more details about it?</p>\n\n<p>Is it like this?</p>\n\n<pre><code>for c, d in enumerate(data):\n    product_id = d['_id']\n    category_id = d['category_id'] \n    prod_to_category[product_id] = category_id\n    for e, pic in enumerate(d['imgs']):\n        picture = imread(io.BytesIO(pic['picture']))\n</code></pre>\n\n<p>Thank you!</p>",
          "rawMarkdown": "Hi , @SpruceMoose, thanks for your advice. I think you are right , I am using decode  after checking my code. But I am confused about what's the meaning of decode_images, could you please show more details about it?\n\nIs it like this?\n\n    for c, d in enumerate(data):\n        product_id = d['_id']\n        category_id = d['category_id'] \n        prod_to_category[product_id] = category_id\n        for e, pic in enumerate(d['imgs']):\n            picture = imread(io.BytesIO(pic['picture']))\n\nThank you!",
          "votes": 1
        },
        {
          "id": 226101,
          "postDate": "2017-10-01T03:04:34.877Z",
          "content": "<p>decode_images will simplify life for you because it allows for making a single prediction for a given product id (1 image per product id). What it does is merge the images in product dictionaries with 2-4 images. You will have to resize afterwards. I have a working version that predicts the category id given an unmerged image. However, many products have a different prediction for each image. I'll think about doing mini-ensembles when I get a better model.</p>",
          "rawMarkdown": "decode_images will simplify life for you because it allows for making a single prediction for a given product id (1 image per product id). What it does is merge the images in product dictionaries with 2-4 images. You will have to resize afterwards. I have a working version that predicts the category id given an unmerged image. However, many products have a different prediction for each image. I'll think about doing mini-ensembles when I get a better model."
        },
        {
          "id": 226106,
          "postDate": "2017-10-01T03:35:29.297Z",
          "content": "<p>Just wondering how do you modify resnet 18 to accept smaller image?</p>",
          "rawMarkdown": "Just wondering how do you modify resnet 18 to accept smaller image?"
        },
        {
          "id": 226107,
          "postDate": "2017-10-01T03:39:35.570Z",
          "content": "<p>I resized the input size.</p>",
          "rawMarkdown": "I resized the input size."
        },
        {
          "id": 226112,
          "postDate": "2017-10-01T04:03:37.637Z",
          "content": "<p>@zhangsongwei I meant resize the image that is returned by decode_images if needed (was in my case)</p>\n\n<p>@CSAdu I use cv2.resize. In Keras, maybe just change input_size. In TF, I think you modify the first layer's filters/dimensions to match your input data. In PyTorch, I could be wrong but it seems you just need to care about the number of channels:</p>\n\n<pre><code>in_channels, height, width = in_shape\n\n    self.conv1 = nn.Conv2d(in_channels, 64, kernel_size=7, stride=2, padding=3, bias=False)\n    self.bn1  = nn.BatchNorm2d(64)\n</code></pre>",
          "rawMarkdown": "@zhangsongwei I meant resize the image that is returned by decode_images if needed (was in my case)\n\n@CSAdu I use cv2.resize. In Keras, maybe just change input_size. In TF, I think you modify the first layer's filters/dimensions to match your input data. In PyTorch, I could be wrong but it seems you just need to care about the number of channels:\n\n    in_channels, height, width = in_shape\n\n        self.conv1 = nn.Conv2d(in_channels, 64, kernel_size=7, stride=2, padding=3, bias=False)\n        self.bn1  = nn.BatchNorm2d(64)"
        },
        {
          "id": 226113,
          "postDate": "2017-10-01T04:12:16.813Z",
          "content": "<p>@zhangsongwei @SpriceMoose So what you guys mean just increase picture size to 224?</p>",
          "rawMarkdown": "@zhangsongwei @SpriceMoose So what you guys mean just increase picture size to 224?"
        },
        {
          "id": 226124,
          "postDate": "2017-10-01T04:57:15.970Z",
          "content": "<p>Yes , it seems that multiple workers could solve it.</p>",
          "rawMarkdown": "Yes , it seems that multiple workers could solve it."
        },
        {
          "id": 226127,
          "postDate": "2017-10-01T05:04:24.287Z",
          "content": "<p>Depends on what you're using for transfer learning</p>",
          "rawMarkdown": "Depends on what you're using for transfer learning"
        },
        {
          "id": 226933,
          "postDate": "2017-10-03T11:51:36.617Z",
          "content": "<p>@Tim Joseph</p>\n\n<p>Could you please show more details about your setup? I renew my setup , but it still takes about 6 hours per epoch.(Resnet18  1080Ti  all images)</p>",
          "rawMarkdown": "@Tim Joseph\n\nCould you please show more details about your setup? I renew my setup , but it still takes about 6 hours per epoch.(Resnet18  1080Ti  all images)"
        },
        {
          "id": 226973,
          "postDate": "2017-10-03T13:27:06.100Z",
          "content": "<p>Sorry, I will not discuss my setup here, since it is not based on any kernel and I am not planing to make it public until the competition is over.\nIf I was using any of the kernels I would probably just put the images into folders and then use them like you normally would do. </p>\n\n<p>But I think 6 hours is good already. I gave your wrong information in my post above by mistake. I only run 1 epoch on 9 million images which takes 2 hours. If you are using 15 million images it should take 2 hours 30 min. If you want to keep your current setup you have to identify the bottleneck first and then based on this change your setup!</p>",
          "rawMarkdown": "Sorry, I will not discuss my setup here, since it is not based on any kernel and I am not planing to make it public until the competition is over.\nIf I was using any of the kernels I would probably just put the images into folders and then use them like you normally would do. \n\nBut I think 6 hours is good already. I gave your wrong information in my post above by mistake. I only run 1 epoch on 9 million images which takes 2 hours. If you are using 15 million images it should take 2 hours 30 min. If you want to keep your current setup you have to identify the bottleneck first and then based on this change your setup!",
          "votes": -1
        }
      ]
    },
    {
      "id": 223542,
      "postDate": "2017-09-22T14:16:49.120Z",
      "content": "<p>Which architeture are you using? Mine is InceptionV3 2 epochs. It took 48h.  GTX 1080TI.</p>",
      "rawMarkdown": "Which architeture are you using? Mine is InceptionV3 2 epochs. It took 48h.  GTX 1080TI.",
      "replies": [
        {
          "id": 223548,
          "postDate": "2017-09-22T14:38:36.867Z",
          "content": "<p>GTX 1070 16G RAM CPU i7, on my machine.  I use googlenet, and  after know your time, I feel better in my heart.</p>",
          "rawMarkdown": "GTX 1070 16G RAM CPU i7, on my machine.  I use googlenet, and  after know your time, I feel better in my heart.",
          "votes": 1
        }
      ]
    },
    {
      "id": 223350,
      "postDate": "2017-09-22T01:09:14.493Z",
      "content": "<p>GTX 1070    16G RAM   CPU i7,  on my machine</p>",
      "rawMarkdown": "GTX 1070    16G RAM   CPU i7,  on my machine"
    },
    {
      "id": 223332,
      "postDate": "2017-09-21T22:35:23.877Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 224967,
      "author_name": "Joy",
      "author_url": "",
      "post_date": "2017-09-28T01:52:50.160000",
      "content": "<p>waiting for your result</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 226344,
      "author_name": "huiqin",
      "author_url": "",
      "post_date": "2017-10-02T00:08:21.650000",
      "content": "<p>My solution is based on  cxflow framework with Xception model and 128 image size, batch_size=300. Using two 1080ti gpus to train , it is  over 10 hours per epoch.  Very slow.  : (</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 225946,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2017-09-30T13:58:12.287000",
      "content": "<p>I've used SE-ResNet-50 with a 5270-neuron classification layer on Keras + TensorFlow on a single 1080 Ti, batch size 256, training/validation split is 80%-20%, input image size 160x160. It takes ~6 hours to do a single epoch (with 8 workers). </p>\n\n<p>This uses my BSON generator for Keras that I posted in Kernels, but with additional data augmentation, train.bson stored on SSD drive. </p>\n\n<p>Currently I'm creating predictions for the test set using 10 crops per image and this takes about 18 hours... I should try to optimize this a bit. ;-)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 226029,
          "author_name": "Brian Luo",
          "author_url": "",
          "post_date": "2017-09-30T19:19:50.877000",
          "content": "<p>Thanks for sharing, your BSON generator is really helpful! Do you train from scratch or fine-tuning on SE-ResNet-50? I tried to fine-tuning on Xception, it took me 6 hours on a single epoch. (1080, 32RAM, I7, batch size 32) Is that normal?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226030,
          "author_name": "Brian Luo",
          "author_url": "",
          "post_date": "2017-09-30T19:25:40.793000",
          "content": "<p>But from scratch, it will take me more than 30 hours (now it takes me 17 hours on half of the training set)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226144,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-10-01T06:52:21.337000",
          "content": "<p>By the way, how did you do with 8 workers?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226147,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-10-01T07:07:57.920000",
          "content": "<p>By the way , how did you do with 8 workers?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226183,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2017-10-01T10:00:33.613000",
          "content": "<p>I did not train from scratch, fine-tuning only (on the last block of the network). It might help to eventually fine-tune <em>all</em> the layers but with a tiny learning rate, but I'd save that for until the network already gets a good score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226201,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-10-01T12:31:59.527000",
          "content": "<p>Hi Analog, I use your BSONIterator class to fine-tune ResNet18 with the same database(the same machine, the same batch size). It takes about 33 hours for one epoch. I profile the program and find the lock code scope  takes about 70%  of the total time. </p>\n\n<p>Is it normal? Is there any way to speed up the training process?</p>\n\n<p>And as a green hand, I don't how to train it on 8 workers as you,  could you share more details?</p>\n\n<p>Best regard. </p>\n\n<hr>\n\n<pre><code>        with self.lock:\n            image_row = self.images_df.iloc[j]\n            product_id = image_row[\"product_id\"]\n            offset_row = self.offsets_df.loc[product_id]\n            # Read this product's data from the BSON file.\n            self.file.seek(offset_row[\"offset\"])\n            item_data = self.file.read(offset_row[\"length\"])\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 226238,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2017-10-01T14:06:35.890000",
          "content": "<p>It makes sense that it spends most of the time there, since disk I/O is slow (especially since I think you mentioned you're not using an SSD). If it takes 33 hours for one epoch, then your computer spends most of its time generating the batches, and your GPU is mostly waiting (not computing anything).</p>\n\n<p>To use 8 workers, just add the parameter <code>workers=8</code> to the <code>model.fit_generator()</code> call.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226246,
          "author_name": "Vinh Nguyen",
          "author_url": "",
          "post_date": "2017-10-01T14:42:37.473000",
          "content": "<p>hi @zhangsongwei which program did you use to profile the code, thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226351,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-10-02T01:01:18.450000",
          "content": "<p>Just compute the time.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 225860,
      "author_name": "SpruceMoose",
      "author_url": "",
      "post_date": "2017-09-30T06:44:13.263000",
      "content": "<p>Using the random access kernel, I can train 1 epoch in 10-11hours with the normal resolutions and train-time augmentations (random rotate, flip, and transpose). batch size is 96 for training and 32 for validation. 98-2% split match the lb +- 10^-3</p>",
      "votes": 0,
      "replies": [
        {
          "id": 225862,
          "author_name": "Five Years",
          "author_url": "",
          "post_date": "2017-09-30T06:47:38.307000",
          "content": "<p>Oh, thank you, and what's your hardware?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 225867,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2017-09-30T07:34:20.207000",
          "content": "<p>1080 (unfortunately not Ti), 64 GB of ram, i7, and ~3TB of storage</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 225868,
          "author_name": "Five Years",
          "author_url": "",
          "post_date": "2017-09-30T07:38:37.580000",
          "content": "<p>16GB ram is so sad</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 225879,
          "author_name": "Tim Joseph",
          "author_url": "",
          "post_date": "2017-09-30T08:38:08.203000",
          "content": "<p>16GB ram is what I have, too. 32GB would be perfect, but 16GB is just enough!</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 226007,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2017-09-30T17:05:41.883000",
          "content": "<p>Yeah 16 is more than enough for this competition if you're using random access</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226114,
          "author_name": "Five Years",
          "author_url": "",
          "post_date": "2017-10-01T04:22:57.740000",
          "content": "<p>Yes, i just use this method</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 230990,
          "author_name": "Juan Pizarro",
          "author_url": "",
          "post_date": "2017-10-13T11:17:41.147000",
          "content": "<p>SpruceMoose how many images You use for train and validation set?</p>\n\n<p>You train all layers?</p>\n\n<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 225769,
      "author_name": "Will",
      "author_url": "",
      "post_date": "2017-09-30T01:27:26.097000",
      "content": "<p>My net is ResNet18, and a epoch took about one day.</p>\n\n<p>Mine is 1080TI. So sad!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 225865,
          "author_name": "Five Years",
          "author_url": "",
          "post_date": "2017-09-30T07:27:07.013000",
          "content": "<p>It's ok, I use googlenet, the  net is faster.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 225880,
          "author_name": "Tim Joseph",
          "author_url": "",
          "post_date": "2017-09-30T08:38:39.060000",
          "content": "<p>Something must be wrong with your setup. 1 Epoch takes 2 hours for me on my 1080TI with resnet-18</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226087,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-10-01T00:40:53.557000",
          "content": "<p>@Tim Joseph</p>\n\n<p>I used Human's BSON generator for Keras that  posted in his Kernels(<a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson</a>), but without additional data augmentation. And I put my data into ResNet-18 using pytorch. My batch size is 256.</p>\n\n<p>I tried to find out what happened, but didn't succeed. Could you please show more details about you setup?</p>\n\n<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226089,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2017-10-01T00:55:45.253000",
          "content": "<p>@zhangsongwei does that mean you're using decode without decode_images? (the combining part) If so, I recommend combining product images into one and then resizing. It speeds things up and doesn't seem to hurt acc.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226095,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-10-01T02:26:02.630000",
          "content": "<p>Hi , @SpruceMoose, thanks for your advice. I think you are right , I am using decode  after checking my code. But I am confused about what's the meaning of decode_images, could you please show more details about it?</p>\n\n<p>Is it like this?</p>\n\n<pre><code>for c, d in enumerate(data):\n    product_id = d['_id']\n    category_id = d['category_id'] \n    prod_to_category[product_id] = category_id\n    for e, pic in enumerate(d['imgs']):\n        picture = imread(io.BytesIO(pic['picture']))\n</code></pre>\n\n<p>Thank you!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 226101,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2017-10-01T03:04:34.877000",
          "content": "<p>decode_images will simplify life for you because it allows for making a single prediction for a given product id (1 image per product id). What it does is merge the images in product dictionaries with 2-4 images. You will have to resize afterwards. I have a working version that predicts the category id given an unmerged image. However, many products have a different prediction for each image. I'll think about doing mini-ensembles when I get a better model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226106,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2017-10-01T03:35:29.297000",
          "content": "<p>Just wondering how do you modify resnet 18 to accept smaller image?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226107,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-10-01T03:39:35.570000",
          "content": "<p>I resized the input size.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226112,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2017-10-01T04:03:37.637000",
          "content": "<p>@zhangsongwei I meant resize the image that is returned by decode_images if needed (was in my case)</p>\n\n<p>@CSAdu I use cv2.resize. In Keras, maybe just change input_size. In TF, I think you modify the first layer's filters/dimensions to match your input data. In PyTorch, I could be wrong but it seems you just need to care about the number of channels:</p>\n\n<pre><code>in_channels, height, width = in_shape\n\n    self.conv1 = nn.Conv2d(in_channels, 64, kernel_size=7, stride=2, padding=3, bias=False)\n    self.bn1  = nn.BatchNorm2d(64)\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226113,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2017-10-01T04:12:16.813000",
          "content": "<p>@zhangsongwei @SpriceMoose So what you guys mean just increase picture size to 224?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226124,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-10-01T04:57:15.970000",
          "content": "<p>Yes , it seems that multiple workers could solve it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226127,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2017-10-01T05:04:24.287000",
          "content": "<p>Depends on what you're using for transfer learning</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226933,
          "author_name": "Will",
          "author_url": "",
          "post_date": "2017-10-03T11:51:36.617000",
          "content": "<p>@Tim Joseph</p>\n\n<p>Could you please show more details about your setup? I renew my setup , but it still takes about 6 hours per epoch.(Resnet18  1080Ti  all images)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 226973,
          "author_name": "Tim Joseph",
          "author_url": "",
          "post_date": "2017-10-03T13:27:06.100000",
          "content": "<p>Sorry, I will not discuss my setup here, since it is not based on any kernel and I am not planing to make it public until the competition is over.\nIf I was using any of the kernels I would probably just put the images into folders and then use them like you normally would do. </p>\n\n<p>But I think 6 hours is good already. I gave your wrong information in my post above by mistake. I only run 1 epoch on 9 million images which takes 2 hours. If you are using 15 million images it should take 2 hours 30 min. If you want to keep your current setup you have to identify the bottleneck first and then based on this change your setup!</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 223542,
      "author_name": "Aloisio Dourado",
      "author_url": "",
      "post_date": "2017-09-22T14:16:49.120000",
      "content": "<p>Which architeture are you using? Mine is InceptionV3 2 epochs. It took 48h.  GTX 1080TI.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 223548,
          "author_name": "Five Years",
          "author_url": "",
          "post_date": "2017-09-22T14:38:36.867000",
          "content": "<p>GTX 1070 16G RAM CPU i7, on my machine.  I use googlenet, and  after know your time, I feel better in my heart.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 223350,
      "author_name": "Five Years",
      "author_url": "",
      "post_date": "2017-09-22T01:09:14.493000",
      "content": "<p>GTX 1070    16G RAM   CPU i7,  on my machine</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 223332,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-09-21T22:35:23.877000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "223135": "I trained all image and 2 epoch cost about 1.5 day",
    "224967": "waiting for your result",
    "226344": "My solution is based on  cxflow framework with Xception model and 128 image size, batch_size=300. Using two 1080ti gpus to train , it is  over 10 hours per epoch.  Very slow.  : (",
    "225946": "I've used SE-ResNet-50 with a 5270-neuron classification layer on Keras + TensorFlow on a single 1080 Ti, batch size 256, training/validation split is 80%-20%, input image size 160x160. It takes ~6 hours to do a single epoch (with 8 workers). \n\nThis uses my BSON generator for Keras that I posted in Kernels, but with additional data augmentation, train.bson stored on SSD drive. \n\nCurrently I'm creating predictions for the test set using 10 crops per image and this takes about 18 hours... I should try to optimize this a bit. ;-)",
    "225860": "Using the random access kernel, I can train 1 epoch in 10-11hours with the normal resolutions and train-time augmentations (random rotate, flip, and transpose). batch size is 96 for training and 32 for validation. 98-2% split match the lb +- 10^-3",
    "225769": "My net is ResNet18, and a epoch took about one day.\n\nMine is 1080TI. So sad!",
    "223542": "Which architeture are you using? Mine is InceptionV3 2 epochs. It took 48h.  GTX 1080TI.",
    "223350": "GTX 1070    16G RAM   CPU i7,  on my machine",
    "223332": ""
  }
}