{
  "id": 40498,
  "title": "pytorch starter kit here!",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/40498",
  "author_name": "",
  "post_date": "2017-10-03T15:34:21.260391200Z",
  "votes": 62,
  "comment_count": 63,
  "views": 0,
  "content": "<p>** important **\nplease refer to @Vladimir Iglovikov at <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652</a></p>\n\n<p>please train with long epoch to get better results . My models results are from training only up to 4 to 5 epoch. </p>\n\n<hr>\n\n<p>This is a pytorch starter kit. It is not completed yet. I use ideas from the kernels and discussion contributions. Thanks to kagglers, especially @Bruno G. do Amaral, @Human Analog, @Adam Blazek for code and discussion.</p>\n\n<p>Download: \n<a href=\"https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing\">https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing</a></p>\n\n<hr>\n\n<p>version.10-03:</p>\n\n<ul>\n<li>initial version for early experiments. see \"trainer.py\"</li>\n</ul>\n\n<hr>\n\n<p>version.10-07:</p>\n\n<ul>\n<li><p>LB 0.63358(160x160 single center crop) or 0.63839(180x180) or  0.64356 (180x180 resize to 160x160) for se-resnet50 trained on 160x160 crops.</p></li>\n<li><p>LB 0.66708 for same se-resnet50 if train=224x224 crops from 256, test = resize to 224</p></li>\n<li><p>LB 0.68939 for same se-resnet50 if train=180x180, test = 180x180 (weight initialised from 224 se-resnet50 of above)</p></li>\n<li><p>LB 0.69565 for inception3 if train=180x180, test = 180x180 (be careful on how to control the lr, batch size, momentum, level of augmentation,etc. I will have a writeup on this later)</p></li>\n<li><p>support various network like resnet, resnext, inception v3, etc. see trainer_xxx.py</p></li>\n<li><p>support gradient accumulation </p></li>\n</ul>\n\n<p>note: code is messy and dirty. It may not be backward- compatible.</p>\n\n<hr>\n\n<p>version.10-17 (latest):</p>\n\n<ul>\n<li><p>advance version. implement multi-gpu support in training. see \"trainer_excited_resnext101_32x4d.py\"</p></li>\n<li><p>LB 0.71064 for se-resnext101_32x4d (single crop 180/180)</p></li>\n<li><p>support xception, se-xception, resnext101_32x4d, se-resnext101_32x4d, inceptionv4, inception_resnetv2 </p></li>\n<li><p>folder \"10-17/senet_conversion\" contains conversion tools for converting caffe model to pytorch  </p></li>\n<li><p>note: </p>\n\n<p><em>1.</em> may not be backward compatible </p>\n\n<p><em>2.</em> old training scripts are found in \"dummy-00/<strong>temp</strong>/old_trainer\"</p></li>\n</ul>\n\n<hr>\n\n<p><br>\nnext version:</p>\n\n<ul>\n<li><p>focal loss for class imbalance</p></li>\n<li><p>super large batch size to speed up training</p></li>\n<li><p>efficient multi-crop for testing</p></li>\n<li><p>divide test/train data into easy and difficult set for faster testing/training.  class balancing</p></li>\n</ul>\n\n<hr>\n\n<p>Note: </p>\n\n<p>My software and thread will be constantly update. Also, ther great works:</p>\n\n<ul>\n<li><p><a href=\"https://www.kaggle.com/blazeka/multi-gpu-tensorflow-convnet-0-65/notebook\">https://www.kaggle.com/blazeka/multi-gpu-tensorflow-convnet-0-65/notebook</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40715\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40715</a></p></li>\n</ul>",
  "messages": [
    {
      "id": "227035",
      "postDate": "10/03/2017 15:34:21",
      "content": "<p>** important **\nplease refer to @Vladimir Iglovikov at <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652</a></p>\n\n<p>please train with long epoch to get better results . My models results are from training only up to 4 to 5 epoch. </p>\n\n<hr>\n\n<p>This is a pytorch starter kit. It is not completed yet. I use ideas from the kernels and discussion contributions. Thanks to kagglers, especially @Bruno G. do Amaral, @Human Analog, @Adam Blazek for code and discussion.</p>\n\n<p>Download: \n<a href=\"https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing\">https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing</a></p>\n\n<hr>\n\n<p>version.10-03:</p>\n\n<ul>\n<li>initial version for early experiments. see \"trainer.py\"</li>\n</ul>\n\n<hr>\n\n<p>version.10-07:</p>\n\n<ul>\n<li><p>LB 0.63358(160x160 single center crop) or 0.63839(180x180) or  0.64356 (180x180 resize to 160x160) for se-resnet50 trained on 160x160 crops.</p></li>\n<li><p>LB 0.66708 for same se-resnet50 if train=224x224 crops from 256, test = resize to 224</p></li>\n<li><p>LB 0.68939 for same se-resnet50 if train=180x180, test = 180x180 (weight initialised from 224 se-resnet50 of above)</p></li>\n<li><p>LB 0.69565 for inception3 if train=180x180, test = 180x180 (be careful on how to control the lr, batch size, momentum, level of augmentation,etc. I will have a writeup on this later)</p></li>\n<li><p>support various network like resnet, resnext, inception v3, etc. see trainer_xxx.py</p></li>\n<li><p>support gradient accumulation </p></li>\n</ul>\n\n<p>note: code is messy and dirty. It may not be backward- compatible.</p>\n\n<hr>\n\n<p>version.10-17 (latest):</p>\n\n<ul>\n<li><p>advance version. implement multi-gpu support in training. see \"trainer_excited_resnext101_32x4d.py\"</p></li>\n<li><p>LB 0.71064 for se-resnext101_32x4d (single crop 180/180)</p></li>\n<li><p>support xception, se-xception, resnext101_32x4d, se-resnext101_32x4d, inceptionv4, inception_resnetv2 </p></li>\n<li><p>folder \"10-17/senet_conversion\" contains conversion tools for converting caffe model to pytorch  </p></li>\n<li><p>note: </p>\n\n<p><em>1.</em> may not be backward compatible </p>\n\n<p><em>2.</em> old training scripts are found in \"dummy-00/<strong>temp</strong>/old_trainer\"</p></li>\n</ul>\n\n<hr>\n\n<p><br>\nnext version:</p>\n\n<ul>\n<li><p>focal loss for class imbalance</p></li>\n<li><p>super large batch size to speed up training</p></li>\n<li><p>efficient multi-crop for testing</p></li>\n<li><p>divide test/train data into easy and difficult set for faster testing/training.  class balancing</p></li>\n</ul>\n\n<hr>\n\n<p>Note: </p>\n\n<p>My software and thread will be constantly update. Also, ther great works:</p>\n\n<ul>\n<li><p><a href=\"https://www.kaggle.com/blazeka/multi-gpu-tensorflow-convnet-0-65/notebook\">https://www.kaggle.com/blazeka/multi-gpu-tensorflow-convnet-0-65/notebook</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40715\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40715</a></p></li>\n</ul>",
      "rawMarkdown": "** important **\nplease refer to @Vladimir Iglovikov at https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652\n\nplease train with long epoch to get better results . My models results are from training only up to 4 to 5 epoch. \n\n------\n\nThis is a pytorch starter kit. It is not completed yet. I use ideas from the kernels and discussion contributions. Thanks to kagglers, especially @Bruno G. do Amaral, @Human Analog, @Adam Blazek for code and discussion.\n\n\nDownload: \nhttps://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing\n\n---------------\nversion.10-03:\n\n- initial version for early experiments. see \"trainer.py\"\n\n---------------\n\nversion.10-07:\n\n- LB 0.63358(160x160 single center crop) or 0.63839(180x180) or  0.64356 (180x180 resize to 160x160) for se-resnet50 trained on 160x160 crops.\n\n- LB 0.66708 for same se-resnet50 if train=224x224 crops from 256, test = resize to 224\n\n- LB 0.68939 for same se-resnet50 if train=180x180, test = 180x180 (weight initialised from 224 se-resnet50 of above)\n\n-  LB 0.69565 for inception3 if train=180x180, test = 180x180 (be careful on how to control the lr, batch size, momentum, level of augmentation,etc. I will have a writeup on this later)\n\n-  support various network like resnet, resnext, inception v3, etc. see trainer_xxx.py\n\n- support gradient accumulation \n\nnote: code is messy and dirty. It may not be backward- compatible.\n\n---------------\n\nversion.10-17 (latest):\n\n- advance version. implement multi-gpu support in training. see \"trainer_excited_resnext101_32x4d.py\"\n\n- LB 0.71064 for se-resnext101_32x4d (single crop 180/180)\n\n- support xception, se-xception, resnext101_32x4d, se-resnext101_32x4d, inceptionv4, inception_resnetv2 \n\n-  folder \"10-17/senet_conversion\" contains conversion tools for converting caffe model to pytorch  \n\n- note: \n\n     *1.* may not be backward compatible \n\n     *2.* old training scripts are found in \"dummy-00/__temp__/old_trainer\"\n\n---------------     \n\n<br>\nnext version:\n \n- focal loss for class imbalance\n\n- super large batch size to speed up training\n\n- efficient multi-crop for testing\n\n- divide test/train data into easy and difficult set for faster testing/training.  class balancing\n\n\n---------------\n\nNote: \n\nMy software and thread will be constantly update. Also, ther great works:\n\n- https://www.kaggle.com/blazeka/multi-gpu-tensorflow-convnet-0-65/notebook\n\n- https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40715",
      "votes": null
    },
    {
      "id": "227081",
      "postDate": "10/03/2017 17:19:22",
      "content": "<p>Welcome Heng in this new competition. I am quite sure that this one is also going to be very tight competitive one because of your generous contributions towards the solution. Keep it up. However, I think, in general, this competition will give an advantage to those who have access to more GPU resources. Otherwise, it will be interesting to see how people come up with ingenious ideas to deal with this large dataset. Good luck for a gold medal this time!</p>",
      "rawMarkdown": "Welcome Heng in this new competition. I am quite sure that this one is also going to be very tight competitive one because of your generous contributions towards the solution. Keep it up. However, I think, in general, this competition will give an advantage to those who have access to more GPU resources. Otherwise, it will be interesting to see how people come up with ingenious ideas to deal with this large dataset. Good luck for a gold medal this time!",
      "votes": null
    },
    {
      "id": "227387",
      "postDate": "10/04/2017 08:37:07",
      "content": "<p>Trained more iterations to get LB 58.6% (12 hr).  initial observations:</p>\n\n<ul>\n<li><p>Too much layers are frozen. Next step is to retrain/finetune with more parameters. </p></li>\n<li><p>I estimate the results of single crop resnet50 is around LB = 62 to 65%</p></li>\n<li><p>I estimate the leader board top score by the kagglers to be near to 82% at the end of the competition. </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/227387/7469/loss.png\" alt=\"enter image description here\" title=\"\"></p></li>\n</ul>",
      "rawMarkdown": "Trained more iterations to get LB 58.6% (12 hr).  initial observations:\n\n- Too much layers are frozen. Next step is to retrain/finetune with more parameters. \n\n- I estimate the results of single crop resnet50 is around LB = 62 to 65%\n\n- I estimate the leader board top score by the kagglers to be near to 82% at the end of the competition. \n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/227387/7469/loss.png",
      "votes": null
    },
    {
      "id": "227862",
      "postDate": "10/05/2017 10:06:35",
      "content": "<p>** experiments on pretrain/freezing, learning rates. etc **</p>\n\n<p>baseline results for resnet50 LB = 0.62978 :</p>\n\n<ol>\n<li><p>train all layers (no freezing)</p></li>\n<li><p>rates:</p>\n\n<ul><li><p>0.01 for epoch 1,2</p></li>\n<li><p>0.001 for epoch 3</p></li>\n<li><p>0.0001 till 3.25</p></li></ul>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/227862/7495/loss_full.png\" alt=\"enter image description here\" title=\"\"></p></li>\n</ol>",
      "rawMarkdown": "** experiments on pretrain/freezing, learning rates. etc **\n\n\nbaseline results for resnet50 LB = 0.62978 :\n\n1. train all layers (no freezing)\n\n2. rates:\n\n - 0.01 for epoch 1,2\n\n -  0.001 for epoch 3\n\n -  0.0001 till 3.25\n\n\n   ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/227862/7495/loss_full.png",
      "votes": null
    },
    {
      "id": "227923",
      "postDate": "10/05/2017 13:36:52",
      "content": "<p>Hi Heng, good to see you in this competition again.\nIs there a reason you don't use github instead of Google Drive for your code?</p>",
      "rawMarkdown": "Hi Heng, good to see you in this competition again.\nIs there a reason you don't use github instead of Google Drive for your code?",
      "votes": null
    },
    {
      "id": "227928",
      "postDate": "10/05/2017 13:41:50",
      "content": "<p>i have an github account too. But i find it difficult to update because my code organisation is messy and file size limit. Maybe i will try to learn the pycharm github integration  or use some friendly github gui tool one day.</p>",
      "rawMarkdown": "i have an github account too. But i find it difficult to update because my code organisation is messy and file size limit. Maybe i will try to learn the pycharm github integration  or use some friendly github gui tool one day.",
      "votes": null
    },
    {
      "id": "227935",
      "postDate": "10/05/2017 13:57:29",
      "content": "<p>I think for public projects like yours it would be much better to split your project into source code and files. So you can still host the big files on Google Drive, but everything related to source code on github.\nThere are three major advantages: 1. Everyone can easily read the code without downloading everything. 2. If you use readable commit messages you create an implicit change log. 3. Issue tracker + Pull requests can make this a community effort.</p>\n\n<p>That said, I appreciate your effort no matter where all the stuff is hosted ;D</p>",
      "rawMarkdown": "I think for public projects like yours it would be much better to split your project into source code and files. So you can still host the big files on Google Drive, but everything related to source code on github.\nThere are three major advantages: 1. Everyone can easily read the code without downloading everything. 2. If you use readable commit messages you create an implicit change log. 3. Issue tracker + Pull requests can make this a community effort.\n\nThat said, I appreciate your effort no matter where all the stuff is hosted ;D",
      "votes": null
    },
    {
      "id": "227946",
      "postDate": "10/05/2017 14:09:20",
      "content": "<p>Good suggestion! I will get it done when i have time :)</p>",
      "rawMarkdown": "Good suggestion! I will get it done when i have time :)",
      "votes": null
    },
    {
      "id": "228236",
      "postDate": "10/06/2017 07:14:11",
      "content": "<p>** experiments on input sizes, network strides. etc **</p>\n\n<p>224x224 (bilinear up-scale) input vs 160x160 input for resnet50. \n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228236/7507/Slide2.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "** experiments on input sizes, network strides. etc **\n\n224x224 (bilinear up-scale) input vs 160x160 input for resnet50. \n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228236/7507/Slide2.png",
      "votes": null
    },
    {
      "id": "228238",
      "postDate": "10/06/2017 07:17:46",
      "content": "<p>To get top results, you should take care of the issues:</p>\n\n<ol>\n<li><p>balancing class  </p></li>\n<li><p>small objects</p></li>\n<li><p>how to train efficiency (extract information efficiency, not all data provide \"useful\" information) </p></li>\n<li><p>break down the problem into simpler ones (e.g. special handling for minority class)</p></li>\n</ol>\n\n<p>... to be updated ...</p>",
      "rawMarkdown": "To get top results, you should take care of the issues:\n\n 1. balancing class  \n\n 2. small objects\n\n 3.  how to train efficiency (extract information efficiency, not all data provide \"useful\" information) \n\n 4. break down the problem into simpler ones (e.g. special handling for minority class)\n\n... to be updated ...",
      "votes": null
    },
    {
      "id": "228255",
      "postDate": "10/06/2017 08:08:56",
      "content": "<p>** experiments on different network structure **</p>\n\n<p>...this is still in progress ...</p>\n\n<p>based on initial results it seems that any of the large network (those that obtained about 20% top1 error on imagenet like resnext, inceptionv3,  resnet, etc) can get about 67to 69% for single crop on cdiscount image. With proper\n ensemble, 70 to 72% is easily obtainable. Care has to be taken in augmentation, input resolution, learning rate, batch size, etc. I estimate top 50 ranks results &gt;72% at the end of the competitions if kagglers have enough gpu.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228255/7517/large_input_nets2.png\" alt=\"enter image description here\" title=\"\">\n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228255/7536/batch_size1.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "** experiments on different network structure **\n\n...this is still in progress ...\n\nbased on initial results it seems that any of the large network (those that obtained about 20% top1 error on imagenet like resnext, inceptionv3,  resnet, etc) can get about 67to 69% for single crop on cdiscount image. With proper\n ensemble, 70 to 72% is easily obtainable. Care has to be taken in augmentation, input resolution, learning rate, batch size, etc. I estimate top 50 ranks results &gt;72% at the end of the competitions if kagglers have enough gpu.\n\n ![enter image description here][1]\n ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228255/7517/large_input_nets2.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228255/7536/batch_size1.png",
      "votes": null
    },
    {
      "id": "228593",
      "postDate": "10/07/2017 06:42:07",
      "content": "<p>It seems that data augmentation can improve balancing class.</p>",
      "rawMarkdown": "It seems that data augmentation can improve balancing class.",
      "votes": null
    },
    {
      "id": "228661",
      "postDate": "10/07/2017 12:05:50",
      "content": "<p>compare resnet50 and se-resnet50</p>\n\n<p>(there is a typo mistake in the figure below. the LB results should be flipped, i.e. SE-resnet is better)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228661/7515/se-resnet50.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "compare resnet50 and se-resnet50\n\n(there is a typo mistake in the figure below. the LB results should be flipped, i.e. SE-resnet is better)\n\n![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228661/7515/se-resnet50.png",
      "votes": null
    },
    {
      "id": "228684",
      "postDate": "10/07/2017 13:40:41",
      "content": "<p>** experiments on augmentation **</p>\n\n<p>... to be updated ...</p>\n\n<p>seems that crop 160x160 from 180x180 is not a good way to augment. I suspect downsize 180x180 to 160x160 gives better results. But I do not have time or extra gpu to do this experiments. anyone  has done this comparison before?</p>",
      "rawMarkdown": "** experiments on augmentation **\n\n... to be updated ...\n\nseems that crop 160x160 from 180x180 is not a good way to augment. I suspect downsize 180x180 to 160x160 gives better results. But I do not have time or extra gpu to do this experiments. anyone  has done this comparison before?",
      "votes": null
    },
    {
      "id": "228735",
      "postDate": "10/07/2017 17:52:09",
      "content": "<p>Heng, how did you create pretrain_convert_table.py and what weights are you loading?</p>",
      "rawMarkdown": "Heng, how did you create pretrain_convert_table.py and what weights are you loading?",
      "votes": null
    },
    {
      "id": "228777",
      "postDate": "10/07/2017 20:27:58",
      "content": "<p>this is for resnet50. the weights are actaully the same as the default pytorch model zoo. i reorganize resnet for future experiments and the naming changed. i manually create the mapping for the model zoo keys to my keys. actually you can just use the default pytorch model zoo and ignore my resnet_xx.py</p>",
      "rawMarkdown": "this is for resnet50. the weights are actaully the same as the default pytorch model zoo. i reorganize resnet for future experiments and the naming changed. i manually create the mapping for the model zoo keys to my keys. actually you can just use the default pytorch model zoo and ignore my resnet_xx.py",
      "votes": null
    },
    {
      "id": "228890",
      "postDate": "10/08/2017 07:26:43",
      "content": "<p>Hi, Heng, I use similar way to load data as you, converting bson to files. </p>\n\n<p>The loading process is very fast at first(0.3s, CPU 30%, RAM 10% of 30G), but some iterations later, it becomes very slow(6s, CPU 5%, RAM 10% of 30G).</p>\n\n<p>I set the CPU mode to performance, and try different num_workers(0,4,8,12 …), but the situation stays the same.</p>\n\n<p>Here is my code:</p>\n\n<pre><code>from torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\n\n\nclass xiongDataset(Dataset):\n    def __init__(self, csv_file, root_dir, transform=None):\n        self.train_names=[]\n        self.root_dir=root_dir\n        self.transform = transform\n        train_images = pd.read_csv(csv_file)\n        train_ids = list(train_images['product_id'])\n        train_idxs = list(train_images['img_idx'])\n        self.labels = list(train_images['category_idx'])\n        num_train = len(train_images)\n       for i in range(num_train):\n           train_name = '{}_{}.jpg'.format(train_ids[i],train_idxs[i])\n           self.train_names.append(train_name)\n\n   def __len__(self):\n       return len(self.train_names)\n\n   def __getitem__(self, idx):\n       img = cv2.imread(self.root_dir + self.train_names[idx])\n       label = self.labels[idx]\n       if self.transform is not None:\n            img = self.transform(img)\n\n        return img,label\n\ntrain_data = xiongDataset('../train_images.csv','../train/',transform=transforms.ToTensor())\n\n\n\ndata_loader= DataLoader(train_data,batch_size=256,shuffle=False,num_workers=0)\n</code></pre>",
      "rawMarkdown": "Hi, Heng, I use similar way to load data as you, converting bson to files. \n\nThe loading process is very fast at first(0.3s, CPU 30%, RAM 10% of 30G), but some iterations later, it becomes very slow(6s, CPU 5%, RAM 10% of 30G).\n\nI set the CPU mode to performance, and try different num_workers(0,4,8,12 …), but the situation stays the same.\n\nHere is my code:\n   \n    from torch.utils.data import Dataset, DataLoader\n    from torchvision import transforms\n\n\n    class xiongDataset(Dataset):\n        def __init__(self, csv_file, root_dir, transform=None):\n            self.train_names=[]\n            self.root_dir=root_dir\n            self.transform = transform\n            train_images = pd.read_csv(csv_file)\n            train_ids = list(train_images['product_id'])\n            train_idxs = list(train_images['img_idx'])\n            self.labels = list(train_images['category_idx'])\n            num_train = len(train_images)\n           for i in range(num_train):\n               train_name = '{}_{}.jpg'.format(train_ids[i],train_idxs[i])\n               self.train_names.append(train_name)\n\n       def __len__(self):\n           return len(self.train_names)\n\n       def __getitem__(self, idx):\n           img = cv2.imread(self.root_dir + self.train_names[idx])\n           label = self.labels[idx]\n           if self.transform is not None:\n                img = self.transform(img)\n\n            return img,label\n\n    train_data = xiongDataset('../train_images.csv','../train/',transform=transforms.ToTensor())\n\n\n\n    data_loader= DataLoader(train_data,batch_size=256,shuffle=False,num_workers=0)",
      "votes": null
    },
    {
      "id": "228891",
      "postDate": "10/08/2017 07:32:19",
      "content": "<p>i use ssd drive (solid state drive). if you are using regular hard disk, you need a better random access database format like leveldb, hdfs i think.</p>\n\n<p>when i use  regular hard disk, i face the same problem you said.</p>",
      "rawMarkdown": "i use ssd drive (solid state drive). if you are using regular hard disk, you need a better random access database format like leveldb, hdfs i think.\n\nwhen i use  regular hard disk, i face the same problem you said.",
      "votes": null
    },
    {
      "id": "228892",
      "postDate": "10/08/2017 07:51:44",
      "content": "<p>Thank you :)</p>",
      "rawMarkdown": "Thank you :)",
      "votes": null
    },
    {
      "id": "228972",
      "postDate": "10/08/2017 12:42:45",
      "content": "<p>Thank you for your advice. I do as what you said, and  make it. Now, loading data only takes less than 0.08 seconds, and one iteration takes 0.3 seconds(batch size=256, Resnet18, a 1080Ti).</p>",
      "rawMarkdown": "Thank you for your advice. I do as what you said, and  make it. Now, loading data only takes less than 0.08 seconds, and one iteration takes 0.3 seconds(batch size=256, Resnet18, a 1080Ti).",
      "votes": null
    },
    {
      "id": "229345",
      "postDate": "10/09/2017 12:28:08",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "229987",
      "postDate": "10/11/2017 00:47:45",
      "content": "<p>Hi, Heng, have you tried multi-gpu? I also train with pytorch, and I set <code>model = nn.DataParallel(model, device_ids=[0, 1])</code>, but there is several errors <a href=\"https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492\">https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492</a></p>\n\n<p>And then I tried your way <code>os.environ[\"CUDA_VISIBLE_DEVICES\"] = '0,1'</code>(I have 2 GPUs), but only one GPU works.</p>",
      "rawMarkdown": "Hi, Heng, have you tried multi-gpu? I also train with pytorch, and I set ```model = nn.DataParallel(model, device_ids=[0, 1])```, but there is several errors https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492\n\nAnd then I tried your way ```os.environ[\"CUDA_VISIBLE_DEVICES\"] = '0,1'```(I have 2 GPUs), but only one GPU works.",
      "votes": null
    },
    {
      "id": "229995",
      "postDate": "10/11/2017 01:28:44",
      "content": "<p>haven't try yet. multi-gpu development would come later</p>",
      "rawMarkdown": "haven't try yet. multi-gpu development would come later",
      "votes": null
    },
    {
      "id": "230220",
      "postDate": "10/11/2017 14:36:31",
      "content": "<p>I have encountered same problem by using ssd. Probably it occurred due to GC, so adding a line <code>gc.disable()</code> in code is a way (I am scared that it may cause some bad effect...).</p>",
      "rawMarkdown": "I have encountered same problem by using ssd. Probably it occurred due to GC, so adding a line `gc.disable()` in code is a way (I am scared that it may cause some bad effect...).",
      "votes": null
    },
    {
      "id": "230350",
      "postDate": "10/11/2017 19:30:48",
      "content": "<p>How many GPUs do you have? I remember you have 4 Titan X? Just want to compare the running time.</p>",
      "rawMarkdown": "How many GPUs do you have? I remember you have 4 Titan X? Just want to compare the running time.",
      "votes": null
    },
    {
      "id": "230391",
      "postDate": "10/11/2017 21:25:24",
      "content": "<p>Hello Heng,</p>\n\n<p>Thanks a lot for sharing your scripts and models so generously! I am struggling to break through 50% acc so going through your scripts will be very helpful to find the issues in my pipeline.</p>\n\n<p>I have a few questions on your train &amp; valid split:</p>\n\n<ul>\n<li>You only used 5K samples for valid. How did you choose them? how do you ensure that the class weights are respected?</li>\n<li>There are 10k less samples in your train vs total labeled.  Did you choose to drop 5K samples on purpose?</li>\n</ul>",
      "rawMarkdown": "Hello Heng,\n\nThanks a lot for sharing your scripts and models so generously! I am struggling to break through 50% acc so going through your scripts will be very helpful to find the issues in my pipeline.\n\nI have a few questions on your train &amp; valid split:\n\n-  You only used 5K samples for valid. How did you choose them? how do you ensure that the class weights are respected?\n- There are 10k less samples in your train vs total labeled.  Did you choose to drop 5K samples on purpose?",
      "votes": null
    },
    {
      "id": "230393",
      "postDate": "10/11/2017 21:37:14",
      "content": "<p>validation set is not 5k. it is 50k.</p>\n\n<p>the 5k is used in the training iteration is a subset of validation to make \"visualisation\" fast.</p>\n\n<p>thay are choosen randomly. (not the best way)</p>\n\n<p>I will come back to data sampling later.  now i am still at early stage of experiments</p>",
      "rawMarkdown": "validation set is not 5k. it is 50k.\n\nthe 5k is used in the training iteration is a subset of validation to make \"visualisation\" fast.\n\nthay are choosen randomly. (not the best way)\n\nI will come back to data sampling later.  now i am still at early stage of experiments",
      "votes": null
    },
    {
      "id": "230397",
      "postDate": "10/11/2017 21:50:04",
      "content": "<p>i am using 1 pascal titanx and 3 1080 ti. but each gpu is use to train one model. i haven't train across multiple gpu yet.  i am still \"exploring\" different models</p>",
      "rawMarkdown": "i am using 1 pascal titanx and 3 1080 ti. but each gpu is use to train one model. i haven't train across multiple gpu yet.  i am still \"exploring\" different models",
      "votes": null
    },
    {
      "id": "230576",
      "postDate": "10/12/2017 08:54:40",
      "content": "<p>Can't find pretrained_file:  /root/share/data/models/pytorch/imagenet/resenet/resnet50-19c8e357.pth</p>",
      "rawMarkdown": "Can't find pretrained_file:  /root/share/data/models/pytorch/imagenet/resenet/resnet50-19c8e357.pth",
      "votes": null
    },
    {
      "id": "230577",
      "postDate": "10/12/2017 08:56:13",
      "content": "<p><a href=\"https://www.google.com.sg/search?q=resnet50-19c8e357.pth&amp;oq=resnet50-19c8e357.pth&amp;aqs=chrome..69i57.1903j0j4&amp;sourceid=chrome&amp;ie=UTF-8\">https://www.google.com.sg/search?q=resnet50-19c8e357.pth&amp;oq=resnet50-19c8e357.pth&amp;aqs=chrome..69i57.1903j0j4&amp;sourceid=chrome&amp;ie=UTF-8</a></p>",
      "rawMarkdown": "https://www.google.com.sg/search?q=resnet50-19c8e357.pth&amp;oq=resnet50-19c8e357.pth&amp;aqs=chrome..69i57.1903j0j4&amp;sourceid=chrome&amp;ie=UTF-8",
      "votes": null
    },
    {
      "id": "230596",
      "postDate": "10/12/2017 09:52:49",
      "content": "<p>How to support multi-gpu ? According  to <a href=\"https://discuss.pytorch.org/t/does-it-support-multi-gpu-card-on-a-single-node/75/7\">https://discuss.pytorch.org/t/does-it-support-multi-gpu-card-on-a-single-node/75/7</a> this shouldn't be too difficult ?</p>",
      "rawMarkdown": "How to support multi-gpu ? According  to [https://discuss.pytorch.org/t/does-it-support-multi-gpu-card-on-a-single-node/75/7][1] this shouldn't be too difficult ?\n\n\n  [1]: https://discuss.pytorch.org/t/does-it-support-multi-gpu-card-on-a-single-node/75/7",
      "votes": null
    },
    {
      "id": "230603",
      "postDate": "10/12/2017 10:24:25",
      "content": "<p>Have you tried multi-gpu? I tried two gpus, but there are several errors <a href=\"https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492\">https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492</a></p>",
      "rawMarkdown": "Have you tried multi-gpu? I tried two gpus, but there are several errors https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492",
      "votes": null
    },
    {
      "id": "231533",
      "postDate": "10/15/2017 09:13:01",
      "content": "<p>Is the 10-7 release the lastest code?</p>\n\n<pre><code>probs  = F.softmax(logits)\nloss = F.cross_entropy(logits, labels)\n</code></pre>\n\n<p>This seems wrong, since crossentropy already includes log_softmax and NLLLoss!</p>",
      "rawMarkdown": "Is the 10-7 release the lastest code?\n\n    probs  = F.softmax(logits)\n    loss = F.cross_entropy(logits, labels)\n\nThis seems wrong, since crossentropy already includes log_softmax and NLLLoss!",
      "votes": null
    },
    {
      "id": "231535",
      "postDate": "10/15/2017 09:16:20",
      "content": "<p>should be correct</p>\n\n<p>loss = F.cross_entropy(logits, labels)</p>\n\n<p>the input is logits, not probs</p>\n\n<p>the loss statement is not dependent on the prob statement.</p>",
      "rawMarkdown": "should be correct\n\nloss = F.cross_entropy(logits, labels)\n\nthe input is logits, not probs\n\nthe loss statement is not dependent on the prob statement.",
      "votes": null
    },
    {
      "id": "231537",
      "postDate": "10/15/2017 09:20:23",
      "content": "<p>Oh wow. You are totally right, so stupid &lt;--</p>",
      "rawMarkdown": "Oh wow. You are totally right, so stupid &lt;--",
      "votes": null
    },
    {
      "id": "231550",
      "postDate": "10/15/2017 10:37:48",
      "content": "<p>I'm curious about your model, how do you train your model, how do you predict the image in one crop</p>",
      "rawMarkdown": "I'm curious about your model, how do you train your model, how do you predict the image in one crop",
      "votes": null
    },
    {
      "id": "232860",
      "postDate": "10/18/2017 16:08:33",
      "content": "<p>results of se-resnext101. This is trained with the latest version 10-17 of the pytorch starter kit above. You need multi-gpu to use this (or it would be really slow).  LB 0.71064 for se-resnext101_32x4d (single crop 180/180)</p>\n\n<p>I try several sets of parameters, etc and finally it works! if the parameters are not correct the loss won't drop. If the parameters are correct, you will see the loss drop slowly over several epoch.</p>\n\n<p>if se-resnext101 works, senet should be better: <a href=\"https://github.com/hujie-frank/SENet\">https://github.com/hujie-frank/SENet</a></p>\n\n<p>my next plan would be senet or larger (like se-resnext269)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/232860/7697/se-resnext101.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "results of se-resnext101. This is trained with the latest version 10-17 of the pytorch starter kit above. You need multi-gpu to use this (or it would be really slow).  LB 0.71064 for se-resnext101_32x4d (single crop 180/180)\n\nI try several sets of parameters, etc and finally it works! if the parameters are not correct the loss won't drop. If the parameters are correct, you will see the loss drop slowly over several epoch.\n\nif se-resnext101 works, senet should be better: https://github.com/hujie-frank/SENet\n\nmy next plan would be senet or larger (like se-resnext269)\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/232860/7697/se-resnext101.png",
      "votes": null
    },
    {
      "id": "232919",
      "postDate": "10/18/2017 19:37:17",
      "content": "<p>Do you use the weights from resneXt pretrained on imagenet or did you map the weights from the original se-resneXt to your own pytorch implementation?</p>",
      "rawMarkdown": "Do you use the weights from resneXt pretrained on imagenet or did you map the weights from the original se-resneXt to your own pytorch implementation?",
      "votes": null
    },
    {
      "id": "233048",
      "postDate": "10/19/2017 03:34:54",
      "content": "<p>imagenet pretrained weights for resnext and se-resnext can be downloaded. i think there are no difference if you start off with \"resnext imagenet weights --&gt;  resnext weights --&gt; se-resnext weights\" or \"se-resnext imagenet weights  --&gt; se-resnext weights\" .</p>\n\n<p>for the graph above, i am using \"resnext imagenet weights --&gt;  resnext weights --&gt; se-resnext weights\".</p>\n\n<p>but for the new experiment on senet (which is se-resnext152) i am using \"se-resnext imagenet weights  --&gt; se-resnext weights\" directly.</p>\n\n<p>both seems to work well</p>",
      "rawMarkdown": "imagenet pretrained weights for resnext and se-resnext can be downloaded. i think there are no difference if you start off with \"resnext imagenet weights --&gt;  resnext weights --&gt; se-resnext weights\" or \"se-resnext imagenet weights  --&gt; se-resnext weights\" .\n\nfor the graph above, i am using \"resnext imagenet weights --&gt;  resnext weights --&gt; se-resnext weights\".\n\nbut for the new experiment on senet (which is se-resnext152) i am using \"se-resnext imagenet weights  --&gt; se-resnext weights\" directly.\n\nboth seems to work well",
      "votes": null
    },
    {
      "id": "233073",
      "postDate": "10/19/2017 05:23:22",
      "content": "<p>It seems that pytorch only has resnet imagenet weights. Did you only use that to fine-tune resnext or se-resnext?</p>",
      "rawMarkdown": "It seems that pytorch only has resnet imagenet weights. Did you only use that to fine-tune resnext or se-resnext?",
      "votes": null
    },
    {
      "id": "233074",
      "postDate": "10/19/2017 05:34:34",
      "content": "<p>I write a converter to convert from caffe to pytorch</p>",
      "rawMarkdown": "I write a converter to convert from caffe to pytorch",
      "votes": null
    },
    {
      "id": "233091",
      "postDate": "10/19/2017 07:36:15",
      "content": "<p>i added a conversion to for snet (se-resnext152) in folder \"10-17/senet_conversion\". you can modify this for se-resnext101</p>",
      "rawMarkdown": "i added a conversion to for snet (se-resnext152) in folder \"10-17/senet_conversion\". you can modify this for se-resnext101",
      "votes": null
    },
    {
      "id": "233106",
      "postDate": "10/19/2017 08:19:40",
      "content": "<p>what about the performance of the se-resnext?</p>",
      "rawMarkdown": "what about the performance of the se-resnext?",
      "votes": null
    },
    {
      "id": "233115",
      "postDate": "10/19/2017 08:56:28",
      "content": "<p>how is your optimazer? sgd? or any other?</p>",
      "rawMarkdown": "how is your optimazer? sgd? or any other?",
      "votes": null
    },
    {
      "id": "234459",
      "postDate": "10/23/2017 11:37:28",
      "content": "<p>My linux rig reboots few minutes after I run the script using 2 GPUs. Any idea on the cause and how to resolve it? Thanks!</p>\n\n<p>And thanks again for sharing the scripts!</p>",
      "rawMarkdown": "My linux rig reboots few minutes after I run the script using 2 GPUs. Any idea on the cause and how to resolve it? Thanks!\n\nAnd thanks again for sharing the scripts!",
      "votes": null
    },
    {
      "id": "234474",
      "postDate": "10/23/2017 12:22:01",
      "content": "<p>Maybe it's overheating?</p>",
      "rawMarkdown": "Maybe it's overheating?",
      "votes": null
    },
    {
      "id": "234518",
      "postDate": "10/23/2017 14:26:48",
      "content": "<p>Thanks for your reply Tim. The GPUs are running 2 scripts separately and seems fine. In this case do you think overheating is still a probable cause? Thanks!</p>",
      "rawMarkdown": "Thanks for your reply Tim. The GPUs are running 2 scripts separately and seems fine. In this case do you think overheating is still a probable cause? Thanks!",
      "votes": null
    },
    {
      "id": "234524",
      "postDate": "10/23/2017 14:37:24",
      "content": "<p>Hi Heng,\nFirst thank you for sharing all the knowledge and information. I enjoy reading your posts.\nI was wondering, why in your do_submit() code, do you split the test set?</p>",
      "rawMarkdown": "Hi Heng,\nFirst thank you for sharing all the knowledge and information. I enjoy reading your posts.\nI was wondering, why in your do_submit() code, do you split the test set?",
      "votes": null
    },
    {
      "id": "234526",
      "postDate": "10/23/2017 14:43:25",
      "content": "<p>I can run in 4 separate machines. In case of running in a single machine, i don't have to start from zero if the testing stop due to some error. Lastly, the saved npy files are smaller , making them faster to read and write</p>",
      "rawMarkdown": "I can run in 4 separate machines. In case of running in a single machine, i don't have to start from zero if the testing stop due to some error. Lastly, the saved npy files are smaller , making them faster to read and write",
      "votes": null
    },
    {
      "id": "234532",
      "postDate": "10/23/2017 14:57:20",
      "content": "<p>Just monitor it with nvidia-smi. Or look in your logs what could have caused the error!</p>",
      "rawMarkdown": "Just monitor it with nvidia-smi. Or look in your logs what could have caused the error!",
      "votes": null
    },
    {
      "id": "237449",
      "postDate": "10/30/2017 10:44:39",
      "content": "<p>Hi Heng, any chance that you could further share how to get <code>xception.keras.convert.pth</code>? Thanks!</p>",
      "rawMarkdown": "Hi Heng, any chance that you could further share how to get `xception.keras.convert.pth`? Thanks!",
      "votes": null
    },
    {
      "id": "237452",
      "postDate": "10/30/2017 10:45:53",
      "content": "<p>It works by reducing the batch size. Maybe it's really due to overheating. Thanks!</p>",
      "rawMarkdown": "It works by reducing the batch size. Maybe it's really due to overheating. Thanks!",
      "votes": null
    },
    {
      "id": "237466",
      "postDate": "10/30/2017 11:29:21",
      "content": "<p>Oh, did you also check whether you have a memory leak? Higher batch size means more memory is needed (I am speaking about main memory, not GPU memory here)</p>",
      "rawMarkdown": "Oh, did you also check whether you have a memory leak? Higher batch size means more memory is needed (I am speaking about main memory, not GPU memory here)",
      "votes": null
    },
    {
      "id": "237987",
      "postDate": "10/31/2017 13:22:52",
      "content": "<p>Oh may I know how I can check this?</p>",
      "rawMarkdown": "Oh may I know how I can check this?",
      "votes": null
    },
    {
      "id": "242305",
      "postDate": "11/11/2017 05:30:23",
      "content": "<p>Wondering if you find where to got the xception pretrained model?</p>",
      "rawMarkdown": "Wondering if you find where to got the xception pretrained model?",
      "votes": null
    },
    {
      "id": "242727",
      "postDate": "11/12/2017 12:54:58",
      "content": "<p>Sorry no luck...</p>",
      "rawMarkdown": "Sorry no luck...",
      "votes": null
    },
    {
      "id": "242733",
      "postDate": "11/12/2017 13:13:33",
      "content": "<p>xception.keras.convert.pth is converted by myself. However,  conversion from keras to pytorch is not correct. I don't know why, but could be due  to difference in padding, ceil/floor in max pooling.</p>\n\n<p>nevertheless, the wrong converted weights seems to be usable for initialization for cdisount image training.</p>",
      "rawMarkdown": "xception.keras.convert.pth is converted by myself. However,  conversion from keras to pytorch is not correct. I don't know why, but could be due  to difference in padding, ceil/floor in max pooling.\n\nnevertheless, the wrong converted weights seems to be usable for initialization for cdisount image training.",
      "votes": null
    },
    {
      "id": "242968",
      "postDate": "11/13/2017 04:05:20",
      "content": "<p>Have you check this <a href=\"https://github.com/ysh329/deep-learning-model-convertor\">https://github.com/ysh329/deep-learning-model-convertor</a></p>",
      "rawMarkdown": "Have you check this https://github.com/ysh329/deep-learning-model-convertor",
      "votes": null
    },
    {
      "id": "249030",
      "postDate": "11/27/2017 15:18:08",
      "content": "<p>How did you guys perform the train/val split?</p>",
      "rawMarkdown": "How did you guys perform the train/val split?",
      "votes": null
    },
    {
      "id": "251376",
      "postDate": "12/01/2017 04:29:05",
      "content": "<p>Has anyone had this error training the Resnet101?</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/251376/7996/Screen%20Shot%202017-11-30%20at%2010.28.28%20PM.png\" alt=\"problem\" title=\"\"></p>",
      "rawMarkdown": "Has anyone had this error training the Resnet101?\n\n![problem][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/251376/7996/Screen%20Shot%202017-11-30%20at%2010.28.28%20PM.png",
      "votes": null
    },
    {
      "id": "251377",
      "postDate": "12/01/2017 04:39:08",
      "content": "<p>what is src in your code?</p>",
      "rawMarkdown": "what is src in your code?",
      "votes": null
    },
    {
      "id": "251630",
      "postDate": "12/01/2017 13:59:19",
      "content": "<hr>\n\n<p>Nevermind it looks like some of my images are missing.</p>",
      "rawMarkdown": "Nevermind it looks like some of my images are missing.",
      "votes": null
    },
    {
      "id": "327935",
      "postDate": "05/13/2018 01:02:58",
      "content": "<p>Hi, Heng. Did you save the data(.bson or .jpg)? The  data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~</p>",
      "rawMarkdown": "Hi, Heng. Did you save the data(.bson or .jpg)? The  data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~",
      "votes": null
    },
    {
      "id": "2162194",
      "postDate": "02/28/2023 03:37:00",
      "content": "<p>this is very handy, thank you for your work!</p>",
      "rawMarkdown": "this is very handy, thank you for your work!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2162194,
      "author_name": "ianmoonee",
      "author_url": "",
      "post_date": "02/28/2023 03:37:00",
      "content": "<p>this is very handy, thank you for your work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 227081,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "10/03/2017 17:19:22",
      "content": "<p>Welcome Heng in this new competition. I am quite sure that this one is also going to be very tight competitive one because of your generous contributions towards the solution. Keep it up. However, I think, in general, this competition will give an advantage to those who have access to more GPU resources. Otherwise, it will be interesting to see how people come up with ingenious ideas to deal with this large dataset. Good luck for a gold medal this time!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 227387,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/04/2017 08:37:07",
      "content": "<p>Trained more iterations to get LB 58.6% (12 hr).  initial observations:</p>\n\n<ul>\n<li><p>Too much layers are frozen. Next step is to retrain/finetune with more parameters. </p></li>\n<li><p>I estimate the results of single crop resnet50 is around LB = 62 to 65%</p></li>\n<li><p>I estimate the leader board top score by the kagglers to be near to 82% at the end of the competition. </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/227387/7469/loss.png\" alt=\"enter image description here\" title=\"\"></p></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 227862,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/05/2017 10:06:35",
      "content": "<p>** experiments on pretrain/freezing, learning rates. etc **</p>\n\n<p>baseline results for resnet50 LB = 0.62978 :</p>\n\n<ol>\n<li><p>train all layers (no freezing)</p></li>\n<li><p>rates:</p>\n\n<ul><li><p>0.01 for epoch 1,2</p></li>\n<li><p>0.001 for epoch 3</p></li>\n<li><p>0.0001 till 3.25</p></li></ul>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/227862/7495/loss_full.png\" alt=\"enter image description here\" title=\"\"></p></li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 227923,
      "author_name": "timjoseph",
      "author_url": "",
      "post_date": "10/05/2017 13:36:52",
      "content": "<p>Hi Heng, good to see you in this competition again.\nIs there a reason you don't use github instead of Google Drive for your code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 227928,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/05/2017 13:41:50",
          "content": "<p>i have an github account too. But i find it difficult to update because my code organisation is messy and file size limit. Maybe i will try to learn the pycharm github integration  or use some friendly github gui tool one day.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 227935,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/05/2017 13:57:29",
          "content": "<p>I think for public projects like yours it would be much better to split your project into source code and files. So you can still host the big files on Google Drive, but everything related to source code on github.\nThere are three major advantages: 1. Everyone can easily read the code without downloading everything. 2. If you use readable commit messages you create an implicit change log. 3. Issue tracker + Pull requests can make this a community effort.</p>\n\n<p>That said, I appreciate your effort no matter where all the stuff is hosted ;D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 227946,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/05/2017 14:09:20",
          "content": "<p>Good suggestion! I will get it done when i have time :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 228236,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/06/2017 07:14:11",
      "content": "<p>** experiments on input sizes, network strides. etc **</p>\n\n<p>224x224 (bilinear up-scale) input vs 160x160 input for resnet50. \n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228236/7507/Slide2.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 228238,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/06/2017 07:17:46",
      "content": "<p>To get top results, you should take care of the issues:</p>\n\n<ol>\n<li><p>balancing class  </p></li>\n<li><p>small objects</p></li>\n<li><p>how to train efficiency (extract information efficiency, not all data provide \"useful\" information) </p></li>\n<li><p>break down the problem into simpler ones (e.g. special handling for minority class)</p></li>\n</ol>\n\n<p>... to be updated ...</p>",
      "votes": null,
      "replies": [
        {
          "id": 228593,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "10/07/2017 06:42:07",
          "content": "<p>It seems that data augmentation can improve balancing class.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 228255,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/06/2017 08:08:56",
      "content": "<p>** experiments on different network structure **</p>\n\n<p>...this is still in progress ...</p>\n\n<p>based on initial results it seems that any of the large network (those that obtained about 20% top1 error on imagenet like resnext, inceptionv3,  resnet, etc) can get about 67to 69% for single crop on cdiscount image. With proper\n ensemble, 70 to 72% is easily obtainable. Care has to be taken in augmentation, input resolution, learning rate, batch size, etc. I estimate top 50 ranks results &gt;72% at the end of the competitions if kagglers have enough gpu.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228255/7517/large_input_nets2.png\" alt=\"enter image description here\" title=\"\">\n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228255/7536/batch_size1.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 228661,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/07/2017 12:05:50",
      "content": "<p>compare resnet50 and se-resnet50</p>\n\n<p>(there is a typo mistake in the figure below. the LB results should be flipped, i.e. SE-resnet is better)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/228661/7515/se-resnet50.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 228684,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/07/2017 13:40:41",
      "content": "<p>** experiments on augmentation **</p>\n\n<p>... to be updated ...</p>\n\n<p>seems that crop 160x160 from 180x180 is not a good way to augment. I suspect downsize 180x180 to 160x160 gives better results. But I do not have time or extra gpu to do this experiments. anyone  has done this comparison before?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 228735,
      "author_name": "timjoseph",
      "author_url": "",
      "post_date": "10/07/2017 17:52:09",
      "content": "<p>Heng, how did you create pretrain_convert_table.py and what weights are you loading?</p>",
      "votes": null,
      "replies": [
        {
          "id": 228777,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/07/2017 20:27:58",
          "content": "<p>this is for resnet50. the weights are actaully the same as the default pytorch model zoo. i reorganize resnet for future experiments and the naming changed. i manually create the mapping for the model zoo keys to my keys. actually you can just use the default pytorch model zoo and ignore my resnet_xx.py</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228892,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/08/2017 07:51:44",
          "content": "<p>Thank you :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 228890,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "10/08/2017 07:26:43",
      "content": "<p>Hi, Heng, I use similar way to load data as you, converting bson to files. </p>\n\n<p>The loading process is very fast at first(0.3s, CPU 30%, RAM 10% of 30G), but some iterations later, it becomes very slow(6s, CPU 5%, RAM 10% of 30G).</p>\n\n<p>I set the CPU mode to performance, and try different num_workers(0,4,8,12 …), but the situation stays the same.</p>\n\n<p>Here is my code:</p>\n\n<pre><code>from torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\n\n\nclass xiongDataset(Dataset):\n    def __init__(self, csv_file, root_dir, transform=None):\n        self.train_names=[]\n        self.root_dir=root_dir\n        self.transform = transform\n        train_images = pd.read_csv(csv_file)\n        train_ids = list(train_images['product_id'])\n        train_idxs = list(train_images['img_idx'])\n        self.labels = list(train_images['category_idx'])\n        num_train = len(train_images)\n       for i in range(num_train):\n           train_name = '{}_{}.jpg'.format(train_ids[i],train_idxs[i])\n           self.train_names.append(train_name)\n\n   def __len__(self):\n       return len(self.train_names)\n\n   def __getitem__(self, idx):\n       img = cv2.imread(self.root_dir + self.train_names[idx])\n       label = self.labels[idx]\n       if self.transform is not None:\n            img = self.transform(img)\n\n        return img,label\n\ntrain_data = xiongDataset('../train_images.csv','../train/',transform=transforms.ToTensor())\n\n\n\ndata_loader= DataLoader(train_data,batch_size=256,shuffle=False,num_workers=0)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 228891,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/08/2017 07:32:19",
          "content": "<p>i use ssd drive (solid state drive). if you are using regular hard disk, you need a better random access database format like leveldb, hdfs i think.</p>\n\n<p>when i use  regular hard disk, i face the same problem you said.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228972,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "10/08/2017 12:42:45",
          "content": "<p>Thank you for your advice. I do as what you said, and  make it. Now, loading data only takes less than 0.08 seconds, and one iteration takes 0.3 seconds(batch size=256, Resnet18, a 1080Ti).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230220,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "10/11/2017 14:36:31",
          "content": "<p>I have encountered same problem by using ssd. Probably it occurred due to GC, so adding a line <code>gc.disable()</code> in code is a way (I am scared that it may cause some bad effect...).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 229345,
      "author_name": "erayonler",
      "author_url": "",
      "post_date": "10/09/2017 12:28:08",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 229987,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "10/11/2017 00:47:45",
      "content": "<p>Hi, Heng, have you tried multi-gpu? I also train with pytorch, and I set <code>model = nn.DataParallel(model, device_ids=[0, 1])</code>, but there is several errors <a href=\"https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492\">https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492</a></p>\n\n<p>And then I tried your way <code>os.environ[\"CUDA_VISIBLE_DEVICES\"] = '0,1'</code>(I have 2 GPUs), but only one GPU works.</p>",
      "votes": null,
      "replies": [
        {
          "id": 229995,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 01:28:44",
          "content": "<p>haven't try yet. multi-gpu development would come later</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 230350,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "10/11/2017 19:30:48",
      "content": "<p>How many GPUs do you have? I remember you have 4 Titan X? Just want to compare the running time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 230397,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 21:50:04",
          "content": "<p>i am using 1 pascal titanx and 3 1080 ti. but each gpu is use to train one model. i haven't train across multiple gpu yet.  i am still \"exploring\" different models</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 230391,
      "author_name": "lamdang",
      "author_url": "",
      "post_date": "10/11/2017 21:25:24",
      "content": "<p>Hello Heng,</p>\n\n<p>Thanks a lot for sharing your scripts and models so generously! I am struggling to break through 50% acc so going through your scripts will be very helpful to find the issues in my pipeline.</p>\n\n<p>I have a few questions on your train &amp; valid split:</p>\n\n<ul>\n<li>You only used 5K samples for valid. How did you choose them? how do you ensure that the class weights are respected?</li>\n<li>There are 10k less samples in your train vs total labeled.  Did you choose to drop 5K samples on purpose?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 230393,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 21:37:14",
          "content": "<p>validation set is not 5k. it is 50k.</p>\n\n<p>the 5k is used in the training iteration is a subset of validation to make \"visualisation\" fast.</p>\n\n<p>thay are choosen randomly. (not the best way)</p>\n\n<p>I will come back to data sampling later.  now i am still at early stage of experiments</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 230576,
      "author_name": "yanchao727",
      "author_url": "",
      "post_date": "10/12/2017 08:54:40",
      "content": "<p>Can't find pretrained_file:  /root/share/data/models/pytorch/imagenet/resenet/resnet50-19c8e357.pth</p>",
      "votes": null,
      "replies": [
        {
          "id": 230577,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/12/2017 08:56:13",
          "content": "<p><a href=\"https://www.google.com.sg/search?q=resnet50-19c8e357.pth&amp;oq=resnet50-19c8e357.pth&amp;aqs=chrome..69i57.1903j0j4&amp;sourceid=chrome&amp;ie=UTF-8\">https://www.google.com.sg/search?q=resnet50-19c8e357.pth&amp;oq=resnet50-19c8e357.pth&amp;aqs=chrome..69i57.1903j0j4&amp;sourceid=chrome&amp;ie=UTF-8</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 230596,
      "author_name": "yanchao727",
      "author_url": "",
      "post_date": "10/12/2017 09:52:49",
      "content": "<p>How to support multi-gpu ? According  to <a href=\"https://discuss.pytorch.org/t/does-it-support-multi-gpu-card-on-a-single-node/75/7\">https://discuss.pytorch.org/t/does-it-support-multi-gpu-card-on-a-single-node/75/7</a> this shouldn't be too difficult ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 230603,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "10/12/2017 10:24:25",
          "content": "<p>Have you tried multi-gpu? I tried two gpus, but there are several errors <a href=\"https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492\">https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 231533,
      "author_name": "timjoseph",
      "author_url": "",
      "post_date": "10/15/2017 09:13:01",
      "content": "<p>Is the 10-7 release the lastest code?</p>\n\n<pre><code>probs  = F.softmax(logits)\nloss = F.cross_entropy(logits, labels)\n</code></pre>\n\n<p>This seems wrong, since crossentropy already includes log_softmax and NLLLoss!</p>",
      "votes": null,
      "replies": [
        {
          "id": 231535,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/15/2017 09:16:20",
          "content": "<p>should be correct</p>\n\n<p>loss = F.cross_entropy(logits, labels)</p>\n\n<p>the input is logits, not probs</p>\n\n<p>the loss statement is not dependent on the prob statement.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231537,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/15/2017 09:20:23",
          "content": "<p>Oh wow. You are totally right, so stupid &lt;--</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 231550,
      "author_name": "ifighting",
      "author_url": "",
      "post_date": "10/15/2017 10:37:48",
      "content": "<p>I'm curious about your model, how do you train your model, how do you predict the image in one crop</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 232860,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/18/2017 16:08:33",
      "content": "<p>results of se-resnext101. This is trained with the latest version 10-17 of the pytorch starter kit above. You need multi-gpu to use this (or it would be really slow).  LB 0.71064 for se-resnext101_32x4d (single crop 180/180)</p>\n\n<p>I try several sets of parameters, etc and finally it works! if the parameters are not correct the loss won't drop. If the parameters are correct, you will see the loss drop slowly over several epoch.</p>\n\n<p>if se-resnext101 works, senet should be better: <a href=\"https://github.com/hujie-frank/SENet\">https://github.com/hujie-frank/SENet</a></p>\n\n<p>my next plan would be senet or larger (like se-resnext269)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/232860/7697/se-resnext101.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 232919,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/18/2017 19:37:17",
          "content": "<p>Do you use the weights from resneXt pretrained on imagenet or did you map the weights from the original se-resneXt to your own pytorch implementation?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233048,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/19/2017 03:34:54",
          "content": "<p>imagenet pretrained weights for resnext and se-resnext can be downloaded. i think there are no difference if you start off with \"resnext imagenet weights --&gt;  resnext weights --&gt; se-resnext weights\" or \"se-resnext imagenet weights  --&gt; se-resnext weights\" .</p>\n\n<p>for the graph above, i am using \"resnext imagenet weights --&gt;  resnext weights --&gt; se-resnext weights\".</p>\n\n<p>but for the new experiment on senet (which is se-resnext152) i am using \"se-resnext imagenet weights  --&gt; se-resnext weights\" directly.</p>\n\n<p>both seems to work well</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233073,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "10/19/2017 05:23:22",
          "content": "<p>It seems that pytorch only has resnet imagenet weights. Did you only use that to fine-tune resnext or se-resnext?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233074,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/19/2017 05:34:34",
          "content": "<p>I write a converter to convert from caffe to pytorch</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233091,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/19/2017 07:36:15",
          "content": "<p>i added a conversion to for snet (se-resnext152) in folder \"10-17/senet_conversion\". you can modify this for se-resnext101</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233106,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "10/19/2017 08:19:40",
          "content": "<p>what about the performance of the se-resnext?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233115,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "10/19/2017 08:56:28",
          "content": "<p>how is your optimazer? sgd? or any other?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 234459,
      "author_name": "terryli",
      "author_url": "",
      "post_date": "10/23/2017 11:37:28",
      "content": "<p>My linux rig reboots few minutes after I run the script using 2 GPUs. Any idea on the cause and how to resolve it? Thanks!</p>\n\n<p>And thanks again for sharing the scripts!</p>",
      "votes": null,
      "replies": [
        {
          "id": 234474,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/23/2017 12:22:01",
          "content": "<p>Maybe it's overheating?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234518,
          "author_name": "terryli",
          "author_url": "",
          "post_date": "10/23/2017 14:26:48",
          "content": "<p>Thanks for your reply Tim. The GPUs are running 2 scripts separately and seems fine. In this case do you think overheating is still a probable cause? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234532,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/23/2017 14:57:20",
          "content": "<p>Just monitor it with nvidia-smi. Or look in your logs what could have caused the error!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 237452,
          "author_name": "terryli",
          "author_url": "",
          "post_date": "10/30/2017 10:45:53",
          "content": "<p>It works by reducing the batch size. Maybe it's really due to overheating. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 237466,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/30/2017 11:29:21",
          "content": "<p>Oh, did you also check whether you have a memory leak? Higher batch size means more memory is needed (I am speaking about main memory, not GPU memory here)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 237987,
          "author_name": "terryli",
          "author_url": "",
          "post_date": "10/31/2017 13:22:52",
          "content": "<p>Oh may I know how I can check this?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 234524,
      "author_name": "burgalon",
      "author_url": "",
      "post_date": "10/23/2017 14:37:24",
      "content": "<p>Hi Heng,\nFirst thank you for sharing all the knowledge and information. I enjoy reading your posts.\nI was wondering, why in your do_submit() code, do you split the test set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 234526,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/23/2017 14:43:25",
          "content": "<p>I can run in 4 separate machines. In case of running in a single machine, i don't have to start from zero if the testing stop due to some error. Lastly, the saved npy files are smaller , making them faster to read and write</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 237449,
      "author_name": "terryli",
      "author_url": "",
      "post_date": "10/30/2017 10:44:39",
      "content": "<p>Hi Heng, any chance that you could further share how to get <code>xception.keras.convert.pth</code>? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 242305,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "11/11/2017 05:30:23",
          "content": "<p>Wondering if you find where to got the xception pretrained model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242727,
          "author_name": "terryli",
          "author_url": "",
          "post_date": "11/12/2017 12:54:58",
          "content": "<p>Sorry no luck...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 242733,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/12/2017 13:13:33",
      "content": "<p>xception.keras.convert.pth is converted by myself. However,  conversion from keras to pytorch is not correct. I don't know why, but could be due  to difference in padding, ceil/floor in max pooling.</p>\n\n<p>nevertheless, the wrong converted weights seems to be usable for initialization for cdisount image training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 242968,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "11/13/2017 04:05:20",
          "content": "<p>Have you check this <a href=\"https://github.com/ysh329/deep-learning-model-convertor\">https://github.com/ysh329/deep-learning-model-convertor</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 249030,
      "author_name": "skinish",
      "author_url": "",
      "post_date": "11/27/2017 15:18:08",
      "content": "<p>How did you guys perform the train/val split?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 251376,
      "author_name": "stevenknguyen",
      "author_url": "",
      "post_date": "12/01/2017 04:29:05",
      "content": "<p>Has anyone had this error training the Resnet101?</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/251376/7996/Screen%20Shot%202017-11-30%20at%2010.28.28%20PM.png\" alt=\"problem\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 251377,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "12/01/2017 04:39:08",
          "content": "<p>what is src in your code?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 251630,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "12/01/2017 13:59:19",
          "content": "<hr>\n\n<p>Nevermind it looks like some of my images are missing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327935,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "05/13/2018 01:02:58",
      "content": "<p>Hi, Heng. Did you save the data(.bson or .jpg)? The  data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "227035": "** important **\nplease refer to @Vladimir Iglovikov at https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652\n\nplease train with long epoch to get better results . My models results are from training only up to 4 to 5 epoch. \n\n------\n\nThis is a pytorch starter kit. It is not completed yet. I use ideas from the kernels and discussion contributions. Thanks to kagglers, especially @Bruno G. do Amaral, @Human Analog, @Adam Blazek for code and discussion.\n\n\nDownload: \nhttps://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing\n\n---------------\nversion.10-03:\n\n- initial version for early experiments. see \"trainer.py\"\n\n---------------\n\nversion.10-07:\n\n- LB 0.63358(160x160 single center crop) or 0.63839(180x180) or  0.64356 (180x180 resize to 160x160) for se-resnet50 trained on 160x160 crops.\n\n- LB 0.66708 for same se-resnet50 if train=224x224 crops from 256, test = resize to 224\n\n- LB 0.68939 for same se-resnet50 if train=180x180, test = 180x180 (weight initialised from 224 se-resnet50 of above)\n\n-  LB 0.69565 for inception3 if train=180x180, test = 180x180 (be careful on how to control the lr, batch size, momentum, level of augmentation,etc. I will have a writeup on this later)\n\n-  support various network like resnet, resnext, inception v3, etc. see trainer_xxx.py\n\n- support gradient accumulation \n\nnote: code is messy and dirty. It may not be backward- compatible.\n\n---------------\n\nversion.10-17 (latest):\n\n- advance version. implement multi-gpu support in training. see \"trainer_excited_resnext101_32x4d.py\"\n\n- LB 0.71064 for se-resnext101_32x4d (single crop 180/180)\n\n- support xception, se-xception, resnext101_32x4d, se-resnext101_32x4d, inceptionv4, inception_resnetv2 \n\n-  folder \"10-17/senet_conversion\" contains conversion tools for converting caffe model to pytorch  \n\n- note: \n\n     *1.* may not be backward compatible \n\n     *2.* old training scripts are found in \"dummy-00/__temp__/old_trainer\"\n\n---------------     \n\n<br>\nnext version:\n \n- focal loss for class imbalance\n\n- super large batch size to speed up training\n\n- efficient multi-crop for testing\n\n- divide test/train data into easy and difficult set for faster testing/training.  class balancing\n\n\n---------------\n\nNote: \n\nMy software and thread will be constantly update. Also, ther great works:\n\n- https://www.kaggle.com/blazeka/multi-gpu-tensorflow-convnet-0-65/notebook\n\n- https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40715",
    "227081": "Welcome Heng in this new competition. I am quite sure that this one is also going to be very tight competitive one because of your generous contributions towards the solution. Keep it up. However, I think, in general, this competition will give an advantage to those who have access to more GPU resources. Otherwise, it will be interesting to see how people come up with ingenious ideas to deal with this large dataset. Good luck for a gold medal this time!",
    "227387": "Trained more iterations to get LB 58.6% (12 hr).  initial observations:\n\n- Too much layers are frozen. Next step is to retrain/finetune with more parameters. \n\n- I estimate the results of single crop resnet50 is around LB = 62 to 65%\n\n- I estimate the leader board top score by the kagglers to be near to 82% at the end of the competition. \n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/227387/7469/loss.png",
    "227862": "** experiments on pretrain/freezing, learning rates. etc **\n\n\nbaseline results for resnet50 LB = 0.62978 :\n\n1. train all layers (no freezing)\n\n2. rates:\n\n - 0.01 for epoch 1,2\n\n -  0.001 for epoch 3\n\n -  0.0001 till 3.25\n\n\n   ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/227862/7495/loss_full.png",
    "227923": "Hi Heng, good to see you in this competition again.\nIs there a reason you don't use github instead of Google Drive for your code?",
    "227928": "i have an github account too. But i find it difficult to update because my code organisation is messy and file size limit. Maybe i will try to learn the pycharm github integration  or use some friendly github gui tool one day.",
    "227935": "I think for public projects like yours it would be much better to split your project into source code and files. So you can still host the big files on Google Drive, but everything related to source code on github.\nThere are three major advantages: 1. Everyone can easily read the code without downloading everything. 2. If you use readable commit messages you create an implicit change log. 3. Issue tracker + Pull requests can make this a community effort.\n\nThat said, I appreciate your effort no matter where all the stuff is hosted ;D",
    "227946": "Good suggestion! I will get it done when i have time :)",
    "228236": "** experiments on input sizes, network strides. etc **\n\n224x224 (bilinear up-scale) input vs 160x160 input for resnet50. \n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228236/7507/Slide2.png",
    "228238": "To get top results, you should take care of the issues:\n\n 1. balancing class  \n\n 2. small objects\n\n 3.  how to train efficiency (extract information efficiency, not all data provide \"useful\" information) \n\n 4. break down the problem into simpler ones (e.g. special handling for minority class)\n\n... to be updated ...",
    "228255": "** experiments on different network structure **\n\n...this is still in progress ...\n\nbased on initial results it seems that any of the large network (those that obtained about 20% top1 error on imagenet like resnext, inceptionv3,  resnet, etc) can get about 67to 69% for single crop on cdiscount image. With proper\n ensemble, 70 to 72% is easily obtainable. Care has to be taken in augmentation, input resolution, learning rate, batch size, etc. I estimate top 50 ranks results &gt;72% at the end of the competitions if kagglers have enough gpu.\n\n ![enter image description here][1]\n ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228255/7517/large_input_nets2.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228255/7536/batch_size1.png",
    "228593": "It seems that data augmentation can improve balancing class.",
    "228661": "compare resnet50 and se-resnet50\n\n(there is a typo mistake in the figure below. the LB results should be flipped, i.e. SE-resnet is better)\n\n![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/228661/7515/se-resnet50.png",
    "228684": "** experiments on augmentation **\n\n... to be updated ...\n\nseems that crop 160x160 from 180x180 is not a good way to augment. I suspect downsize 180x180 to 160x160 gives better results. But I do not have time or extra gpu to do this experiments. anyone  has done this comparison before?",
    "228735": "Heng, how did you create pretrain_convert_table.py and what weights are you loading?",
    "228777": "this is for resnet50. the weights are actaully the same as the default pytorch model zoo. i reorganize resnet for future experiments and the naming changed. i manually create the mapping for the model zoo keys to my keys. actually you can just use the default pytorch model zoo and ignore my resnet_xx.py",
    "228890": "Hi, Heng, I use similar way to load data as you, converting bson to files. \n\nThe loading process is very fast at first(0.3s, CPU 30%, RAM 10% of 30G), but some iterations later, it becomes very slow(6s, CPU 5%, RAM 10% of 30G).\n\nI set the CPU mode to performance, and try different num_workers(0,4,8,12 …), but the situation stays the same.\n\nHere is my code:\n   \n    from torch.utils.data import Dataset, DataLoader\n    from torchvision import transforms\n\n\n    class xiongDataset(Dataset):\n        def __init__(self, csv_file, root_dir, transform=None):\n            self.train_names=[]\n            self.root_dir=root_dir\n            self.transform = transform\n            train_images = pd.read_csv(csv_file)\n            train_ids = list(train_images['product_id'])\n            train_idxs = list(train_images['img_idx'])\n            self.labels = list(train_images['category_idx'])\n            num_train = len(train_images)\n           for i in range(num_train):\n               train_name = '{}_{}.jpg'.format(train_ids[i],train_idxs[i])\n               self.train_names.append(train_name)\n\n       def __len__(self):\n           return len(self.train_names)\n\n       def __getitem__(self, idx):\n           img = cv2.imread(self.root_dir + self.train_names[idx])\n           label = self.labels[idx]\n           if self.transform is not None:\n                img = self.transform(img)\n\n            return img,label\n\n    train_data = xiongDataset('../train_images.csv','../train/',transform=transforms.ToTensor())\n\n\n\n    data_loader= DataLoader(train_data,batch_size=256,shuffle=False,num_workers=0)",
    "228891": "i use ssd drive (solid state drive). if you are using regular hard disk, you need a better random access database format like leveldb, hdfs i think.\n\nwhen i use  regular hard disk, i face the same problem you said.",
    "228892": "Thank you :)",
    "228972": "Thank you for your advice. I do as what you said, and  make it. Now, loading data only takes less than 0.08 seconds, and one iteration takes 0.3 seconds(batch size=256, Resnet18, a 1080Ti).",
    "229345": "Thanks for sharing",
    "229987": "Hi, Heng, have you tried multi-gpu? I also train with pytorch, and I set ```model = nn.DataParallel(model, device_ids=[0, 1])```, but there is several errors https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492\n\nAnd then I tried your way ```os.environ[\"CUDA_VISIBLE_DEVICES\"] = '0,1'```(I have 2 GPUs), but only one GPU works.",
    "229995": "haven't try yet. multi-gpu development would come later",
    "230220": "I have encountered same problem by using ssd. Probably it occurred due to GC, so adding a line `gc.disable()` in code is a way (I am scared that it may cause some bad effect...).",
    "230350": "How many GPUs do you have? I remember you have 4 Titan X? Just want to compare the running time.",
    "230391": "Hello Heng,\n\nThanks a lot for sharing your scripts and models so generously! I am struggling to break through 50% acc so going through your scripts will be very helpful to find the issues in my pipeline.\n\nI have a few questions on your train &amp; valid split:\n\n-  You only used 5K samples for valid. How did you choose them? how do you ensure that the class weights are respected?\n- There are 10k less samples in your train vs total labeled.  Did you choose to drop 5K samples on purpose?",
    "230393": "validation set is not 5k. it is 50k.\n\nthe 5k is used in the training iteration is a subset of validation to make \"visualisation\" fast.\n\nthay are choosen randomly. (not the best way)\n\nI will come back to data sampling later.  now i am still at early stage of experiments",
    "230397": "i am using 1 pascal titanx and 3 1080 ti. but each gpu is use to train one model. i haven't train across multiple gpu yet.  i am still \"exploring\" different models",
    "230576": "Can't find pretrained_file:  /root/share/data/models/pytorch/imagenet/resenet/resnet50-19c8e357.pth",
    "230577": "https://www.google.com.sg/search?q=resnet50-19c8e357.pth&amp;oq=resnet50-19c8e357.pth&amp;aqs=chrome..69i57.1903j0j4&amp;sourceid=chrome&amp;ie=UTF-8",
    "230596": "How to support multi-gpu ? According  to [https://discuss.pytorch.org/t/does-it-support-multi-gpu-card-on-a-single-node/75/7][1] this shouldn't be too difficult ?\n\n\n  [1]: https://discuss.pytorch.org/t/does-it-support-multi-gpu-card-on-a-single-node/75/7",
    "230603": "Have you tried multi-gpu? I tried two gpus, but there are several errors https://discuss.pytorch.org/t/multi-gpu-error-the-accuracy-of-one-gpu-is-much-lower-than-two-gpus/8492",
    "231533": "Is the 10-7 release the lastest code?\n\n    probs  = F.softmax(logits)\n    loss = F.cross_entropy(logits, labels)\n\nThis seems wrong, since crossentropy already includes log_softmax and NLLLoss!",
    "231535": "should be correct\n\nloss = F.cross_entropy(logits, labels)\n\nthe input is logits, not probs\n\nthe loss statement is not dependent on the prob statement.",
    "231537": "Oh wow. You are totally right, so stupid &lt;--",
    "231550": "I'm curious about your model, how do you train your model, how do you predict the image in one crop",
    "232860": "results of se-resnext101. This is trained with the latest version 10-17 of the pytorch starter kit above. You need multi-gpu to use this (or it would be really slow).  LB 0.71064 for se-resnext101_32x4d (single crop 180/180)\n\nI try several sets of parameters, etc and finally it works! if the parameters are not correct the loss won't drop. If the parameters are correct, you will see the loss drop slowly over several epoch.\n\nif se-resnext101 works, senet should be better: https://github.com/hujie-frank/SENet\n\nmy next plan would be senet or larger (like se-resnext269)\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/232860/7697/se-resnext101.png",
    "232919": "Do you use the weights from resneXt pretrained on imagenet or did you map the weights from the original se-resneXt to your own pytorch implementation?",
    "233048": "imagenet pretrained weights for resnext and se-resnext can be downloaded. i think there are no difference if you start off with \"resnext imagenet weights --&gt;  resnext weights --&gt; se-resnext weights\" or \"se-resnext imagenet weights  --&gt; se-resnext weights\" .\n\nfor the graph above, i am using \"resnext imagenet weights --&gt;  resnext weights --&gt; se-resnext weights\".\n\nbut for the new experiment on senet (which is se-resnext152) i am using \"se-resnext imagenet weights  --&gt; se-resnext weights\" directly.\n\nboth seems to work well",
    "233073": "It seems that pytorch only has resnet imagenet weights. Did you only use that to fine-tune resnext or se-resnext?",
    "233074": "I write a converter to convert from caffe to pytorch",
    "233091": "i added a conversion to for snet (se-resnext152) in folder \"10-17/senet_conversion\". you can modify this for se-resnext101",
    "233106": "what about the performance of the se-resnext?",
    "233115": "how is your optimazer? sgd? or any other?",
    "234459": "My linux rig reboots few minutes after I run the script using 2 GPUs. Any idea on the cause and how to resolve it? Thanks!\n\nAnd thanks again for sharing the scripts!",
    "234474": "Maybe it's overheating?",
    "234518": "Thanks for your reply Tim. The GPUs are running 2 scripts separately and seems fine. In this case do you think overheating is still a probable cause? Thanks!",
    "234524": "Hi Heng,\nFirst thank you for sharing all the knowledge and information. I enjoy reading your posts.\nI was wondering, why in your do_submit() code, do you split the test set?",
    "234526": "I can run in 4 separate machines. In case of running in a single machine, i don't have to start from zero if the testing stop due to some error. Lastly, the saved npy files are smaller , making them faster to read and write",
    "234532": "Just monitor it with nvidia-smi. Or look in your logs what could have caused the error!",
    "237449": "Hi Heng, any chance that you could further share how to get `xception.keras.convert.pth`? Thanks!",
    "237452": "It works by reducing the batch size. Maybe it's really due to overheating. Thanks!",
    "237466": "Oh, did you also check whether you have a memory leak? Higher batch size means more memory is needed (I am speaking about main memory, not GPU memory here)",
    "237987": "Oh may I know how I can check this?",
    "242305": "Wondering if you find where to got the xception pretrained model?",
    "242727": "Sorry no luck...",
    "242733": "xception.keras.convert.pth is converted by myself. However,  conversion from keras to pytorch is not correct. I don't know why, but could be due  to difference in padding, ceil/floor in max pooling.\n\nnevertheless, the wrong converted weights seems to be usable for initialization for cdisount image training.",
    "242968": "Have you check this https://github.com/ysh329/deep-learning-model-convertor",
    "249030": "How did you guys perform the train/val split?",
    "251376": "Has anyone had this error training the Resnet101?\n\n![problem][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/251376/7996/Screen%20Shot%202017-11-30%20at%2010.28.28%20PM.png",
    "251377": "what is src in your code?",
    "251630": "Nevermind it looks like some of my images are missing.",
    "327935": "Hi, Heng. Did you save the data(.bson or .jpg)? The  data has been removed after the competition, but I still want to use several images(Non-commercial). Could you please share it? Thank you~",
    "2162194": "this is very handy, thank you for your work!"
  },
  "source": "meta"
}