{
  "id": 41021,
  "title": "Release of trained models (Cdiscount model zoo)",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/41021",
  "author_name": "",
  "post_date": "2017-10-11T16:59:21.959542200Z",
  "votes": 63,
  "comment_count": 52,
  "views": 0,
  "content": "<p>** important **\nplease refer to @Vladimir Iglovikov at <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652</a></p>\n\n<p>please train with long epoch to get better results . My models results are from training only up to 4 to 5 epoch. </p>\n\n<hr>\n\n<p>you can download my trained models at the share drive:</p>\n\n<p><a href=\"https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE\">https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE</a></p>\n\n<p>(go to the \"release\" folder)</p>\n\n<p>performance:</p>\n\n<ol>\n<li><p>SE-InceptionV3 : LB 0.69673 (single 180/180 crop)</p></li>\n<li><p>InceptionV3 : LB 0.69565 (single 180/180 crop)</p></li>\n<li><p>Xception : LB 0.69422 (single 180/180 crop)</p></li>\n</ol>\n\n<p>compare with other models:</p>\n\n<ul>\n<li><p>SE-Resnet50 : LB 0.68939 (single 180/180 crop)</p></li>\n<li><p>SE-resnext101_32x4d : LB 0.71064  (single crop 180/180)</p></li>\n<li><p>InceptionV4 :  in progress</p></li>\n<li><p>Inception_ResnetV2 :  in progress</p></li>\n<li><p>Resnet152 :  in progress</p></li>\n<li><p>ResDrop269 (Stochastic Depth Resnet) : in progress</p></li>\n</ul>\n\n<h2>*Note: all results here are trained only for limited number of epoch. if they are trained longer, i expect some slight improvement e.g. +0.005.</h2>\n\n<p>[what to do with it?]</p>\n\n<ol>\n<li><p>use it to finetune your new models of different scale</p></li>\n<li><p>perform test-time augmentation of multi-crops and scales</p></li>\n<li><p>extend the model, e.g. use more inception or fc layers</p></li>\n</ol>\n\n<p>You can refer to the vgg paper \"Very Deep Convolutional Networks for Large-Scale Image Recognition\" -Karen Simonyan, Andrew Zisserman, Arxiv 2014. If you google for the presentation slides of ilsvrc, you can find more tricks.</p>\n\n<p>Yet another paper mentions \"top-k pooling\" for inference. please refer to \"<a href=\"http://ww.dahua.me/publications/dhl17_polynet.pdf\">http://ww.dahua.me/publications/dhl17_polynet.pdf</a>\", see section.5</p>\n\n<p>I expect multi-crop results of maybe LB 0.71 to 0.72, but I haven't got enough time and resources to try. e.g. 144 crops per image is a nightmare for me.</p>\n\n<p>if you got good (or bad) results, I would appreciate you can report your results here so that me and other can repeat (or avoid) it.</p>\n\n<p>good luck!</p>\n\n<hr>\n\n<p>[training details]</p>\n\n<p>These models are created using my starter kit at: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a></p>\n\n<ol>\n<li><p>refer to \"training_code\" for details of augmentation and data for train split. \nThese files are from my starter kit. They should contain enough information for your future work. If there is anything unclear, please ask at  the forum or refer to the start kit.</p></li>\n<li><p>refer to \"demo_code\" for details of how to apply model to test an image. e.g. how to normalize input image with mean and std values. \"label_to_cat_id\" maps the class label to the cat_id.</p></li>\n</ol>",
  "messages": [
    {
      "id": "230288",
      "postDate": "10/11/2017 16:59:21",
      "content": "<p>** important **\nplease refer to @Vladimir Iglovikov at <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652</a></p>\n\n<p>please train with long epoch to get better results . My models results are from training only up to 4 to 5 epoch. </p>\n\n<hr>\n\n<p>you can download my trained models at the share drive:</p>\n\n<p><a href=\"https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE\">https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE</a></p>\n\n<p>(go to the \"release\" folder)</p>\n\n<p>performance:</p>\n\n<ol>\n<li><p>SE-InceptionV3 : LB 0.69673 (single 180/180 crop)</p></li>\n<li><p>InceptionV3 : LB 0.69565 (single 180/180 crop)</p></li>\n<li><p>Xception : LB 0.69422 (single 180/180 crop)</p></li>\n</ol>\n\n<p>compare with other models:</p>\n\n<ul>\n<li><p>SE-Resnet50 : LB 0.68939 (single 180/180 crop)</p></li>\n<li><p>SE-resnext101_32x4d : LB 0.71064  (single crop 180/180)</p></li>\n<li><p>InceptionV4 :  in progress</p></li>\n<li><p>Inception_ResnetV2 :  in progress</p></li>\n<li><p>Resnet152 :  in progress</p></li>\n<li><p>ResDrop269 (Stochastic Depth Resnet) : in progress</p></li>\n</ul>\n\n<h2>*Note: all results here are trained only for limited number of epoch. if they are trained longer, i expect some slight improvement e.g. +0.005.</h2>\n\n<p>[what to do with it?]</p>\n\n<ol>\n<li><p>use it to finetune your new models of different scale</p></li>\n<li><p>perform test-time augmentation of multi-crops and scales</p></li>\n<li><p>extend the model, e.g. use more inception or fc layers</p></li>\n</ol>\n\n<p>You can refer to the vgg paper \"Very Deep Convolutional Networks for Large-Scale Image Recognition\" -Karen Simonyan, Andrew Zisserman, Arxiv 2014. If you google for the presentation slides of ilsvrc, you can find more tricks.</p>\n\n<p>Yet another paper mentions \"top-k pooling\" for inference. please refer to \"<a href=\"http://ww.dahua.me/publications/dhl17_polynet.pdf\">http://ww.dahua.me/publications/dhl17_polynet.pdf</a>\", see section.5</p>\n\n<p>I expect multi-crop results of maybe LB 0.71 to 0.72, but I haven't got enough time and resources to try. e.g. 144 crops per image is a nightmare for me.</p>\n\n<p>if you got good (or bad) results, I would appreciate you can report your results here so that me and other can repeat (or avoid) it.</p>\n\n<p>good luck!</p>\n\n<hr>\n\n<p>[training details]</p>\n\n<p>These models are created using my starter kit at: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a></p>\n\n<ol>\n<li><p>refer to \"training_code\" for details of augmentation and data for train split. \nThese files are from my starter kit. They should contain enough information for your future work. If there is anything unclear, please ask at  the forum or refer to the start kit.</p></li>\n<li><p>refer to \"demo_code\" for details of how to apply model to test an image. e.g. how to normalize input image with mean and std values. \"label_to_cat_id\" maps the class label to the cat_id.</p></li>\n</ol>",
      "rawMarkdown": "** important **\nplease refer to @Vladimir Iglovikov at https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652\n\nplease train with long epoch to get better results . My models results are from training only up to 4 to 5 epoch. \n\n------\n\nyou can download my trained models at the share drive:\n\nhttps://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE\n\n(go to the \"release\" folder)\n\nperformance:\n\n1. SE-InceptionV3 : LB 0.69673 (single 180/180 crop)\n\n2. InceptionV3 : LB 0.69565 (single 180/180 crop)\n\n3.  Xception : LB 0.69422 (single 180/180 crop)\n\n\ncompare with other models:\n\n -  SE-Resnet50 : LB 0.68939 (single 180/180 crop)\n\n -  SE-resnext101_32x4d : LB 0.71064  (single crop 180/180)\n\n -  InceptionV4 :  in progress\n\n -  Inception_ResnetV2 :  in progress\n\n -  Resnet152 :  in progress\n\n - ResDrop269 (Stochastic Depth Resnet) : in progress\n\n*Note: all results here are trained only for limited number of epoch. if they are trained longer, i expect some slight improvement e.g. +0.005.\n----\n\n[what to do with it?]\n\n1.  use it to finetune your new models of different scale\n\n2. perform test-time augmentation of multi-crops and scales\n\n3. extend the model, e.g. use more inception or fc layers\n\nYou can refer to the vgg paper \"Very Deep Convolutional Networks for Large-Scale Image Recognition\" -Karen Simonyan, Andrew Zisserman, Arxiv 2014. If you google for the presentation slides of ilsvrc, you can find more tricks.\n\nYet another paper mentions \"top-k pooling\" for inference. please refer to \"http://ww.dahua.me/publications/dhl17_polynet.pdf\", see section.5\n\nI expect multi-crop results of maybe LB 0.71 to 0.72, but I haven't got enough time and resources to try. e.g. 144 crops per image is a nightmare for me.\n\nif you got good (or bad) results, I would appreciate you can report your results here so that me and other can repeat (or avoid) it.\n\ngood luck!\n\n----\n\n[training details]\n\nThese models are created using my starter kit at: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\n\n1. refer to \"training_code\" for details of augmentation and data for train split. \nThese files are from my starter kit. They should contain enough information for your future work. If there is anything unclear, please ask at  the forum or refer to the start kit.\n\n\n2. refer to \"demo_code\" for details of how to apply model to test an image. e.g. how to normalize input image with mean and std values. \"label_to_cat_id\" maps the class label to the cat_id.",
      "votes": null
    },
    {
      "id": "230321",
      "postDate": "10/11/2017 18:06:47",
      "content": "<p>Here is a contribution to keras-tf people:\n<a href=\"https://drive.google.com/drive/folders/0B6ZtD5YbyCsJVHM4OXl1Y2stQW8?usp=sharing\">https://drive.google.com/drive/folders/0B6ZtD5YbyCsJVHM4OXl1Y2stQW8?usp=sharing</a></p>\n\n<p>You can find Xception and Inception v3 models and a pickle storing class order. On validation the models reach upper 66 (67)%. Images are preprocessed by model's default: x/255. and ((x / 255.) - 0.5) * 2.).</p>",
      "rawMarkdown": "Here is a contribution to keras-tf people:\nhttps://drive.google.com/drive/folders/0B6ZtD5YbyCsJVHM4OXl1Y2stQW8?usp=sharing\n\nYou can find Xception and Inception v3 models and a pickle storing class order. On validation the models reach upper 66 (67)%. Images are preprocessed by model's default: x/255. and ((x / 255.) - 0.5) * 2.).",
      "votes": null
    },
    {
      "id": "230401",
      "postDate": "10/11/2017 21:58:11",
      "content": "<p>Thank for sharing</p>\n\n<p>What was the strategy of training you used?</p>\n\n<ul>\n<li>epochs</li>\n<li>all images?</li>\n<li>crop or downsample</li>\n<li>how much time take an epoch</li>\n<li>learning rate</li>\n<li>...</li>\n</ul>\n\n<p>thanks</p>",
      "rawMarkdown": "Thank for sharing\n\nWhat was the strategy of training you used?\n\n- epochs\n- all images?\n- crop or downsample\n- how much time take an epoch\n- learning rate\n- ...\n\nthanks",
      "votes": null
    },
    {
      "id": "230417",
      "postDate": "10/11/2017 23:00:48",
      "content": "<p>inception v3: 12 crops verus 144 crops. See google paper: <a href=\"https://arxiv.org/pdf/1602.07261.pdf\">https://arxiv.org/pdf/1602.07261.pdf</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/230417/7606/crops.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "inception v3: 12 crops verus 144 crops. See google paper: https://arxiv.org/pdf/1602.07261.pdf\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/230417/7606/crops.png",
      "votes": null
    },
    {
      "id": "230421",
      "postDate": "10/11/2017 23:15:39",
      "content": "<p>This is pretty thorough. Thanks for sharing the model. Learning a lot from your posts in this competition.</p>",
      "rawMarkdown": "This is pretty thorough. Thanks for sharing the model. Learning a lot from your posts in this competition.",
      "votes": null
    },
    {
      "id": "230445",
      "postDate": "10/12/2017 00:15:28",
      "content": "<p><a href=\"http://image-net.org/challenges/talks/2016/Hikvision_at_ImageNet_2016.pdf\">http://image-net.org/challenges/talks/2016/Hikvision_at_ImageNet_2016.pdf</a> \n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/230445/7607/crops1.png\" alt=\"enter image description here\" title=\"\">.. </p>",
      "rawMarkdown": "http://image-net.org/challenges/talks/2016/Hikvision_at_ImageNet_2016.pdf \n ![enter image description here][1].. \n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/230445/7607/crops1.png",
      "votes": null
    },
    {
      "id": "230652",
      "postDate": "10/12/2017 11:50:49",
      "content": "<p>Thanks for sharing!\nLoading the wight in Keras leads to error: ''You are trying to load a weight file containing 81 layers into a model with 80 layers.''. Can you share with us your architecture (layers, input size, batch size etc). </p>",
      "rawMarkdown": "Thanks for sharing!\nLoading the wight in Keras leads to error: ''You are trying to load a weight file containing 81 layers into a model with 80 layers.''. Can you share with us your architecture (layers, input size, batch size etc).",
      "votes": null
    },
    {
      "id": "230667",
      "postDate": "10/12/2017 12:40:36",
      "content": "<p>This is very helpful. Thanks</p>",
      "rawMarkdown": "This is very helpful. Thanks",
      "votes": null
    },
    {
      "id": "230725",
      "postDate": "10/12/2017 15:41:31",
      "content": "<p>@Beri, </p>\n\n<p>are you using <code>model.load_weights()</code> ? These hdf5-s contain the architecture as well. Should have model loaded if you do <code>from keras.models import load_model; model = load_weights(\"somefile.hdf5\")</code></p>",
      "rawMarkdown": "Beri, \n\nare you using `model.load_weights()` ? These hdf5-s contain the architecture as well. Should have model loaded if you do `from keras.models import load_model; model = load_weights(\"somefile.hdf5\")`",
      "votes": null
    },
    {
      "id": "230730",
      "postDate": "10/12/2017 15:55:44",
      "content": "<p>Yes\nI am loading the model from keras.applications (Xception) and then setting the wieght</p>",
      "rawMarkdown": "Yes\nI am loading the model from keras.applications (Xception) and then setting the wieght",
      "votes": null
    },
    {
      "id": "230735",
      "postDate": "10/12/2017 16:13:17",
      "content": "<p>@JuanPizarro</p>\n\n<p>This kernal should answer most of your questions:\n<a href=\"https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\">https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights</a></p>\n\n<p>About training: \nI evaluate on 10k images (from val set) every 2 million images. \nTo reach this performance I need 50 million images.\nSee the learning curves:\n<img src=\"https://image.ibb.co/dpXWdw/tf_xception_learning.jpg\" alt=\"Image\" title=\"\"></p>",
      "rawMarkdown": "JuanPizarro\n\nThis kernal should answer most of your questions:\nhttps://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\n\nAbout training: \nI evaluate on 10k images (from val set) every 2 million images. \nTo reach this performance I need 50 million images.\nSee the learning curves:\n![Image][1]\n\n\n  [1]: https://image.ibb.co/dpXWdw/tf_xception_learning.jpg",
      "votes": null
    },
    {
      "id": "231050",
      "postDate": "10/13/2017 15:20:12",
      "content": "<p>resnext101_32x4d in progress of training but showing promising results</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/231050/7636/in_progress_resnxt101.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "resnext101_32x4d in progress of training but showing promising results\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/231050/7636/in_progress_resnxt101.png",
      "votes": null
    },
    {
      "id": "231187",
      "postDate": "10/13/2017 20:39:51",
      "content": "<p>How long does rexnext take for an epoch? (How many images are in one epoch?)</p>",
      "rawMarkdown": "How long does rexnext take for an epoch? (How many images are in one epoch?)",
      "votes": null
    },
    {
      "id": "231236",
      "postDate": "10/14/2017 02:48:26",
      "content": "<p>For a batch of 256 images ( parallel over 2 gpu), 1000 iterations takes 35 min. This gives a training rate of 1000*256/35 images per minute.</p>\n\n<p>I am also trying the SE-Resnext101, which seems to give better results</p>\n\n<p>Note: after i verify the multi-gpu code is correct, i will releasr the code later.</p>\n\n<pre><code>        # one iteration update  -------------\n        images = Variable(images).cuda()\n        labels = Variable(labels).cuda()\n\n        if NUM_CUDA_DEVICES!=0:\n            logits = torch.nn.DataParallel(net)(images) #use default\n        else:\n            logits = net(images)\n\n        probs = F.softmax(logits) \n        loss  = F.cross_entropy(logits, labels)\n        acc   = top_accuracy(probs, labels, top_k=(1,))\n</code></pre>",
      "rawMarkdown": "For a batch of 256 images ( parallel over 2 gpu), 1000 iterations takes 35 min. This gives a training rate of 1000*256/35 images per minute.\n\nI am also trying the SE-Resnext101, which seems to give better results\n\nNote: after i verify the multi-gpu code is correct, i will releasr the code later.\n\n    \n            # one iteration update  -------------\n            images = Variable(images).cuda()\n            labels = Variable(labels).cuda()\n\n            if NUM_CUDA_DEVICES!=0:\n                logits = torch.nn.DataParallel(net)(images) #use default\n            else:\n                logits = net(images)\n\n            probs = F.softmax(logits) \n            loss  = F.cross_entropy(logits, labels)\n            acc   = top_accuracy(probs, labels, top_k=(1,))",
      "votes": null
    },
    {
      "id": "231262",
      "postDate": "10/14/2017 06:39:14",
      "content": "<p>Does nn.dataparallel work on your 1080ti machine? I still experience a known bug :/\nAnd what's your learning rate schedule/optimizer? I can for some reason not reproduce your results in the sense that I will take several learning rate decreases for me and many more epochs to reach 60%...</p>",
      "rawMarkdown": "Does nn.dataparallel work on your 1080ti machine? I still experience a known bug :/\nAnd what's your learning rate schedule/optimizer? I can for some reason not reproduce your results in the sense that I will take several learning rate decreases for me and many more epochs to reach 60%...",
      "votes": null
    },
    {
      "id": "231390",
      "postDate": "10/14/2017 17:09:31",
      "content": "<p>my resnext101-180 is initialised from resnext101-224 (see <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780</a>). I  trained  resnext101-224 to about 0.62 validation accuracy within 1.0 epoch (by increasing batch size=16x32, rate=0.01) and then switch to resnext101-180 ( batch size= 4x64, rate =0.01).</p>\n\n<p>My learning rate, batch size, momentum is tuned by hand. I am not sure if such hand tuned learning hyper-parameters will optimum or sub-optimum results.</p>",
      "rawMarkdown": "my resnext101-180 is initialised from resnext101-224 (see https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780). I  trained  resnext101-224 to about 0.62 validation accuracy within 1.0 epoch (by increasing batch size=16x32, rate=0.01) and then switch to resnext101-180 ( batch size= 4x64, rate =0.01).\n\nMy learning rate, batch size, momentum is tuned by hand. I am not sure if such hand tuned learning hyper-parameters will optimum or sub-optimum results.",
      "votes": null
    },
    {
      "id": "231393",
      "postDate": "10/14/2017 17:14:28",
      "content": "<p>The problem is: I cannot get a network to more than 52% accuracy with a constant learning rate, batch size and momentum. From your numbers I would assume that even with everything fixed I should be able to hit more than 52%...</p>",
      "rawMarkdown": "The problem is: I cannot get a network to more than 52% accuracy with a constant learning rate, batch size and momentum. From your numbers I would assume that even with everything fixed I should be able to hit more than 52%...",
      "votes": null
    },
    {
      "id": "231395",
      "postDate": "10/14/2017 17:28:12",
      "content": "<p>i suggest you can start off with se-resnet50 (or resnet50) on <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a>.</p>\n\n<p>keep rate =0.01, use imagenet pretrained, batch = 2x224 (or 2x160). one epoch should give you 0.57 (or 0.55). With this, you can check your setup is correct or not. You can monitor your accuracy or loss graph at every 30 min. They should match my graphs.</p>\n\n<p>This should be a good baseline. Then if you change network, the new graphs should be better</p>",
      "rawMarkdown": "i suggest you can start off with se-resnet50 (or resnet50) on https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498.\n\nkeep rate =0.01, use imagenet pretrained, batch = 2x224 (or 2x160). one epoch should give you 0.57 (or 0.55). With this, you can check your setup is correct or not. You can monitor your accuracy or loss graph at every 30 min. They should match my graphs.\n\nThis should be a good baseline. Then if you change network, the new graphs should be better",
      "votes": null
    },
    {
      "id": "231399",
      "postDate": "10/14/2017 17:39:53",
      "content": "<p>Made some modification to se-inception3. Here are the results. The weights of the modified model is initisalised from the original model above. I use learning rate start from 0.1, 0.05, 0.01, 0.005, 0.0025. Batch size = 4*128. This gives LB 0.69809 (single crop 180/180)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/231399/7638/new_inception3.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/231399/7639/new_inception3_1.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "Made some modification to se-inception3. Here are the results. The weights of the modified model is initisalised from the original model above. I use learning rate start from 0.1, 0.05, 0.01, 0.005, 0.0025. Batch size = 4*128. This gives LB 0.69809 (single crop 180/180)\n\n  ![enter image description here][1]\n\n   ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/231399/7638/new_inception3.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/231399/7639/new_inception3_1.png",
      "votes": null
    },
    {
      "id": "231402",
      "postDate": "10/14/2017 17:45:36",
      "content": "<p>When you say 2x224 you mean a batch of size 2?</p>",
      "rawMarkdown": "When you say 2x224 you mean a batch of size 2?",
      "votes": null
    },
    {
      "id": "231403",
      "postDate": "10/14/2017 17:46:52",
      "content": "<p>batch size =224</p>\n\n<p>accumulation =2</p>\n\n<p>effective batch size = 2x224</p>",
      "rawMarkdown": "batch size =224\n\naccumulation =2\n\neffective batch size = 2x224",
      "votes": null
    },
    {
      "id": "231404",
      "postDate": "10/14/2017 17:55:53",
      "content": "<p>Ahhhh, thanks. I was always wondering what you mean :)</p>",
      "rawMarkdown": "Ahhhh, thanks. I was always wondering what you mean :)",
      "votes": null
    },
    {
      "id": "231479",
      "postDate": "10/15/2017 03:23:41",
      "content": "<p>xception-180 results</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/231479/7641/xception-180.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "xception-180 results\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/231479/7641/xception-180.png",
      "votes": null
    },
    {
      "id": "231713",
      "postDate": "10/15/2017 18:51:20",
      "content": "<p>One more question: This resnext is initialized from imagenet weight right?</p>",
      "rawMarkdown": "One more question: This resnext is initialized from imagenet weight right?",
      "votes": null
    },
    {
      "id": "231740",
      "postDate": "10/15/2017 21:08:12",
      "content": "<p>yes</p>",
      "rawMarkdown": "yes",
      "votes": null
    },
    {
      "id": "231916",
      "postDate": "10/16/2017 12:06:45",
      "content": "<p>I try to replicate your results with my own implementation, but for me it is a little hard to understand your starter-kit code. Can you just confirm or correct:</p>\n\n<ol>\n<li>Use resnet-50 pretrained from torchvision on imagenet. Does not use any pretrained weights on CDiscount.</li>\n<li>1 epoch = ~15 million images and use image size 180</li>\n<li>SGD with lr=0.01, momentum=0.9, weight_decay=0.0005</li>\n<li>Batch-size=160 and accumulate 2 batches. Effective learning rate is lr/accumulation, so effective learning rate for this network is 0.01/2 = 0.005</li>\n<li>Keep all parameters constant in the first epoch to get ~0.55 accuracy.</li>\n</ol>",
      "rawMarkdown": "I try to replicate your results with my own implementation, but for me it is a little hard to understand your starter-kit code. Can you just confirm or correct:\n\n 1.  Use resnet-50 pretrained from torchvision on imagenet. Does not use any pretrained weights on CDiscount.\n 2. 1 epoch = ~15 million images and use image size 180\n 3. SGD with lr=0.01, momentum=0.9, weight_decay=0.0005\n 4. Batch-size=160 and accumulate 2 batches. Effective learning rate is lr/accumulation, so effective learning rate for this network is 0.01/2 = 0.005\n 5. Keep all parameters constant in the first epoch to get ~0.55 accuracy.",
      "votes": null
    },
    {
      "id": "231943",
      "postDate": "10/16/2017 13:31:30",
      "content": "<p>\"1. Use resnet-50 pretrained from torchvision on imagenet. Does not use any pretrained weights on CDiscount.\"</p>\n\n<p>you can use imagenet pretrained model from torchvision for start. </p>\n\n<p>I started off with imagenet pretrained model. But during my development, i change parameters for different experiments and i reused the previously trained weights as initialization for each new experiment. </p>\n\n<p>The SE-Resnet50 : LB 0.68939 (single 180/180 crop) model is the results of numerous experiments</p>\n\n<p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p>\n\n<p>\"2. 1 epoch = ~15 million images and use image size 180\"</p>\n\n<p>i use the train_id_v0_7019896 file. it has 7019896  products and 12283645 images. 1 epoch = 12283645 images. image size is 180 (but training images are perturbed by scale, shift, rotate change)</p>\n\n<p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p>\n\n<p>\"3. SGD with lr=0.01, momentum=0.9, weight_decay=0.0005\"</p>\n\n<p>this are not fixed values. they are changed when i think the loss does not improved. Roughly, lr is changed from 0.01,0.001,0.0001, momentum=0.9, 0.5,0.1, weight_decay=0.0005,0.0001.</p>\n\n<p>You can use this values for starting. </p>\n\n<p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p>\n\n<p>\"4. Batch-size=160 and accumulate 2 batches. Effective learning rate is lr/accumulation, so effective learning rate for this network is 0.01/2 = 0.005\"</p>\n\n<p>No. effective learning rate is still 0.01. In pytorch, you first set gradient=0. After one batch:</p>\n\n<p>gradient = sum/160</p>\n\n<p>Then you make an accumulation:</p>\n\n<p>gradient = sum/160 + sum/160 = 2*sum/160</p>\n\n<p>when you update the weights</p>\n\n<p>weights += -rate/2 * 2*sum/160  =  -rate*sum/160 </p>\n\n<p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p>\n\n<p>\"5.  Keep all parameters constant in the first epoch to get ~0.55 accuracy.\"</p>\n\n<p>If you keep the parameters fixed, you can get 0.55 accuracy. This is my first experiment. In my later experiment,  i change my parameters as training progress.</p>\n\n<p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p>\n\n<p>Maybe it is easier to understand what affects the rate of loss and make adjustments from there. From my experiments,  the are:</p>\n\n<ol>\n<li><p>type of augmentation</p></li>\n<li><p>learning rate (and batch size)</p></li>\n<li><p>momentum</p></li>\n</ol>\n\n<p>These three are the most important factors.</p>",
      "rawMarkdown": "\"1. Use resnet-50 pretrained from torchvision on imagenet. Does not use any pretrained weights on CDiscount.\"\n\nyou can use imagenet pretrained model from torchvision for start. \n\nI started off with imagenet pretrained model. But during my development, i change parameters for different experiments and i reused the previously trained weights as initialization for each new experiment. \n\nThe SE-Resnet50 : LB 0.68939 (single 180/180 crop) model is the results of numerous experiments\n\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\n\"2. 1 epoch = ~15 million images and use image size 180\"\n\ni use the train_id_v0_7019896 file. it has 7019896  products and 12283645 images. 1 epoch = 12283645 images. image size is 180 (but training images are perturbed by scale, shift, rotate change)\n\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\n\"3. SGD with lr=0.01, momentum=0.9, weight_decay=0.0005\"\n\nthis are not fixed values. they are changed when i think the loss does not improved. Roughly, lr is changed from 0.01,0.001,0.0001, momentum=0.9, 0.5,0.1, weight_decay=0.0005,0.0001.\n\nYou can use this values for starting. \n\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\n\"4. Batch-size=160 and accumulate 2 batches. Effective learning rate is lr/accumulation, so effective learning rate for this network is 0.01/2 = 0.005\"\n\nNo. effective learning rate is still 0.01. In pytorch, you first set gradient=0. After one batch:\n\ngradient = sum/160\n\nThen you make an accumulation:\n\ngradient = sum/160 + sum/160 = 2*sum/160\n\nwhen you update the weights\n\nweights += -rate/2 * 2*sum/160  =  -rate*sum/160 \n\n\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\n\"5.  Keep all parameters constant in the first epoch to get ~0.55 accuracy.\"\n\nIf you keep the parameters fixed, you can get 0.55 accuracy. This is my first experiment. In my later experiment,  i change my parameters as training progress.\n\n\n\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nMaybe it is easier to understand what affects the rate of loss and make adjustments from there. From my experiments,  the are:\n\n1. type of augmentation\n\n2. learning rate (and batch size)\n\n3. momentum\n\nThese three are the most important factors.",
      "votes": null
    },
    {
      "id": "231949",
      "postDate": "10/16/2017 13:45:13",
      "content": "<p>I create a new log at my sharedrive. Go to 10-03/results/resnet50-full-00\n<a href=\"https://drive.google.com/open?id=0B_DICebvRE-kYnItWXZvS0RFblU\">https://drive.google.com/open?id=0B_DICebvRE-kYnItWXZvS0RFblU</a></p>\n\n<p>you can see the training logs:</p>\n\n<p>log.train-0.txt : 0 to 16 k iterations</p>\n\n<p>log.train-1.txt: 16  to 41 k </p>\n\n<p>log.train-2.txt: 41 to 66, 66 to 79 k</p>\n\n<p>the learning rate is also recorded. this results gives about LB 0.63. This model uses 160 crops from 180 for training (this may not be the best way according to my later experiments) </p>",
      "rawMarkdown": "I create a new log at my sharedrive. Go to 10-03/results/resnet50-full-00\nhttps://drive.google.com/open?id=0B_DICebvRE-kYnItWXZvS0RFblU\n\n\nyou can see the training logs:\n\nlog.train-0.txt : 0 to 16 k iterations\n\nlog.train-1.txt: 16  to 41 k \n\nlog.train-2.txt: 41 to 66, 66 to 79 k\n\nthe learning rate is also recorded. this results gives about LB 0.63. This model uses 160 crops from 180 for training (this may not be the best way according to my later experiments)",
      "votes": null
    },
    {
      "id": "231954",
      "postDate": "10/16/2017 14:01:00",
      "content": "<p>Ah alright. Thank you very much for explaining. The log is and empty file by the way!</p>",
      "rawMarkdown": "Ah alright. Thank you very much for explaining. The log is and empty file by the way!",
      "votes": null
    },
    {
      "id": "231958",
      "postDate": "10/16/2017 14:04:49",
      "content": "<p>check this new link</p>\n\n<p><a href=\"https://drive.google.com/open?id=0B_DICebvRE-kaG5xNWtlVEdsMW8\">https://drive.google.com/open?id=0B_DICebvRE-kaG5xNWtlVEdsMW8</a></p>",
      "rawMarkdown": "check this new link\n\nhttps://drive.google.com/open?id=0B_DICebvRE-kaG5xNWtlVEdsMW8",
      "votes": null
    },
    {
      "id": "231970",
      "postDate": "10/16/2017 14:31:32",
      "content": "<p>do you ever try dpn or densenet?</p>",
      "rawMarkdown": "do you ever try dpn or densenet?",
      "votes": null
    },
    {
      "id": "231972",
      "postDate": "10/16/2017 14:35:08",
      "content": "<p>i tried dpn for 0.1 epoch but it is very slow. I did not try densenet. </p>\n\n<p>if you got time and resource, i suggest you to try because:</p>\n\n<ol>\n<li><p>ensemble different kind of net gives good results</p></li>\n<li><p>when you are forming team later, having a \"rare\" kind of network would have an advantage</p></li>\n</ol>",
      "rawMarkdown": "i tried dpn for 0.1 epoch but it is very slow. I did not try densenet. \n\nif you got time and resource, i suggest you to try because:\n\n1. ensemble different kind of net gives good results\n\n2. when you are forming team later, having a \"rare\" kind of network would have an advantage",
      "votes": null
    },
    {
      "id": "232171",
      "postDate": "10/17/2017 01:48:18",
      "content": "<p>yep, now i am trying the dpn and densenet</p>",
      "rawMarkdown": "yep, now i am trying the dpn and densenet",
      "votes": null
    },
    {
      "id": "232174",
      "postDate": "10/17/2017 01:53:50",
      "content": "<p>for your reference:  my first 10k iteratiopns using dpn96 using 224 as input. </p>\n\n<p>good luck!</p>\n\n<pre><code>** dataset setting **\ntrain_dataset.split = train_id_v0_5655916\nvalid_dataset.split = valid_id_v0_5000\nlen(train_dataset)  = 9898228\nlen(valid_dataset)  = 8785\nlen(train_loader)   = 309319\nlen(valid_loader)   = 275\nbatch_size  = 32\niter_accum  = 16\nbatch_size*iter_accum  = 512\n\n\n** start training here! **\noptimizer=&lt;torch.optim.sgd.SGD object at 0x7f32209351d0&gt;\nLR=Step Learning Rates\nrates=[' 0.0100']\nsteps=['      0']\n\nrate   iter   epoch  | valid_loss/acc | train_loss/acc | batch_loss/acc |  time   \n-------------------------------------------------------------------------------------\n0.0000   12.0 k   0.31  | 2.2947  0.5511 | 0.0000  0.0000 | 0.0000  0.0000 |     1 min \n0.0100   13.0 k   0.36  | 2.0813  0.5779 | 2.1591  0.5587 | 3.0312  0.4688 |   150 min \n0.0100   14.0 k   0.41  | 2.0427  0.5855 | 2.1341  0.5711 | 2.6260  0.5312 |   297 min\n</code></pre>",
      "rawMarkdown": "for your reference:  my first 10k iteratiopns using dpn96 using 224 as input. \n\ngood luck!\n\n \n\n    ** dataset setting **\n\ttrain_dataset.split = train_id_v0_5655916\n\tvalid_dataset.split = valid_id_v0_5000\n\tlen(train_dataset)  = 9898228\n\tlen(valid_dataset)  = 8785\n\tlen(train_loader)   = 309319\n\tlen(valid_loader)   = 275\n\tbatch_size  = 32\n\titer_accum  = 16\n\tbatch_size*iter_accum  = 512\n\n\n    ** start training here! **\n    optimizer=",
      "votes": null
    },
    {
      "id": "233398",
      "postDate": "10/20/2017 04:23:50",
      "content": "<p>it seems that you train the deep model very fast, how many epoch did you train for a deep model, etc se-resnext-50?</p>",
      "rawMarkdown": "it seems that you train the deep model very fast, how many epoch did you train for a deep model, etc se-resnext-50?",
      "votes": null
    },
    {
      "id": "233859",
      "postDate": "10/21/2017 15:24:12",
      "content": "<p>After you hit public score of 0.75 and above, you can consider pesudo lable learning of weak/semi supervised learning. Google and other cvpr paper shows that deep network still would work if there is about 25 to 20 percent label noise. Very large noisy train data is still effective</p>",
      "rawMarkdown": "After you hit public score of 0.75 and above, you can consider pesudo lable learning of weak/semi supervised learning. Google and other cvpr paper shows that deep network still would work if there is about 25 to 20 percent label noise. Very large noisy train data is still effective",
      "votes": null
    },
    {
      "id": "237644",
      "postDate": "10/30/2017 17:22:21",
      "content": "<p>Thanks so much for pre-trained weights and the kernel. They are immensely useful, especially for people with weaker hardware like me. I successfully ran the xception model, but when I try to load the inception model I get an error </p>\n\n<pre><code>You are trying to load a weight file containing 190 layers into a model with 189 layers.\n</code></pre>\n\n<p>How did you set up the inception model?</p>",
      "rawMarkdown": "Thanks so much for pre-trained weights and the kernel. They are immensely useful, especially for people with weaker hardware like me. I successfully ran the xception model, but when I try to load the inception model I get an error \n\n    You are trying to load a weight file containing 190 layers into a model with 189 layers.\n\nHow did you set up the inception model?",
      "votes": null
    },
    {
      "id": "237649",
      "postDate": "10/30/2017 17:31:15",
      "content": "<p>Hello bbrant \nin my case the inception net has an extra Dense(3k, relu) before final classification.</p>\n\n<p>Try loading models directly: </p>\n\n<pre><code>from keras.models import load_model\nmodel = load_model(\"/path/to/model.hdf5\")\nmodel.summary()\nmodel.fit(...)\n...\n</code></pre>\n\n<p>The files should already have arhitecture in them. </p>",
      "rawMarkdown": "Hello bbrant \nin my case the inception net has an extra Dense(3k, relu) before final classification.\n\nTry loading models directly: \n\n    from keras.models import load_model\n    model = load_model(\"/path/to/model.hdf5\")\n    model.summary()\n    model.fit(...)\n    ...\n\nThe files should already have arhitecture in them.",
      "votes": null
    },
    {
      "id": "237663",
      "postDate": "10/30/2017 17:50:45",
      "content": "<p>Ahh, thanks. I wasn't aware that the files contained the architecture and that the load_model function existed. I am new to data science. Thanks so much for your help and the models! </p>",
      "rawMarkdown": "Ahh, thanks. I wasn't aware that the files contained the architecture and that the load_model function existed. I am new to data science. Thanks so much for your help and the models!",
      "votes": null
    },
    {
      "id": "241803",
      "postDate": "11/09/2017 20:04:17",
      "content": "<p>Hi, \nThe file <em>Xception.Py</em> contain the line \n<code>import pyinn to P</code> <br>\nWhat kind of package need to be installed to keep it working?</p>",
      "rawMarkdown": "Hi, \nThe file *Xception.Py* contain the line \n`import pyinn to P`  \nWhat kind of package need to be installed to keep it working?",
      "votes": null
    },
    {
      "id": "241899",
      "postDate": "11/10/2017 04:19:30",
      "content": "<p><a href=\"https://github.com/szagoruyko/pyinn\">https://github.com/szagoruyko/pyinn</a></p>",
      "rawMarkdown": "https://github.com/szagoruyko/pyinn",
      "votes": null
    },
    {
      "id": "242188",
      "postDate": "11/10/2017 19:27:19",
      "content": "<p>@Heng CherKeng,\nThanks for the insight. How large is your validation set, and may I ask you to share confusion matrix for all classes in this competition (any top performing model)?</p>",
      "rawMarkdown": "Heng CherKeng,\nThanks for the insight. How large is your validation set, and may I ask you to share confusion matrix for all classes in this competition (any top performing model)?",
      "votes": null
    },
    {
      "id": "243481",
      "postDate": "11/14/2017 06:30:35",
      "content": "<p>Hi, Thanks for sharing.\nAny plan to put the model also of SE-resnext101_32x4d ?</p>",
      "rawMarkdown": "Hi, Thanks for sharing.\nAny plan to put the model also of SE-resnext101_32x4d ?",
      "votes": null
    },
    {
      "id": "243483",
      "postDate": "11/14/2017 06:32:59",
      "content": "<p>sorry. there is no plan to release SE-resnext101_32x4d.</p>",
      "rawMarkdown": "sorry. there is no plan to release SE-resnext101_32x4d.",
      "votes": null
    },
    {
      "id": "246824",
      "postDate": "11/21/2017 20:16:39",
      "content": "<p>I am using Caffe on 1080 ti with cudnn enabled. Everything looks fine. But the speed I am getting with the same resnext models is almost like one third of the speed you are reporting. Also I cannot go above 10 batch size for 180X180 inputs. Are you training end to end? I wonder why is there so much difference in the batch size and speed between caffe and pytorch. And for my time benchmarking, I am not taking data layer time into account. So memory read write is also not the reason.</p>",
      "rawMarkdown": "I am using Caffe on 1080 ti with cudnn enabled. Everything looks fine. But the speed I am getting with the same resnext models is almost like one third of the speed you are reporting. Also I cannot go above 10 batch size for 180X180 inputs. Are you training end to end? I wonder why is there so much difference in the batch size and speed between caffe and pytorch. And for my time benchmarking, I am not taking data layer time into account. So memory read write is also not the reason.",
      "votes": null
    },
    {
      "id": "249029",
      "postDate": "11/27/2017 15:17:58",
      "content": "<p>How did you guys perform the train/val split?</p>",
      "rawMarkdown": "How did you guys perform the train/val split?",
      "votes": null
    },
    {
      "id": "254908",
      "postDate": "12/07/2017 21:08:34",
      "content": "<p>@Heng CherKeng Thanks for sharing. I am using the inception and xception pre-trained models you provided. Would it be possible that you could provide the initial logs of these two models please?</p>",
      "rawMarkdown": "Heng CherKeng Thanks for sharing. I am using the inception and xception pre-trained models you provided. Would it be possible that you could provide the initial logs of these two models please?",
      "votes": null
    },
    {
      "id": "255137",
      "postDate": "12/08/2017 12:14:55",
      "content": "<p>Can you give me some insight on how to train the networks? </p>\n\n<p>I tried training th header and then the whole model but training the whole model just went super wrong. I followed the instructions in here <a href=\"https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\">https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights</a> but removed that 'middle train' as I don't know what I am training in there. </p>",
      "rawMarkdown": "Can you give me some insight on how to train the networks? \n\nI tried training th header and then the whole model but training the whole model just went super wrong. I followed the instructions in here https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights but removed that 'middle train' as I don't know what I am training in there.",
      "votes": null
    },
    {
      "id": "267273",
      "postDate": "01/10/2018 23:56:52",
      "content": "<p>@Heng CherKeng, thanks for sharing this was interesting to read!</p>\n\n<p>I am new to kaggle and wanted to try this competition.</p>\n\n<p>So does the model zoo mean that I can use any of your models?\nAlso, I can't find any train/test data in the Data section of the competition.\nAnother thing: can someone share the submission file to see how it looks like or is it prohibited?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Heng CherKeng, thanks for sharing this was interesting to read!\n\nI am new to kaggle and wanted to try this competition.\n\nSo does the model zoo mean that I can use any of your models?\nAlso, I can't find any train/test data in the Data section of the competition.\nAnother thing: can someone share the submission file to see how it looks like or is it prohibited?\n\nThanks!",
      "votes": null
    },
    {
      "id": "267413",
      "postDate": "01/11/2018 09:54:50",
      "content": "<p>@Heng CherKeng, can you explain how to use this model zoo?\nI have found data, installed pytorch and I don't get what to do next.\nI've tried to run demo.py from your zoo folder, though it didn't work out for me.</p>",
      "rawMarkdown": "Heng CherKeng, can you explain how to use this model zoo?\nI have found data, installed pytorch and I don't get what to do next.\nI've tried to run demo.py from your zoo folder, though it didn't work out for me.",
      "votes": null
    },
    {
      "id": "302337",
      "postDate": "03/23/2018 23:13:13",
      "content": "<p>Thanks a lot for sharing, I learned a lot from your code. Now that the competition is over would you share your code for other models?</p>",
      "rawMarkdown": "Thanks a lot for sharing, I learned a lot from your code. Now that the competition is over would you share your code for other models?",
      "votes": null
    },
    {
      "id": "302650",
      "postDate": "03/24/2018 14:23:25",
      "content": "<p>Ehsan, </p>\n\n<p><a href=\"https://github.com/miha-skalic/convolutedPredictions_Cdiscount\">https://github.com/miha-skalic/convolutedPredictions_Cdiscount</a></p>\n\n<p>this repository might be of interest to you. It contains the code of the whole team.</p>",
      "rawMarkdown": "Ehsan, \n\nhttps://github.com/miha-skalic/convolutedPredictions_Cdiscount\n\nthis repository might be of interest to you. It contains the code of the whole team.",
      "votes": null
    },
    {
      "id": "303255",
      "postDate": "03/25/2018 21:22:56",
      "content": "<p>Thank you so much. </p>",
      "rawMarkdown": "Thank you so much.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 230321,
      "author_name": "mihaskalic",
      "author_url": "",
      "post_date": "10/11/2017 18:06:47",
      "content": "<p>Here is a contribution to keras-tf people:\n<a href=\"https://drive.google.com/drive/folders/0B6ZtD5YbyCsJVHM4OXl1Y2stQW8?usp=sharing\">https://drive.google.com/drive/folders/0B6ZtD5YbyCsJVHM4OXl1Y2stQW8?usp=sharing</a></p>\n\n<p>You can find Xception and Inception v3 models and a pickle storing class order. On validation the models reach upper 66 (67)%. Images are preprocessed by model's default: x/255. and ((x / 255.) - 0.5) * 2.).</p>",
      "votes": null,
      "replies": [
        {
          "id": 230401,
          "author_name": "jpizarrom",
          "author_url": "",
          "post_date": "10/11/2017 21:58:11",
          "content": "<p>Thank for sharing</p>\n\n<p>What was the strategy of training you used?</p>\n\n<ul>\n<li>epochs</li>\n<li>all images?</li>\n<li>crop or downsample</li>\n<li>how much time take an epoch</li>\n<li>learning rate</li>\n<li>...</li>\n</ul>\n\n<p>thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230652,
          "author_name": "elberi",
          "author_url": "",
          "post_date": "10/12/2017 11:50:49",
          "content": "<p>Thanks for sharing!\nLoading the wight in Keras leads to error: ''You are trying to load a weight file containing 81 layers into a model with 80 layers.''. Can you share with us your architecture (layers, input size, batch size etc). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230725,
          "author_name": "mihaskalic",
          "author_url": "",
          "post_date": "10/12/2017 15:41:31",
          "content": "<p>@Beri, </p>\n\n<p>are you using <code>model.load_weights()</code> ? These hdf5-s contain the architecture as well. Should have model loaded if you do <code>from keras.models import load_model; model = load_weights(\"somefile.hdf5\")</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230730,
          "author_name": "elberi",
          "author_url": "",
          "post_date": "10/12/2017 15:55:44",
          "content": "<p>Yes\nI am loading the model from keras.applications (Xception) and then setting the wieght</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230735,
          "author_name": "mihaskalic",
          "author_url": "",
          "post_date": "10/12/2017 16:13:17",
          "content": "<p>@JuanPizarro</p>\n\n<p>This kernal should answer most of your questions:\n<a href=\"https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\">https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights</a></p>\n\n<p>About training: \nI evaluate on 10k images (from val set) every 2 million images. \nTo reach this performance I need 50 million images.\nSee the learning curves:\n<img src=\"https://image.ibb.co/dpXWdw/tf_xception_learning.jpg\" alt=\"Image\" title=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 237644,
          "author_name": "bbrandt",
          "author_url": "",
          "post_date": "10/30/2017 17:22:21",
          "content": "<p>Thanks so much for pre-trained weights and the kernel. They are immensely useful, especially for people with weaker hardware like me. I successfully ran the xception model, but when I try to load the inception model I get an error </p>\n\n<pre><code>You are trying to load a weight file containing 190 layers into a model with 189 layers.\n</code></pre>\n\n<p>How did you set up the inception model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 237649,
          "author_name": "mihaskalic",
          "author_url": "",
          "post_date": "10/30/2017 17:31:15",
          "content": "<p>Hello bbrant \nin my case the inception net has an extra Dense(3k, relu) before final classification.</p>\n\n<p>Try loading models directly: </p>\n\n<pre><code>from keras.models import load_model\nmodel = load_model(\"/path/to/model.hdf5\")\nmodel.summary()\nmodel.fit(...)\n...\n</code></pre>\n\n<p>The files should already have arhitecture in them. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 237663,
          "author_name": "bbrandt",
          "author_url": "",
          "post_date": "10/30/2017 17:50:45",
          "content": "<p>Ahh, thanks. I wasn't aware that the files contained the architecture and that the load_model function existed. I am new to data science. Thanks so much for your help and the models! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 230417,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/11/2017 23:00:48",
      "content": "<p>inception v3: 12 crops verus 144 crops. See google paper: <a href=\"https://arxiv.org/pdf/1602.07261.pdf\">https://arxiv.org/pdf/1602.07261.pdf</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/230417/7606/crops.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 230421,
      "author_name": "dongxu027",
      "author_url": "",
      "post_date": "10/11/2017 23:15:39",
      "content": "<p>This is pretty thorough. Thanks for sharing the model. Learning a lot from your posts in this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 230445,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/12/2017 00:15:28",
      "content": "<p><a href=\"http://image-net.org/challenges/talks/2016/Hikvision_at_ImageNet_2016.pdf\">http://image-net.org/challenges/talks/2016/Hikvision_at_ImageNet_2016.pdf</a> \n <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/230445/7607/crops1.png\" alt=\"enter image description here\" title=\"\">.. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 230667,
      "author_name": "ashayshah7",
      "author_url": "",
      "post_date": "10/12/2017 12:40:36",
      "content": "<p>This is very helpful. Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 231050,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/13/2017 15:20:12",
      "content": "<p>resnext101_32x4d in progress of training but showing promising results</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/231050/7636/in_progress_resnxt101.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 231187,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/13/2017 20:39:51",
          "content": "<p>How long does rexnext take for an epoch? (How many images are in one epoch?)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231236,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/14/2017 02:48:26",
          "content": "<p>For a batch of 256 images ( parallel over 2 gpu), 1000 iterations takes 35 min. This gives a training rate of 1000*256/35 images per minute.</p>\n\n<p>I am also trying the SE-Resnext101, which seems to give better results</p>\n\n<p>Note: after i verify the multi-gpu code is correct, i will releasr the code later.</p>\n\n<pre><code>        # one iteration update  -------------\n        images = Variable(images).cuda()\n        labels = Variable(labels).cuda()\n\n        if NUM_CUDA_DEVICES!=0:\n            logits = torch.nn.DataParallel(net)(images) #use default\n        else:\n            logits = net(images)\n\n        probs = F.softmax(logits) \n        loss  = F.cross_entropy(logits, labels)\n        acc   = top_accuracy(probs, labels, top_k=(1,))\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231262,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/14/2017 06:39:14",
          "content": "<p>Does nn.dataparallel work on your 1080ti machine? I still experience a known bug :/\nAnd what's your learning rate schedule/optimizer? I can for some reason not reproduce your results in the sense that I will take several learning rate decreases for me and many more epochs to reach 60%...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231390,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/14/2017 17:09:31",
          "content": "<p>my resnext101-180 is initialised from resnext101-224 (see <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780</a>). I  trained  resnext101-224 to about 0.62 validation accuracy within 1.0 epoch (by increasing batch size=16x32, rate=0.01) and then switch to resnext101-180 ( batch size= 4x64, rate =0.01).</p>\n\n<p>My learning rate, batch size, momentum is tuned by hand. I am not sure if such hand tuned learning hyper-parameters will optimum or sub-optimum results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231393,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/14/2017 17:14:28",
          "content": "<p>The problem is: I cannot get a network to more than 52% accuracy with a constant learning rate, batch size and momentum. From your numbers I would assume that even with everything fixed I should be able to hit more than 52%...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231395,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/14/2017 17:28:12",
          "content": "<p>i suggest you can start off with se-resnet50 (or resnet50) on <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a>.</p>\n\n<p>keep rate =0.01, use imagenet pretrained, batch = 2x224 (or 2x160). one epoch should give you 0.57 (or 0.55). With this, you can check your setup is correct or not. You can monitor your accuracy or loss graph at every 30 min. They should match my graphs.</p>\n\n<p>This should be a good baseline. Then if you change network, the new graphs should be better</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231402,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/14/2017 17:45:36",
          "content": "<p>When you say 2x224 you mean a batch of size 2?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231403,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/14/2017 17:46:52",
          "content": "<p>batch size =224</p>\n\n<p>accumulation =2</p>\n\n<p>effective batch size = 2x224</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231404,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/14/2017 17:55:53",
          "content": "<p>Ahhhh, thanks. I was always wondering what you mean :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231713,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/15/2017 18:51:20",
          "content": "<p>One more question: This resnext is initialized from imagenet weight right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231740,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/15/2017 21:08:12",
          "content": "<p>yes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231916,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/16/2017 12:06:45",
          "content": "<p>I try to replicate your results with my own implementation, but for me it is a little hard to understand your starter-kit code. Can you just confirm or correct:</p>\n\n<ol>\n<li>Use resnet-50 pretrained from torchvision on imagenet. Does not use any pretrained weights on CDiscount.</li>\n<li>1 epoch = ~15 million images and use image size 180</li>\n<li>SGD with lr=0.01, momentum=0.9, weight_decay=0.0005</li>\n<li>Batch-size=160 and accumulate 2 batches. Effective learning rate is lr/accumulation, so effective learning rate for this network is 0.01/2 = 0.005</li>\n<li>Keep all parameters constant in the first epoch to get ~0.55 accuracy.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231943,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/16/2017 13:31:30",
          "content": "<p>\"1. Use resnet-50 pretrained from torchvision on imagenet. Does not use any pretrained weights on CDiscount.\"</p>\n\n<p>you can use imagenet pretrained model from torchvision for start. </p>\n\n<p>I started off with imagenet pretrained model. But during my development, i change parameters for different experiments and i reused the previously trained weights as initialization for each new experiment. </p>\n\n<p>The SE-Resnet50 : LB 0.68939 (single 180/180 crop) model is the results of numerous experiments</p>\n\n<p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p>\n\n<p>\"2. 1 epoch = ~15 million images and use image size 180\"</p>\n\n<p>i use the train_id_v0_7019896 file. it has 7019896  products and 12283645 images. 1 epoch = 12283645 images. image size is 180 (but training images are perturbed by scale, shift, rotate change)</p>\n\n<p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p>\n\n<p>\"3. SGD with lr=0.01, momentum=0.9, weight_decay=0.0005\"</p>\n\n<p>this are not fixed values. they are changed when i think the loss does not improved. Roughly, lr is changed from 0.01,0.001,0.0001, momentum=0.9, 0.5,0.1, weight_decay=0.0005,0.0001.</p>\n\n<p>You can use this values for starting. </p>\n\n<p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p>\n\n<p>\"4. Batch-size=160 and accumulate 2 batches. Effective learning rate is lr/accumulation, so effective learning rate for this network is 0.01/2 = 0.005\"</p>\n\n<p>No. effective learning rate is still 0.01. In pytorch, you first set gradient=0. After one batch:</p>\n\n<p>gradient = sum/160</p>\n\n<p>Then you make an accumulation:</p>\n\n<p>gradient = sum/160 + sum/160 = 2*sum/160</p>\n\n<p>when you update the weights</p>\n\n<p>weights += -rate/2 * 2*sum/160  =  -rate*sum/160 </p>\n\n<p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p>\n\n<p>\"5.  Keep all parameters constant in the first epoch to get ~0.55 accuracy.\"</p>\n\n<p>If you keep the parameters fixed, you can get 0.55 accuracy. This is my first experiment. In my later experiment,  i change my parameters as training progress.</p>\n\n<p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p>\n\n<p>Maybe it is easier to understand what affects the rate of loss and make adjustments from there. From my experiments,  the are:</p>\n\n<ol>\n<li><p>type of augmentation</p></li>\n<li><p>learning rate (and batch size)</p></li>\n<li><p>momentum</p></li>\n</ol>\n\n<p>These three are the most important factors.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231949,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/16/2017 13:45:13",
          "content": "<p>I create a new log at my sharedrive. Go to 10-03/results/resnet50-full-00\n<a href=\"https://drive.google.com/open?id=0B_DICebvRE-kYnItWXZvS0RFblU\">https://drive.google.com/open?id=0B_DICebvRE-kYnItWXZvS0RFblU</a></p>\n\n<p>you can see the training logs:</p>\n\n<p>log.train-0.txt : 0 to 16 k iterations</p>\n\n<p>log.train-1.txt: 16  to 41 k </p>\n\n<p>log.train-2.txt: 41 to 66, 66 to 79 k</p>\n\n<p>the learning rate is also recorded. this results gives about LB 0.63. This model uses 160 crops from 180 for training (this may not be the best way according to my later experiments) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231954,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/16/2017 14:01:00",
          "content": "<p>Ah alright. Thank you very much for explaining. The log is and empty file by the way!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231958,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/16/2017 14:04:49",
          "content": "<p>check this new link</p>\n\n<p><a href=\"https://drive.google.com/open?id=0B_DICebvRE-kaG5xNWtlVEdsMW8\">https://drive.google.com/open?id=0B_DICebvRE-kaG5xNWtlVEdsMW8</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 246824,
          "author_name": "ankitsinghamaan",
          "author_url": "",
          "post_date": "11/21/2017 20:16:39",
          "content": "<p>I am using Caffe on 1080 ti with cudnn enabled. Everything looks fine. But the speed I am getting with the same resnext models is almost like one third of the speed you are reporting. Also I cannot go above 10 batch size for 180X180 inputs. Are you training end to end? I wonder why is there so much difference in the batch size and speed between caffe and pytorch. And for my time benchmarking, I am not taking data layer time into account. So memory read write is also not the reason.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 231399,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/14/2017 17:39:53",
      "content": "<p>Made some modification to se-inception3. Here are the results. The weights of the modified model is initisalised from the original model above. I use learning rate start from 0.1, 0.05, 0.01, 0.005, 0.0025. Batch size = 4*128. This gives LB 0.69809 (single crop 180/180)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/231399/7638/new_inception3.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/231399/7639/new_inception3_1.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 231479,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/15/2017 03:23:41",
      "content": "<p>xception-180 results</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/231479/7641/xception-180.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 231970,
      "author_name": "ifighting",
      "author_url": "",
      "post_date": "10/16/2017 14:31:32",
      "content": "<p>do you ever try dpn or densenet?</p>",
      "votes": null,
      "replies": [
        {
          "id": 231972,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/16/2017 14:35:08",
          "content": "<p>i tried dpn for 0.1 epoch but it is very slow. I did not try densenet. </p>\n\n<p>if you got time and resource, i suggest you to try because:</p>\n\n<ol>\n<li><p>ensemble different kind of net gives good results</p></li>\n<li><p>when you are forming team later, having a \"rare\" kind of network would have an advantage</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 232171,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "10/17/2017 01:48:18",
          "content": "<p>yep, now i am trying the dpn and densenet</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 232174,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/17/2017 01:53:50",
          "content": "<p>for your reference:  my first 10k iteratiopns using dpn96 using 224 as input. </p>\n\n<p>good luck!</p>\n\n<pre><code>** dataset setting **\ntrain_dataset.split = train_id_v0_5655916\nvalid_dataset.split = valid_id_v0_5000\nlen(train_dataset)  = 9898228\nlen(valid_dataset)  = 8785\nlen(train_loader)   = 309319\nlen(valid_loader)   = 275\nbatch_size  = 32\niter_accum  = 16\nbatch_size*iter_accum  = 512\n\n\n** start training here! **\noptimizer=&lt;torch.optim.sgd.SGD object at 0x7f32209351d0&gt;\nLR=Step Learning Rates\nrates=[' 0.0100']\nsteps=['      0']\n\nrate   iter   epoch  | valid_loss/acc | train_loss/acc | batch_loss/acc |  time   \n-------------------------------------------------------------------------------------\n0.0000   12.0 k   0.31  | 2.2947  0.5511 | 0.0000  0.0000 | 0.0000  0.0000 |     1 min \n0.0100   13.0 k   0.36  | 2.0813  0.5779 | 2.1591  0.5587 | 3.0312  0.4688 |   150 min \n0.0100   14.0 k   0.41  | 2.0427  0.5855 | 2.1341  0.5711 | 2.6260  0.5312 |   297 min\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233398,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "10/20/2017 04:23:50",
          "content": "<p>it seems that you train the deep model very fast, how many epoch did you train for a deep model, etc se-resnext-50?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 233859,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/21/2017 15:24:12",
          "content": "<p>After you hit public score of 0.75 and above, you can consider pesudo lable learning of weak/semi supervised learning. Google and other cvpr paper shows that deep network still would work if there is about 25 to 20 percent label noise. Very large noisy train data is still effective</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 241803,
      "author_name": "algila",
      "author_url": "",
      "post_date": "11/09/2017 20:04:17",
      "content": "<p>Hi, \nThe file <em>Xception.Py</em> contain the line \n<code>import pyinn to P</code> <br>\nWhat kind of package need to be installed to keep it working?</p>",
      "votes": null,
      "replies": [
        {
          "id": 241899,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/10/2017 04:19:30",
          "content": "<p><a href=\"https://github.com/szagoruyko/pyinn\">https://github.com/szagoruyko/pyinn</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 242188,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "11/10/2017 19:27:19",
      "content": "<p>@Heng CherKeng,\nThanks for the insight. How large is your validation set, and may I ask you to share confusion matrix for all classes in this competition (any top performing model)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 243481,
      "author_name": "algila",
      "author_url": "",
      "post_date": "11/14/2017 06:30:35",
      "content": "<p>Hi, Thanks for sharing.\nAny plan to put the model also of SE-resnext101_32x4d ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 243483,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/14/2017 06:32:59",
          "content": "<p>sorry. there is no plan to release SE-resnext101_32x4d.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 249029,
      "author_name": "skinish",
      "author_url": "",
      "post_date": "11/27/2017 15:17:58",
      "content": "<p>How did you guys perform the train/val split?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 254908,
      "author_name": "panatopos",
      "author_url": "",
      "post_date": "12/07/2017 21:08:34",
      "content": "<p>@Heng CherKeng Thanks for sharing. I am using the inception and xception pre-trained models you provided. Would it be possible that you could provide the initial logs of these two models please?</p>",
      "votes": null,
      "replies": [
        {
          "id": 255137,
          "author_name": "skinish",
          "author_url": "",
          "post_date": "12/08/2017 12:14:55",
          "content": "<p>Can you give me some insight on how to train the networks? </p>\n\n<p>I tried training th header and then the whole model but training the whole model just went super wrong. I followed the instructions in here <a href=\"https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\">https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights</a> but removed that 'middle train' as I don't know what I am training in there. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 267273,
      "author_name": "nate123",
      "author_url": "",
      "post_date": "01/10/2018 23:56:52",
      "content": "<p>@Heng CherKeng, thanks for sharing this was interesting to read!</p>\n\n<p>I am new to kaggle and wanted to try this competition.</p>\n\n<p>So does the model zoo mean that I can use any of your models?\nAlso, I can't find any train/test data in the Data section of the competition.\nAnother thing: can someone share the submission file to see how it looks like or is it prohibited?</p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267413,
      "author_name": "nate123",
      "author_url": "",
      "post_date": "01/11/2018 09:54:50",
      "content": "<p>@Heng CherKeng, can you explain how to use this model zoo?\nI have found data, installed pytorch and I don't get what to do next.\nI've tried to run demo.py from your zoo folder, though it didn't work out for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302337,
      "author_name": "ehsanfathi",
      "author_url": "",
      "post_date": "03/23/2018 23:13:13",
      "content": "<p>Thanks a lot for sharing, I learned a lot from your code. Now that the competition is over would you share your code for other models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 302650,
          "author_name": "mihaskalic",
          "author_url": "",
          "post_date": "03/24/2018 14:23:25",
          "content": "<p>Ehsan, </p>\n\n<p><a href=\"https://github.com/miha-skalic/convolutedPredictions_Cdiscount\">https://github.com/miha-skalic/convolutedPredictions_Cdiscount</a></p>\n\n<p>this repository might be of interest to you. It contains the code of the whole team.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 303255,
          "author_name": "ehsanfathi",
          "author_url": "",
          "post_date": "03/25/2018 21:22:56",
          "content": "<p>Thank you so much. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "230288": "** important **\nplease refer to @Vladimir Iglovikov at https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652\n\nplease train with long epoch to get better results . My models results are from training only up to 4 to 5 epoch. \n\n------\n\nyou can download my trained models at the share drive:\n\nhttps://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE\n\n(go to the \"release\" folder)\n\nperformance:\n\n1. SE-InceptionV3 : LB 0.69673 (single 180/180 crop)\n\n2. InceptionV3 : LB 0.69565 (single 180/180 crop)\n\n3.  Xception : LB 0.69422 (single 180/180 crop)\n\n\ncompare with other models:\n\n -  SE-Resnet50 : LB 0.68939 (single 180/180 crop)\n\n -  SE-resnext101_32x4d : LB 0.71064  (single crop 180/180)\n\n -  InceptionV4 :  in progress\n\n -  Inception_ResnetV2 :  in progress\n\n -  Resnet152 :  in progress\n\n - ResDrop269 (Stochastic Depth Resnet) : in progress\n\n*Note: all results here are trained only for limited number of epoch. if they are trained longer, i expect some slight improvement e.g. +0.005.\n----\n\n[what to do with it?]\n\n1.  use it to finetune your new models of different scale\n\n2. perform test-time augmentation of multi-crops and scales\n\n3. extend the model, e.g. use more inception or fc layers\n\nYou can refer to the vgg paper \"Very Deep Convolutional Networks for Large-Scale Image Recognition\" -Karen Simonyan, Andrew Zisserman, Arxiv 2014. If you google for the presentation slides of ilsvrc, you can find more tricks.\n\nYet another paper mentions \"top-k pooling\" for inference. please refer to \"http://ww.dahua.me/publications/dhl17_polynet.pdf\", see section.5\n\nI expect multi-crop results of maybe LB 0.71 to 0.72, but I haven't got enough time and resources to try. e.g. 144 crops per image is a nightmare for me.\n\nif you got good (or bad) results, I would appreciate you can report your results here so that me and other can repeat (or avoid) it.\n\ngood luck!\n\n----\n\n[training details]\n\nThese models are created using my starter kit at: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\n\n1. refer to \"training_code\" for details of augmentation and data for train split. \nThese files are from my starter kit. They should contain enough information for your future work. If there is anything unclear, please ask at  the forum or refer to the start kit.\n\n\n2. refer to \"demo_code\" for details of how to apply model to test an image. e.g. how to normalize input image with mean and std values. \"label_to_cat_id\" maps the class label to the cat_id.",
    "230321": "Here is a contribution to keras-tf people:\nhttps://drive.google.com/drive/folders/0B6ZtD5YbyCsJVHM4OXl1Y2stQW8?usp=sharing\n\nYou can find Xception and Inception v3 models and a pickle storing class order. On validation the models reach upper 66 (67)%. Images are preprocessed by model's default: x/255. and ((x / 255.) - 0.5) * 2.).",
    "230401": "Thank for sharing\n\nWhat was the strategy of training you used?\n\n- epochs\n- all images?\n- crop or downsample\n- how much time take an epoch\n- learning rate\n- ...\n\nthanks",
    "230417": "inception v3: 12 crops verus 144 crops. See google paper: https://arxiv.org/pdf/1602.07261.pdf\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/230417/7606/crops.png",
    "230421": "This is pretty thorough. Thanks for sharing the model. Learning a lot from your posts in this competition.",
    "230445": "http://image-net.org/challenges/talks/2016/Hikvision_at_ImageNet_2016.pdf \n ![enter image description here][1].. \n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/230445/7607/crops1.png",
    "230652": "Thanks for sharing!\nLoading the wight in Keras leads to error: ''You are trying to load a weight file containing 81 layers into a model with 80 layers.''. Can you share with us your architecture (layers, input size, batch size etc).",
    "230667": "This is very helpful. Thanks",
    "230725": "Beri, \n\nare you using `model.load_weights()` ? These hdf5-s contain the architecture as well. Should have model loaded if you do `from keras.models import load_model; model = load_weights(\"somefile.hdf5\")`",
    "230730": "Yes\nI am loading the model from keras.applications (Xception) and then setting the wieght",
    "230735": "JuanPizarro\n\nThis kernal should answer most of your questions:\nhttps://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\n\nAbout training: \nI evaluate on 10k images (from val set) every 2 million images. \nTo reach this performance I need 50 million images.\nSee the learning curves:\n![Image][1]\n\n\n  [1]: https://image.ibb.co/dpXWdw/tf_xception_learning.jpg",
    "231050": "resnext101_32x4d in progress of training but showing promising results\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/231050/7636/in_progress_resnxt101.png",
    "231187": "How long does rexnext take for an epoch? (How many images are in one epoch?)",
    "231236": "For a batch of 256 images ( parallel over 2 gpu), 1000 iterations takes 35 min. This gives a training rate of 1000*256/35 images per minute.\n\nI am also trying the SE-Resnext101, which seems to give better results\n\nNote: after i verify the multi-gpu code is correct, i will releasr the code later.\n\n    \n            # one iteration update  -------------\n            images = Variable(images).cuda()\n            labels = Variable(labels).cuda()\n\n            if NUM_CUDA_DEVICES!=0:\n                logits = torch.nn.DataParallel(net)(images) #use default\n            else:\n                logits = net(images)\n\n            probs = F.softmax(logits) \n            loss  = F.cross_entropy(logits, labels)\n            acc   = top_accuracy(probs, labels, top_k=(1,))",
    "231262": "Does nn.dataparallel work on your 1080ti machine? I still experience a known bug :/\nAnd what's your learning rate schedule/optimizer? I can for some reason not reproduce your results in the sense that I will take several learning rate decreases for me and many more epochs to reach 60%...",
    "231390": "my resnext101-180 is initialised from resnext101-224 (see https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780). I  trained  resnext101-224 to about 0.62 validation accuracy within 1.0 epoch (by increasing batch size=16x32, rate=0.01) and then switch to resnext101-180 ( batch size= 4x64, rate =0.01).\n\nMy learning rate, batch size, momentum is tuned by hand. I am not sure if such hand tuned learning hyper-parameters will optimum or sub-optimum results.",
    "231393": "The problem is: I cannot get a network to more than 52% accuracy with a constant learning rate, batch size and momentum. From your numbers I would assume that even with everything fixed I should be able to hit more than 52%...",
    "231395": "i suggest you can start off with se-resnet50 (or resnet50) on https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498.\n\nkeep rate =0.01, use imagenet pretrained, batch = 2x224 (or 2x160). one epoch should give you 0.57 (or 0.55). With this, you can check your setup is correct or not. You can monitor your accuracy or loss graph at every 30 min. They should match my graphs.\n\nThis should be a good baseline. Then if you change network, the new graphs should be better",
    "231399": "Made some modification to se-inception3. Here are the results. The weights of the modified model is initisalised from the original model above. I use learning rate start from 0.1, 0.05, 0.01, 0.005, 0.0025. Batch size = 4*128. This gives LB 0.69809 (single crop 180/180)\n\n  ![enter image description here][1]\n\n   ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/231399/7638/new_inception3.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/231399/7639/new_inception3_1.png",
    "231402": "When you say 2x224 you mean a batch of size 2?",
    "231403": "batch size =224\n\naccumulation =2\n\neffective batch size = 2x224",
    "231404": "Ahhhh, thanks. I was always wondering what you mean :)",
    "231479": "xception-180 results\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/231479/7641/xception-180.png",
    "231713": "One more question: This resnext is initialized from imagenet weight right?",
    "231740": "yes",
    "231916": "I try to replicate your results with my own implementation, but for me it is a little hard to understand your starter-kit code. Can you just confirm or correct:\n\n 1.  Use resnet-50 pretrained from torchvision on imagenet. Does not use any pretrained weights on CDiscount.\n 2. 1 epoch = ~15 million images and use image size 180\n 3. SGD with lr=0.01, momentum=0.9, weight_decay=0.0005\n 4. Batch-size=160 and accumulate 2 batches. Effective learning rate is lr/accumulation, so effective learning rate for this network is 0.01/2 = 0.005\n 5. Keep all parameters constant in the first epoch to get ~0.55 accuracy.",
    "231943": "\"1. Use resnet-50 pretrained from torchvision on imagenet. Does not use any pretrained weights on CDiscount.\"\n\nyou can use imagenet pretrained model from torchvision for start. \n\nI started off with imagenet pretrained model. But during my development, i change parameters for different experiments and i reused the previously trained weights as initialization for each new experiment. \n\nThe SE-Resnet50 : LB 0.68939 (single 180/180 crop) model is the results of numerous experiments\n\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\n\"2. 1 epoch = ~15 million images and use image size 180\"\n\ni use the train_id_v0_7019896 file. it has 7019896  products and 12283645 images. 1 epoch = 12283645 images. image size is 180 (but training images are perturbed by scale, shift, rotate change)\n\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\n\"3. SGD with lr=0.01, momentum=0.9, weight_decay=0.0005\"\n\nthis are not fixed values. they are changed when i think the loss does not improved. Roughly, lr is changed from 0.01,0.001,0.0001, momentum=0.9, 0.5,0.1, weight_decay=0.0005,0.0001.\n\nYou can use this values for starting. \n\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\n\"4. Batch-size=160 and accumulate 2 batches. Effective learning rate is lr/accumulation, so effective learning rate for this network is 0.01/2 = 0.005\"\n\nNo. effective learning rate is still 0.01. In pytorch, you first set gradient=0. After one batch:\n\ngradient = sum/160\n\nThen you make an accumulation:\n\ngradient = sum/160 + sum/160 = 2*sum/160\n\nwhen you update the weights\n\nweights += -rate/2 * 2*sum/160  =  -rate*sum/160 \n\n\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\n\"5.  Keep all parameters constant in the first epoch to get ~0.55 accuracy.\"\n\nIf you keep the parameters fixed, you can get 0.55 accuracy. This is my first experiment. In my later experiment,  i change my parameters as training progress.\n\n\n\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nMaybe it is easier to understand what affects the rate of loss and make adjustments from there. From my experiments,  the are:\n\n1. type of augmentation\n\n2. learning rate (and batch size)\n\n3. momentum\n\nThese three are the most important factors.",
    "231949": "I create a new log at my sharedrive. Go to 10-03/results/resnet50-full-00\nhttps://drive.google.com/open?id=0B_DICebvRE-kYnItWXZvS0RFblU\n\n\nyou can see the training logs:\n\nlog.train-0.txt : 0 to 16 k iterations\n\nlog.train-1.txt: 16  to 41 k \n\nlog.train-2.txt: 41 to 66, 66 to 79 k\n\nthe learning rate is also recorded. this results gives about LB 0.63. This model uses 160 crops from 180 for training (this may not be the best way according to my later experiments)",
    "231954": "Ah alright. Thank you very much for explaining. The log is and empty file by the way!",
    "231958": "check this new link\n\nhttps://drive.google.com/open?id=0B_DICebvRE-kaG5xNWtlVEdsMW8",
    "231970": "do you ever try dpn or densenet?",
    "231972": "i tried dpn for 0.1 epoch but it is very slow. I did not try densenet. \n\nif you got time and resource, i suggest you to try because:\n\n1. ensemble different kind of net gives good results\n\n2. when you are forming team later, having a \"rare\" kind of network would have an advantage",
    "232171": "yep, now i am trying the dpn and densenet",
    "232174": "for your reference:  my first 10k iteratiopns using dpn96 using 224 as input. \n\ngood luck!\n\n \n\n    ** dataset setting **\n\ttrain_dataset.split = train_id_v0_5655916\n\tvalid_dataset.split = valid_id_v0_5000\n\tlen(train_dataset)  = 9898228\n\tlen(valid_dataset)  = 8785\n\tlen(train_loader)   = 309319\n\tlen(valid_loader)   = 275\n\tbatch_size  = 32\n\titer_accum  = 16\n\tbatch_size*iter_accum  = 512\n\n\n    ** start training here! **\n    optimizer=",
    "233398": "it seems that you train the deep model very fast, how many epoch did you train for a deep model, etc se-resnext-50?",
    "233859": "After you hit public score of 0.75 and above, you can consider pesudo lable learning of weak/semi supervised learning. Google and other cvpr paper shows that deep network still would work if there is about 25 to 20 percent label noise. Very large noisy train data is still effective",
    "237644": "Thanks so much for pre-trained weights and the kernel. They are immensely useful, especially for people with weaker hardware like me. I successfully ran the xception model, but when I try to load the inception model I get an error \n\n    You are trying to load a weight file containing 190 layers into a model with 189 layers.\n\nHow did you set up the inception model?",
    "237649": "Hello bbrant \nin my case the inception net has an extra Dense(3k, relu) before final classification.\n\nTry loading models directly: \n\n    from keras.models import load_model\n    model = load_model(\"/path/to/model.hdf5\")\n    model.summary()\n    model.fit(...)\n    ...\n\nThe files should already have arhitecture in them.",
    "237663": "Ahh, thanks. I wasn't aware that the files contained the architecture and that the load_model function existed. I am new to data science. Thanks so much for your help and the models!",
    "241803": "Hi, \nThe file *Xception.Py* contain the line \n`import pyinn to P`  \nWhat kind of package need to be installed to keep it working?",
    "241899": "https://github.com/szagoruyko/pyinn",
    "242188": "Heng CherKeng,\nThanks for the insight. How large is your validation set, and may I ask you to share confusion matrix for all classes in this competition (any top performing model)?",
    "243481": "Hi, Thanks for sharing.\nAny plan to put the model also of SE-resnext101_32x4d ?",
    "243483": "sorry. there is no plan to release SE-resnext101_32x4d.",
    "246824": "I am using Caffe on 1080 ti with cudnn enabled. Everything looks fine. But the speed I am getting with the same resnext models is almost like one third of the speed you are reporting. Also I cannot go above 10 batch size for 180X180 inputs. Are you training end to end? I wonder why is there so much difference in the batch size and speed between caffe and pytorch. And for my time benchmarking, I am not taking data layer time into account. So memory read write is also not the reason.",
    "249029": "How did you guys perform the train/val split?",
    "254908": "Heng CherKeng Thanks for sharing. I am using the inception and xception pre-trained models you provided. Would it be possible that you could provide the initial logs of these two models please?",
    "255137": "Can you give me some insight on how to train the networks? \n\nI tried training th header and then the whole model but training the whole model just went super wrong. I followed the instructions in here https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights but removed that 'middle train' as I don't know what I am training in there.",
    "267273": "Heng CherKeng, thanks for sharing this was interesting to read!\n\nI am new to kaggle and wanted to try this competition.\n\nSo does the model zoo mean that I can use any of your models?\nAlso, I can't find any train/test data in the Data section of the competition.\nAnother thing: can someone share the submission file to see how it looks like or is it prohibited?\n\nThanks!",
    "267413": "Heng CherKeng, can you explain how to use this model zoo?\nI have found data, installed pytorch and I don't get what to do next.\nI've tried to run demo.py from your zoo folder, though it didn't work out for me.",
    "302337": "Thanks a lot for sharing, I learned a lot from your code. Now that the competition is over would you share your code for other models?",
    "302650": "Ehsan, \n\nhttps://github.com/miha-skalic/convolutedPredictions_Cdiscount\n\nthis repository might be of interest to you. It contains the code of the whole team.",
    "303255": "Thank you so much."
  },
  "source": "meta"
}