{
  "id": 41652,
  "title": "Single model performance",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/41652",
  "author_name": "",
  "post_date": "2017-10-22T01:11:31.800916500Z",
  "votes": 43,
  "comment_count": 91,
  "views": 0,
  "content": "<p>I am not familiar with this type of the train set and this number of target classes. I hope this thread will help to figure out what models perform better and which of them worth considering.</p>\n\n<p>To start:</p>\n\n<p>Resnet 101, no bagging, no TTA</p>\n\n<ul>\n<li>val_loss = 1.18 </li>\n<li>val_acc = 0.744 </li>\n<li>LB = 0.7437</li>\n</ul>",
  "messages": [
    {
      "id": "233986",
      "postDate": "10/22/2017 01:11:31",
      "content": "<p>I am not familiar with this type of the train set and this number of target classes. I hope this thread will help to figure out what models perform better and which of them worth considering.</p>\n\n<p>To start:</p>\n\n<p>Resnet 101, no bagging, no TTA</p>\n\n<ul>\n<li>val_loss = 1.18 </li>\n<li>val_acc = 0.744 </li>\n<li>LB = 0.7437</li>\n</ul>",
      "rawMarkdown": "I am not familiar with this type of the train set and this number of target classes. I hope this thread will help to figure out what models perform better and which of them worth considering.\n\nTo start:\n\nResnet 101, no bagging, no TTA\n\n - val_loss = 1.18 \n - val_acc = 0.744 \n - LB = 0.7437",
      "votes": null
    },
    {
      "id": "233991",
      "postDate": "10/22/2017 01:25:43",
      "content": "<p>thanks a lot!</p>",
      "rawMarkdown": "thanks a lot!",
      "votes": null
    },
    {
      "id": "233996",
      "postDate": "10/22/2017 01:48:17",
      "content": "<p>That's quite high score for single model! How long did you train for this and which GPU did you use?</p>\n\n<p>Due to limited resources, I'm using images scaled down to 90x90 and simple Xception, no TTA</p>\n\n<ul>\n<li><p>val_acc = .6904 (for each image) </p></li>\n<li><p>LB= .7092 (averaged softmax output for all images in same product id)</p></li>\n</ul>\n\n<p>I'm trying many other approaches like using top 2000 frequent items, using category 1,2 information, etc, but nothing beats this baseline yet. Sad to see much better results with basic model, but thanks for the info.</p>",
      "rawMarkdown": "That's quite high score for single model! How long did you train for this and which GPU did you use?\n\nDue to limited resources, I'm using images scaled down to 90x90 and simple Xception, no TTA\n\n - val_acc = .6904 (for each image) \n\n - LB= .7092 (averaged softmax output for all images in same product id)\n\nI'm trying many other approaches like using top 2000 frequent items, using category 1,2 information, etc, but nothing beats this baseline yet. Sad to see much better results with basic model, but thanks for the info.",
      "votes": null
    },
    {
      "id": "233997",
      "postDate": "10/22/2017 01:51:12",
      "content": "<p>4 x GTX 1080 Ti</p>\n\n<p>18 epochs x 4.5 hours</p>\n\n<p>I have 0.69 after 7th epoch</p>",
      "rawMarkdown": "4 x GTX 1080 Ti\n\n18 epochs x 4.5 hours\n\nI have 0.69 after 7th epoch",
      "votes": null
    },
    {
      "id": "234000",
      "postDate": "10/22/2017 01:58:43",
      "content": "<p>seems that we need long epoches!</p>",
      "rawMarkdown": "seems that we need long epoches!",
      "votes": null
    },
    {
      "id": "234005",
      "postDate": "10/22/2017 02:40:41",
      "content": "<p>Amazing score. Did you do training without pretrained weight?</p>\n\n<p>My results: \ninception_v3, 160 crop, val loss: 1.78~, val acc: 0.66~, lb: 0.673~</p>",
      "rawMarkdown": "Amazing score. Did you do training without pretrained weight?\n\nMy results: \ninception_v3, 160 crop, val loss: 1.78~, val acc: 0.66~, lb: 0.673~",
      "votes": null
    },
    {
      "id": "234006",
      "postDate": "10/22/2017 02:41:18",
      "content": "<p>i am also using 4 x GTX 1080 Ti. For resnext101 (about same flops as your resnet101 ),  1 epoch=12283645 images takes about 1.3 hr. I am using 180x180 as input. I suppose your input size is not 180x180?</p>",
      "rawMarkdown": "i am also using 4 x GTX 1080 Ti. For resnext101 (about same flops as your resnet101 ),  1 epoch=12283645 images takes about 1.3 hr. I am using 180x180 as input. I suppose your input size is not 180x180?",
      "votes": null
    },
    {
      "id": "234007",
      "postDate": "10/22/2017 02:43:55",
      "content": "<p>what does TTA mean?</p>",
      "rawMarkdown": "what does TTA mean?",
      "votes": null
    },
    {
      "id": "234010",
      "postDate": "10/22/2017 02:52:06",
      "content": "<p>I use 160x160 crops.</p>\n\n<p>1.3hrs per epoch.... It looks like I am doing something wrong.</p>\n\n<p>It may happen that my SSD is not fast enough and I need to buy faster one.</p>",
      "rawMarkdown": "I use 160x160 crops.\n\n1.3hrs per epoch.... It looks like I am doing something wrong.\n\nIt may happen that my SSD is not fast enough and I need to buy faster one.",
      "votes": null
    },
    {
      "id": "234012",
      "postDate": "10/22/2017 02:53:57",
      "content": "<p>TTA= test time augmentation (e.g. multi crops, resize in test time)</p>",
      "rawMarkdown": "TTA= test time augmentation (e.g. multi crops, resize in test time)",
      "votes": null
    },
    {
      "id": "234013",
      "postDate": "10/22/2017 02:56:01",
      "content": "<p>Test-time augmentation. In my case, I can get extra 0.01 just with 12-way TTA(flip, shift up/down, etc) with exact same weight - almost free cake :)</p>",
      "rawMarkdown": "Test-time augmentation. In my case, I can get extra 0.01 just with 12-way TTA(flip, shift up/down, etc) with exact same weight - almost free cake :)",
      "votes": null
    },
    {
      "id": "234019",
      "postDate": "10/22/2017 03:31:19",
      "content": "<p>I used pre-trained. I would be also interested how long will it take to get the same result without pre-trained, but it may be too computationally expensive.</p>",
      "rawMarkdown": "I used pre-trained. I would be also interested how long will it take to get the same result without pre-trained, but it may be too computationally expensive.",
      "votes": null
    },
    {
      "id": "234026",
      "postDate": "10/22/2017 03:51:23",
      "content": "<p>I used resnext-101, with image centor croped to 128X128, used pre-trained. It can get 0.59 val _acc after 1 epoches. I used sgd optimizer. Training is going...</p>",
      "rawMarkdown": "I used resnext-101, with image centor croped to 128X128, used pre-trained. It can get 0.59 val _acc after 1 epoches. I used sgd optimizer. Training is going...",
      "votes": null
    },
    {
      "id": "234032",
      "postDate": "10/22/2017 04:10:23",
      "content": "<p>i am sorry that i made a mistake. 1 epoch=12283645 images takes about 13 hr, with  180x180 as input. So your timing may be reasonable. </p>\n\n<p>Also speed of resnext and resnet are not the same even if they have the same flops, <a href=\"https://github.com/Cadene/pretrained-models.pytorch/issues/8\">https://github.com/Cadene/pretrained-models.pytorch/issues/8</a></p>\n\n<p>I suspect it is due to group convolution in pytorch implementation ,\n<a href=\"https://discuss.pytorch.org/t/does-pytorch-optimize-the-group-parameter-in-convs/3060\">https://discuss.pytorch.org/t/does-pytorch-optimize-the-group-parameter-in-convs/3060</a></p>\n\n<p>What’s New in cuDNN 7?\nGrouped Convolutions for models such as ResNeXt and Xception and CTC (Connectionist Temporal Classification) loss layer for temporal classification</p>\n\n<p><a href=\"https://developer.nvidia.com/cudnn\">https://developer.nvidia.com/cudnn</a></p>",
      "rawMarkdown": "i am sorry that i made a mistake. 1 epoch=12283645 images takes about 13 hr, with  180x180 as input. So your timing may be reasonable. \n\nAlso speed of resnext and resnet are not the same even if they have the same flops, https://github.com/Cadene/pretrained-models.pytorch/issues/8\n\nI suspect it is due to group convolution in pytorch implementation ,\nhttps://discuss.pytorch.org/t/does-pytorch-optimize-the-group-parameter-in-convs/3060\n\nWhat’s New in cuDNN 7?\nGrouped Convolutions for models such as ResNeXt and Xception and CTC (Connectionist Temporal Classification) loss layer for temporal classification\n\nhttps://developer.nvidia.com/cudnn",
      "votes": null
    },
    {
      "id": "234047",
      "postDate": "10/22/2017 05:45:20",
      "content": "<p>I'm just wondering, how many layers are actually training in your 101 layers net? For me, training one more block of Xception takes me 2 more hours for one epoch. (single GTX 1080)</p>",
      "rawMarkdown": "I'm just wondering, how many layers are actually training in your 101 layers net? For me, training one more block of Xception takes me 2 more hours for one epoch. (single GTX 1080)",
      "votes": null
    },
    {
      "id": "234061",
      "postDate": "10/22/2017 07:06:31",
      "content": "<p>I use single ResNet 50 and my current LB-score is 0.45813, local validation score is 0.4914.</p>",
      "rawMarkdown": "I use single ResNet 50 and my current LB-score is 0.45813, local validation score is 0.4914.",
      "votes": null
    },
    {
      "id": "234093",
      "postDate": "10/22/2017 09:31:06",
      "content": "<p>SE-ResNet-50:</p>\n\n<ul>\n<li>finetuning about 7 epochs (about 16 hours per epoch on 2x1080)</li>\n<li>train augmentation: 161x161 random crops + horizontal flips</li>\n<li>val single crop accuracy: 70.96</li>\n<li>val 10 crops accuracy: 71.79</li>\n<li>LB 10 crops accuracy: 71.67</li>\n</ul>",
      "rawMarkdown": "SE-ResNet-50:\n\n - finetuning about 7 epochs (about 16 hours per epoch on 2x1080)\n - train augmentation: 161x161 random crops + horizontal flips\n - val single crop accuracy: 70.96\n - val 10 crops accuracy: 71.79\n - LB 10 crops accuracy: 71.67",
      "votes": null
    },
    {
      "id": "234099",
      "postDate": "10/22/2017 09:56:47",
      "content": "<p>thanks for the information! What is the learning rate you used for finetunning?</p>",
      "rawMarkdown": "thanks for the information! What is the learning rate you used for finetunning?",
      "votes": null
    },
    {
      "id": "234101",
      "postDate": "10/22/2017 10:10:07",
      "content": "<p>What's your learning rate schedule/optimize?</p>",
      "rawMarkdown": "What's your learning rate schedule/optimize?",
      "votes": null
    },
    {
      "id": "234102",
      "postDate": "10/22/2017 10:11:05",
      "content": "<p>I am using SGD with Nesterov momentum (0.9) and batch_size=256</p>\n\n<ul>\n<li>epochs 1..5: 0.001</li>\n<li>epoch 6: 0.0001</li>\n<li>epoch 7: 0.00001</li>\n</ul>",
      "rawMarkdown": "I am using SGD with Nesterov momentum (0.9) and batch_size=256\n\n - epochs 1..5: 0.001\n - epoch 6: 0.0001\n - epoch 7: 0.00001",
      "votes": null
    },
    {
      "id": "234106",
      "postDate": "10/22/2017 10:17:56",
      "content": "<p>Grouped Convs seem to be supported in the current master of Pytorch: <a href=\"https://github.com/pytorch/pytorch/pull/3057\">https://github.com/pytorch/pytorch/pull/3057</a></p>",
      "rawMarkdown": "Grouped Convs seem to be supported in the current master of Pytorch: https://github.com/pytorch/pytorch/pull/3057",
      "votes": null
    },
    {
      "id": "234111",
      "postDate": "10/22/2017 10:43:25",
      "content": "<p>i think this is spatial, where group = in_planes. It improve mobilenet but not Resnext, where  group=32 or 64. </p>",
      "rawMarkdown": "i think this is spatial, where group = in_planes. It improve mobilenet but not Resnext, where  group=32 or 64.",
      "votes": null
    },
    {
      "id": "234115",
      "postDate": "10/22/2017 10:49:21",
      "content": "<p>@Vladimir Iglovikov</p>\n\n<p>I rewrite my code an i get the following results:</p>\n\n<ul>\n<li><p>System: 4x1080Ti/pytorch</p></li>\n<li><p>input =160x160</p></li>\n<li><p>network = resnet101</p></li>\n<li><p>batch = 256</p></li>\n<li><p>time for 1000 iterations = 1000*256 images = 7 min.</p></li>\n</ul>\n\n<p>if one epoch = 12283645 images, i am getting time per epoch = 336 min = 5.6 hr.</p>\n\n<p>how many images are there in one epoch for your timing of 4.5hr? Thanks a lot!</p>\n\n<p>(I am trying to see if my system has been setup correctly. This is the first time i use multi-gpu training intensively)</p>",
      "rawMarkdown": "Vladimir Iglovikov\n \nI rewrite my code an i get the following results:\n\n -  System: 4x1080Ti/pytorch\n\n -  input =160x160\n\n -  network = resnet101\n\n -  batch = 256\n\n -  time for 1000 iterations = 1000*256 images = 7 min.\n\nif one epoch = 12283645 images, i am getting time per epoch = 336 min = 5.6 hr.\n\nhow many images are there in one epoch for your timing of 4.5hr? Thanks a lot!\n\n(I am trying to see if my system has been setup correctly. This is the first time i use multi-gpu training intensively)",
      "votes": null
    },
    {
      "id": "234137",
      "postDate": "10/22/2017 11:56:17",
      "content": "<p>one epoch 4.5 hours? why so fast?</p>",
      "rawMarkdown": "one epoch 4.5 hours? why so fast?",
      "votes": null
    },
    {
      "id": "234151",
      "postDate": "10/22/2017 12:49:19",
      "content": "<p>Resnet-18: val accuracy 0.63 trained for 6 epochs,  simple SGD, lr 0.1 -&gt; 0.01  -&gt; 0.001 , change lr after 2 epochs, no data augmentation no crop, 180x180, just change FC layer to get the 5270 classes and train</p>",
      "rawMarkdown": "Resnet-18: val accuracy 0.63 trained for 6 epochs,  simple SGD, lr 0.1 -&gt; 0.01  -&gt; 0.001 , change lr after 2 epochs, no data augmentation no crop, 180x180, just change FC layer to get the 5270 classes and train",
      "votes": null
    },
    {
      "id": "234175",
      "postDate": "10/22/2017 14:21:12",
      "content": "<p>One epoch - 11738624 Images. Rest are used for validation. </p>\n\n<p>Batch_size = 512</p>",
      "rawMarkdown": "One epoch - 11738624 Images. Rest are used for validation. \n\nBatch_size = 512",
      "votes": null
    },
    {
      "id": "234204",
      "postDate": "10/22/2017 16:45:04",
      "content": "<p>Under the same setting (batch=512), i am able to get training iteration of 40,526  images per minute. For an epoch of 11,738,624 it would take 282 min or 4.7 hr, which is close to your results. Thanks again!</p>",
      "rawMarkdown": "Under the same setting (batch=512), i am able to get training iteration of 40,526  images per minute. For an epoch of 11,738,624 it would take 282 min or 4.7 hr, which is close to your results. Thanks again!",
      "votes": null
    },
    {
      "id": "234294",
      "postDate": "10/22/2017 22:25:40",
      "content": "<p>Some useful info in this thread. Thanks all. I'm curious what people are settling on for training augmentation? I adapted some training code used in previous challenges last week and set some models training. I've got one model moving past the mid 60s in validation error and another that got stuck in the mid 50s. </p>\n\n<p>The model that got stuck was high capacity but I had colour/saturation augmentation enabled. The one that's still chugging along and doing much better only had random crop/scale and rotation. I'm guessing the prevalence of almost constant black or white backgrounds in most images means that adding variability to that requires much more learning than leaving it be.</p>\n\n<p>Next angle is to see if disabling the random crop and just leaving h-flip and a bit of rotation on does better... basically assuming that the preprocessing of the competition dataset leaves things relatively well centred and similarly scaled across train and test datasets. Has anyone hit close to .7 or better with almost no augmentation?</p>",
      "rawMarkdown": "Some useful info in this thread. Thanks all. I'm curious what people are settling on for training augmentation? I adapted some training code used in previous challenges last week and set some models training. I've got one model moving past the mid 60s in validation error and another that got stuck in the mid 50s. \n\nThe model that got stuck was high capacity but I had colour/saturation augmentation enabled. The one that's still chugging along and doing much better only had random crop/scale and rotation. I'm guessing the prevalence of almost constant black or white backgrounds in most images means that adding variability to that requires much more learning than leaving it be.\n\nNext angle is to see if disabling the random crop and just leaving h-flip and a bit of rotation on does better... basically assuming that the preprocessing of the competition dataset leaves things relatively well centred and similarly scaled across train and test datasets. Has anyone hit close to .7 or better with almost no augmentation?",
      "votes": null
    },
    {
      "id": "234309",
      "postDate": "10/22/2017 23:06:07",
      "content": "<p>Are you training all params in your net?</p>",
      "rawMarkdown": "Are you training all params in your net?",
      "votes": null
    },
    {
      "id": "234997",
      "postDate": "10/24/2017 14:54:34",
      "content": "<p>here is a possible Cdiscount in-house baseline results:\n<a href=\"https://www.linkedin.com/in/zhiwei-li-81b444126/\">https://www.linkedin.com/in/zhiwei-li-81b444126/</a></p>\n\n<p>Apr 2017 – Sep 2017  Employment Duration 6 mos</p>\n\n<p>LocationBordeaux Area, France</p>\n\n<p>Large Scale Image Classification for e-business product categorization</p>\n\n<p>• Research work on different Deep Learning architectures like Inception, ResNet etc.</p>\n\n<p>• Worked on 20 million images in roughly 9000 unbalanced class for product categories</p>\n\n<p>• Implemented transfer learning based on Inception V3 model using TensorFlow and Keras</p>\n\n<p>• Designed a quick model construction pipeline, raised ~60 times efficacy, achieved 72.3% accuracy</p>\n\n<p>• This model will be used to improve the categorization ability of current model</p>",
      "rawMarkdown": "here is a possible Cdiscount in-house baseline results:\nhttps://www.linkedin.com/in/zhiwei-li-81b444126/\n\nApr 2017 – Sep 2017  Employment Duration 6 mos\n\nLocationBordeaux Area, France\n\nLarge Scale Image Classification for e-business product categorization\n\n• Research work on different Deep Learning architectures like Inception, ResNet etc.\n\n• Worked on 20 million images in roughly 9000 unbalanced class for product categories\n\n• Implemented transfer learning based on Inception V3 model using TensorFlow and Keras\n\n• Designed a quick model construction pipeline, raised ~60 times efficacy, achieved 72.3% accuracy\n\n• This model will be used to improve the categorization ability of current model",
      "votes": null
    },
    {
      "id": "235034",
      "postDate": "10/24/2017 16:40:10",
      "content": "<blockquote>\n  <p>9000 unbalanced classes </p>\n</blockquote>\n\n<p>This probably makes the problem harder.</p>",
      "rawMarkdown": "&gt; 9000 unbalanced classes \n\nThis probably makes the problem harder.",
      "votes": null
    },
    {
      "id": "235081",
      "postDate": "10/24/2017 17:44:20",
      "content": "<p>&gt; This probably makes the problem harder.</p>\n\n<p>Or you just caught a resume 'embellishment' gone wrong.. <br>\nThat's linkedin.</p>",
      "rawMarkdown": "&gt; This probably makes the problem harder.\n\n\n\n\nOr you just caught a resume 'embellishment' gone wrong..  \nThat's linkedin.",
      "votes": null
    },
    {
      "id": "235885",
      "postDate": "10/26/2017 08:40:11",
      "content": "<p>when i test with renet v2 101 one epoch (train:9 val:1)  <br>\n : lr rate 0.01, <br>\n : aug flip  <br>\n : size 180*180  <br>\n :  rmsprop optimizer  <br>\n : init with tensorflow slim renet 101  <br>\n : no weight decay, no batch norm decay  <br>\n : batch size 48 <br></p>\n\n<p>train acc graph movi around 60%, but\nvaldation set give me acc 0.1%</p>\n\n<p>what makes my model so overfitting?\n - 5270 1*1 conv layer?\ni refenrece following model\n - <a href=\"https://github.com/tensorflow/models/tree/master/research/slim\">https://github.com/tensorflow/models/tree/master/research/slim</a></p>",
      "rawMarkdown": "when i test with renet v2 101 one epoch (train:9 val:1)  <br>\n : lr rate 0.01, <br>\n : aug flip  <br>\n : size 180*180  <br>\n :  rmsprop optimizer  <br>\n : init with tensorflow slim renet 101  <br>\n : no weight decay, no batch norm decay  <br>\n : batch size 48 <br>\n\ntrain acc graph movi around 60%, but\nvaldation set give me acc 0.1%\n\nwhat makes my model so overfitting?\n - 5270 1*1 conv layer?\ni refenrece following model\n - https://github.com/tensorflow/models/tree/master/research/slim",
      "votes": null
    },
    {
      "id": "235892",
      "postDate": "10/26/2017 08:48:24",
      "content": "<p>you can use sgd, rmsprop is not recommended</p>",
      "rawMarkdown": "you can use sgd, rmsprop is not recommended",
      "votes": null
    },
    {
      "id": "235896",
      "postDate": "10/26/2017 08:55:53",
      "content": "<p>why rmsprop is not recommended?</p>",
      "rawMarkdown": "why rmsprop is not recommended?",
      "votes": null
    },
    {
      "id": "235902",
      "postDate": "10/26/2017 09:08:27",
      "content": "<p>RMSProp should not make much of difference imho. I think you have bug in your code with that much of a difference in train/val!</p>",
      "rawMarkdown": "RMSProp should not make much of difference imho. I think you have bug in your code with that much of a difference in train/val!",
      "votes": null
    },
    {
      "id": "235941",
      "postDate": "10/26/2017 10:55:59",
      "content": "<p>@DaeYoungPark</p>\n\n<p>refer to: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/42069\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/42069</a></p>",
      "rawMarkdown": "DaeYoungPark\n \nrefer to: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/42069",
      "votes": null
    },
    {
      "id": "236060",
      "postDate": "10/26/2017 16:20:38",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    },
    {
      "id": "236164",
      "postDate": "10/26/2017 19:11:46",
      "content": "<p>Its a useful source and am new to kaggle hope to learn something new</p>",
      "rawMarkdown": "Its a useful source and am new to kaggle hope to learn something new",
      "votes": null
    },
    {
      "id": "236285",
      "postDate": "10/27/2017 01:27:06",
      "content": "<p>Thank you for tour refer!!</p>",
      "rawMarkdown": "Thank you for tour refer!!",
      "votes": null
    },
    {
      "id": "236443",
      "postDate": "10/27/2017 11:14:47",
      "content": "<p>Does 'val single crop accuracy: 70.96' mean that you got a validation accuracy of 70.96% after always training with the same 161x161 crop?</p>\n\n<p>Does 'val 10 crops accuracy: 71.79'  mean that the accuracy went up to 71.79 when you used 10 different crops per image?</p>\n\n<p>The conclusion being that augmentation helps.</p>\n\n<p>Sorry to ask. I am still getting used to the shorthand that people use.</p>\n\n<p>LB - is the score you got on the leader board, right?</p>\n\n<p>And during validation do you take a middle crop or resize from 180 to 161?</p>",
      "rawMarkdown": "Does 'val single crop accuracy: 70.96' mean that you got a validation accuracy of 70.96% after always training with the same 161x161 crop?\n\nDoes 'val 10 crops accuracy: 71.79'  mean that the accuracy went up to 71.79 when you used 10 different crops per image?\n\nThe conclusion being that augmentation helps.\n\nSorry to ask. I am still getting used to the shorthand that people use.\n\nLB - is the score you got on the leader board, right?\n\nAnd during validation do you take a middle crop or resize from 180 to 161?",
      "votes": null
    },
    {
      "id": "236444",
      "postDate": "10/27/2017 11:15:32",
      "content": "<p>You must have a bug. What are you doing for pre-processing?</p>",
      "rawMarkdown": "You must have a bug. What are you doing for pre-processing?",
      "votes": null
    },
    {
      "id": "237360",
      "postDate": "10/30/2017 06:21:15",
      "content": "<p>From discussions, it seems that the highest score we get is 0.70 with single model. Did you tried any other tricks?</p>\n\n<p>Did you make use the other two categories info? </p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "From discussions, it seems that the highest score we get is 0.70 with single model. Did you tried any other tricks?\n\nDid you make use the other two categories info? \n\nThanks!",
      "votes": null
    },
    {
      "id": "237844",
      "postDate": "10/31/2017 04:44:58",
      "content": "<p>How did you adjust your learning rate?\nThanks!</p>",
      "rawMarkdown": "How did you adjust your learning rate?\nThanks!",
      "votes": null
    },
    {
      "id": "238069",
      "postDate": "10/31/2017 15:25:41",
      "content": "<p>Wondering how do you modify resnet101 to accept smaller images. Seems diffrent people has diffrent strategies</p>",
      "rawMarkdown": "Wondering how do you modify resnet101 to accept smaller images. Seems diffrent people has diffrent strategies",
      "votes": null
    },
    {
      "id": "238242",
      "postDate": "10/31/2017 23:31:53",
      "content": "<p>Resnet50, 180x180 input, make use of the other two categories info, 6 epoch, no TTA:\nval_acc: 0.692\nLB: 0.705</p>",
      "rawMarkdown": "Resnet50, 180x180 input, make use of the other two categories info, 6 epoch, no TTA:\nval_acc: 0.692\nLB: 0.705",
      "votes": null
    },
    {
      "id": "238325",
      "postDate": "11/01/2017 06:46:30",
      "content": "<p>What's the meaning of the other two categories info? Did you train multi classifier?</p>",
      "rawMarkdown": "What's the meaning of the other two categories info? Did you train multi classifier?",
      "votes": null
    },
    {
      "id": "238623",
      "postDate": "11/01/2017 19:05:37",
      "content": "<p>In addition to the categories we are interested in, I also train the model to predict the categories of the father levels, like animal - mammal - cat, where cat is our target.</p>",
      "rawMarkdown": "In addition to the categories we are interested in, I also train the model to predict the categories of the father levels, like animal - mammal - cat, where cat is our target.",
      "votes": null
    },
    {
      "id": "238775",
      "postDate": "11/02/2017 06:10:18",
      "content": "<p>Are you using the father levels as auxiliary classifier?</p>",
      "rawMarkdown": "Are you using the father levels as auxiliary classifier?",
      "votes": null
    },
    {
      "id": "238776",
      "postDate": "11/02/2017 06:12:05",
      "content": "<p>regarding hierarchical categorization, one can consider the yolo9000 approach (wordTree): <a href=\"https://arxiv.org/abs/1612.08242\">https://arxiv.org/abs/1612.08242</a> </p>",
      "rawMarkdown": "regarding hierarchical categorization, one can consider the yolo9000 approach (wordTree): https://arxiv.org/abs/1612.08242",
      "votes": null
    },
    {
      "id": "242245",
      "postDate": "11/10/2017 22:30:23",
      "content": "<p>Xception, 180x180 \nval acc: 0.7277\naverage images for product LB: 0.73785.    </p>",
      "rawMarkdown": "Xception, 180x180 \nval acc: 0.7277\naverage images for product LB: 0.73785.",
      "votes": null
    },
    {
      "id": "242271",
      "postDate": "11/11/2017 01:28:21",
      "content": "<p>average images for product，you predict the prob of every image and average them?</p>",
      "rawMarkdown": "average images for product，you predict the prob of every image and average them?",
      "votes": null
    },
    {
      "id": "242282",
      "postDate": "11/11/2017 02:42:55",
      "content": "<p>Yes, for test data, for each product I predict the prob of every image and average them.  </p>",
      "rawMarkdown": "Yes, for test data, for each product I predict the prob of every image and average them.",
      "votes": null
    },
    {
      "id": "242303",
      "postDate": "11/11/2017 05:24:16",
      "content": "<p>I'm also using Xception now. But I only got LB: 0.68 for about 15 epoch. Do you train all parameters in the model? Which augmentation strategy do you use? Thanks a lot!</p>",
      "rawMarkdown": "I'm also using Xception now. But I only got LB: 0.68 for about 15 epoch. Do you train all parameters in the model? Which augmentation strategy do you use? Thanks a lot!",
      "votes": null
    },
    {
      "id": "242306",
      "postDate": "11/11/2017 05:32:00",
      "content": "<p>Do you use pretrained model?</p>",
      "rawMarkdown": "Do you use pretrained model?",
      "votes": null
    },
    {
      "id": "242335",
      "postDate": "11/11/2017 09:26:11",
      "content": "<p>i'm still using a single resnet-18 because of very limited resources. 160x160 random crop in training, from scratch, lr 0.1 -&gt; 0.01 -&gt; 0.001 -&gt; 0.0005 -&gt; 0.0001, changing lr every few epochs, SGD with momentum 0.9, 192 images batches. I have modified the network in the last few conv layers. LB 0.69. inference with only 1 center 160x160 crop.</p>",
      "rawMarkdown": "i'm still using a single resnet-18 because of very limited resources. 160x160 random crop in training, from scratch, lr 0.1 -&gt; 0.01 -&gt; 0.001 -&gt; 0.0005 -&gt; 0.0001, changing lr every few epochs, SGD with momentum 0.9, 192 images batches. I have modified the network in the last few conv layers. LB 0.69. inference with only 1 center 160x160 crop.",
      "votes": null
    },
    {
      "id": "242363",
      "postDate": "11/11/2017 10:45:11",
      "content": "<p>How did you modify the conv layers? Seems like a pretty good score for resnet-18!\nAny data augmentation? We are still looking for the right kind...</p>",
      "rawMarkdown": "How did you modify the conv layers? Seems like a pretty good score for resnet-18!\nAny data augmentation? We are still looking for the right kind...",
      "votes": null
    },
    {
      "id": "242370",
      "postDate": "11/11/2017 11:19:51",
      "content": "<p>it seems to me that there's a kind of flaw in vanilla resnet-18 in the final few stages, if the final classifier has to distinguish among 5270 classes. The final FC layer is just a matrix, it basically can do 2 things: rotate a vector and project or embed the vector in the target space. Embedding in higher dimensional space does not change the dimension of the initial subspace, it cannot give higher resolution. Resnet 18 has a average pooling before the FC, there is too much of a bottlenck before FC. So vanilla resnet 18 cannot get good class resolution beyond a 10-20% of the full set.  Unfortunately time is running and i have very old metal.... i'm working on a completely different back-end for my little resnet 18  but not sure if i can make it in time. </p>",
      "rawMarkdown": "it seems to me that there's a kind of flaw in vanilla resnet-18 in the final few stages, if the final classifier has to distinguish among 5270 classes. The final FC layer is just a matrix, it basically can do 2 things: rotate a vector and project or embed the vector in the target space. Embedding in higher dimensional space does not change the dimension of the initial subspace, it cannot give higher resolution. Resnet 18 has a average pooling before the FC, there is too much of a bottlenck before FC. So vanilla resnet 18 cannot get good class resolution beyond a 10-20% of the full set.  Unfortunately time is running and i have very old metal.... i'm working on a completely different back-end for my little resnet 18  but not sure if i can make it in time.",
      "votes": null
    },
    {
      "id": "242525",
      "postDate": "11/11/2017 19:56:11",
      "content": "<p>@CSAdu: Yes I use Keras pre-trained model.</p>",
      "rawMarkdown": "CSAdu: Yes I use Keras pre-trained model.",
      "votes": null
    },
    {
      "id": "242526",
      "postDate": "11/11/2017 19:58:09",
      "content": "<p>@Brian Luo: Yes I train all parameters. I use small augmentation to get 0.68 then continue training without augmentation.</p>",
      "rawMarkdown": "Brian Luo: Yes I train all parameters. I use small augmentation to get 0.68 then continue training without augmentation.",
      "votes": null
    },
    {
      "id": "242761",
      "postDate": "11/12/2017 14:56:23",
      "content": "<p>So do you use pretrained model? Where you find your pretrained model. I start to train the SE-Resnet50 but seems after 1 epoch I only got 0.50, it this validation accuracy close to your data after 1 epoch?</p>",
      "rawMarkdown": "So do you use pretrained model? Where you find your pretrained model. I start to train the SE-Resnet50 but seems after 1 epoch I only got 0.50, it this validation accuracy close to your data after 1 epoch?",
      "votes": null
    },
    {
      "id": "242964",
      "postDate": "11/13/2017 03:50:10",
      "content": "<p>Thanks for your reply. How many epoch do you use to get 0.73?</p>",
      "rawMarkdown": "Thanks for your reply. How many epoch do you use to get 0.73?",
      "votes": null
    },
    {
      "id": "243159",
      "postDate": "11/13/2017 13:30:15",
      "content": "<p>It takes me about 60 epochs to get 0.68 and another 100 epochs to get 0.73, each epoch is 1/10 train data.</p>",
      "rawMarkdown": "It takes me about 60 epochs to get 0.68 and another 100 epochs to get 0.73, each epoch is 1/10 train data.",
      "votes": null
    },
    {
      "id": "245805",
      "postDate": "11/19/2017 19:16:05",
      "content": "<p>Are you using the wordTree approach? I'm not sure how to apply this to our task. Do you know how to predict conditional probabilities for each intermediate node?</p>",
      "rawMarkdown": "Are you using the wordTree approach? I'm not sure how to apply this to our task. Do you know how to predict conditional probabilities for each intermediate node?",
      "votes": null
    },
    {
      "id": "245946",
      "postDate": "11/20/2017 06:28:59",
      "content": "<p>My resnet50 only got to about 63% accuracy before it started massively overfitting</p>\n\n<p>EDIT: went back to an earlier epoch and changed the data augmentation, so it got to 65% (with room for improvement) after a couple of epochs with more overfitting than some of my other models, but not so much as before where the validation accuracy was actually decreasing</p>",
      "rawMarkdown": "My resnet50 only got to about 63% accuracy before it started massively overfitting\n\nEDIT: went back to an earlier epoch and changed the data augmentation, so it got to 65% (with room for improvement) after a couple of epochs with more overfitting than some of my other models, but not so much as before where the validation accuracy was actually decreasing",
      "votes": null
    },
    {
      "id": "246516",
      "postDate": "11/21/2017 08:19:04",
      "content": "<p>Did you use a globalaveragePooling layer just before the last FC layers (which gives 2048--&gt;5270)?\nOr some other poolling layers  which gives X&gt;2048 ---&gt;5270 at the top?\nThanks</p>",
      "rawMarkdown": "Did you use a globalaveragePooling layer just before the last FC layers (which gives 2048--&gt;5270)?\nOr some other poolling layers  which gives X&gt;2048 ---&gt;5270 at the top?\nThanks",
      "votes": null
    },
    {
      "id": "246544",
      "postDate": "11/21/2017 09:07:49",
      "content": "<p>Just GlobalAveragePooling.</p>",
      "rawMarkdown": "Just GlobalAveragePooling.",
      "votes": null
    },
    {
      "id": "247083",
      "postDate": "11/22/2017 09:25:34",
      "content": "<p>inception resnet v2, tune keras pretrained model. Train top layers for 2 epoch, then train all parameters for 12 epoch. \n180x180, single image val acc: 0.7305,  LB: 0.74564(average prediction for all images of a product.)</p>",
      "rawMarkdown": "inception resnet v2, tune keras pretrained model. Train top layers for 2 epoch, then train all parameters for 12 epoch. \n180x180, single image val acc: 0.7305,  LB: 0.74564(average prediction for all images of a product.)",
      "votes": null
    },
    {
      "id": "247212",
      "postDate": "11/22/2017 15:45:43",
      "content": "<p>How long you need to train one epoch for inception resnet v2?</p>",
      "rawMarkdown": "How long you need to train one epoch for inception resnet v2?",
      "votes": null
    },
    {
      "id": "247386",
      "postDate": "11/22/2017 23:20:26",
      "content": "<p>It takes me about 9 hours to train all parameters for each epoch with AMD 1950X, 4 gtx 1080ti and Samsung 960 pro. My GPUs usage is not 100%, I think there is still space to improve speed of I/O of my code.</p>",
      "rawMarkdown": "It takes me about 9 hours to train all parameters for each epoch with AMD 1950X, 4 gtx 1080ti and Samsung 960 pro. My GPUs usage is not 100%, I think there is still space to improve speed of I/O of my code.",
      "votes": null
    },
    {
      "id": "248763",
      "postDate": "11/26/2017 22:49:09",
      "content": "<p>Wondering what is your best single model and its accuracy?</p>",
      "rawMarkdown": "Wondering what is your best single model and its accuracy?",
      "votes": null
    },
    {
      "id": "248797",
      "postDate": "11/27/2017 01:08:15",
      "content": "<p>My best model gets about 74.8% on the leaderboard by itself. I only have two or three real models trained and ensembling currently isn't getting me very big gains.</p>",
      "rawMarkdown": "My best model gets about 74.8% on the leaderboard by itself. I only have two or three real models trained and ensembling currently isn't getting me very big gains.",
      "votes": null
    },
    {
      "id": "249023",
      "postDate": "11/27/2017 15:00:58",
      "content": "<p>Thanks for your share. I have some questions:</p>\n\n<ol>\n<li>Do you train your model like it is described in here? <a href=\"https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\">https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights</a></li>\n<li>Sorry, I am new at image processing. What does 180x180 mean? that the images were reshaped to 180x180 dimensions?</li>\n</ol>",
      "rawMarkdown": "Thanks for your share. I have some questions:\n\n 1. Do you train your model like it is described in here? https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\n 2. Sorry, I am new at image processing. What does 180x180 mean? that the images were reshaped to 180x180 dimensions?",
      "votes": null
    },
    {
      "id": "249027",
      "postDate": "11/27/2017 15:08:40",
      "content": "<p>How did you guys perform the train/val split?</p>",
      "rawMarkdown": "How did you guys perform the train/val split?",
      "votes": null
    },
    {
      "id": "249228",
      "postDate": "11/28/2017 01:21:34",
      "content": "<p>You are welcome.</p>\n\n<ol>\n<li>I use this data generator <a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson</a>, and training is similar to <a href=\"https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\">https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights</a>  </li>\n<li>180x180 is the original size of the images. I didn't reshape them.</li>\n</ol>",
      "rawMarkdown": "You are welcome.\n\n1. I use this data generator https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson, and training is similar to https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights  \n2. 180x180 is the original size of the images. I didn't reshape them.",
      "votes": null
    },
    {
      "id": "249377",
      "postDate": "11/28/2017 09:52:34",
      "content": "<p>thanks for the details; do you have any tips / ideas why it \"doesn't work\" for me? I tried \"Inception_v3\" from Keras, after one epoch (about 80% of Train data) I was getting only 20% accuracy (on Validation only 13%) !!!</p>\n\n<p>I use my own Generator, 512 Batch size, random shuffle of the data .... ,aybe I need to try with the Generator from the link</p>",
      "rawMarkdown": "thanks for the details; do you have any tips / ideas why it \"doesn't work\" for me? I tried \"Inception_v3\" from Keras, after one epoch (about 80% of Train data) I was getting only 20% accuracy (on Validation only 13%) !!!\n\nI use my own Generator, 512 Batch size, random shuffle of the data .... ,aybe I need to try with the Generator from the link",
      "votes": null
    },
    {
      "id": "249405",
      "postDate": "11/28/2017 11:41:04",
      "content": "<p>My guess: </p>\n\n<ol>\n<li><p>Did you use the keras application build in preprocess function? \nfrom keras.applications.inception_v3 import preprocess_input,</p></li>\n<li><p>What optimizer did you use, can you try 'adam' ?</p></li>\n</ol>",
      "rawMarkdown": "My guess: \n\n1.  Did you use the keras application build in preprocess function? \nfrom keras.applications.inception_v3 import preprocess_input,\n\n2. What optimizer did you use, can you try 'adam' ?",
      "votes": null
    },
    {
      "id": "249406",
      "postDate": "11/28/2017 11:44:35",
      "content": "<p>I didn't use the function itself (think it was throwing an error?) but did:\nimg = img / 255.\nimg = img - 0.5\n(I think there was one more step .. basically I copied the 'tf' mode)</p>\n\n<p>I used rmsprop, will try with adam in the evening ...</p>\n\n<p>btw. I used custom Input shape (180,180,3) - could that also be the problem?</p>",
      "rawMarkdown": "I didn't use the function itself (think it was throwing an error?) but did:\nimg = img / 255.\nimg = img - 0.5\n(I think there was one more step .. basically I copied the 'tf' mode)\n\nI used rmsprop, will try with adam in the evening ...\n\nbtw. I used custom Input shape (180,180,3) - could that also be the problem?",
      "votes": null
    },
    {
      "id": "249679",
      "postDate": "11/29/2017 01:14:17",
      "content": "<p>Usually, my val_acc is higher than acc when acc is lower than 60%.  I was using (img/255-0.5)*2.0 and can get 40% with 10% data(with inception resnet v2 and xception). Then I switch to the keras preprocess and can get 50% with 10% data(with resnet101 or resnet 151).  I also use (180,180,3).  What is your batch size? </p>",
      "rawMarkdown": "Usually, my val_acc is higher than acc when acc is lower than 60%.  I was using (img/255-0.5)*2.0 and can get 40% with 10% data(with inception resnet v2 and xception). Then I switch to the keras preprocess and can get 50% with 10% data(with resnet101 or resnet 151).  I also use (180,180,3).  What is your batch size?",
      "votes": null
    },
    {
      "id": "249765",
      "postDate": "11/29/2017 07:05:49",
      "content": "<p>thank you, with some corrections (guess the preprocessing was wrong) I got 50% after around 1.5 epoch (around 12M Images), an improvement but still no so great</p>\n\n<p>I'm using batch size = 512, this time it was with adam optimizer (all default settings), will test with Xception or the pretrained weights from Heng CherKeng</p>",
      "rawMarkdown": "thank you, with some corrections (guess the preprocessing was wrong) I got 50% after around 1.5 epoch (around 12M Images), an improvement but still no so great\n\nI'm using batch size = 512, this time it was with adam optimizer (all default settings), will test with Xception or the pretrained weights from Heng CherKeng",
      "votes": null
    },
    {
      "id": "249807",
      "postDate": "11/29/2017 08:28:49",
      "content": "<p>Hi @iFighting, @minu, @Heng CherKeng,\nIt seems like you guys are using MXNet?\nWhat are you impressions comparing to PyTorch?\nDoes MXNet facilitates specific aspects? Were you able to attain better results training with MXNet?</p>",
      "rawMarkdown": "Hi @iFighting, @minu, @Heng CherKeng,\nIt seems like you guys are using MXNet?\nWhat are you impressions comparing to PyTorch?\nDoes MXNet facilitates specific aspects? Were you able to attain better results training with MXNet?",
      "votes": null
    },
    {
      "id": "249816",
      "postDate": "11/29/2017 08:50:38",
      "content": "<p>i am using pytorch only for this competition.  I use caffe, mxnet, tensorflow before. I think their results are all about the same (because they use the \"same formula\" and based on nvidia cudnn). Some DL framework are faster. But overall, I think pytorch is fast and easy to debug and experiment with new structure.</p>",
      "rawMarkdown": "i am using pytorch only for this competition.  I use caffe, mxnet, tensorflow before. I think their results are all about the same (because they use the \"same formula\" and based on nvidia cudnn). Some DL framework are faster. But overall, I think pytorch is fast and easy to debug and experiment with new structure.",
      "votes": null
    },
    {
      "id": "250062",
      "postDate": "11/29/2017 17:43:48",
      "content": "<p>what is the 'tf' mode? Can you guys point me where to read more about retraining and transfer learning with Xception (or Inception, resnet, etc)? :)</p>",
      "rawMarkdown": "what is the 'tf' mode? Can you guys point me where to read more about retraining and transfer learning with Xception (or Inception, resnet, etc)? :)",
      "votes": null
    },
    {
      "id": "250103",
      "postDate": "11/29/2017 19:49:01",
      "content": "<p>I assume 'tf' stands for tensorflow, but might be wrong :)\nwhen you check the code in <a href=\"https://github.com/fchollet/keras/blob/master/keras/applications/imagenet_utils.py\">https://github.com/fchollet/keras/blob/master/keras/applications/imagenet_utils.py</a> you will see, but basically it's (since 2 days ago, before it was the transformation mentioned by Enhao above): <br>\nx /= 127.5\nx -= 1.</p>\n\n<p>for a tutorial you can check out this keras blog: <a href=\"https://blog.keras.io/building-powerful-image-classification-models-using-very-little-data.html\">https://blog.keras.io/building-powerful-image-classification-models-using-very-little-data.html</a></p>",
      "rawMarkdown": "I assume 'tf' stands for tensorflow, but might be wrong :)\nwhen you check the code in https://github.com/fchollet/keras/blob/master/keras/applications/imagenet_utils.py you will see, but basically it's (since 2 days ago, before it was the transformation mentioned by Enhao above):  \nx /= 127.5\nx -= 1.\n\nfor a tutorial you can check out this keras blog: https://blog.keras.io/building-powerful-image-classification-models-using-very-little-data.html",
      "votes": null
    },
    {
      "id": "250479",
      "postDate": "11/30/2017 03:38:05",
      "content": "<p>Update: Resnet50 gets  0.71450 on the leaderboard after 10 epochs</p>",
      "rawMarkdown": "Update: Resnet50 gets  0.71450 on the leaderboard after 10 epochs",
      "votes": null
    },
    {
      "id": "250497",
      "postDate": "11/30/2017 04:03:18",
      "content": "<p>Just wondering what you have done to boost from 65% to 71%?</p>",
      "rawMarkdown": "Just wondering what you have done to boost from 65% to 71%?",
      "votes": null
    },
    {
      "id": "250520",
      "postDate": "11/30/2017 04:34:56",
      "content": "<ul>\n<li>65% was my validation accuracy at two epochs, while 0.7134 was my validation accuracy at ten epochs</li>\n<li>for the last epoch I trained without data augmentation to put the training and testing conditions in harmony</li>\n<li>leaderboard accuracy is a bit higher since it involves more than one image for each product being predicted</li>\n</ul>",
      "rawMarkdown": "* 65% was my validation accuracy at two epochs, while 0.7134 was my validation accuracy at ten epochs\n* for the last epoch I trained without data augmentation to put the training and testing conditions in harmony\n* leaderboard accuracy is a bit higher since it involves more than one image for each product being predicted",
      "votes": null
    },
    {
      "id": "250545",
      "postDate": "11/30/2017 05:07:04",
      "content": "<p>Thanks! And 65% at 2 epoch is quite high I believed. Would you mind if you can share how do you achieve this? My resnet 101 only have about 58% after two epoch with 0.001 lr</p>",
      "rawMarkdown": "Thanks! And 65% at 2 epoch is quite high I believed. Would you mind if you can share how do you achieve this? My resnet 101 only have about 58% after two epoch with 0.001 lr",
      "votes": null
    },
    {
      "id": "250607",
      "postDate": "11/30/2017 06:20:52",
      "content": "<p>yes, i use mxnet</p>",
      "rawMarkdown": "yes, i use mxnet",
      "votes": null
    },
    {
      "id": "250633",
      "postDate": "11/30/2017 06:40:21",
      "content": "<p>I'm just familiar with MXNet. and MXNet is efficient in Multi-GPU.</p>",
      "rawMarkdown": "I'm just familiar with MXNet. and MXNet is efficient in Multi-GPU.",
      "votes": null
    },
    {
      "id": "252273",
      "postDate": "12/02/2017 16:49:51",
      "content": "<p>What's different between (img/255-0.5)*2.0 and keras preprocess (x /= 127.5 x -= 1)? I think they're the same thing, why you get higher result with keras preprocess?</p>",
      "rawMarkdown": "What's different between (img/255-0.5)*2.0 and keras preprocess (x /= 127.5 x -= 1)? I think they're the same thing, why you get higher result with keras preprocess?",
      "votes": null
    },
    {
      "id": "252279",
      "postDate": "12/02/2017 16:57:55",
      "content": "<p>yes, most likely the main difference was coming from the adam optimizer (I did both changes at once)</p>",
      "rawMarkdown": "yes, most likely the main difference was coming from the adam optimizer (I did both changes at once)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 233991,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/22/2017 01:25:43",
      "content": "<p>thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 233996,
      "author_name": "jandjenter",
      "author_url": "",
      "post_date": "10/22/2017 01:48:17",
      "content": "<p>That's quite high score for single model! How long did you train for this and which GPU did you use?</p>\n\n<p>Due to limited resources, I'm using images scaled down to 90x90 and simple Xception, no TTA</p>\n\n<ul>\n<li><p>val_acc = .6904 (for each image) </p></li>\n<li><p>LB= .7092 (averaged softmax output for all images in same product id)</p></li>\n</ul>\n\n<p>I'm trying many other approaches like using top 2000 frequent items, using category 1,2 information, etc, but nothing beats this baseline yet. Sad to see much better results with basic model, but thanks for the info.</p>",
      "votes": null,
      "replies": [
        {
          "id": 233997,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "10/22/2017 01:51:12",
          "content": "<p>4 x GTX 1080 Ti</p>\n\n<p>18 epochs x 4.5 hours</p>\n\n<p>I have 0.69 after 7th epoch</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234000,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/22/2017 01:58:43",
          "content": "<p>seems that we need long epoches!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234006,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/22/2017 02:41:18",
          "content": "<p>i am also using 4 x GTX 1080 Ti. For resnext101 (about same flops as your resnet101 ),  1 epoch=12283645 images takes about 1.3 hr. I am using 180x180 as input. I suppose your input size is not 180x180?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234007,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "10/22/2017 02:43:55",
          "content": "<p>what does TTA mean?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234010,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "10/22/2017 02:52:06",
          "content": "<p>I use 160x160 crops.</p>\n\n<p>1.3hrs per epoch.... It looks like I am doing something wrong.</p>\n\n<p>It may happen that my SSD is not fast enough and I need to buy faster one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234012,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/22/2017 02:53:57",
          "content": "<p>TTA= test time augmentation (e.g. multi crops, resize in test time)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234013,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "10/22/2017 02:56:01",
          "content": "<p>Test-time augmentation. In my case, I can get extra 0.01 just with 12-way TTA(flip, shift up/down, etc) with exact same weight - almost free cake :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234032,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/22/2017 04:10:23",
          "content": "<p>i am sorry that i made a mistake. 1 epoch=12283645 images takes about 13 hr, with  180x180 as input. So your timing may be reasonable. </p>\n\n<p>Also speed of resnext and resnet are not the same even if they have the same flops, <a href=\"https://github.com/Cadene/pretrained-models.pytorch/issues/8\">https://github.com/Cadene/pretrained-models.pytorch/issues/8</a></p>\n\n<p>I suspect it is due to group convolution in pytorch implementation ,\n<a href=\"https://discuss.pytorch.org/t/does-pytorch-optimize-the-group-parameter-in-convs/3060\">https://discuss.pytorch.org/t/does-pytorch-optimize-the-group-parameter-in-convs/3060</a></p>\n\n<p>What’s New in cuDNN 7?\nGrouped Convolutions for models such as ResNeXt and Xception and CTC (Connectionist Temporal Classification) loss layer for temporal classification</p>\n\n<p><a href=\"https://developer.nvidia.com/cudnn\">https://developer.nvidia.com/cudnn</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234106,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/22/2017 10:17:56",
          "content": "<p>Grouped Convs seem to be supported in the current master of Pytorch: <a href=\"https://github.com/pytorch/pytorch/pull/3057\">https://github.com/pytorch/pytorch/pull/3057</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234111,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/22/2017 10:43:25",
          "content": "<p>i think this is spatial, where group = in_planes. It improve mobilenet but not Resnext, where  group=32 or 64. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234115,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/22/2017 10:49:21",
          "content": "<p>@Vladimir Iglovikov</p>\n\n<p>I rewrite my code an i get the following results:</p>\n\n<ul>\n<li><p>System: 4x1080Ti/pytorch</p></li>\n<li><p>input =160x160</p></li>\n<li><p>network = resnet101</p></li>\n<li><p>batch = 256</p></li>\n<li><p>time for 1000 iterations = 1000*256 images = 7 min.</p></li>\n</ul>\n\n<p>if one epoch = 12283645 images, i am getting time per epoch = 336 min = 5.6 hr.</p>\n\n<p>how many images are there in one epoch for your timing of 4.5hr? Thanks a lot!</p>\n\n<p>(I am trying to see if my system has been setup correctly. This is the first time i use multi-gpu training intensively)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234137,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "10/22/2017 11:56:17",
          "content": "<p>one epoch 4.5 hours? why so fast?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234175,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "10/22/2017 14:21:12",
          "content": "<p>One epoch - 11738624 Images. Rest are used for validation. </p>\n\n<p>Batch_size = 512</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234204,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/22/2017 16:45:04",
          "content": "<p>Under the same setting (batch=512), i am able to get training iteration of 40,526  images per minute. For an epoch of 11,738,624 it would take 282 min or 4.7 hr, which is close to your results. Thanks again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234309,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/22/2017 23:06:07",
          "content": "<p>Are you training all params in your net?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 234005,
      "author_name": "lyakaap",
      "author_url": "",
      "post_date": "10/22/2017 02:40:41",
      "content": "<p>Amazing score. Did you do training without pretrained weight?</p>\n\n<p>My results: \ninception_v3, 160 crop, val loss: 1.78~, val acc: 0.66~, lb: 0.673~</p>",
      "votes": null,
      "replies": [
        {
          "id": 234019,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "10/22/2017 03:31:19",
          "content": "<p>I used pre-trained. I would be also interested how long will it take to get the same result without pre-trained, but it may be too computationally expensive.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234026,
          "author_name": "heroxrq",
          "author_url": "",
          "post_date": "10/22/2017 03:51:23",
          "content": "<p>I used resnext-101, with image centor croped to 128X128, used pre-trained. It can get 0.59 val _acc after 1 epoches. I used sgd optimizer. Training is going...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 234047,
      "author_name": "brianlzm",
      "author_url": "",
      "post_date": "10/22/2017 05:45:20",
      "content": "<p>I'm just wondering, how many layers are actually training in your 101 layers net? For me, training one more block of Xception takes me 2 more hours for one epoch. (single GTX 1080)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 234061,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "10/22/2017 07:06:31",
      "content": "<p>I use single ResNet 50 and my current LB-score is 0.45813, local validation score is 0.4914.</p>",
      "votes": null,
      "replies": [
        {
          "id": 236444,
          "author_name": "jinkos",
          "author_url": "",
          "post_date": "10/27/2017 11:15:32",
          "content": "<p>You must have a bug. What are you doing for pre-processing?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 234093,
      "author_name": "davletag",
      "author_url": "",
      "post_date": "10/22/2017 09:31:06",
      "content": "<p>SE-ResNet-50:</p>\n\n<ul>\n<li>finetuning about 7 epochs (about 16 hours per epoch on 2x1080)</li>\n<li>train augmentation: 161x161 random crops + horizontal flips</li>\n<li>val single crop accuracy: 70.96</li>\n<li>val 10 crops accuracy: 71.79</li>\n<li>LB 10 crops accuracy: 71.67</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 234099,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/22/2017 09:56:47",
          "content": "<p>thanks for the information! What is the learning rate you used for finetunning?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234102,
          "author_name": "davletag",
          "author_url": "",
          "post_date": "10/22/2017 10:11:05",
          "content": "<p>I am using SGD with Nesterov momentum (0.9) and batch_size=256</p>\n\n<ul>\n<li>epochs 1..5: 0.001</li>\n<li>epoch 6: 0.0001</li>\n<li>epoch 7: 0.00001</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 236443,
          "author_name": "jinkos",
          "author_url": "",
          "post_date": "10/27/2017 11:14:47",
          "content": "<p>Does 'val single crop accuracy: 70.96' mean that you got a validation accuracy of 70.96% after always training with the same 161x161 crop?</p>\n\n<p>Does 'val 10 crops accuracy: 71.79'  mean that the accuracy went up to 71.79 when you used 10 different crops per image?</p>\n\n<p>The conclusion being that augmentation helps.</p>\n\n<p>Sorry to ask. I am still getting used to the shorthand that people use.</p>\n\n<p>LB - is the score you got on the leader board, right?</p>\n\n<p>And during validation do you take a middle crop or resize from 180 to 161?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242761,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "11/12/2017 14:56:23",
          "content": "<p>So do you use pretrained model? Where you find your pretrained model. I start to train the SE-Resnet50 but seems after 1 epoch I only got 0.50, it this validation accuracy close to your data after 1 epoch?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 234101,
      "author_name": "timjoseph",
      "author_url": "",
      "post_date": "10/22/2017 10:10:07",
      "content": "<p>What's your learning rate schedule/optimize?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 234151,
      "author_name": "andreanobile",
      "author_url": "",
      "post_date": "10/22/2017 12:49:19",
      "content": "<p>Resnet-18: val accuracy 0.63 trained for 6 epochs,  simple SGD, lr 0.1 -&gt; 0.01  -&gt; 0.001 , change lr after 2 epochs, no data augmentation no crop, 180x180, just change FC layer to get the 5270 classes and train</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 234294,
      "author_name": "rwightman",
      "author_url": "",
      "post_date": "10/22/2017 22:25:40",
      "content": "<p>Some useful info in this thread. Thanks all. I'm curious what people are settling on for training augmentation? I adapted some training code used in previous challenges last week and set some models training. I've got one model moving past the mid 60s in validation error and another that got stuck in the mid 50s. </p>\n\n<p>The model that got stuck was high capacity but I had colour/saturation augmentation enabled. The one that's still chugging along and doing much better only had random crop/scale and rotation. I'm guessing the prevalence of almost constant black or white backgrounds in most images means that adding variability to that requires much more learning than leaving it be.</p>\n\n<p>Next angle is to see if disabling the random crop and just leaving h-flip and a bit of rotation on does better... basically assuming that the preprocessing of the competition dataset leaves things relatively well centred and similarly scaled across train and test datasets. Has anyone hit close to .7 or better with almost no augmentation?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 234997,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/24/2017 14:54:34",
      "content": "<p>here is a possible Cdiscount in-house baseline results:\n<a href=\"https://www.linkedin.com/in/zhiwei-li-81b444126/\">https://www.linkedin.com/in/zhiwei-li-81b444126/</a></p>\n\n<p>Apr 2017 – Sep 2017  Employment Duration 6 mos</p>\n\n<p>LocationBordeaux Area, France</p>\n\n<p>Large Scale Image Classification for e-business product categorization</p>\n\n<p>• Research work on different Deep Learning architectures like Inception, ResNet etc.</p>\n\n<p>• Worked on 20 million images in roughly 9000 unbalanced class for product categories</p>\n\n<p>• Implemented transfer learning based on Inception V3 model using TensorFlow and Keras</p>\n\n<p>• Designed a quick model construction pipeline, raised ~60 times efficacy, achieved 72.3% accuracy</p>\n\n<p>• This model will be used to improve the categorization ability of current model</p>",
      "votes": null,
      "replies": [
        {
          "id": 235034,
          "author_name": "andreaslup",
          "author_url": "",
          "post_date": "10/24/2017 16:40:10",
          "content": "<blockquote>\n  <p>9000 unbalanced classes </p>\n</blockquote>\n\n<p>This probably makes the problem harder.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 235081,
          "author_name": "fernandoconstantino",
          "author_url": "",
          "post_date": "10/24/2017 17:44:20",
          "content": "<p>&gt; This probably makes the problem harder.</p>\n\n<p>Or you just caught a resume 'embellishment' gone wrong.. <br>\nThat's linkedin.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 235885,
      "author_name": "dyaprk",
      "author_url": "",
      "post_date": "10/26/2017 08:40:11",
      "content": "<p>when i test with renet v2 101 one epoch (train:9 val:1)  <br>\n : lr rate 0.01, <br>\n : aug flip  <br>\n : size 180*180  <br>\n :  rmsprop optimizer  <br>\n : init with tensorflow slim renet 101  <br>\n : no weight decay, no batch norm decay  <br>\n : batch size 48 <br></p>\n\n<p>train acc graph movi around 60%, but\nvaldation set give me acc 0.1%</p>\n\n<p>what makes my model so overfitting?\n - 5270 1*1 conv layer?\ni refenrece following model\n - <a href=\"https://github.com/tensorflow/models/tree/master/research/slim\">https://github.com/tensorflow/models/tree/master/research/slim</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 235892,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "10/26/2017 08:48:24",
          "content": "<p>you can use sgd, rmsprop is not recommended</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 235896,
          "author_name": "dyaprk",
          "author_url": "",
          "post_date": "10/26/2017 08:55:53",
          "content": "<p>why rmsprop is not recommended?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 235902,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/26/2017 09:08:27",
          "content": "<p>RMSProp should not make much of difference imho. I think you have bug in your code with that much of a difference in train/val!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 235941,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/26/2017 10:55:59",
          "content": "<p>@DaeYoungPark</p>\n\n<p>refer to: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/42069\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/42069</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 236285,
          "author_name": "dyaprk",
          "author_url": "",
          "post_date": "10/27/2017 01:27:06",
          "content": "<p>Thank you for tour refer!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 236060,
      "author_name": "nelsonzhao",
      "author_url": "",
      "post_date": "10/26/2017 16:20:38",
      "content": "<p>Thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 236164,
      "author_name": "owaisqayum",
      "author_url": "",
      "post_date": "10/26/2017 19:11:46",
      "content": "<p>Its a useful source and am new to kaggle hope to learn something new</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 237360,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "10/30/2017 06:21:15",
      "content": "<p>From discussions, it seems that the highest score we get is 0.70 with single model. Did you tried any other tricks?</p>\n\n<p>Did you make use the other two categories info? </p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 237844,
      "author_name": "owruby",
      "author_url": "",
      "post_date": "10/31/2017 04:44:58",
      "content": "<p>How did you adjust your learning rate?\nThanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 238069,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "10/31/2017 15:25:41",
      "content": "<p>Wondering how do you modify resnet101 to accept smaller images. Seems diffrent people has diffrent strategies</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 238242,
      "author_name": "ybw9000",
      "author_url": "",
      "post_date": "10/31/2017 23:31:53",
      "content": "<p>Resnet50, 180x180 input, make use of the other two categories info, 6 epoch, no TTA:\nval_acc: 0.692\nLB: 0.705</p>",
      "votes": null,
      "replies": [
        {
          "id": 238325,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "11/01/2017 06:46:30",
          "content": "<p>What's the meaning of the other two categories info? Did you train multi classifier?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 238623,
          "author_name": "ybw9000",
          "author_url": "",
          "post_date": "11/01/2017 19:05:37",
          "content": "<p>In addition to the categories we are interested in, I also train the model to predict the categories of the father levels, like animal - mammal - cat, where cat is our target.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 238775,
          "author_name": "zhangsongwei",
          "author_url": "",
          "post_date": "11/02/2017 06:10:18",
          "content": "<p>Are you using the father levels as auxiliary classifier?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 238776,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/02/2017 06:12:05",
          "content": "<p>regarding hierarchical categorization, one can consider the yolo9000 approach (wordTree): <a href=\"https://arxiv.org/abs/1612.08242\">https://arxiv.org/abs/1612.08242</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 245805,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "11/19/2017 19:16:05",
          "content": "<p>Are you using the wordTree approach? I'm not sure how to apply this to our task. Do you know how to predict conditional probabilities for each intermediate node?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 242245,
      "author_name": "blackarrow3542",
      "author_url": "",
      "post_date": "11/10/2017 22:30:23",
      "content": "<p>Xception, 180x180 \nval acc: 0.7277\naverage images for product LB: 0.73785.    </p>",
      "votes": null,
      "replies": [
        {
          "id": 242271,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "11/11/2017 01:28:21",
          "content": "<p>average images for product，you predict the prob of every image and average them?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242282,
          "author_name": "blackarrow3542",
          "author_url": "",
          "post_date": "11/11/2017 02:42:55",
          "content": "<p>Yes, for test data, for each product I predict the prob of every image and average them.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242303,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "11/11/2017 05:24:16",
          "content": "<p>I'm also using Xception now. But I only got LB: 0.68 for about 15 epoch. Do you train all parameters in the model? Which augmentation strategy do you use? Thanks a lot!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242306,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "11/11/2017 05:32:00",
          "content": "<p>Do you use pretrained model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242525,
          "author_name": "blackarrow3542",
          "author_url": "",
          "post_date": "11/11/2017 19:56:11",
          "content": "<p>@CSAdu: Yes I use Keras pre-trained model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242526,
          "author_name": "blackarrow3542",
          "author_url": "",
          "post_date": "11/11/2017 19:58:09",
          "content": "<p>@Brian Luo: Yes I train all parameters. I use small augmentation to get 0.68 then continue training without augmentation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242964,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "11/13/2017 03:50:10",
          "content": "<p>Thanks for your reply. How many epoch do you use to get 0.73?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 243159,
          "author_name": "blackarrow3542",
          "author_url": "",
          "post_date": "11/13/2017 13:30:15",
          "content": "<p>It takes me about 60 epochs to get 0.68 and another 100 epochs to get 0.73, each epoch is 1/10 train data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 246516,
          "author_name": "pierretisseur",
          "author_url": "",
          "post_date": "11/21/2017 08:19:04",
          "content": "<p>Did you use a globalaveragePooling layer just before the last FC layers (which gives 2048--&gt;5270)?\nOr some other poolling layers  which gives X&gt;2048 ---&gt;5270 at the top?\nThanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 246544,
          "author_name": "blackarrow3542",
          "author_url": "",
          "post_date": "11/21/2017 09:07:49",
          "content": "<p>Just GlobalAveragePooling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 247083,
          "author_name": "blackarrow3542",
          "author_url": "",
          "post_date": "11/22/2017 09:25:34",
          "content": "<p>inception resnet v2, tune keras pretrained model. Train top layers for 2 epoch, then train all parameters for 12 epoch. \n180x180, single image val acc: 0.7305,  LB: 0.74564(average prediction for all images of a product.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 247212,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "11/22/2017 15:45:43",
          "content": "<p>How long you need to train one epoch for inception resnet v2?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 247386,
          "author_name": "blackarrow3542",
          "author_url": "",
          "post_date": "11/22/2017 23:20:26",
          "content": "<p>It takes me about 9 hours to train all parameters for each epoch with AMD 1950X, 4 gtx 1080ti and Samsung 960 pro. My GPUs usage is not 100%, I think there is still space to improve speed of I/O of my code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 249023,
          "author_name": "skinish",
          "author_url": "",
          "post_date": "11/27/2017 15:00:58",
          "content": "<p>Thanks for your share. I have some questions:</p>\n\n<ol>\n<li>Do you train your model like it is described in here? <a href=\"https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\">https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights</a></li>\n<li>Sorry, I am new at image processing. What does 180x180 mean? that the images were reshaped to 180x180 dimensions?</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 249228,
          "author_name": "blackarrow3542",
          "author_url": "",
          "post_date": "11/28/2017 01:21:34",
          "content": "<p>You are welcome.</p>\n\n<ol>\n<li>I use this data generator <a href=\"https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson\">https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson</a>, and training is similar to <a href=\"https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\">https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights</a>  </li>\n<li>180x180 is the original size of the images. I didn't reshape them.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 249377,
          "author_name": "steelrose",
          "author_url": "",
          "post_date": "11/28/2017 09:52:34",
          "content": "<p>thanks for the details; do you have any tips / ideas why it \"doesn't work\" for me? I tried \"Inception_v3\" from Keras, after one epoch (about 80% of Train data) I was getting only 20% accuracy (on Validation only 13%) !!!</p>\n\n<p>I use my own Generator, 512 Batch size, random shuffle of the data .... ,aybe I need to try with the Generator from the link</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 249405,
          "author_name": "blackarrow3542",
          "author_url": "",
          "post_date": "11/28/2017 11:41:04",
          "content": "<p>My guess: </p>\n\n<ol>\n<li><p>Did you use the keras application build in preprocess function? \nfrom keras.applications.inception_v3 import preprocess_input,</p></li>\n<li><p>What optimizer did you use, can you try 'adam' ?</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 249406,
          "author_name": "steelrose",
          "author_url": "",
          "post_date": "11/28/2017 11:44:35",
          "content": "<p>I didn't use the function itself (think it was throwing an error?) but did:\nimg = img / 255.\nimg = img - 0.5\n(I think there was one more step .. basically I copied the 'tf' mode)</p>\n\n<p>I used rmsprop, will try with adam in the evening ...</p>\n\n<p>btw. I used custom Input shape (180,180,3) - could that also be the problem?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 249679,
          "author_name": "blackarrow3542",
          "author_url": "",
          "post_date": "11/29/2017 01:14:17",
          "content": "<p>Usually, my val_acc is higher than acc when acc is lower than 60%.  I was using (img/255-0.5)*2.0 and can get 40% with 10% data(with inception resnet v2 and xception). Then I switch to the keras preprocess and can get 50% with 10% data(with resnet101 or resnet 151).  I also use (180,180,3).  What is your batch size? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 249765,
          "author_name": "steelrose",
          "author_url": "",
          "post_date": "11/29/2017 07:05:49",
          "content": "<p>thank you, with some corrections (guess the preprocessing was wrong) I got 50% after around 1.5 epoch (around 12M Images), an improvement but still no so great</p>\n\n<p>I'm using batch size = 512, this time it was with adam optimizer (all default settings), will test with Xception or the pretrained weights from Heng CherKeng</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250062,
          "author_name": "skinish",
          "author_url": "",
          "post_date": "11/29/2017 17:43:48",
          "content": "<p>what is the 'tf' mode? Can you guys point me where to read more about retraining and transfer learning with Xception (or Inception, resnet, etc)? :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250103,
          "author_name": "steelrose",
          "author_url": "",
          "post_date": "11/29/2017 19:49:01",
          "content": "<p>I assume 'tf' stands for tensorflow, but might be wrong :)\nwhen you check the code in <a href=\"https://github.com/fchollet/keras/blob/master/keras/applications/imagenet_utils.py\">https://github.com/fchollet/keras/blob/master/keras/applications/imagenet_utils.py</a> you will see, but basically it's (since 2 days ago, before it was the transformation mentioned by Enhao above): <br>\nx /= 127.5\nx -= 1.</p>\n\n<p>for a tutorial you can check out this keras blog: <a href=\"https://blog.keras.io/building-powerful-image-classification-models-using-very-little-data.html\">https://blog.keras.io/building-powerful-image-classification-models-using-very-little-data.html</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 252273,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "12/02/2017 16:49:51",
          "content": "<p>What's different between (img/255-0.5)*2.0 and keras preprocess (x /= 127.5 x -= 1)? I think they're the same thing, why you get higher result with keras preprocess?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 252279,
          "author_name": "steelrose",
          "author_url": "",
          "post_date": "12/02/2017 16:57:55",
          "content": "<p>yes, most likely the main difference was coming from the adam optimizer (I did both changes at once)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 242335,
      "author_name": "andreanobile",
      "author_url": "",
      "post_date": "11/11/2017 09:26:11",
      "content": "<p>i'm still using a single resnet-18 because of very limited resources. 160x160 random crop in training, from scratch, lr 0.1 -&gt; 0.01 -&gt; 0.001 -&gt; 0.0005 -&gt; 0.0001, changing lr every few epochs, SGD with momentum 0.9, 192 images batches. I have modified the network in the last few conv layers. LB 0.69. inference with only 1 center 160x160 crop.</p>",
      "votes": null,
      "replies": [
        {
          "id": 242363,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "11/11/2017 10:45:11",
          "content": "<p>How did you modify the conv layers? Seems like a pretty good score for resnet-18!\nAny data augmentation? We are still looking for the right kind...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242370,
          "author_name": "andreanobile",
          "author_url": "",
          "post_date": "11/11/2017 11:19:51",
          "content": "<p>it seems to me that there's a kind of flaw in vanilla resnet-18 in the final few stages, if the final classifier has to distinguish among 5270 classes. The final FC layer is just a matrix, it basically can do 2 things: rotate a vector and project or embed the vector in the target space. Embedding in higher dimensional space does not change the dimension of the initial subspace, it cannot give higher resolution. Resnet 18 has a average pooling before the FC, there is too much of a bottlenck before FC. So vanilla resnet 18 cannot get good class resolution beyond a 10-20% of the full set.  Unfortunately time is running and i have very old metal.... i'm working on a completely different back-end for my little resnet 18  but not sure if i can make it in time. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 245946,
      "author_name": "eachshadow",
      "author_url": "",
      "post_date": "11/20/2017 06:28:59",
      "content": "<p>My resnet50 only got to about 63% accuracy before it started massively overfitting</p>\n\n<p>EDIT: went back to an earlier epoch and changed the data augmentation, so it got to 65% (with room for improvement) after a couple of epochs with more overfitting than some of my other models, but not so much as before where the validation accuracy was actually decreasing</p>",
      "votes": null,
      "replies": [
        {
          "id": 248763,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "11/26/2017 22:49:09",
          "content": "<p>Wondering what is your best single model and its accuracy?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 248797,
          "author_name": "eachshadow",
          "author_url": "",
          "post_date": "11/27/2017 01:08:15",
          "content": "<p>My best model gets about 74.8% on the leaderboard by itself. I only have two or three real models trained and ensembling currently isn't getting me very big gains.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250479,
          "author_name": "eachshadow",
          "author_url": "",
          "post_date": "11/30/2017 03:38:05",
          "content": "<p>Update: Resnet50 gets  0.71450 on the leaderboard after 10 epochs</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250497,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "11/30/2017 04:03:18",
          "content": "<p>Just wondering what you have done to boost from 65% to 71%?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250520,
          "author_name": "eachshadow",
          "author_url": "",
          "post_date": "11/30/2017 04:34:56",
          "content": "<ul>\n<li>65% was my validation accuracy at two epochs, while 0.7134 was my validation accuracy at ten epochs</li>\n<li>for the last epoch I trained without data augmentation to put the training and testing conditions in harmony</li>\n<li>leaderboard accuracy is a bit higher since it involves more than one image for each product being predicted</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250545,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "11/30/2017 05:07:04",
          "content": "<p>Thanks! And 65% at 2 epoch is quite high I believed. Would you mind if you can share how do you achieve this? My resnet 101 only have about 58% after two epoch with 0.001 lr</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 249027,
      "author_name": "skinish",
      "author_url": "",
      "post_date": "11/27/2017 15:08:40",
      "content": "<p>How did you guys perform the train/val split?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 249807,
      "author_name": "burgalon",
      "author_url": "",
      "post_date": "11/29/2017 08:28:49",
      "content": "<p>Hi @iFighting, @minu, @Heng CherKeng,\nIt seems like you guys are using MXNet?\nWhat are you impressions comparing to PyTorch?\nDoes MXNet facilitates specific aspects? Were you able to attain better results training with MXNet?</p>",
      "votes": null,
      "replies": [
        {
          "id": 249816,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/29/2017 08:50:38",
          "content": "<p>i am using pytorch only for this competition.  I use caffe, mxnet, tensorflow before. I think their results are all about the same (because they use the \"same formula\" and based on nvidia cudnn). Some DL framework are faster. But overall, I think pytorch is fast and easy to debug and experiment with new structure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250607,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "11/30/2017 06:20:52",
          "content": "<p>yes, i use mxnet</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 250633,
          "author_name": "mwbyeon",
          "author_url": "",
          "post_date": "11/30/2017 06:40:21",
          "content": "<p>I'm just familiar with MXNet. and MXNet is efficient in Multi-GPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "233986": "I am not familiar with this type of the train set and this number of target classes. I hope this thread will help to figure out what models perform better and which of them worth considering.\n\nTo start:\n\nResnet 101, no bagging, no TTA\n\n - val_loss = 1.18 \n - val_acc = 0.744 \n - LB = 0.7437",
    "233991": "thanks a lot!",
    "233996": "That's quite high score for single model! How long did you train for this and which GPU did you use?\n\nDue to limited resources, I'm using images scaled down to 90x90 and simple Xception, no TTA\n\n - val_acc = .6904 (for each image) \n\n - LB= .7092 (averaged softmax output for all images in same product id)\n\nI'm trying many other approaches like using top 2000 frequent items, using category 1,2 information, etc, but nothing beats this baseline yet. Sad to see much better results with basic model, but thanks for the info.",
    "233997": "4 x GTX 1080 Ti\n\n18 epochs x 4.5 hours\n\nI have 0.69 after 7th epoch",
    "234000": "seems that we need long epoches!",
    "234005": "Amazing score. Did you do training without pretrained weight?\n\nMy results: \ninception_v3, 160 crop, val loss: 1.78~, val acc: 0.66~, lb: 0.673~",
    "234006": "i am also using 4 x GTX 1080 Ti. For resnext101 (about same flops as your resnet101 ),  1 epoch=12283645 images takes about 1.3 hr. I am using 180x180 as input. I suppose your input size is not 180x180?",
    "234007": "what does TTA mean?",
    "234010": "I use 160x160 crops.\n\n1.3hrs per epoch.... It looks like I am doing something wrong.\n\nIt may happen that my SSD is not fast enough and I need to buy faster one.",
    "234012": "TTA= test time augmentation (e.g. multi crops, resize in test time)",
    "234013": "Test-time augmentation. In my case, I can get extra 0.01 just with 12-way TTA(flip, shift up/down, etc) with exact same weight - almost free cake :)",
    "234019": "I used pre-trained. I would be also interested how long will it take to get the same result without pre-trained, but it may be too computationally expensive.",
    "234026": "I used resnext-101, with image centor croped to 128X128, used pre-trained. It can get 0.59 val _acc after 1 epoches. I used sgd optimizer. Training is going...",
    "234032": "i am sorry that i made a mistake. 1 epoch=12283645 images takes about 13 hr, with  180x180 as input. So your timing may be reasonable. \n\nAlso speed of resnext and resnet are not the same even if they have the same flops, https://github.com/Cadene/pretrained-models.pytorch/issues/8\n\nI suspect it is due to group convolution in pytorch implementation ,\nhttps://discuss.pytorch.org/t/does-pytorch-optimize-the-group-parameter-in-convs/3060\n\nWhat’s New in cuDNN 7?\nGrouped Convolutions for models such as ResNeXt and Xception and CTC (Connectionist Temporal Classification) loss layer for temporal classification\n\nhttps://developer.nvidia.com/cudnn",
    "234047": "I'm just wondering, how many layers are actually training in your 101 layers net? For me, training one more block of Xception takes me 2 more hours for one epoch. (single GTX 1080)",
    "234061": "I use single ResNet 50 and my current LB-score is 0.45813, local validation score is 0.4914.",
    "234093": "SE-ResNet-50:\n\n - finetuning about 7 epochs (about 16 hours per epoch on 2x1080)\n - train augmentation: 161x161 random crops + horizontal flips\n - val single crop accuracy: 70.96\n - val 10 crops accuracy: 71.79\n - LB 10 crops accuracy: 71.67",
    "234099": "thanks for the information! What is the learning rate you used for finetunning?",
    "234101": "What's your learning rate schedule/optimize?",
    "234102": "I am using SGD with Nesterov momentum (0.9) and batch_size=256\n\n - epochs 1..5: 0.001\n - epoch 6: 0.0001\n - epoch 7: 0.00001",
    "234106": "Grouped Convs seem to be supported in the current master of Pytorch: https://github.com/pytorch/pytorch/pull/3057",
    "234111": "i think this is spatial, where group = in_planes. It improve mobilenet but not Resnext, where  group=32 or 64.",
    "234115": "Vladimir Iglovikov\n \nI rewrite my code an i get the following results:\n\n -  System: 4x1080Ti/pytorch\n\n -  input =160x160\n\n -  network = resnet101\n\n -  batch = 256\n\n -  time for 1000 iterations = 1000*256 images = 7 min.\n\nif one epoch = 12283645 images, i am getting time per epoch = 336 min = 5.6 hr.\n\nhow many images are there in one epoch for your timing of 4.5hr? Thanks a lot!\n\n(I am trying to see if my system has been setup correctly. This is the first time i use multi-gpu training intensively)",
    "234137": "one epoch 4.5 hours? why so fast?",
    "234151": "Resnet-18: val accuracy 0.63 trained for 6 epochs,  simple SGD, lr 0.1 -&gt; 0.01  -&gt; 0.001 , change lr after 2 epochs, no data augmentation no crop, 180x180, just change FC layer to get the 5270 classes and train",
    "234175": "One epoch - 11738624 Images. Rest are used for validation. \n\nBatch_size = 512",
    "234204": "Under the same setting (batch=512), i am able to get training iteration of 40,526  images per minute. For an epoch of 11,738,624 it would take 282 min or 4.7 hr, which is close to your results. Thanks again!",
    "234294": "Some useful info in this thread. Thanks all. I'm curious what people are settling on for training augmentation? I adapted some training code used in previous challenges last week and set some models training. I've got one model moving past the mid 60s in validation error and another that got stuck in the mid 50s. \n\nThe model that got stuck was high capacity but I had colour/saturation augmentation enabled. The one that's still chugging along and doing much better only had random crop/scale and rotation. I'm guessing the prevalence of almost constant black or white backgrounds in most images means that adding variability to that requires much more learning than leaving it be.\n\nNext angle is to see if disabling the random crop and just leaving h-flip and a bit of rotation on does better... basically assuming that the preprocessing of the competition dataset leaves things relatively well centred and similarly scaled across train and test datasets. Has anyone hit close to .7 or better with almost no augmentation?",
    "234309": "Are you training all params in your net?",
    "234997": "here is a possible Cdiscount in-house baseline results:\nhttps://www.linkedin.com/in/zhiwei-li-81b444126/\n\nApr 2017 – Sep 2017  Employment Duration 6 mos\n\nLocationBordeaux Area, France\n\nLarge Scale Image Classification for e-business product categorization\n\n• Research work on different Deep Learning architectures like Inception, ResNet etc.\n\n• Worked on 20 million images in roughly 9000 unbalanced class for product categories\n\n• Implemented transfer learning based on Inception V3 model using TensorFlow and Keras\n\n• Designed a quick model construction pipeline, raised ~60 times efficacy, achieved 72.3% accuracy\n\n• This model will be used to improve the categorization ability of current model",
    "235034": "&gt; 9000 unbalanced classes \n\nThis probably makes the problem harder.",
    "235081": "&gt; This probably makes the problem harder.\n\n\n\n\nOr you just caught a resume 'embellishment' gone wrong..  \nThat's linkedin.",
    "235885": "when i test with renet v2 101 one epoch (train:9 val:1)  <br>\n : lr rate 0.01, <br>\n : aug flip  <br>\n : size 180*180  <br>\n :  rmsprop optimizer  <br>\n : init with tensorflow slim renet 101  <br>\n : no weight decay, no batch norm decay  <br>\n : batch size 48 <br>\n\ntrain acc graph movi around 60%, but\nvaldation set give me acc 0.1%\n\nwhat makes my model so overfitting?\n - 5270 1*1 conv layer?\ni refenrece following model\n - https://github.com/tensorflow/models/tree/master/research/slim",
    "235892": "you can use sgd, rmsprop is not recommended",
    "235896": "why rmsprop is not recommended?",
    "235902": "RMSProp should not make much of difference imho. I think you have bug in your code with that much of a difference in train/val!",
    "235941": "DaeYoungPark\n \nrefer to: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/42069",
    "236060": "Thanks a lot!",
    "236164": "Its a useful source and am new to kaggle hope to learn something new",
    "236285": "Thank you for tour refer!!",
    "236443": "Does 'val single crop accuracy: 70.96' mean that you got a validation accuracy of 70.96% after always training with the same 161x161 crop?\n\nDoes 'val 10 crops accuracy: 71.79'  mean that the accuracy went up to 71.79 when you used 10 different crops per image?\n\nThe conclusion being that augmentation helps.\n\nSorry to ask. I am still getting used to the shorthand that people use.\n\nLB - is the score you got on the leader board, right?\n\nAnd during validation do you take a middle crop or resize from 180 to 161?",
    "236444": "You must have a bug. What are you doing for pre-processing?",
    "237360": "From discussions, it seems that the highest score we get is 0.70 with single model. Did you tried any other tricks?\n\nDid you make use the other two categories info? \n\nThanks!",
    "237844": "How did you adjust your learning rate?\nThanks!",
    "238069": "Wondering how do you modify resnet101 to accept smaller images. Seems diffrent people has diffrent strategies",
    "238242": "Resnet50, 180x180 input, make use of the other two categories info, 6 epoch, no TTA:\nval_acc: 0.692\nLB: 0.705",
    "238325": "What's the meaning of the other two categories info? Did you train multi classifier?",
    "238623": "In addition to the categories we are interested in, I also train the model to predict the categories of the father levels, like animal - mammal - cat, where cat is our target.",
    "238775": "Are you using the father levels as auxiliary classifier?",
    "238776": "regarding hierarchical categorization, one can consider the yolo9000 approach (wordTree): https://arxiv.org/abs/1612.08242",
    "242245": "Xception, 180x180 \nval acc: 0.7277\naverage images for product LB: 0.73785.",
    "242271": "average images for product，you predict the prob of every image and average them?",
    "242282": "Yes, for test data, for each product I predict the prob of every image and average them.",
    "242303": "I'm also using Xception now. But I only got LB: 0.68 for about 15 epoch. Do you train all parameters in the model? Which augmentation strategy do you use? Thanks a lot!",
    "242306": "Do you use pretrained model?",
    "242335": "i'm still using a single resnet-18 because of very limited resources. 160x160 random crop in training, from scratch, lr 0.1 -&gt; 0.01 -&gt; 0.001 -&gt; 0.0005 -&gt; 0.0001, changing lr every few epochs, SGD with momentum 0.9, 192 images batches. I have modified the network in the last few conv layers. LB 0.69. inference with only 1 center 160x160 crop.",
    "242363": "How did you modify the conv layers? Seems like a pretty good score for resnet-18!\nAny data augmentation? We are still looking for the right kind...",
    "242370": "it seems to me that there's a kind of flaw in vanilla resnet-18 in the final few stages, if the final classifier has to distinguish among 5270 classes. The final FC layer is just a matrix, it basically can do 2 things: rotate a vector and project or embed the vector in the target space. Embedding in higher dimensional space does not change the dimension of the initial subspace, it cannot give higher resolution. Resnet 18 has a average pooling before the FC, there is too much of a bottlenck before FC. So vanilla resnet 18 cannot get good class resolution beyond a 10-20% of the full set.  Unfortunately time is running and i have very old metal.... i'm working on a completely different back-end for my little resnet 18  but not sure if i can make it in time.",
    "242525": "CSAdu: Yes I use Keras pre-trained model.",
    "242526": "Brian Luo: Yes I train all parameters. I use small augmentation to get 0.68 then continue training without augmentation.",
    "242761": "So do you use pretrained model? Where you find your pretrained model. I start to train the SE-Resnet50 but seems after 1 epoch I only got 0.50, it this validation accuracy close to your data after 1 epoch?",
    "242964": "Thanks for your reply. How many epoch do you use to get 0.73?",
    "243159": "It takes me about 60 epochs to get 0.68 and another 100 epochs to get 0.73, each epoch is 1/10 train data.",
    "245805": "Are you using the wordTree approach? I'm not sure how to apply this to our task. Do you know how to predict conditional probabilities for each intermediate node?",
    "245946": "My resnet50 only got to about 63% accuracy before it started massively overfitting\n\nEDIT: went back to an earlier epoch and changed the data augmentation, so it got to 65% (with room for improvement) after a couple of epochs with more overfitting than some of my other models, but not so much as before where the validation accuracy was actually decreasing",
    "246516": "Did you use a globalaveragePooling layer just before the last FC layers (which gives 2048--&gt;5270)?\nOr some other poolling layers  which gives X&gt;2048 ---&gt;5270 at the top?\nThanks",
    "246544": "Just GlobalAveragePooling.",
    "247083": "inception resnet v2, tune keras pretrained model. Train top layers for 2 epoch, then train all parameters for 12 epoch. \n180x180, single image val acc: 0.7305,  LB: 0.74564(average prediction for all images of a product.)",
    "247212": "How long you need to train one epoch for inception resnet v2?",
    "247386": "It takes me about 9 hours to train all parameters for each epoch with AMD 1950X, 4 gtx 1080ti and Samsung 960 pro. My GPUs usage is not 100%, I think there is still space to improve speed of I/O of my code.",
    "248763": "Wondering what is your best single model and its accuracy?",
    "248797": "My best model gets about 74.8% on the leaderboard by itself. I only have two or three real models trained and ensembling currently isn't getting me very big gains.",
    "249023": "Thanks for your share. I have some questions:\n\n 1. Do you train your model like it is described in here? https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights\n 2. Sorry, I am new at image processing. What does 180x180 mean? that the images were reshaped to 180x180 dimensions?",
    "249027": "How did you guys perform the train/val split?",
    "249228": "You are welcome.\n\n1. I use this data generator https://www.kaggle.com/humananalog/keras-generator-for-reading-directly-from-bson, and training is similar to https://www.kaggle.com/mihaskalic/keras-xception-model-0-68-on-pl-weights  \n2. 180x180 is the original size of the images. I didn't reshape them.",
    "249377": "thanks for the details; do you have any tips / ideas why it \"doesn't work\" for me? I tried \"Inception_v3\" from Keras, after one epoch (about 80% of Train data) I was getting only 20% accuracy (on Validation only 13%) !!!\n\nI use my own Generator, 512 Batch size, random shuffle of the data .... ,aybe I need to try with the Generator from the link",
    "249405": "My guess: \n\n1.  Did you use the keras application build in preprocess function? \nfrom keras.applications.inception_v3 import preprocess_input,\n\n2. What optimizer did you use, can you try 'adam' ?",
    "249406": "I didn't use the function itself (think it was throwing an error?) but did:\nimg = img / 255.\nimg = img - 0.5\n(I think there was one more step .. basically I copied the 'tf' mode)\n\nI used rmsprop, will try with adam in the evening ...\n\nbtw. I used custom Input shape (180,180,3) - could that also be the problem?",
    "249679": "Usually, my val_acc is higher than acc when acc is lower than 60%.  I was using (img/255-0.5)*2.0 and can get 40% with 10% data(with inception resnet v2 and xception). Then I switch to the keras preprocess and can get 50% with 10% data(with resnet101 or resnet 151).  I also use (180,180,3).  What is your batch size?",
    "249765": "thank you, with some corrections (guess the preprocessing was wrong) I got 50% after around 1.5 epoch (around 12M Images), an improvement but still no so great\n\nI'm using batch size = 512, this time it was with adam optimizer (all default settings), will test with Xception or the pretrained weights from Heng CherKeng",
    "249807": "Hi @iFighting, @minu, @Heng CherKeng,\nIt seems like you guys are using MXNet?\nWhat are you impressions comparing to PyTorch?\nDoes MXNet facilitates specific aspects? Were you able to attain better results training with MXNet?",
    "249816": "i am using pytorch only for this competition.  I use caffe, mxnet, tensorflow before. I think their results are all about the same (because they use the \"same formula\" and based on nvidia cudnn). Some DL framework are faster. But overall, I think pytorch is fast and easy to debug and experiment with new structure.",
    "250062": "what is the 'tf' mode? Can you guys point me where to read more about retraining and transfer learning with Xception (or Inception, resnet, etc)? :)",
    "250103": "I assume 'tf' stands for tensorflow, but might be wrong :)\nwhen you check the code in https://github.com/fchollet/keras/blob/master/keras/applications/imagenet_utils.py you will see, but basically it's (since 2 days ago, before it was the transformation mentioned by Enhao above):  \nx /= 127.5\nx -= 1.\n\nfor a tutorial you can check out this keras blog: https://blog.keras.io/building-powerful-image-classification-models-using-very-little-data.html",
    "250479": "Update: Resnet50 gets  0.71450 on the leaderboard after 10 epochs",
    "250497": "Just wondering what you have done to boost from 65% to 71%?",
    "250520": "* 65% was my validation accuracy at two epochs, while 0.7134 was my validation accuracy at ten epochs\n* for the last epoch I trained without data augmentation to put the training and testing conditions in harmony\n* leaderboard accuracy is a bit higher since it involves more than one image for each product being predicted",
    "250545": "Thanks! And 65% at 2 epoch is quite high I believed. Would you mind if you can share how do you achieve this? My resnet 101 only have about 58% after two epoch with 0.001 lr",
    "250607": "yes, i use mxnet",
    "250633": "I'm just familiar with MXNet. and MXNet is efficient in Multi-GPU.",
    "252273": "What's different between (img/255-0.5)*2.0 and keras preprocess (x /= 127.5 x -= 1)? I think they're the same thing, why you get higher result with keras preprocess?",
    "252279": "yes, most likely the main difference was coming from the adam optimizer (I did both changes at once)"
  },
  "source": "meta"
}