{
  "id": 42069,
  "title": "repeating resnet101 LB=0.74 results",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/42069",
  "author_name": "",
  "post_date": "2017-10-26T10:44:14.674374200Z",
  "votes": 25,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Thanks to @Vladimir Iglovikov for this discussion thread at <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652</a>.</p>\n\n<p>Here, I try to repeat his results:</p>\n\n<ul>\n<li>val_loss = 1.18</li>\n<li>val_acc = 0.744</li>\n<li>LB = 0.7437 (0.69 after 7th epoch)</li>\n</ul>\n\n<p>.</p>\n\n<ul>\n<li>4 x GTX 1080 Ti</li>\n<li>18 epochs x 4.5 hours</li>\n<li>One epoch - 11738624 images</li>\n<li>batch size = 512</li>\n</ul>\n\n<p>So far below is my results (LB 0.704). It seems that i have overfitting and my results are not as good as @Vladimir Iglovikov. You may want to try other train parameters and better augmentation. I need to review my training pipeline and try second time.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/235935/7750/results.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>My code is based on pytorch. It runs at 4 hr 27 min per 11.2 million images. </p>\n\n<p>You can download my code and trained  model at : <a href=\"https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing\">https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing</a> (see folder \"10-26\").</p>\n\n<p>You should be able to run it. The imagenet pretrain model is from pytorch model zoo. Please also refer to: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a></p>",
  "messages": [
    {
      "id": "235935",
      "postDate": "10/26/2017 10:44:14",
      "content": "<p>Thanks to @Vladimir Iglovikov for this discussion thread at <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652</a>.</p>\n\n<p>Here, I try to repeat his results:</p>\n\n<ul>\n<li>val_loss = 1.18</li>\n<li>val_acc = 0.744</li>\n<li>LB = 0.7437 (0.69 after 7th epoch)</li>\n</ul>\n\n<p>.</p>\n\n<ul>\n<li>4 x GTX 1080 Ti</li>\n<li>18 epochs x 4.5 hours</li>\n<li>One epoch - 11738624 images</li>\n<li>batch size = 512</li>\n</ul>\n\n<p>So far below is my results (LB 0.704). It seems that i have overfitting and my results are not as good as @Vladimir Iglovikov. You may want to try other train parameters and better augmentation. I need to review my training pipeline and try second time.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/235935/7750/results.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>My code is based on pytorch. It runs at 4 hr 27 min per 11.2 million images. </p>\n\n<p>You can download my code and trained  model at : <a href=\"https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing\">https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing</a> (see folder \"10-26\").</p>\n\n<p>You should be able to run it. The imagenet pretrain model is from pytorch model zoo. Please also refer to: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a></p>",
      "rawMarkdown": "Thanks to @Vladimir Iglovikov for this discussion thread at https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652.\n\nHere, I try to repeat his results:\n\n  - val_loss = 1.18\n  - val_acc = 0.744\n  - LB = 0.7437 (0.69 after 7th epoch)\n\n.\n\n\n  - 4 x GTX 1080 Ti\n  - 18 epochs x 4.5 hours\n  - One epoch - 11738624 images\n  - batch size = 512\n\nSo far below is my results (LB 0.704). It seems that i have overfitting and my results are not as good as @Vladimir Iglovikov. You may want to try other train parameters and better augmentation. I need to review my training pipeline and try second time.\n\n  ![enter image description here][1]\n\nMy code is based on pytorch. It runs at 4 hr 27 min per 11.2 million images. \n\nYou can download my code and trained  model at : https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing (see folder \"10-26\").\n\nYou should be able to run it. The imagenet pretrain model is from pytorch model zoo. Please also refer to: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/235935/7750/results.png",
      "votes": null
    },
    {
      "id": "236128",
      "postDate": "10/26/2017 18:23:15",
      "content": "<p>Thanks! I am not familiar with pythorch, so I could not find this in your code: can you tell how you managed to change the default picture size from 224X224 to 180X180 or 160X160 in your case and still use the pre training weigths? Did you change some thing else besides the input layer? I am using a prerained model of resnet-101 in Keras, and can't make it work when change the input size to less than 224X224.</p>",
      "rawMarkdown": "Thanks! I am not familiar with pythorch, so I could not find this in your code: can you tell how you managed to change the default picture size from 224X224 to 180X180 or 160X160 in your case and still use the pre training weigths? Did you change some thing else besides the input layer? I am using a prerained model of resnet-101 in Keras, and can't make it work when change the input size to less than 224X224.",
      "votes": null
    },
    {
      "id": "236157",
      "postDate": "10/26/2017 19:02:11",
      "content": "<p>Pytorch does not check for input size in convolutional layers afaik, just like it should be. A convolution can be applied to arbitrary image size, so there is no need to change the sizes. Since most architectures nowadays use a global pooling layers at the end, the fully connected layers receive an input of the size of number of channels of the last convolutional layer, which is independent of feature map dimensions.</p>",
      "rawMarkdown": "Pytorch does not check for input size in convolutional layers afaik, just like it should be. A convolution can be applied to arbitrary image size, so there is no need to change the sizes. Since most architectures nowadays use a global pooling layers at the end, the fully connected layers receive an input of the size of number of channels of the last convolutional layer, which is independent of feature map dimensions.",
      "votes": null
    },
    {
      "id": "236167",
      "postDate": "10/26/2017 19:40:24",
      "content": "<p>The problem occurs in the MaxPooling2D just after the first batch normalization.</p>",
      "rawMarkdown": "The problem occurs in the MaxPooling2D just after the first batch normalization.",
      "votes": null
    },
    {
      "id": "237111",
      "postDate": "10/29/2017 13:58:45",
      "content": "<p>Thanks again CK! Any chance that you can also share the <code>time_to_str</code>? ;P</p>",
      "rawMarkdown": "Thanks again CK! Any chance that you can also share the `time_to_str`? ;P",
      "votes": null
    },
    {
      "id": "237118",
      "postDate": "10/29/2017 14:20:37",
      "content": "<pre><code>def time_to_str(t):\nt  = int(t)\nhr = t//60\nmin = t%60\nreturn '%2d hr %02d min'%(hr,min)\n</code></pre>",
      "rawMarkdown": "def time_to_str(t):\n    t  = int(t)\n    hr = t//60\n    min = t%60\n    return '%2d hr %02d min'%(hr,min)",
      "votes": null
    },
    {
      "id": "237521",
      "postDate": "10/30/2017 13:55:43",
      "content": "<p>After trying several days, it seems impossible for me to get 0.74. My best results is about LB 0.71 (vanilla resnet101  image prediction), but this is after lots of  efforts in parameter tuning. </p>\n\n<p>But I believe that it is possible to get about 0.73 to 0.74 with single model like resnet101. Anyone has any suggestions that this can be done?</p>\n\n<p>I already try different augmentation, droupout, different batch size, learning rate, momentum.</p>\n\n<p>If there is no special tricks to get 0.74, then i suspect my multi-gpu training code could be buggy.</p>",
      "rawMarkdown": "After trying several days, it seems impossible for me to get 0.74. My best results is about LB 0.71 (vanilla resnet101  image prediction), but this is after lots of  efforts in parameter tuning. \n\nBut I believe that it is possible to get about 0.73 to 0.74 with single model like resnet101. Anyone has any suggestions that this can be done?\n\nI already try different augmentation, droupout, different batch size, learning rate, momentum.\n\nIf there is no special tricks to get 0.74, then i suspect my multi-gpu training code could be buggy.",
      "votes": null
    },
    {
      "id": "237540",
      "postDate": "10/30/2017 14:47:59",
      "content": "<p>Try ensemble of different folds from CV, or ensemble of different checkpoints of the model (last 3?). That will still be one model but a bit augmented. ;)</p>",
      "rawMarkdown": "Try ensemble of different folds from CV, or ensemble of different checkpoints of the model (last 3?). That will still be one model but a bit augmented. ;)",
      "votes": null
    },
    {
      "id": "237544",
      "postDate": "10/30/2017 14:52:46",
      "content": "<p>thanks for the reply. I know that will improve the LB score. But I am looking for a solution for single model (no ensemble of any kind) with no image  augmentation. I am wondering how to do it.</p>",
      "rawMarkdown": "thanks for the reply. I know that will improve the LB score. But I am looking for a solution for single model (no ensemble of any kind) with no image  augmentation. I am wondering how to do it.",
      "votes": null
    },
    {
      "id": "237742",
      "postDate": "10/30/2017 20:57:56",
      "content": "<p>Perhaps your classifier design is slightly different from the model that got 0.74. By this I mean the layer(s) that is/are placed on top of the ResNet-101 feature extraction layers.</p>",
      "rawMarkdown": "Perhaps your classifier design is slightly different from the model that got 0.74. By this I mean the layer(s) that is/are placed on top of the ResNet-101 feature extraction layers.",
      "votes": null
    },
    {
      "id": "237842",
      "postDate": "10/31/2017 04:04:09",
      "content": "<p>I don't know about that @Human Analog, all these pretrained models are highly optimized architectures.Adding conv layers to already existing pre-trained architectures will only decrease performance. I may be wrong though</p>",
      "rawMarkdown": "I don't know about that @Human Analog, all these pretrained models are highly optimized architectures.Adding conv layers to already existing pre-trained architectures will only decrease performance. I may be wrong though",
      "votes": null
    },
    {
      "id": "237909",
      "postDate": "10/31/2017 09:37:16",
      "content": "<p>I said nothing about conv layers. ;-) But when you take a pretrained network such as ResNet-101, you have to add your own classifier on top (for the 5270 classes). There are different ways to do this. You can connect a fully-connected layer directly to the last convolution layer, or you can first use global average pooling and then use a fully-connected layer, or you can use multiple fully-connected layers, or many other schemes. So just because two people say they both use ResNet-101 does not mean they also use the same classifier design on top of that network... and that is where the difference in score may be coming from.</p>",
      "rawMarkdown": "I said nothing about conv layers. ;-) But when you take a pretrained network such as ResNet-101, you have to add your own classifier on top (for the 5270 classes). There are different ways to do this. You can connect a fully-connected layer directly to the last convolution layer, or you can first use global average pooling and then use a fully-connected layer, or you can use multiple fully-connected layers, or many other schemes. So just because two people say they both use ResNet-101 does not mean they also use the same classifier design on top of that network... and that is where the difference in score may be coming from.",
      "votes": null
    },
    {
      "id": "237926",
      "postDate": "10/31/2017 10:31:27",
      "content": "<p>.</p>",
      "rawMarkdown": ".",
      "votes": null
    },
    {
      "id": "238333",
      "postDate": "11/01/2017 07:01:08",
      "content": "<p>I think maybe he makes use of the other two categories info. That is to say, he maybe divide the data into not 5270 classes first.</p>",
      "rawMarkdown": "I think maybe he makes use of the other two categories info. That is to say, he maybe divide the data into not 5270 classes first.",
      "votes": null
    },
    {
      "id": "241869",
      "postDate": "11/10/2017 02:00:15",
      "content": "<p>I've tried your code and commented out the following code:</p>\n\n<pre><code>optimizer = optim.SGD([ # iter_accum==1\n    {'params': net.layer0.parameters(), 'lr': 0.001},\n    {'params': net.layer1.parameters(), 'lr': 0.001},\n    {'params': net.layer2.parameters(), 'lr': 0.001},\n    {'params': net.layer3.parameters(), 'lr': 0.001},\n    {'params': net.layer4.parameters(), 'lr': 0.001},\n    {'params': net.fc.parameters(),     'lr': 0.01},\n</code></pre>\n\n<p>but it seems not working , the loss is continuing to go up, should I do other things before trying the code ?\nMy output:</p>\n\n<pre><code>   rate   iter(k)  epoch   num(m)  | valid_loss/acc | train_loss/acc | batch_loss/acc |  time\n------------------------------------------------------------------------------------------------\n0.00000  243.0 k  10.13  122.79 | 1.1624  0.7296 | 0.0000  0.0000 | 0.0000  0.0000 |  0 hr 07 min\n0.00100  244.0 k  10.17  123.31 | 1.2460  0.7103 | 1.3729  0.6854 | 1.0941  0.7168 |  1 hr 00 min\n0.00100  245.0 k  10.21  123.82 | 1.2573  0.7089 | 1.2933  0.6953 | 1.1591  0.7227 |  1 hr 53 min\n0.00100  246.0 k  10.26  124.33 | 1.2772  0.7047 | 1.3211  0.6937 | 1.1279  0.7539 |  2 hr 46 min\n0.00100  247.0 k  10.30  124.84 | 1.2723  0.7060 | 1.3415  0.6931 | 1.2441  0.7090 |  3 hr 39 min\n0.00100  248.0 k  10.34  125.35 | 1.2756  0.7053 | 1.3496  0.6882 | 1.3224  0.7070 |  4 hr 32 min\n0.00100  249.0 k  10.38  125.87 | 1.2828  0.7036 | 1.3322  0.6984 | 1.4751  0.6504 |  5 hr 25 min\n0.00100  250.0 k  10.42  126.38 | 1.2861  0.7038 | 1.3474  0.6920 | 1.2393  0.7012 |  6 hr 18 min\n0.00100  251.0 k  10.47  126.89 | 1.2878  0.7031 | 1.3415  0.6977 | 1.4016  0.6797 |  7 hr 10 min\n0.00100  252.0 k  10.51  127.40 | 1.2865  0.7039 | 1.3785  0.6859 | 1.2657  0.6953 |  8 hr 03 min\n0.00100  253.0 k  10.55  127.91 | 1.2992  0.7012 | 1.3416  0.6922 | 1.3070  0.7012 |  8 hr 56 min\n0.00100  254.0 k  10.59  128.43 | 1.3002  0.6994 | 1.3434  0.6984 | 1.3961  0.6660 |  9 hr 49 min\n0.00100  255.0 k  10.64  128.94 | 1.2917  0.7025 | 1.3414  0.6944 | 1.4535  0.6758 | 10 hr 42 min\n0.00100  256.0 k  10.68  129.45 | 1.2977  0.7019 | 1.3330  0.6901 | 1.1098  0.7324 | 11 hr 35 min\n0.00100  257.0 k  10.72  129.96 | 1.2938  0.7029 | 1.3711  0.6858 | 1.2438  0.7148 | 12 hr 28 min\n0.00100  258.0 k  10.76  130.47 | 1.3145  0.6984 | 1.3411  0.6898 | 1.4022  0.6797 | 13 hr 21 min\n0.00100  259.0 k  10.80  130.99 | 1.2982  0.7012 | 1.3481  0.6931 | 1.4801  0.6602 | 14 hr 14 min\n0.00100  260.0 k  10.85  131.50 | 1.2986  0.7016 | 1.3390  0.6902 | 1.3132  0.6875 | 15 hr 07 min\n0.00100  261.0 k  10.89  132.01 | 1.3142  0.6981 | 1.3408  0.6945 | 1.3962  0.6934 | 15 hr 59 min\n0.00100  262.0 k  10.93  132.52 | 1.3067  0.7007 | 1.3624  0.6867 | 1.4176  0.6699 | 16 hr 52 min\n0.00100  263.0 k  10.97  133.03 | 1.3073  0.7002 | 1.3651  0.6900 | 1.3471  0.6758 | 17 hr 45 min\n0.00100  264.0 k  11.02  133.55 | 1.3006  0.7019 | 1.3382  0.6921 | 1.3121  0.7070 | 18 hr 38 min\n0.00100  265.0 k  11.06  134.06 | 1.3100  0.6993 | 1.3512  0.6952 | 1.3898  0.6953 | 19 hr 31 min\n0.00100  266.0 k  11.10  134.57 | 1.3145  0.6994 | 1.3197  0.6980 | 1.4913  0.6895 | 20 hr 24 min\n0.00100  267.0 k  11.14  135.08 | 1.3053  0.7000 | 1.3225  0.6900 | 1.2124  0.7285 | 21 hr 17 min\n0.00100  268.0 k  11.18  135.59 | 1.3041  0.7015 | 1.2959  0.7011 | 1.2908  0.6855 | 22 hr 10 min\n0.00100  269.0 k  11.23  136.11 | 1.3076  0.7005 | 1.2619  0.7072 | 1.2847  0.7070 | 23 hr 03 min\n0.00100  270.0 k  11.27  136.62 | 1.3160  0.6994 | 1.2833  0.7049 | 1.2478  0.7266 | 23 hr 56 min\n</code></pre>",
      "rawMarkdown": "I've tried your code and commented out the following code:\n\n    optimizer = optim.SGD([ # iter_accum==1\n        {'params': net.layer0.parameters(), 'lr': 0.001},\n        {'params': net.layer1.parameters(), 'lr': 0.001},\n        {'params': net.layer2.parameters(), 'lr': 0.001},\n        {'params': net.layer3.parameters(), 'lr': 0.001},\n        {'params': net.layer4.parameters(), 'lr': 0.001},\n        {'params': net.fc.parameters(),     'lr': 0.01},\n\nbut it seems not working , the loss is continuing to go up, should I do other things before trying the code ?\nMy output:\n\n       rate   iter(k)  epoch   num(m)  | valid_loss/acc | train_loss/acc | batch_loss/acc |  time\n    ------------------------------------------------------------------------------------------------\n    0.00000  243.0 k  10.13  122.79 | 1.1624  0.7296 | 0.0000  0.0000 | 0.0000  0.0000 |  0 hr 07 min\n    0.00100  244.0 k  10.17  123.31 | 1.2460  0.7103 | 1.3729  0.6854 | 1.0941  0.7168 |  1 hr 00 min\n    0.00100  245.0 k  10.21  123.82 | 1.2573  0.7089 | 1.2933  0.6953 | 1.1591  0.7227 |  1 hr 53 min\n    0.00100  246.0 k  10.26  124.33 | 1.2772  0.7047 | 1.3211  0.6937 | 1.1279  0.7539 |  2 hr 46 min\n    0.00100  247.0 k  10.30  124.84 | 1.2723  0.7060 | 1.3415  0.6931 | 1.2441  0.7090 |  3 hr 39 min\n    0.00100  248.0 k  10.34  125.35 | 1.2756  0.7053 | 1.3496  0.6882 | 1.3224  0.7070 |  4 hr 32 min\n    0.00100  249.0 k  10.38  125.87 | 1.2828  0.7036 | 1.3322  0.6984 | 1.4751  0.6504 |  5 hr 25 min\n    0.00100  250.0 k  10.42  126.38 | 1.2861  0.7038 | 1.3474  0.6920 | 1.2393  0.7012 |  6 hr 18 min\n    0.00100  251.0 k  10.47  126.89 | 1.2878  0.7031 | 1.3415  0.6977 | 1.4016  0.6797 |  7 hr 10 min\n    0.00100  252.0 k  10.51  127.40 | 1.2865  0.7039 | 1.3785  0.6859 | 1.2657  0.6953 |  8 hr 03 min\n    0.00100  253.0 k  10.55  127.91 | 1.2992  0.7012 | 1.3416  0.6922 | 1.3070  0.7012 |  8 hr 56 min\n    0.00100  254.0 k  10.59  128.43 | 1.3002  0.6994 | 1.3434  0.6984 | 1.3961  0.6660 |  9 hr 49 min\n    0.00100  255.0 k  10.64  128.94 | 1.2917  0.7025 | 1.3414  0.6944 | 1.4535  0.6758 | 10 hr 42 min\n    0.00100  256.0 k  10.68  129.45 | 1.2977  0.7019 | 1.3330  0.6901 | 1.1098  0.7324 | 11 hr 35 min\n    0.00100  257.0 k  10.72  129.96 | 1.2938  0.7029 | 1.3711  0.6858 | 1.2438  0.7148 | 12 hr 28 min\n    0.00100  258.0 k  10.76  130.47 | 1.3145  0.6984 | 1.3411  0.6898 | 1.4022  0.6797 | 13 hr 21 min\n    0.00100  259.0 k  10.80  130.99 | 1.2982  0.7012 | 1.3481  0.6931 | 1.4801  0.6602 | 14 hr 14 min\n    0.00100  260.0 k  10.85  131.50 | 1.2986  0.7016 | 1.3390  0.6902 | 1.3132  0.6875 | 15 hr 07 min\n    0.00100  261.0 k  10.89  132.01 | 1.3142  0.6981 | 1.3408  0.6945 | 1.3962  0.6934 | 15 hr 59 min\n    0.00100  262.0 k  10.93  132.52 | 1.3067  0.7007 | 1.3624  0.6867 | 1.4176  0.6699 | 16 hr 52 min\n    0.00100  263.0 k  10.97  133.03 | 1.3073  0.7002 | 1.3651  0.6900 | 1.3471  0.6758 | 17 hr 45 min\n    0.00100  264.0 k  11.02  133.55 | 1.3006  0.7019 | 1.3382  0.6921 | 1.3121  0.7070 | 18 hr 38 min\n    0.00100  265.0 k  11.06  134.06 | 1.3100  0.6993 | 1.3512  0.6952 | 1.3898  0.6953 | 19 hr 31 min\n    0.00100  266.0 k  11.10  134.57 | 1.3145  0.6994 | 1.3197  0.6980 | 1.4913  0.6895 | 20 hr 24 min\n    0.00100  267.0 k  11.14  135.08 | 1.3053  0.7000 | 1.3225  0.6900 | 1.2124  0.7285 | 21 hr 17 min\n    0.00100  268.0 k  11.18  135.59 | 1.3041  0.7015 | 1.2959  0.7011 | 1.2908  0.6855 | 22 hr 10 min\n    0.00100  269.0 k  11.23  136.11 | 1.3076  0.7005 | 1.2619  0.7072 | 1.2847  0.7070 | 23 hr 03 min\n    0.00100  270.0 k  11.27  136.62 | 1.3160  0.6994 | 1.2833  0.7049 | 1.2478  0.7266 | 23 hr 56 min",
      "votes": null
    },
    {
      "id": "241900",
      "postDate": "11/10/2017 04:23:48",
      "content": "<p>can you post you code(just the trainer.py) and the whole log file?</p>",
      "rawMarkdown": "can you post you code(just the trainer.py) and the whole log file?",
      "votes": null
    },
    {
      "id": "241902",
      "postDate": "11/10/2017 04:37:16",
      "content": "<p>ok, trainer.py : <a href=\"https://pastebin.com/QTySsgHf\">https://pastebin.com/QTySsgHf</a>\nlog_file: <a href=\"https://pastebin.com/n3AxGrsR\">https://pastebin.com/n3AxGrsR</a></p>",
      "rawMarkdown": "ok, trainer.py : [https://pastebin.com/QTySsgHf][1]\nlog_file: https://pastebin.com/n3AxGrsR\n\n  [1]: https://pastebin.com/QTySsgHf",
      "votes": null
    },
    {
      "id": "241903",
      "postDate": "11/10/2017 04:44:16",
      "content": "<p>note that your train and validation set is not the same as mine.  your results is actually reasonable. the validation score should be close to LB score, which is 0.70.</p>\n\n<p>you can try to use smaller learning rate. But i think the limit is around to 0.70 to 0.71.</p>\n\n<pre><code>if 1:\n    optimizer = optim.SGD(filter(lambda p: p.requires_grad, net.parameters()),\n                          lr=0.0001, momentum=0.9, weight_decay=0.0001) #nesterov=True\n</code></pre>",
      "rawMarkdown": "note that your train and validation set is not the same as mine.  your results is actually reasonable. the validation score should be close to LB score, which is 0.70.\n\nyou can try to use smaller learning rate. But i think the limit is around to 0.70 to 0.71.\n\n    if 1:\n        optimizer = optim.SGD(filter(lambda p: p.requires_grad, net.parameters()),\n                              lr=0.0001, momentum=0.9, weight_decay=0.0001) #nesterov=True",
      "votes": null
    },
    {
      "id": "241905",
      "postDate": "11/10/2017 04:51:07",
      "content": "<p>But my LB commit was 0.64 which was far from 0.70 .\nHaven't you find anything wrong elsewhere ? \nI guess I should take a  close look.</p>",
      "rawMarkdown": "But my LB commit was 0.64 which was far from 0.70 .\nHaven't you find anything wrong elsewhere ? \nI guess I should take a  close look.",
      "votes": null
    },
    {
      "id": "242119",
      "postDate": "11/10/2017 17:03:06",
      "content": "<p>i suspect some bug in your submission code. check the normalization of input at submission. it should be same as training. also note that there are 1 to 4 images per product.</p>",
      "rawMarkdown": "i suspect some bug in your submission code. check the normalization of input at submission. it should be same as training. also note that there are 1 to 4 images per product.",
      "votes": null
    },
    {
      "id": "242955",
      "postDate": "11/13/2017 03:24:03",
      "content": "<p>Oh, can I  use the submission functions in your trainer_resnet50.py code directly ? \nI couldn't find any normalization except \"SequentialSampler\" ?</p>",
      "rawMarkdown": "Oh, can I  use the submission functions in your trainer_resnet50.py code directly ? \nI couldn't find any normalization except \"SequentialSampler\" ?",
      "votes": null
    },
    {
      "id": "244675",
      "postDate": "11/16/2017 17:48:30",
      "content": "<p>I'm having a hard time trying to reproduce your validation score. My best guess is that the problem in image preprocessing, my current code for it looks like this:</p>\n\n<pre><code>mean = [0.485, 0.456, 0.406]\nstd  = [0.229, 0.224, 0.225]\nimg = np.array(img).astype(np.float32)\nimg = fix_center_crop(img, size=(160,160))  \nimg = img/255.\nfor i in range(3):\n    img[:, i] = (img[:, i]-mean[i])/std[i]\nimg = img.transpose((2,0,1))\n</code></pre>\n\n<p>And I'm getting about 0.316 accuracy on validation set.\nAny idea what I'm missing?</p>",
      "rawMarkdown": "I'm having a hard time trying to reproduce your validation score. My best guess is that the problem in image preprocessing, my current code for it looks like this:\n\n    mean = [0.485, 0.456, 0.406]\n    std  = [0.229, 0.224, 0.225]\n    img = np.array(img).astype(np.float32)\n    img = fix_center_crop(img, size=(160,160))  \n    img = img/255.\n    for i in range(3):\n        img[:, i] = (img[:, i]-mean[i])/std[i]\n    img = img.transpose((2,0,1))\n\nAnd I'm getting about 0.316 accuracy on validation set.\nAny idea what I'm missing?",
      "votes": null
    },
    {
      "id": "250145",
      "postDate": "11/29/2017 21:24:41",
      "content": "<p>thank you very much, you seem to use some class \"CDiscountDataset\" but I don't see in in the code you have on google ... can you please share it? (i'm not familiar with pytorch, so not sure what exactly should be there and whether it's basically just the code to read from bson, or also some augmentations etc.)</p>",
      "rawMarkdown": "thank you very much, you seem to use some class \"CDiscountDataset\" but I don't see in in the code you have on google ... can you please share it? (i'm not familiar with pytorch, so not sure what exactly should be there and whether it's basically just the code to read from bson, or also some augmentations etc.)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 236128,
      "author_name": "aloisiodn",
      "author_url": "",
      "post_date": "10/26/2017 18:23:15",
      "content": "<p>Thanks! I am not familiar with pythorch, so I could not find this in your code: can you tell how you managed to change the default picture size from 224X224 to 180X180 or 160X160 in your case and still use the pre training weigths? Did you change some thing else besides the input layer? I am using a prerained model of resnet-101 in Keras, and can't make it work when change the input size to less than 224X224.</p>",
      "votes": null,
      "replies": [
        {
          "id": 236157,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/26/2017 19:02:11",
          "content": "<p>Pytorch does not check for input size in convolutional layers afaik, just like it should be. A convolution can be applied to arbitrary image size, so there is no need to change the sizes. Since most architectures nowadays use a global pooling layers at the end, the fully connected layers receive an input of the size of number of channels of the last convolutional layer, which is independent of feature map dimensions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 236167,
          "author_name": "aloisiodn",
          "author_url": "",
          "post_date": "10/26/2017 19:40:24",
          "content": "<p>The problem occurs in the MaxPooling2D just after the first batch normalization.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 237111,
      "author_name": "terryli",
      "author_url": "",
      "post_date": "10/29/2017 13:58:45",
      "content": "<p>Thanks again CK! Any chance that you can also share the <code>time_to_str</code>? ;P</p>",
      "votes": null,
      "replies": [
        {
          "id": 237118,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/29/2017 14:20:37",
          "content": "<pre><code>def time_to_str(t):\nt  = int(t)\nhr = t//60\nmin = t%60\nreturn '%2d hr %02d min'%(hr,min)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 237521,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/30/2017 13:55:43",
      "content": "<p>After trying several days, it seems impossible for me to get 0.74. My best results is about LB 0.71 (vanilla resnet101  image prediction), but this is after lots of  efforts in parameter tuning. </p>\n\n<p>But I believe that it is possible to get about 0.73 to 0.74 with single model like resnet101. Anyone has any suggestions that this can be done?</p>\n\n<p>I already try different augmentation, droupout, different batch size, learning rate, momentum.</p>\n\n<p>If there is no special tricks to get 0.74, then i suspect my multi-gpu training code could be buggy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 237540,
          "author_name": "mpekalski",
          "author_url": "",
          "post_date": "10/30/2017 14:47:59",
          "content": "<p>Try ensemble of different folds from CV, or ensemble of different checkpoints of the model (last 3?). That will still be one model but a bit augmented. ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 237544,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/30/2017 14:52:46",
          "content": "<p>thanks for the reply. I know that will improve the LB score. But I am looking for a solution for single model (no ensemble of any kind) with no image  augmentation. I am wondering how to do it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 237742,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/30/2017 20:57:56",
          "content": "<p>Perhaps your classifier design is slightly different from the model that got 0.74. By this I mean the layer(s) that is/are placed on top of the ResNet-101 feature extraction layers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 237842,
          "author_name": "rteja1113",
          "author_url": "",
          "post_date": "10/31/2017 04:04:09",
          "content": "<p>I don't know about that @Human Analog, all these pretrained models are highly optimized architectures.Adding conv layers to already existing pre-trained architectures will only decrease performance. I may be wrong though</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 237909,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/31/2017 09:37:16",
          "content": "<p>I said nothing about conv layers. ;-) But when you take a pretrained network such as ResNet-101, you have to add your own classifier on top (for the 5270 classes). There are different ways to do this. You can connect a fully-connected layer directly to the last convolution layer, or you can first use global average pooling and then use a fully-connected layer, or you can use multiple fully-connected layers, or many other schemes. So just because two people say they both use ResNet-101 does not mean they also use the same classifier design on top of that network... and that is where the difference in score may be coming from.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 237926,
      "author_name": "keveri",
      "author_url": "",
      "post_date": "10/31/2017 10:31:27",
      "content": "<p>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 238333,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "11/01/2017 07:01:08",
      "content": "<p>I think maybe he makes use of the other two categories info. That is to say, he maybe divide the data into not 5270 classes first.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 241869,
      "author_name": "yanchao727",
      "author_url": "",
      "post_date": "11/10/2017 02:00:15",
      "content": "<p>I've tried your code and commented out the following code:</p>\n\n<pre><code>optimizer = optim.SGD([ # iter_accum==1\n    {'params': net.layer0.parameters(), 'lr': 0.001},\n    {'params': net.layer1.parameters(), 'lr': 0.001},\n    {'params': net.layer2.parameters(), 'lr': 0.001},\n    {'params': net.layer3.parameters(), 'lr': 0.001},\n    {'params': net.layer4.parameters(), 'lr': 0.001},\n    {'params': net.fc.parameters(),     'lr': 0.01},\n</code></pre>\n\n<p>but it seems not working , the loss is continuing to go up, should I do other things before trying the code ?\nMy output:</p>\n\n<pre><code>   rate   iter(k)  epoch   num(m)  | valid_loss/acc | train_loss/acc | batch_loss/acc |  time\n------------------------------------------------------------------------------------------------\n0.00000  243.0 k  10.13  122.79 | 1.1624  0.7296 | 0.0000  0.0000 | 0.0000  0.0000 |  0 hr 07 min\n0.00100  244.0 k  10.17  123.31 | 1.2460  0.7103 | 1.3729  0.6854 | 1.0941  0.7168 |  1 hr 00 min\n0.00100  245.0 k  10.21  123.82 | 1.2573  0.7089 | 1.2933  0.6953 | 1.1591  0.7227 |  1 hr 53 min\n0.00100  246.0 k  10.26  124.33 | 1.2772  0.7047 | 1.3211  0.6937 | 1.1279  0.7539 |  2 hr 46 min\n0.00100  247.0 k  10.30  124.84 | 1.2723  0.7060 | 1.3415  0.6931 | 1.2441  0.7090 |  3 hr 39 min\n0.00100  248.0 k  10.34  125.35 | 1.2756  0.7053 | 1.3496  0.6882 | 1.3224  0.7070 |  4 hr 32 min\n0.00100  249.0 k  10.38  125.87 | 1.2828  0.7036 | 1.3322  0.6984 | 1.4751  0.6504 |  5 hr 25 min\n0.00100  250.0 k  10.42  126.38 | 1.2861  0.7038 | 1.3474  0.6920 | 1.2393  0.7012 |  6 hr 18 min\n0.00100  251.0 k  10.47  126.89 | 1.2878  0.7031 | 1.3415  0.6977 | 1.4016  0.6797 |  7 hr 10 min\n0.00100  252.0 k  10.51  127.40 | 1.2865  0.7039 | 1.3785  0.6859 | 1.2657  0.6953 |  8 hr 03 min\n0.00100  253.0 k  10.55  127.91 | 1.2992  0.7012 | 1.3416  0.6922 | 1.3070  0.7012 |  8 hr 56 min\n0.00100  254.0 k  10.59  128.43 | 1.3002  0.6994 | 1.3434  0.6984 | 1.3961  0.6660 |  9 hr 49 min\n0.00100  255.0 k  10.64  128.94 | 1.2917  0.7025 | 1.3414  0.6944 | 1.4535  0.6758 | 10 hr 42 min\n0.00100  256.0 k  10.68  129.45 | 1.2977  0.7019 | 1.3330  0.6901 | 1.1098  0.7324 | 11 hr 35 min\n0.00100  257.0 k  10.72  129.96 | 1.2938  0.7029 | 1.3711  0.6858 | 1.2438  0.7148 | 12 hr 28 min\n0.00100  258.0 k  10.76  130.47 | 1.3145  0.6984 | 1.3411  0.6898 | 1.4022  0.6797 | 13 hr 21 min\n0.00100  259.0 k  10.80  130.99 | 1.2982  0.7012 | 1.3481  0.6931 | 1.4801  0.6602 | 14 hr 14 min\n0.00100  260.0 k  10.85  131.50 | 1.2986  0.7016 | 1.3390  0.6902 | 1.3132  0.6875 | 15 hr 07 min\n0.00100  261.0 k  10.89  132.01 | 1.3142  0.6981 | 1.3408  0.6945 | 1.3962  0.6934 | 15 hr 59 min\n0.00100  262.0 k  10.93  132.52 | 1.3067  0.7007 | 1.3624  0.6867 | 1.4176  0.6699 | 16 hr 52 min\n0.00100  263.0 k  10.97  133.03 | 1.3073  0.7002 | 1.3651  0.6900 | 1.3471  0.6758 | 17 hr 45 min\n0.00100  264.0 k  11.02  133.55 | 1.3006  0.7019 | 1.3382  0.6921 | 1.3121  0.7070 | 18 hr 38 min\n0.00100  265.0 k  11.06  134.06 | 1.3100  0.6993 | 1.3512  0.6952 | 1.3898  0.6953 | 19 hr 31 min\n0.00100  266.0 k  11.10  134.57 | 1.3145  0.6994 | 1.3197  0.6980 | 1.4913  0.6895 | 20 hr 24 min\n0.00100  267.0 k  11.14  135.08 | 1.3053  0.7000 | 1.3225  0.6900 | 1.2124  0.7285 | 21 hr 17 min\n0.00100  268.0 k  11.18  135.59 | 1.3041  0.7015 | 1.2959  0.7011 | 1.2908  0.6855 | 22 hr 10 min\n0.00100  269.0 k  11.23  136.11 | 1.3076  0.7005 | 1.2619  0.7072 | 1.2847  0.7070 | 23 hr 03 min\n0.00100  270.0 k  11.27  136.62 | 1.3160  0.6994 | 1.2833  0.7049 | 1.2478  0.7266 | 23 hr 56 min\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 241900,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/10/2017 04:23:48",
          "content": "<p>can you post you code(just the trainer.py) and the whole log file?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 241902,
          "author_name": "yanchao727",
          "author_url": "",
          "post_date": "11/10/2017 04:37:16",
          "content": "<p>ok, trainer.py : <a href=\"https://pastebin.com/QTySsgHf\">https://pastebin.com/QTySsgHf</a>\nlog_file: <a href=\"https://pastebin.com/n3AxGrsR\">https://pastebin.com/n3AxGrsR</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 241903,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/10/2017 04:44:16",
          "content": "<p>note that your train and validation set is not the same as mine.  your results is actually reasonable. the validation score should be close to LB score, which is 0.70.</p>\n\n<p>you can try to use smaller learning rate. But i think the limit is around to 0.70 to 0.71.</p>\n\n<pre><code>if 1:\n    optimizer = optim.SGD(filter(lambda p: p.requires_grad, net.parameters()),\n                          lr=0.0001, momentum=0.9, weight_decay=0.0001) #nesterov=True\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 241905,
          "author_name": "yanchao727",
          "author_url": "",
          "post_date": "11/10/2017 04:51:07",
          "content": "<p>But my LB commit was 0.64 which was far from 0.70 .\nHaven't you find anything wrong elsewhere ? \nI guess I should take a  close look.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242119,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/10/2017 17:03:06",
          "content": "<p>i suspect some bug in your submission code. check the normalization of input at submission. it should be same as training. also note that there are 1 to 4 images per product.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 242955,
          "author_name": "yanchao727",
          "author_url": "",
          "post_date": "11/13/2017 03:24:03",
          "content": "<p>Oh, can I  use the submission functions in your trainer_resnet50.py code directly ? \nI couldn't find any normalization except \"SequentialSampler\" ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 244675,
      "author_name": "zavadskyy",
      "author_url": "",
      "post_date": "11/16/2017 17:48:30",
      "content": "<p>I'm having a hard time trying to reproduce your validation score. My best guess is that the problem in image preprocessing, my current code for it looks like this:</p>\n\n<pre><code>mean = [0.485, 0.456, 0.406]\nstd  = [0.229, 0.224, 0.225]\nimg = np.array(img).astype(np.float32)\nimg = fix_center_crop(img, size=(160,160))  \nimg = img/255.\nfor i in range(3):\n    img[:, i] = (img[:, i]-mean[i])/std[i]\nimg = img.transpose((2,0,1))\n</code></pre>\n\n<p>And I'm getting about 0.316 accuracy on validation set.\nAny idea what I'm missing?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 250145,
      "author_name": "steelrose",
      "author_url": "",
      "post_date": "11/29/2017 21:24:41",
      "content": "<p>thank you very much, you seem to use some class \"CDiscountDataset\" but I don't see in in the code you have on google ... can you please share it? (i'm not familiar with pytorch, so not sure what exactly should be there and whether it's basically just the code to read from bson, or also some augmentations etc.)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "235935": "Thanks to @Vladimir Iglovikov for this discussion thread at https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41652.\n\nHere, I try to repeat his results:\n\n  - val_loss = 1.18\n  - val_acc = 0.744\n  - LB = 0.7437 (0.69 after 7th epoch)\n\n.\n\n\n  - 4 x GTX 1080 Ti\n  - 18 epochs x 4.5 hours\n  - One epoch - 11738624 images\n  - batch size = 512\n\nSo far below is my results (LB 0.704). It seems that i have overfitting and my results are not as good as @Vladimir Iglovikov. You may want to try other train parameters and better augmentation. I need to review my training pipeline and try second time.\n\n  ![enter image description here][1]\n\nMy code is based on pytorch. It runs at 4 hr 27 min per 11.2 million images. \n\nYou can download my code and trained  model at : https://drive.google.com/drive/folders/0B_DICebvRE-kb2dFd2FKX1hfRkE?usp=sharing (see folder \"10-26\").\n\nYou should be able to run it. The imagenet pretrain model is from pytorch model zoo. Please also refer to: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/235935/7750/results.png",
    "236128": "Thanks! I am not familiar with pythorch, so I could not find this in your code: can you tell how you managed to change the default picture size from 224X224 to 180X180 or 160X160 in your case and still use the pre training weigths? Did you change some thing else besides the input layer? I am using a prerained model of resnet-101 in Keras, and can't make it work when change the input size to less than 224X224.",
    "236157": "Pytorch does not check for input size in convolutional layers afaik, just like it should be. A convolution can be applied to arbitrary image size, so there is no need to change the sizes. Since most architectures nowadays use a global pooling layers at the end, the fully connected layers receive an input of the size of number of channels of the last convolutional layer, which is independent of feature map dimensions.",
    "236167": "The problem occurs in the MaxPooling2D just after the first batch normalization.",
    "237111": "Thanks again CK! Any chance that you can also share the `time_to_str`? ;P",
    "237118": "def time_to_str(t):\n    t  = int(t)\n    hr = t//60\n    min = t%60\n    return '%2d hr %02d min'%(hr,min)",
    "237521": "After trying several days, it seems impossible for me to get 0.74. My best results is about LB 0.71 (vanilla resnet101  image prediction), but this is after lots of  efforts in parameter tuning. \n\nBut I believe that it is possible to get about 0.73 to 0.74 with single model like resnet101. Anyone has any suggestions that this can be done?\n\nI already try different augmentation, droupout, different batch size, learning rate, momentum.\n\nIf there is no special tricks to get 0.74, then i suspect my multi-gpu training code could be buggy.",
    "237540": "Try ensemble of different folds from CV, or ensemble of different checkpoints of the model (last 3?). That will still be one model but a bit augmented. ;)",
    "237544": "thanks for the reply. I know that will improve the LB score. But I am looking for a solution for single model (no ensemble of any kind) with no image  augmentation. I am wondering how to do it.",
    "237742": "Perhaps your classifier design is slightly different from the model that got 0.74. By this I mean the layer(s) that is/are placed on top of the ResNet-101 feature extraction layers.",
    "237842": "I don't know about that @Human Analog, all these pretrained models are highly optimized architectures.Adding conv layers to already existing pre-trained architectures will only decrease performance. I may be wrong though",
    "237909": "I said nothing about conv layers. ;-) But when you take a pretrained network such as ResNet-101, you have to add your own classifier on top (for the 5270 classes). There are different ways to do this. You can connect a fully-connected layer directly to the last convolution layer, or you can first use global average pooling and then use a fully-connected layer, or you can use multiple fully-connected layers, or many other schemes. So just because two people say they both use ResNet-101 does not mean they also use the same classifier design on top of that network... and that is where the difference in score may be coming from.",
    "237926": ".",
    "238333": "I think maybe he makes use of the other two categories info. That is to say, he maybe divide the data into not 5270 classes first.",
    "241869": "I've tried your code and commented out the following code:\n\n    optimizer = optim.SGD([ # iter_accum==1\n        {'params': net.layer0.parameters(), 'lr': 0.001},\n        {'params': net.layer1.parameters(), 'lr': 0.001},\n        {'params': net.layer2.parameters(), 'lr': 0.001},\n        {'params': net.layer3.parameters(), 'lr': 0.001},\n        {'params': net.layer4.parameters(), 'lr': 0.001},\n        {'params': net.fc.parameters(),     'lr': 0.01},\n\nbut it seems not working , the loss is continuing to go up, should I do other things before trying the code ?\nMy output:\n\n       rate   iter(k)  epoch   num(m)  | valid_loss/acc | train_loss/acc | batch_loss/acc |  time\n    ------------------------------------------------------------------------------------------------\n    0.00000  243.0 k  10.13  122.79 | 1.1624  0.7296 | 0.0000  0.0000 | 0.0000  0.0000 |  0 hr 07 min\n    0.00100  244.0 k  10.17  123.31 | 1.2460  0.7103 | 1.3729  0.6854 | 1.0941  0.7168 |  1 hr 00 min\n    0.00100  245.0 k  10.21  123.82 | 1.2573  0.7089 | 1.2933  0.6953 | 1.1591  0.7227 |  1 hr 53 min\n    0.00100  246.0 k  10.26  124.33 | 1.2772  0.7047 | 1.3211  0.6937 | 1.1279  0.7539 |  2 hr 46 min\n    0.00100  247.0 k  10.30  124.84 | 1.2723  0.7060 | 1.3415  0.6931 | 1.2441  0.7090 |  3 hr 39 min\n    0.00100  248.0 k  10.34  125.35 | 1.2756  0.7053 | 1.3496  0.6882 | 1.3224  0.7070 |  4 hr 32 min\n    0.00100  249.0 k  10.38  125.87 | 1.2828  0.7036 | 1.3322  0.6984 | 1.4751  0.6504 |  5 hr 25 min\n    0.00100  250.0 k  10.42  126.38 | 1.2861  0.7038 | 1.3474  0.6920 | 1.2393  0.7012 |  6 hr 18 min\n    0.00100  251.0 k  10.47  126.89 | 1.2878  0.7031 | 1.3415  0.6977 | 1.4016  0.6797 |  7 hr 10 min\n    0.00100  252.0 k  10.51  127.40 | 1.2865  0.7039 | 1.3785  0.6859 | 1.2657  0.6953 |  8 hr 03 min\n    0.00100  253.0 k  10.55  127.91 | 1.2992  0.7012 | 1.3416  0.6922 | 1.3070  0.7012 |  8 hr 56 min\n    0.00100  254.0 k  10.59  128.43 | 1.3002  0.6994 | 1.3434  0.6984 | 1.3961  0.6660 |  9 hr 49 min\n    0.00100  255.0 k  10.64  128.94 | 1.2917  0.7025 | 1.3414  0.6944 | 1.4535  0.6758 | 10 hr 42 min\n    0.00100  256.0 k  10.68  129.45 | 1.2977  0.7019 | 1.3330  0.6901 | 1.1098  0.7324 | 11 hr 35 min\n    0.00100  257.0 k  10.72  129.96 | 1.2938  0.7029 | 1.3711  0.6858 | 1.2438  0.7148 | 12 hr 28 min\n    0.00100  258.0 k  10.76  130.47 | 1.3145  0.6984 | 1.3411  0.6898 | 1.4022  0.6797 | 13 hr 21 min\n    0.00100  259.0 k  10.80  130.99 | 1.2982  0.7012 | 1.3481  0.6931 | 1.4801  0.6602 | 14 hr 14 min\n    0.00100  260.0 k  10.85  131.50 | 1.2986  0.7016 | 1.3390  0.6902 | 1.3132  0.6875 | 15 hr 07 min\n    0.00100  261.0 k  10.89  132.01 | 1.3142  0.6981 | 1.3408  0.6945 | 1.3962  0.6934 | 15 hr 59 min\n    0.00100  262.0 k  10.93  132.52 | 1.3067  0.7007 | 1.3624  0.6867 | 1.4176  0.6699 | 16 hr 52 min\n    0.00100  263.0 k  10.97  133.03 | 1.3073  0.7002 | 1.3651  0.6900 | 1.3471  0.6758 | 17 hr 45 min\n    0.00100  264.0 k  11.02  133.55 | 1.3006  0.7019 | 1.3382  0.6921 | 1.3121  0.7070 | 18 hr 38 min\n    0.00100  265.0 k  11.06  134.06 | 1.3100  0.6993 | 1.3512  0.6952 | 1.3898  0.6953 | 19 hr 31 min\n    0.00100  266.0 k  11.10  134.57 | 1.3145  0.6994 | 1.3197  0.6980 | 1.4913  0.6895 | 20 hr 24 min\n    0.00100  267.0 k  11.14  135.08 | 1.3053  0.7000 | 1.3225  0.6900 | 1.2124  0.7285 | 21 hr 17 min\n    0.00100  268.0 k  11.18  135.59 | 1.3041  0.7015 | 1.2959  0.7011 | 1.2908  0.6855 | 22 hr 10 min\n    0.00100  269.0 k  11.23  136.11 | 1.3076  0.7005 | 1.2619  0.7072 | 1.2847  0.7070 | 23 hr 03 min\n    0.00100  270.0 k  11.27  136.62 | 1.3160  0.6994 | 1.2833  0.7049 | 1.2478  0.7266 | 23 hr 56 min",
    "241900": "can you post you code(just the trainer.py) and the whole log file?",
    "241902": "ok, trainer.py : [https://pastebin.com/QTySsgHf][1]\nlog_file: https://pastebin.com/n3AxGrsR\n\n  [1]: https://pastebin.com/QTySsgHf",
    "241903": "note that your train and validation set is not the same as mine.  your results is actually reasonable. the validation score should be close to LB score, which is 0.70.\n\nyou can try to use smaller learning rate. But i think the limit is around to 0.70 to 0.71.\n\n    if 1:\n        optimizer = optim.SGD(filter(lambda p: p.requires_grad, net.parameters()),\n                              lr=0.0001, momentum=0.9, weight_decay=0.0001) #nesterov=True",
    "241905": "But my LB commit was 0.64 which was far from 0.70 .\nHaven't you find anything wrong elsewhere ? \nI guess I should take a  close look.",
    "242119": "i suspect some bug in your submission code. check the normalization of input at submission. it should be same as training. also note that there are 1 to 4 images per product.",
    "242955": "Oh, can I  use the submission functions in your trainer_resnet50.py code directly ? \nI couldn't find any normalization except \"SequentialSampler\" ?",
    "244675": "I'm having a hard time trying to reproduce your validation score. My best guess is that the problem in image preprocessing, my current code for it looks like this:\n\n    mean = [0.485, 0.456, 0.406]\n    std  = [0.229, 0.224, 0.225]\n    img = np.array(img).astype(np.float32)\n    img = fix_center_crop(img, size=(160,160))  \n    img = img/255.\n    for i in range(3):\n        img[:, i] = (img[:, i]-mean[i])/std[i]\n    img = img.transpose((2,0,1))\n\nAnd I'm getting about 0.316 accuracy on validation set.\nAny idea what I'm missing?",
    "250145": "thank you very much, you seem to use some class \"CDiscountDataset\" but I don't see in in the code you have on google ... can you please share it? (i'm not familiar with pytorch, so not sure what exactly should be there and whether it's basically just the code to read from bson, or also some augmentations etc.)"
  },
  "source": "meta"
}