{
  "id": 40934,
  "title": "Reusing model",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/40934",
  "author_name": "",
  "post_date": "2017-10-10T13:34:10.132532900Z",
  "votes": 10,
  "comment_count": 20,
  "views": 0,
  "content": "<p>yet another \"trick\". Also, i find that inception3 on 180x180 gives the best results. LB 0.69565 for train/test = 180x180 (single crop).</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/229785/7587/reuse_11.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/229785/7588/reuse_12.png\" alt=\"enter image description here\" title=\"\"></p>",
  "messages": [
    {
      "id": "229785",
      "postDate": "10/10/2017 13:34:10",
      "content": "<p>yet another \"trick\". Also, i find that inception3 on 180x180 gives the best results. LB 0.69565 for train/test = 180x180 (single crop).</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/229785/7587/reuse_11.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/229785/7588/reuse_12.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "yet another \"trick\". Also, i find that inception3 on 180x180 gives the best results. LB 0.69565 for train/test = 180x180 (single crop).\n\n  ![enter image description here][1]\n\n\n  ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/229785/7587/reuse_11.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/229785/7588/reuse_12.png",
      "votes": null
    },
    {
      "id": "230041",
      "postDate": "10/11/2017 05:42:15",
      "content": "<p>So you train all parameters? </p>",
      "rawMarkdown": "So you train all parameters?",
      "votes": null
    },
    {
      "id": "230044",
      "postDate": "10/11/2017 05:59:43",
      "content": "<p>yes, i train all parameters. </p>\n\n<p>but do note that there are no fixed rules on how to finetune models. it is by trial and error. </p>",
      "rawMarkdown": "yes, i train all parameters. \n\nbut do note that there are no fixed rules on how to finetune models. it is by trial and error.",
      "votes": null
    },
    {
      "id": "230045",
      "postDate": "10/11/2017 06:03:42",
      "content": "<p>How long do you train them in one epoch? I spend lot's of time trying to reduce the time but it didn't work. Now I still need 11 hours to train one epoch and just training few params (1080, not TI). </p>",
      "rawMarkdown": "How long do you train them in one epoch? I spend lot's of time trying to reduce the time but it didn't work. Now I still need 11 hours to train one epoch and just training few params (1080, not TI).",
      "votes": null
    },
    {
      "id": "230050",
      "postDate": "10/11/2017 06:07:42",
      "content": "<p>one epoch is about 10 to 11 hrs (1080 ti). </p>\n\n<p>the trick is to plot your loss and accuracy per xxx iteratrions. Adjust momentum, rate, batch size, augmentation, etc as it goes. you must find ways to make speedup.</p>",
      "rawMarkdown": "one epoch is about 10 to 11 hrs (1080 ti). \n\nthe trick is to plot your loss and accuracy per xxx iteratrions. Adjust momentum, rate, batch size, augmentation, etc as it goes. you must find ways to make speedup.",
      "votes": null
    },
    {
      "id": "230052",
      "postDate": "10/11/2017 06:13:39",
      "content": "<p>Thx, the tricks you said are highly related to the result, but are they really related to the training speed? Now I think I have lot's of ideas to experiment but I can't do it only because of the time. I'm using SSD, multiple works, but the speed is still really slow...</p>",
      "rawMarkdown": "Thx, the tricks you said are highly related to the result, but are they really related to the training speed? Now I think I have lot's of ideas to experiment but I can't do it only because of the time. I'm using SSD, multiple works, but the speed is still really slow...",
      "votes": null
    },
    {
      "id": "230054",
      "postDate": "10/11/2017 06:17:32",
      "content": "<p>Yes. I will have some writeup later. Eg reducing momentum will reduce convergence time by half. You need to estimate the most efficient parameters. Eg best parameters gives accuracy of + 0.0xx per epoch. So you don't have to train too many epoch at first</p>",
      "rawMarkdown": "Yes. I will have some writeup later. Eg reducing momentum will reduce convergence time by half. You need to estimate the most efficient parameters. Eg best parameters gives accuracy of + 0.0xx per epoch. So you don't have to train too many epoch at first",
      "votes": null
    },
    {
      "id": "230055",
      "postDate": "10/11/2017 06:22:13",
      "content": "<p>You're right, but I'm really concern about reducing the training time in each epoch, not the number of epochs. Maybe I should try your code first.</p>",
      "rawMarkdown": "You're right, but I'm really concern about reducing the training time in each epoch, not the number of epochs. Maybe I should try your code first.",
      "votes": null
    },
    {
      "id": "230062",
      "postDate": "10/11/2017 06:44:46",
      "content": "<p>oh. I suspect it is augmentation by cpu that causes the bottleneck. I suggest you set up the following systems to compare timing.</p>\n\n<ol>\n<li><p>create dummy data in gpu. run training iterations using the same dummy data over and over agin. everything is in gpu. this is your gpu throughput speed.</p></li>\n<li><p>create dummy data in RAM. you need to transfer data from cpu to gpu. but there is no augmentation (i.e. no work by cpu except data transfer). this measures your ram to gpu transfer bottleneck.</p></li>\n<li><p>load data from SSD. but no augmentation</p></li>\n<li><p>load data from SSD. do some augmentation.</p></li>\n</ol>",
      "rawMarkdown": "oh. I suspect it is augmentation by cpu that causes the bottleneck. I suggest you set up the following systems to compare timing.\n\n1. create dummy data in gpu. run training iterations using the same dummy data over and over agin. everything is in gpu. this is your gpu throughput speed.\n\n2. create dummy data in RAM. you need to transfer data from cpu to gpu. but there is no augmentation (i.e. no work by cpu except data transfer). this measures your ram to gpu transfer bottleneck.\n\n3. load data from SSD. but no augmentation\n\n4. load data from SSD. do some augmentation.",
      "votes": null
    },
    {
      "id": "230115",
      "postDate": "10/11/2017 10:22:31",
      "content": "<p>Do you initialize SE-Inception3 180x180 from a fine-tuned Inception3 180x180? maybe I misunderstand your image...</p>",
      "rawMarkdown": "Do you initialize SE-Inception3 180x180 from a fine-tuned Inception3 180x180? maybe I misunderstand your image...",
      "votes": null
    },
    {
      "id": "230117",
      "postDate": "10/11/2017 10:23:28",
      "content": "<p>yes., the convolution weights are copied. But the missing SE weights are randomly set.  See:  <a href=\"https://github.com/moskomule/senet.pytorch/blob/master/se_inception.py\">https://github.com/moskomule/senet.pytorch/blob/master/se_inception.py</a></p>",
      "rawMarkdown": "yes., the convolution weights are copied. But the missing SE weights are randomly set.  See:  https://github.com/moskomule/senet.pytorch/blob/master/se_inception.py",
      "votes": null
    },
    {
      "id": "230118",
      "postDate": "10/11/2017 10:26:22",
      "content": "<p>se-inception3</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/230118/7589/se-inception3.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "se-inception3\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/230118/7589/se-inception3.png",
      "votes": null
    },
    {
      "id": "230128",
      "postDate": "10/11/2017 11:00:19",
      "content": "<p>I presume it works just because you train all layers so the SE weights are learned on the fly and the dataset is large enough so they have time to settle to good values.</p>",
      "rawMarkdown": "I presume it works just because you train all layers so the SE weights are learned on the fly and the dataset is large enough so they have time to settle to good values.",
      "votes": null
    },
    {
      "id": "230130",
      "postDate": "10/11/2017 11:06:49",
      "content": "<p>SE actually only performs very little better in my experiments. I am not sure if my results are correct. There is no results of inception3 in the paper (but there are results of other inception).</p>\n\n<p>It works because basically SE operation is:  out = scale*out + residual.</p>\n\n<p>scale = sigmoid( ....)</p>\n\n<p>if i start with random SE weights, scale will be about constant 0.5 at start. (maybe it is better to add a x2 multipiler). then the weights can be adjusted by training later. So your comment is more or less correct.</p>",
      "rawMarkdown": "SE actually only performs very little better in my experiments. I am not sure if my results are correct. There is no results of inception3 in the paper (but there are results of other inception).\n\nIt works because basically SE operation is:  out = scale*out + residual.\n\nscale = sigmoid( ....)\n\nif i start with random SE weights, scale will be about constant 0.5 at start. (maybe it is better to add a x2 multipiler). then the weights can be adjusted by training later. So your comment is more or less correct.",
      "votes": null
    },
    {
      "id": "230253",
      "postDate": "10/11/2017 16:00:55",
      "content": "<p>Also refer to the vgg paper\n<a href=\"https://arxiv.org/pdf/1409.1556.pdf\">https://arxiv.org/pdf/1409.1556.pdf</a></p>",
      "rawMarkdown": "Also refer to the vgg paper\nhttps://arxiv.org/pdf/1409.1556.pdf",
      "votes": null
    },
    {
      "id": "230301",
      "postDate": "10/11/2017 17:17:11",
      "content": "<p>Great idea, thanks.</p>\n\n<p>One thing I would like to ask you. Did you use the RMSprop for optimizer like Inception v3 paper?</p>",
      "rawMarkdown": "Great idea, thanks.\n\nOne thing I would like to ask you. Did you use the RMSprop for optimizer like Inception v3 paper?",
      "votes": null
    },
    {
      "id": "230302",
      "postDate": "10/11/2017 17:17:53",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "230306",
      "postDate": "10/11/2017 17:23:59",
      "content": "<p>No. I use sdg+momentum. i have released the code and trained model. you can try experiments of different optimiser.</p>\n\n<p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a></p>",
      "rawMarkdown": "No. I use sdg+momentum. i have released the code and trained model. you can try experiments of different optimiser.\n\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\n\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498",
      "votes": null
    },
    {
      "id": "230433",
      "postDate": "10/11/2017 23:54:54",
      "content": "<p>Thanks a lot! I will profile my code and see if something are wrong.</p>\n\n<p>By the way, do you have any experiments for training same epoch on SE-net directly and reusing non-SE-net to SE-net? (i.e., training 2epoch on SE-net, or 1 epoch on non-SE-net and than use their weights to train 1 epoch on SE-net)</p>",
      "rawMarkdown": "Thanks a lot! I will profile my code and see if something are wrong.\n\nBy the way, do you have any experiments for training same epoch on SE-net directly and reusing non-SE-net to SE-net? (i.e., training 2epoch on SE-net, or 1 epoch on non-SE-net and than use their weights to train 1 epoch on SE-net)",
      "votes": null
    },
    {
      "id": "230437",
      "postDate": "10/12/2017 00:00:20",
      "content": "<p>my models are released at: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021</a></p>\n\n<p>you can do your experiments from there.</p>",
      "rawMarkdown": "my models are released at: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\n\nyou can do your experiments from there.",
      "votes": null
    },
    {
      "id": "230439",
      "postDate": "10/12/2017 00:02:45",
      "content": "<p>Ok, I will do it. Thanks again!</p>",
      "rawMarkdown": "Ok, I will do it. Thanks again!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 230041,
      "author_name": "brianlzm",
      "author_url": "",
      "post_date": "10/11/2017 05:42:15",
      "content": "<p>So you train all parameters? </p>",
      "votes": null,
      "replies": [
        {
          "id": 230044,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 05:59:43",
          "content": "<p>yes, i train all parameters. </p>\n\n<p>but do note that there are no fixed rules on how to finetune models. it is by trial and error. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230045,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/11/2017 06:03:42",
          "content": "<p>How long do you train them in one epoch? I spend lot's of time trying to reduce the time but it didn't work. Now I still need 11 hours to train one epoch and just training few params (1080, not TI). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230050,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 06:07:42",
          "content": "<p>one epoch is about 10 to 11 hrs (1080 ti). </p>\n\n<p>the trick is to plot your loss and accuracy per xxx iteratrions. Adjust momentum, rate, batch size, augmentation, etc as it goes. you must find ways to make speedup.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230052,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/11/2017 06:13:39",
          "content": "<p>Thx, the tricks you said are highly related to the result, but are they really related to the training speed? Now I think I have lot's of ideas to experiment but I can't do it only because of the time. I'm using SSD, multiple works, but the speed is still really slow...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230054,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 06:17:32",
          "content": "<p>Yes. I will have some writeup later. Eg reducing momentum will reduce convergence time by half. You need to estimate the most efficient parameters. Eg best parameters gives accuracy of + 0.0xx per epoch. So you don't have to train too many epoch at first</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230055,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/11/2017 06:22:13",
          "content": "<p>You're right, but I'm really concern about reducing the training time in each epoch, not the number of epochs. Maybe I should try your code first.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230062,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 06:44:46",
          "content": "<p>oh. I suspect it is augmentation by cpu that causes the bottleneck. I suggest you set up the following systems to compare timing.</p>\n\n<ol>\n<li><p>create dummy data in gpu. run training iterations using the same dummy data over and over agin. everything is in gpu. this is your gpu throughput speed.</p></li>\n<li><p>create dummy data in RAM. you need to transfer data from cpu to gpu. but there is no augmentation (i.e. no work by cpu except data transfer). this measures your ram to gpu transfer bottleneck.</p></li>\n<li><p>load data from SSD. but no augmentation</p></li>\n<li><p>load data from SSD. do some augmentation.</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230433,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/11/2017 23:54:54",
          "content": "<p>Thanks a lot! I will profile my code and see if something are wrong.</p>\n\n<p>By the way, do you have any experiments for training same epoch on SE-net directly and reusing non-SE-net to SE-net? (i.e., training 2epoch on SE-net, or 1 epoch on non-SE-net and than use their weights to train 1 epoch on SE-net)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230437,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/12/2017 00:00:20",
          "content": "<p>my models are released at: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021</a></p>\n\n<p>you can do your experiments from there.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230439,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/12/2017 00:02:45",
          "content": "<p>Ok, I will do it. Thanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 230115,
      "author_name": "vzsolt",
      "author_url": "",
      "post_date": "10/11/2017 10:22:31",
      "content": "<p>Do you initialize SE-Inception3 180x180 from a fine-tuned Inception3 180x180? maybe I misunderstand your image...</p>",
      "votes": null,
      "replies": [
        {
          "id": 230117,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 10:23:28",
          "content": "<p>yes., the convolution weights are copied. But the missing SE weights are randomly set.  See:  <a href=\"https://github.com/moskomule/senet.pytorch/blob/master/se_inception.py\">https://github.com/moskomule/senet.pytorch/blob/master/se_inception.py</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230128,
          "author_name": "vzsolt",
          "author_url": "",
          "post_date": "10/11/2017 11:00:19",
          "content": "<p>I presume it works just because you train all layers so the SE weights are learned on the fly and the dataset is large enough so they have time to settle to good values.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230130,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 11:06:49",
          "content": "<p>SE actually only performs very little better in my experiments. I am not sure if my results are correct. There is no results of inception3 in the paper (but there are results of other inception).</p>\n\n<p>It works because basically SE operation is:  out = scale*out + residual.</p>\n\n<p>scale = sigmoid( ....)</p>\n\n<p>if i start with random SE weights, scale will be about constant 0.5 at start. (maybe it is better to add a x2 multipiler). then the weights can be adjusted by training later. So your comment is more or less correct.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 230118,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/11/2017 10:26:22",
      "content": "<p>se-inception3</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/230118/7589/se-inception3.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 230253,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/11/2017 16:00:55",
      "content": "<p>Also refer to the vgg paper\n<a href=\"https://arxiv.org/pdf/1409.1556.pdf\">https://arxiv.org/pdf/1409.1556.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 230301,
      "author_name": "lyakaap",
      "author_url": "",
      "post_date": "10/11/2017 17:17:11",
      "content": "<p>Great idea, thanks.</p>\n\n<p>One thing I would like to ask you. Did you use the RMSprop for optimizer like Inception v3 paper?</p>",
      "votes": null,
      "replies": [
        {
          "id": 230306,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 17:23:59",
          "content": "<p>No. I use sdg+momentum. i have released the code and trained model. you can try experiments of different optimiser.</p>\n\n<p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 230302,
      "author_name": "",
      "author_url": "",
      "post_date": "10/11/2017 17:17:53",
      "content": "",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "229785": "yet another \"trick\". Also, i find that inception3 on 180x180 gives the best results. LB 0.69565 for train/test = 180x180 (single crop).\n\n  ![enter image description here][1]\n\n\n  ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/229785/7587/reuse_11.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/229785/7588/reuse_12.png",
    "230041": "So you train all parameters?",
    "230044": "yes, i train all parameters. \n\nbut do note that there are no fixed rules on how to finetune models. it is by trial and error.",
    "230045": "How long do you train them in one epoch? I spend lot's of time trying to reduce the time but it didn't work. Now I still need 11 hours to train one epoch and just training few params (1080, not TI).",
    "230050": "one epoch is about 10 to 11 hrs (1080 ti). \n\nthe trick is to plot your loss and accuracy per xxx iteratrions. Adjust momentum, rate, batch size, augmentation, etc as it goes. you must find ways to make speedup.",
    "230052": "Thx, the tricks you said are highly related to the result, but are they really related to the training speed? Now I think I have lot's of ideas to experiment but I can't do it only because of the time. I'm using SSD, multiple works, but the speed is still really slow...",
    "230054": "Yes. I will have some writeup later. Eg reducing momentum will reduce convergence time by half. You need to estimate the most efficient parameters. Eg best parameters gives accuracy of + 0.0xx per epoch. So you don't have to train too many epoch at first",
    "230055": "You're right, but I'm really concern about reducing the training time in each epoch, not the number of epochs. Maybe I should try your code first.",
    "230062": "oh. I suspect it is augmentation by cpu that causes the bottleneck. I suggest you set up the following systems to compare timing.\n\n1. create dummy data in gpu. run training iterations using the same dummy data over and over agin. everything is in gpu. this is your gpu throughput speed.\n\n2. create dummy data in RAM. you need to transfer data from cpu to gpu. but there is no augmentation (i.e. no work by cpu except data transfer). this measures your ram to gpu transfer bottleneck.\n\n3. load data from SSD. but no augmentation\n\n4. load data from SSD. do some augmentation.",
    "230115": "Do you initialize SE-Inception3 180x180 from a fine-tuned Inception3 180x180? maybe I misunderstand your image...",
    "230117": "yes., the convolution weights are copied. But the missing SE weights are randomly set.  See:  https://github.com/moskomule/senet.pytorch/blob/master/se_inception.py",
    "230118": "se-inception3\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/230118/7589/se-inception3.png",
    "230128": "I presume it works just because you train all layers so the SE weights are learned on the fly and the dataset is large enough so they have time to settle to good values.",
    "230130": "SE actually only performs very little better in my experiments. I am not sure if my results are correct. There is no results of inception3 in the paper (but there are results of other inception).\n\nIt works because basically SE operation is:  out = scale*out + residual.\n\nscale = sigmoid( ....)\n\nif i start with random SE weights, scale will be about constant 0.5 at start. (maybe it is better to add a x2 multipiler). then the weights can be adjusted by training later. So your comment is more or less correct.",
    "230253": "Also refer to the vgg paper\nhttps://arxiv.org/pdf/1409.1556.pdf",
    "230301": "Great idea, thanks.\n\nOne thing I would like to ask you. Did you use the RMSprop for optimizer like Inception v3 paper?",
    "230302": "",
    "230306": "No. I use sdg+momentum. i have released the code and trained model. you can try experiments of different optimiser.\n\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\n\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40498",
    "230433": "Thanks a lot! I will profile my code and see if something are wrong.\n\nBy the way, do you have any experiments for training same epoch on SE-net directly and reusing non-SE-net to SE-net? (i.e., training 2epoch on SE-net, or 1 epoch on non-SE-net and than use their weights to train 1 epoch on SE-net)",
    "230437": "my models are released at: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\n\nyou can do your experiments from there.",
    "230439": "Ok, I will do it. Thanks again!"
  },
  "source": "meta"
}