{
  "id": 45709,
  "title": "[ my ensemble solution  ]",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/45709",
  "author_name": "",
  "post_date": "2017-12-14T23:55:10.681602100Z",
  "votes": 49,
  "comment_count": 20,
  "views": 0,
  "content": "<p><strong>Updated! please refer to the ppt and xls in the attachment for full details.</strong></p>\n\n<hr>\n\n<p>I present my solution , which is part of team \"Convoluted predictions\". I submitted an ensemble  of 6 models (combine.0, combine.1, combine.2 combine.5, combine.5-max, combine.5-mix). This ensemble has LB=0.787.</p>\n\n<p>There is one model that is discarded and not used (combine.4, single base network), but nevertheless gives good results.</p>\n\n<p>Credits go to my teammate @Miha Skalic who discover the fast way to train the model in 2 stages. By training on the extracted features, it is much more faster and accurate!</p>\n\n<p>.</p>\n\n<p>First, I introduce combine.2. It uses gated scaling in the combination network fcnet3.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8063/c2.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8064/fcnet3.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>How to train fcnet3:</p>\n\n<ol>\n<li>Train the base networks (se-resent101 and xception) image wise.</li>\n<li>Extract the features for each train image  and store in disk (quantised to np.uint8 and store as memmap)</li>\n<li>train fcnet3 product wise.</li>\n</ol>\n\n<p>Tricks in training:</p>\n\n<ol>\n<li><p>for base networks , use dropout and use augmentation=small shift, scale, rotation + flip + rotate90,180,270. Batch_size = 512 to 768, num_epoches =  about 20. SDG/learning rate=0.01,0.001,0.0001 (manually tuned)</p>\n\n<p>time to train (4x1080Ti) = 5 days per base networks</p></li>\n<li><p>for fcnet3, use dropout and i only use augmentation=flip (due to lack of space in my SSD drive). Batch_size = 4096, num_epoches = about 100.SDG/learning rate=0.01,0.001,0.0001 (manually tuned)</p>\n\n<p>time to train (4x1080Ti) = 1.5 days  </p></li>\n</ol>\n\n<hr>\n\n<p>Other variants (to be updated)</p>\n\n<p>a. other base network</p>\n\n<p>I make combine.0 and combine.1 as shown below. These differs in dropout probability. Initially, i intended to use only one of them in final submission. But surprisingly, using both of them is better than using one. (you note that these are exact same models that i post in another post:  <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021</a>) . With my gated Fcnet3, you can use low performance base network in the range of LB=0.69 to 0.71.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8069/combine.0.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8070/combine.1.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8126/graph1.png\" alt=\"enter image description here\" title=\"\">\n (loss curve for combine.1)</p>\n\n<hr>\n\n<p>b. fcnet1  : max to replace scaling</p>\n\n<p>c. fcnet0 : fcnet1   + mixup augmentation</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8124/graph3.png\" alt=\"enter image description here\" title=\"\">\n   <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8125/graph2.png\" alt=\"enter image description here\" title=\"\"></p>",
  "messages": [
    {
      "id": "257777",
      "postDate": "12/14/2017 23:55:10",
      "content": "<p><strong>Updated! please refer to the ppt and xls in the attachment for full details.</strong></p>\n\n<hr>\n\n<p>I present my solution , which is part of team \"Convoluted predictions\". I submitted an ensemble  of 6 models (combine.0, combine.1, combine.2 combine.5, combine.5-max, combine.5-mix). This ensemble has LB=0.787.</p>\n\n<p>There is one model that is discarded and not used (combine.4, single base network), but nevertheless gives good results.</p>\n\n<p>Credits go to my teammate @Miha Skalic who discover the fast way to train the model in 2 stages. By training on the extracted features, it is much more faster and accurate!</p>\n\n<p>.</p>\n\n<p>First, I introduce combine.2. It uses gated scaling in the combination network fcnet3.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8063/c2.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8064/fcnet3.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>How to train fcnet3:</p>\n\n<ol>\n<li>Train the base networks (se-resent101 and xception) image wise.</li>\n<li>Extract the features for each train image  and store in disk (quantised to np.uint8 and store as memmap)</li>\n<li>train fcnet3 product wise.</li>\n</ol>\n\n<p>Tricks in training:</p>\n\n<ol>\n<li><p>for base networks , use dropout and use augmentation=small shift, scale, rotation + flip + rotate90,180,270. Batch_size = 512 to 768, num_epoches =  about 20. SDG/learning rate=0.01,0.001,0.0001 (manually tuned)</p>\n\n<p>time to train (4x1080Ti) = 5 days per base networks</p></li>\n<li><p>for fcnet3, use dropout and i only use augmentation=flip (due to lack of space in my SSD drive). Batch_size = 4096, num_epoches = about 100.SDG/learning rate=0.01,0.001,0.0001 (manually tuned)</p>\n\n<p>time to train (4x1080Ti) = 1.5 days  </p></li>\n</ol>\n\n<hr>\n\n<p>Other variants (to be updated)</p>\n\n<p>a. other base network</p>\n\n<p>I make combine.0 and combine.1 as shown below. These differs in dropout probability. Initially, i intended to use only one of them in final submission. But surprisingly, using both of them is better than using one. (you note that these are exact same models that i post in another post:  <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021</a>) . With my gated Fcnet3, you can use low performance base network in the range of LB=0.69 to 0.71.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8069/combine.0.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8070/combine.1.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8126/graph1.png\" alt=\"enter image description here\" title=\"\">\n (loss curve for combine.1)</p>\n\n<hr>\n\n<p>b. fcnet1  : max to replace scaling</p>\n\n<p>c. fcnet0 : fcnet1   + mixup augmentation</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8124/graph3.png\" alt=\"enter image description here\" title=\"\">\n   <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/257777/8125/graph2.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "**Updated! please refer to the ppt and xls in the attachment for full details.**\n\n---\n\nI present my solution , which is part of team \"Convoluted predictions\". I submitted an ensemble  of 6 models (combine.0, combine.1, combine.2 combine.5, combine.5-max, combine.5-mix). This ensemble has LB=0.787.\n\n\nThere is one model that is discarded and not used (combine.4, single base network), but nevertheless gives good results.\n\nCredits go to my teammate @Miha Skalic who discover the fast way to train the model in 2 stages. By training on the extracted features, it is much more faster and accurate!\n\n\n.\n\n\nFirst, I introduce combine.2. It uses gated scaling in the combination network fcnet3.\n\n\n ![enter image description here][1]\n\n ![enter image description here][2]\n\nHow to train fcnet3:\n\n 1.  Train the base networks (se-resent101 and xception) image wise.\n 2. Extract the features for each train image  and store in disk (quantised to np.uint8 and store as memmap)\n 3. train fcnet3 product wise.\n\nTricks in training:\n\n1.  for base networks , use dropout and use augmentation=small shift, scale, rotation + flip + rotate90,180,270. Batch_size = 512 to 768, num_epoches =  about 20. SDG/learning rate=0.01,0.001,0.0001 (manually tuned)\n\n       time to train (4x1080Ti) = 5 days per base networks\n\n\n2.  for fcnet3, use dropout and i only use augmentation=flip (due to lack of space in my SSD drive). Batch_size = 4096, num_epoches = about 100.SDG/learning rate=0.01,0.001,0.0001 (manually tuned)\n\n       time to train (4x1080Ti) = 1.5 days  \n\n\n\n---\nOther variants (to be updated)\n\na. other base network\n\nI make combine.0 and combine.1 as shown below. These differs in dropout probability. Initially, i intended to use only one of them in final submission. But surprisingly, using both of them is better than using one. (you note that these are exact same models that i post in another post:  https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021) . With my gated Fcnet3, you can use low performance base network in the range of LB=0.69 to 0.71.\n\n\n ![enter image description here][4]\n\n ![enter image description here][5]\n\n ![enter image description here][8]\n (loss curve for combine.1)\n\n---\n\nb. fcnet1  : max to replace scaling\n\n\n\n\nc. fcnet0 : fcnet1   + mixup augmentation\n\n   ![enter image description here][6]\n   ![enter image description here][7]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8063/c2.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8064/fcnet3.png\n  [3]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8065/fcnet3_curve.png\n  [4]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8069/combine.0.png\n  [5]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8070/combine.1.png\n  [6]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8124/graph3.png\n  [7]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8125/graph2.png\n  [8]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8126/graph1.png",
      "votes": null
    },
    {
      "id": "257780",
      "postDate": "12/15/2017 00:06:45",
      "content": "<p>Hi Heng, Congrats  !!, gained a lot from you in this competition.Thank You\nRegarding Network structure was this network trained jointly ? i.e the net parameters are shared across different images that belong to same product ? </p>",
      "rawMarkdown": "Hi Heng, Congrats  !!, gained a lot from you in this competition.Thank You\nRegarding Network structure was this network trained jointly ? i.e the net parameters are shared across different images that belong to same product ?",
      "votes": null
    },
    {
      "id": "257782",
      "postDate": "12/15/2017 00:12:14",
      "content": "<p>I was trying something similar with base models that had 74.8 % and 72.3 % accuracy, respectively, but I couldn't get my MLP to train.</p>",
      "rawMarkdown": "I was trying something similar with base models that had 74.8 % and 72.3 % accuracy, respectively, but I couldn't get my MLP to train.",
      "votes": null
    },
    {
      "id": "257794",
      "postDate": "12/15/2017 00:28:26",
      "content": "<p>Hi Heng, one more question, the 'x.sum(dim=1)' in Forward Propagation of FcNet3, adds up the vectors of 4 images element-wise and gives one long 8192 length vector right ?</p>",
      "rawMarkdown": "Hi Heng, one more question, the 'x.sum(dim=1)' in Forward Propagation of FcNet3, adds up the vectors of 4 images element-wise and gives one long 8192 length vector right ?",
      "votes": null
    },
    {
      "id": "257800",
      "postDate": "12/15/2017 00:32:14",
      "content": "<p>No. The adding is done for 4 images per product. After adding, one product has 2x2048 (2 base network) dim. This will make results independent of image order, see this paper: <a href=\"https://arxiv.org/abs/1612.00593\">https://arxiv.org/abs/1612.00593</a>.\nThe scaling is a gate to select how much to add from each of the 4 image.</p>\n\n<p>\"PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation\"\n- Charles R. Qi, Hao Su, Kaichun Mo, Leonidas J. Guibas</p>",
      "rawMarkdown": "No. The adding is done for 4 images per product. After adding, one product has 2x2048 (2 base network) dim. This will make results independent of image order, see this paper: https://arxiv.org/abs/1612.00593.\nThe scaling is a gate to select how much to add from each of the 4 image.\n\n\"PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation\"\n- Charles R. Qi, Hao Su, Kaichun Mo, Leonidas J. Guibas",
      "votes": null
    },
    {
      "id": "257842",
      "postDate": "12/15/2017 01:35:14",
      "content": "<p>Hi Heng, you split the train/val set by products or by images?</p>",
      "rawMarkdown": "Hi Heng, you split the train/val set by products or by images?",
      "votes": null
    },
    {
      "id": "257843",
      "postDate": "12/15/2017 01:36:37",
      "content": "<p>by products</p>",
      "rawMarkdown": "by products",
      "votes": null
    },
    {
      "id": "257846",
      "postDate": "12/15/2017 01:40:40",
      "content": "<p>Congratulations 2nd place and thank you for the great solution! \nI'm curious whether \"Mixup\" contributed to LB score.</p>",
      "rawMarkdown": "Congratulations 2nd place and thank you for the great solution! \nI'm curious whether \"Mixup\" contributed to LB score.",
      "votes": null
    },
    {
      "id": "257848",
      "postDate": "12/15/2017 01:43:52",
      "content": "<p>i use mixup as a regularizer (noise in train data). \nmixup did better when used in an ensemble. Alone, results is slightly worst than dropout.\n(i will post some results later)</p>",
      "rawMarkdown": "i use mixup as a regularizer (noise in train data). \nmixup did better when used in an ensemble. Alone, results is slightly worst than dropout.\n(i will post some results later)",
      "votes": null
    },
    {
      "id": "257863",
      "postDate": "12/15/2017 02:18:01",
      "content": "<p>Hi Heng,\nDo you use exactly the same training and validation set in the two stages of your pipeline? May I know the corresponding training accuracies (corrsponds to the 0.7256 and 0.7192) for your finetuning models?</p>",
      "rawMarkdown": "Hi Heng,\nDo you use exactly the same training and validation set in the two stages of your pipeline? May I know the corresponding training accuracies (corrsponds to the 0.7256 and 0.7192) for your finetuning models?",
      "votes": null
    },
    {
      "id": "257870",
      "postDate": "12/15/2017 02:30:55",
      "content": "<p>cool, don't even know the final ensemble has such a great improvement! I just simply average all my models. Thanks for sharing!</p>",
      "rawMarkdown": "cool, don't even know the final ensemble has such a great improvement! I just simply average all my models. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "257879",
      "postDate": "12/15/2017 02:47:09",
      "content": "<p>yes. same training/validation set for both stages.0.7256 and 0.7192 (and 0.7753) are validation accuracy  on the same set. I will post the training accuracy later.</p>",
      "rawMarkdown": "yes. same training/validation set for both stages.0.7256 and 0.7192 (and 0.7753) are validation accuracy  on the same set. I will post the training accuracy later.",
      "votes": null
    },
    {
      "id": "257894",
      "postDate": "12/15/2017 03:16:59",
      "content": "<p>Thanks in advance. I ask this because I did something similar to your two stages but I trained my second stage on the validation set only and did not reuse the training set. I actually did a 5fold CV on the validation set. But I wasn't able to get a decent boost in performance.</p>\n\n<p>I believe what I'm doing is a more usual way of stacking than yours (reusing the training set), but the situation might be a little bit different here.</p>",
      "rawMarkdown": "Thanks in advance. I ask this because I did something similar to your two stages but I trained my second stage on the validation set only and did not reuse the training set. I actually did a 5fold CV on the validation set. But I wasn't able to get a decent boost in performance.\n\nI believe what I'm doing is a more usual way of stacking than yours (reusing the training set), but the situation might be a little bit different here.",
      "votes": null
    },
    {
      "id": "257910",
      "postDate": "12/15/2017 04:37:54",
      "content": "<p>What i am doing is closer to multiview pooling. see: <a href=\"https://arxiv.org/pdf/1505.00880.pdf\">https://arxiv.org/pdf/1505.00880.pdf</a></p>\n\n<p>Becuase there are too much data for end-to-end, so i break into 2 stages. Alternate training should improve results (i.e. train base models and fcnet, by alternatively freezing them) but i haven't got time to try it.</p>\n\n<p><img src=\"http://vis-www.cs.umass.edu/mvcnn/images/mvcnn.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>\"Multi-view Convolutional Neural Networks for 3D Shape Recognition\"- Hang Su Subhransu Maji</p>",
      "rawMarkdown": "What i am doing is closer to multiview pooling. see: https://arxiv.org/pdf/1505.00880.pdf\n\nBecuase there are too much data for end-to-end, so i break into 2 stages. Alternate training should improve results (i.e. train base models and fcnet, by alternatively freezing them) but i haven't got time to try it.\n\n\n ![enter image description here][1]\n\n\"Multi-view Convolutional Neural Networks for 3D Shape Recognition\"- Hang Su Subhransu Maji\n\n\n  [1]: http://vis-www.cs.umass.edu/mvcnn/images/mvcnn.png",
      "votes": null
    },
    {
      "id": "258044",
      "postDate": "12/15/2017 10:26:22",
      "content": "<p>Thank you for all your great contributions!</p>",
      "rawMarkdown": "Thank you for all your great contributions!",
      "votes": null
    },
    {
      "id": "258088",
      "postDate": "12/15/2017 13:06:55",
      "content": "<p>I was also wondering about cross validation, i did not try it for lack of computing resources. One of the issues about it is the bottleneck features are not aligned across fold. Did you use bottleneck features or probability predictions?</p>",
      "rawMarkdown": "I was also wondering about cross validation, i did not try it for lack of computing resources. One of the issues about it is the bottleneck features are not aligned across fold. Did you use bottleneck features or probability predictions?",
      "votes": null
    },
    {
      "id": "258134",
      "postDate": "12/15/2017 15:20:46",
      "content": "<p>Thanks Heng, very interesting. It might worth research a little bit more to see if such a boost mainly comes from the multi-view (in this case the boost will be relatively stable no matter how good you trained 1st stage models), or mainly comes from relatively lower 1st stage model performance (in this case the boost will gradually diminish to zero when you have better 1st stage model).</p>\n\n<p>BTW, can you share the formula of \"quantised to np.uint8\"? I guess some scaling is required right?</p>",
      "rawMarkdown": "Thanks Heng, very interesting. It might worth research a little bit more to see if such a boost mainly comes from the multi-view (in this case the boost will be relatively stable no matter how good you trained 1st stage models), or mainly comes from relatively lower 1st stage model performance (in this case the boost will gradually diminish to zero when you have better 1st stage model).\n\nBTW, can you share the formula of \"quantised to np.uint8\"? I guess some scaling is required right?",
      "votes": null
    },
    {
      "id": "258138",
      "postDate": "12/15/2017 15:36:18",
      "content": "<p>regarding using poorer performance base networks, please refer to combine.0 and combine.1 in the updated post.</p>",
      "rawMarkdown": "regarding using poorer performance base networks, please refer to combine.0 and combine.1 in the updated post.",
      "votes": null
    },
    {
      "id": "258144",
      "postDate": "12/15/2017 15:52:29",
      "content": "<p>Informative. Thanks</p>",
      "rawMarkdown": "Informative. Thanks",
      "votes": null
    },
    {
      "id": "258147",
      "postDate": "12/15/2017 16:00:40",
      "content": "<p>@Lam Dang. I switch validation set for training base network and it worked!  I in fact, i re-trained  a trained network  (dpn92) from my team mate. We use different train/validation set and there are some data overlap.</p>\n\n<p>I use \"features\" and \"not probability\" for stage two.</p>",
      "rawMarkdown": "Lam Dang. I switch validation set for training base network and it worked!  I in fact, i re-trained  a trained network  (dpn92) from my team mate. We use different train/validation set and there are some data overlap.\n\nI use \"features\" and \"not probability\" for stage two.",
      "votes": null
    },
    {
      "id": "258176",
      "postDate": "12/15/2017 16:57:30",
      "content": "<p>Congratulations @Heng Cherkeng and team on coming 2nd in this competition. And thanks so much for your always useful contribution to the community with every competition you have participated in. It really is a pleasure competing with you.</p>\n\n<p>And thank you so much for sharing your team's solution once again. I will find time to read through it later.</p>",
      "rawMarkdown": "Congratulations @Heng Cherkeng and team on coming 2nd in this competition. And thanks so much for your always useful contribution to the community with every competition you have participated in. It really is a pleasure competing with you.\n\nAnd thank you so much for sharing your team's solution once again. I will find time to read through it later.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 257780,
      "author_name": "rteja1113",
      "author_url": "",
      "post_date": "12/15/2017 00:06:45",
      "content": "<p>Hi Heng, Congrats  !!, gained a lot from you in this competition.Thank You\nRegarding Network structure was this network trained jointly ? i.e the net parameters are shared across different images that belong to same product ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 257782,
      "author_name": "eachshadow",
      "author_url": "",
      "post_date": "12/15/2017 00:12:14",
      "content": "<p>I was trying something similar with base models that had 74.8 % and 72.3 % accuracy, respectively, but I couldn't get my MLP to train.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 257794,
      "author_name": "rteja1113",
      "author_url": "",
      "post_date": "12/15/2017 00:28:26",
      "content": "<p>Hi Heng, one more question, the 'x.sum(dim=1)' in Forward Propagation of FcNet3, adds up the vectors of 4 images element-wise and gives one long 8192 length vector right ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 257800,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/15/2017 00:32:14",
          "content": "<p>No. The adding is done for 4 images per product. After adding, one product has 2x2048 (2 base network) dim. This will make results independent of image order, see this paper: <a href=\"https://arxiv.org/abs/1612.00593\">https://arxiv.org/abs/1612.00593</a>.\nThe scaling is a gate to select how much to add from each of the 4 image.</p>\n\n<p>\"PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation\"\n- Charles R. Qi, Hao Su, Kaichun Mo, Leonidas J. Guibas</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 257842,
      "author_name": "tangchuanxin",
      "author_url": "",
      "post_date": "12/15/2017 01:35:14",
      "content": "<p>Hi Heng, you split the train/val set by products or by images?</p>",
      "votes": null,
      "replies": [
        {
          "id": 257843,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/15/2017 01:36:37",
          "content": "<p>by products</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 257846,
      "author_name": "lyakaap",
      "author_url": "",
      "post_date": "12/15/2017 01:40:40",
      "content": "<p>Congratulations 2nd place and thank you for the great solution! \nI'm curious whether \"Mixup\" contributed to LB score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 257848,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/15/2017 01:43:52",
          "content": "<p>i use mixup as a regularizer (noise in train data). \nmixup did better when used in an ensemble. Alone, results is slightly worst than dropout.\n(i will post some results later)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 257863,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "12/15/2017 02:18:01",
      "content": "<p>Hi Heng,\nDo you use exactly the same training and validation set in the two stages of your pipeline? May I know the corresponding training accuracies (corrsponds to the 0.7256 and 0.7192) for your finetuning models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 257879,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/15/2017 02:47:09",
          "content": "<p>yes. same training/validation set for both stages.0.7256 and 0.7192 (and 0.7753) are validation accuracy  on the same set. I will post the training accuracy later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 257894,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "12/15/2017 03:16:59",
          "content": "<p>Thanks in advance. I ask this because I did something similar to your two stages but I trained my second stage on the validation set only and did not reuse the training set. I actually did a 5fold CV on the validation set. But I wasn't able to get a decent boost in performance.</p>\n\n<p>I believe what I'm doing is a more usual way of stacking than yours (reusing the training set), but the situation might be a little bit different here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 257910,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/15/2017 04:37:54",
          "content": "<p>What i am doing is closer to multiview pooling. see: <a href=\"https://arxiv.org/pdf/1505.00880.pdf\">https://arxiv.org/pdf/1505.00880.pdf</a></p>\n\n<p>Becuase there are too much data for end-to-end, so i break into 2 stages. Alternate training should improve results (i.e. train base models and fcnet, by alternatively freezing them) but i haven't got time to try it.</p>\n\n<p><img src=\"http://vis-www.cs.umass.edu/mvcnn/images/mvcnn.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>\"Multi-view Convolutional Neural Networks for 3D Shape Recognition\"- Hang Su Subhransu Maji</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258088,
          "author_name": "lamdang",
          "author_url": "",
          "post_date": "12/15/2017 13:06:55",
          "content": "<p>I was also wondering about cross validation, i did not try it for lack of computing resources. One of the issues about it is the bottleneck features are not aligned across fold. Did you use bottleneck features or probability predictions?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258134,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "12/15/2017 15:20:46",
          "content": "<p>Thanks Heng, very interesting. It might worth research a little bit more to see if such a boost mainly comes from the multi-view (in this case the boost will be relatively stable no matter how good you trained 1st stage models), or mainly comes from relatively lower 1st stage model performance (in this case the boost will gradually diminish to zero when you have better 1st stage model).</p>\n\n<p>BTW, can you share the formula of \"quantised to np.uint8\"? I guess some scaling is required right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258138,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/15/2017 15:36:18",
          "content": "<p>regarding using poorer performance base networks, please refer to combine.0 and combine.1 in the updated post.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258144,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "12/15/2017 15:52:29",
          "content": "<p>Informative. Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258147,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/15/2017 16:00:40",
          "content": "<p>@Lam Dang. I switch validation set for training base network and it worked!  I in fact, i re-trained  a trained network  (dpn92) from my team mate. We use different train/validation set and there are some data overlap.</p>\n\n<p>I use \"features\" and \"not probability\" for stage two.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 257870,
      "author_name": "swkcoding1",
      "author_url": "",
      "post_date": "12/15/2017 02:30:55",
      "content": "<p>cool, don't even know the final ensemble has such a great improvement! I just simply average all my models. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258044,
      "author_name": "laurensdlm",
      "author_url": "",
      "post_date": "12/15/2017 10:26:22",
      "content": "<p>Thank you for all your great contributions!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258176,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "12/15/2017 16:57:30",
      "content": "<p>Congratulations @Heng Cherkeng and team on coming 2nd in this competition. And thanks so much for your always useful contribution to the community with every competition you have participated in. It really is a pleasure competing with you.</p>\n\n<p>And thank you so much for sharing your team's solution once again. I will find time to read through it later.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "257777": "**Updated! please refer to the ppt and xls in the attachment for full details.**\n\n---\n\nI present my solution , which is part of team \"Convoluted predictions\". I submitted an ensemble  of 6 models (combine.0, combine.1, combine.2 combine.5, combine.5-max, combine.5-mix). This ensemble has LB=0.787.\n\n\nThere is one model that is discarded and not used (combine.4, single base network), but nevertheless gives good results.\n\nCredits go to my teammate @Miha Skalic who discover the fast way to train the model in 2 stages. By training on the extracted features, it is much more faster and accurate!\n\n\n.\n\n\nFirst, I introduce combine.2. It uses gated scaling in the combination network fcnet3.\n\n\n ![enter image description here][1]\n\n ![enter image description here][2]\n\nHow to train fcnet3:\n\n 1.  Train the base networks (se-resent101 and xception) image wise.\n 2. Extract the features for each train image  and store in disk (quantised to np.uint8 and store as memmap)\n 3. train fcnet3 product wise.\n\nTricks in training:\n\n1.  for base networks , use dropout and use augmentation=small shift, scale, rotation + flip + rotate90,180,270. Batch_size = 512 to 768, num_epoches =  about 20. SDG/learning rate=0.01,0.001,0.0001 (manually tuned)\n\n       time to train (4x1080Ti) = 5 days per base networks\n\n\n2.  for fcnet3, use dropout and i only use augmentation=flip (due to lack of space in my SSD drive). Batch_size = 4096, num_epoches = about 100.SDG/learning rate=0.01,0.001,0.0001 (manually tuned)\n\n       time to train (4x1080Ti) = 1.5 days  \n\n\n\n---\nOther variants (to be updated)\n\na. other base network\n\nI make combine.0 and combine.1 as shown below. These differs in dropout probability. Initially, i intended to use only one of them in final submission. But surprisingly, using both of them is better than using one. (you note that these are exact same models that i post in another post:  https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/41021) . With my gated Fcnet3, you can use low performance base network in the range of LB=0.69 to 0.71.\n\n\n ![enter image description here][4]\n\n ![enter image description here][5]\n\n ![enter image description here][8]\n (loss curve for combine.1)\n\n---\n\nb. fcnet1  : max to replace scaling\n\n\n\n\nc. fcnet0 : fcnet1   + mixup augmentation\n\n   ![enter image description here][6]\n   ![enter image description here][7]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8063/c2.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8064/fcnet3.png\n  [3]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8065/fcnet3_curve.png\n  [4]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8069/combine.0.png\n  [5]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8070/combine.1.png\n  [6]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8124/graph3.png\n  [7]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8125/graph2.png\n  [8]: https://kaggle2.blob.core.windows.net/forum-message-attachments/257777/8126/graph1.png",
    "257780": "Hi Heng, Congrats  !!, gained a lot from you in this competition.Thank You\nRegarding Network structure was this network trained jointly ? i.e the net parameters are shared across different images that belong to same product ?",
    "257782": "I was trying something similar with base models that had 74.8 % and 72.3 % accuracy, respectively, but I couldn't get my MLP to train.",
    "257794": "Hi Heng, one more question, the 'x.sum(dim=1)' in Forward Propagation of FcNet3, adds up the vectors of 4 images element-wise and gives one long 8192 length vector right ?",
    "257800": "No. The adding is done for 4 images per product. After adding, one product has 2x2048 (2 base network) dim. This will make results independent of image order, see this paper: https://arxiv.org/abs/1612.00593.\nThe scaling is a gate to select how much to add from each of the 4 image.\n\n\"PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation\"\n- Charles R. Qi, Hao Su, Kaichun Mo, Leonidas J. Guibas",
    "257842": "Hi Heng, you split the train/val set by products or by images?",
    "257843": "by products",
    "257846": "Congratulations 2nd place and thank you for the great solution! \nI'm curious whether \"Mixup\" contributed to LB score.",
    "257848": "i use mixup as a regularizer (noise in train data). \nmixup did better when used in an ensemble. Alone, results is slightly worst than dropout.\n(i will post some results later)",
    "257863": "Hi Heng,\nDo you use exactly the same training and validation set in the two stages of your pipeline? May I know the corresponding training accuracies (corrsponds to the 0.7256 and 0.7192) for your finetuning models?",
    "257870": "cool, don't even know the final ensemble has such a great improvement! I just simply average all my models. Thanks for sharing!",
    "257879": "yes. same training/validation set for both stages.0.7256 and 0.7192 (and 0.7753) are validation accuracy  on the same set. I will post the training accuracy later.",
    "257894": "Thanks in advance. I ask this because I did something similar to your two stages but I trained my second stage on the validation set only and did not reuse the training set. I actually did a 5fold CV on the validation set. But I wasn't able to get a decent boost in performance.\n\nI believe what I'm doing is a more usual way of stacking than yours (reusing the training set), but the situation might be a little bit different here.",
    "257910": "What i am doing is closer to multiview pooling. see: https://arxiv.org/pdf/1505.00880.pdf\n\nBecuase there are too much data for end-to-end, so i break into 2 stages. Alternate training should improve results (i.e. train base models and fcnet, by alternatively freezing them) but i haven't got time to try it.\n\n\n ![enter image description here][1]\n\n\"Multi-view Convolutional Neural Networks for 3D Shape Recognition\"- Hang Su Subhransu Maji\n\n\n  [1]: http://vis-www.cs.umass.edu/mvcnn/images/mvcnn.png",
    "258044": "Thank you for all your great contributions!",
    "258088": "I was also wondering about cross validation, i did not try it for lack of computing resources. One of the issues about it is the bottleneck features are not aligned across fold. Did you use bottleneck features or probability predictions?",
    "258134": "Thanks Heng, very interesting. It might worth research a little bit more to see if such a boost mainly comes from the multi-view (in this case the boost will be relatively stable no matter how good you trained 1st stage models), or mainly comes from relatively lower 1st stage model performance (in this case the boost will gradually diminish to zero when you have better 1st stage model).\n\nBTW, can you share the formula of \"quantised to np.uint8\"? I guess some scaling is required right?",
    "258138": "regarding using poorer performance base networks, please refer to combine.0 and combine.1 in the updated post.",
    "258144": "Informative. Thanks",
    "258147": "Lam Dang. I switch validation set for training base network and it worked!  I in fact, i re-trained  a trained network  (dpn92) from my team mate. We use different train/validation set and there are some data overlap.\n\nI use \"features\" and \"not probability\" for stage two.",
    "258176": "Congratulations @Heng Cherkeng and team on coming 2nd in this competition. And thanks so much for your always useful contribution to the community with every competition you have participated in. It really is a pleasure competing with you.\n\nAnd thank you so much for sharing your team's solution once again. I will find time to read through it later."
  },
  "source": "meta"
}