{
  "id": 45863,
  "title": "My brief summary of this competition[0.795]",
  "url": "/competitions/cdiscount-image-classification-challenge/writeups/bestfitting-my-brief-summary-of-this-competition-0",
  "author_name": "",
  "post_date": "2017-12-17T15:50:54.040Z",
  "votes": 134,
  "comment_count": 44,
  "views": 0,
  "content": "<p>It is an interesting competition,and also a tough one.<br><br>\nI faced the following challenges:<br>\n1.Large Dataset with 15+ millions images and 5000+ categories.<br>\n2.Some products and 1-4 images<br>\n3.CD/BOOK are very hard to classify.<br>\n4.I estimated the best overall accuracy will be 0.8 also,there are a lot of methods to choose and there are a large space to improve,so it’s very hard to win.</p>\n\n<p><b></b></p><b>\n\n<h2>My solution summary:</h2>\n\n</b><p><b></b><br>\n1.Making preparations for such a big dataset<br>\n2.Finetuning pretrained models.(0.759/0.757,inception-resnet-v2,0.757/0.756 resnet50).<br>\n3.Make full use of multi-images of a product. (0.772/0.771 inception-resnet-v2 and 0.769/0.768+ Resnet50 models).<br>\n4.Use OCR to add semantics to the models.\n<b></b></p><b>\n\n<h2>Preparing:</h2>\n\n</b><p><b></b><br>\nAt early stage of this competition,I spent 4 days on trying to feed the images to pytorch efficiently,I I found the best way is to buy 512M SSDs.<br>\nI also found that if I want to store prediction results into a file,I must use sparse matrix,I spend 1.5 days on writing codes to make sure I can keep probabilities of  prediction in a small file(~40M)\n<b></b></p><b>\n\n<h2>Finetuning pretrained models.</h2>\n\n</b><p><b></b><br>\nWe can not win such a competition only by downloading pretrained models and finetuning them,we must understand the principles behind every models.<br>\nAfter got some statistics of this data set,I searched and read all kinds of papers related:network structures,the optimizers,the train plan．．．．．．<br>\nThen I started my experiments on Resnet34,I have following result:<br>\n1.The network structure of almost all pretrained model are for imagenet with 1000 labels ,but our dataset have 5270 labels.If we use then directly,there will be some representation bottleneck.<br>\n2.I found train with SGD need more epochs to converge than ADAM.<br></p>\n\n<p>So I modified the head of resnet34,add a 1x1 kernel convolution layer after layer4 of resnet,which increase the channel from 512 to 5270,and then a FC of 5270*5270.<br></p>\n\n<p>And I used Adam with optim.Adam(net.parameters(), lr=lr) and lr plan is:<br></p>\n\n<pre>        lr = 0.0003\n        if epoch &gt; 7:\n            lr = 0.0001\n        if epoch &gt; 9:\n            lr = 0.00005\n        if epoch &gt; 11:\n            lr = 0.00001\n</pre>\n\n<p>When I  trained 11.5 epochs with 160x160 patch random cropped from 180*180 images,I predicted and averaged predictions of the images of product,I could get a score &gt;0.72 on public LB.<br><br>\nThen I thought,why not add a FC layer more?So I added a 5270x5270,and I found that it improved more than 0.5<br></p>\n\n<p>Since I can train 1 epochs in 2.5 hours on a 4x 1080Ti machine,I could do a lot of experiments,I found if the result is not so good on first 1-2 epochs,the result not be good at end,this speed up my experiments.<br>\nI tried but not used in the next stage:<br>\n1.Multi-level categories as multi-task.<br>\n2.Hard Example,Focal Loss,<br>\n3.Dilation<br>\n4.Dropout<br>\n<br></p>\n\n<p>After I got these conclusion,I turned to resnet50,I found resnet50 is better than resnet34: 0.756 on public.<br><br></p>\n\n<p>Because I am busy with my everyday job ,I just let my GPUs keep running,and fine-tuned Resnet101,resnet152,and inceptionresnetv2 and inceptionV4.\nInceptionv4 is a little weak at this stage,others could easily get a score &gt;0.755 on public LB.<br></p>\n\n<p><b></b></p><b>\n\n<h2>Make full use of multi-images of a product</h2>\n\n</b><p><b></b><br>\nLast weeks of this competition,I had more time,and I started analyzing the dataset.<br>\nI trained my models using images of a product separately before,but when we have a look at a product,we will treat the images as a whole.So I concatenate the images into a image,and fine-tune the models.<br><br></p>\n\n<p>That’s to say,We can split the trainset  into 4 parts,products with 1,2,3,4 images,if we finetune products with 4 images ,we concatenate the 4 images into a image.if we finetune products with 3 images ,we concatenate the 3 images into a image. I also finetune the model with product with only 1 images.<br>\nBecause I found that there are some invalid or meaningless images for classification in images for multi-images products,that will be noise for product with only on product.It was easy to finetune,the acc improved a lot after 2-3 epochs,it took no so much time.<br><br></p>\n\n<p>After this step,I got a cluster of model for resnet50,a model trained with all images,4 models for 1,2,3,4 images products.I predicted the product accordingly and ensembled them,I got a big improvement.Now,Almost all the models can get acc score near 0.77 .<br><br></p>\n\n<p>I made several submissions just now,I found that if I ensemble resnet50 models with inception-resnet-v2 models,we can get 0.782/0.781 on  private and public LB.<br><br></p>\n\n<p><b></b></p><b>\n\n<h2>OCR</h2>\n\n</b><p><b></b><br>\nI found that CD and BOOK are very hard to classify,because we need know the meaning of the cover!<br>\nAnd I want to evaluate the performance of all kinds of algorithms OCR related for my next project,so I spent more than a week on it,I use CTPN to extract the boxes containing text,the result was quite good.Then I used CRNN to extract the text from the boxes.But I found that the text are not so good,because the images are small sized and CRNN was trained for English and I have no French dataset to retrain it,I can generate it,but I can not found suitable books without copyright limitation to extract French words.Without good words extracted from CD and BOOK cover,I could not use Word2Vec,so I can not provide semantic meaning to the model. As a last resort,I extracted the features of boxes from CRNN and feed them into my multi-input CNN,the first input is Resnet50 FC features and the second is CRNN features.<br>\nAfter got the prediction from this OCR-Network,I ensembled with the predictions of  Resnet50 ,I make a lot of submissions in last week to confirm I did not over-fit.<br>\nI found that OCR solution  contributed ~0.14% to my final submissions.<br>\n[Edit],I found that if I use OCR on 0.782 model,the result can improve 0.35%,this means that the improvement got from OCR is covered by multi-model ensemble.\n<b></b></p><b>\n\n<h2>Other models</h2>\n\n</b><p><b></b><br>\nI also tried densenet161,densenet169,dpn92,they are much worse than resnet50,since I have trained them,so I just ensembled them to add some diversities.<br><br>\nAfter ensembled the models,the public LB can reach 0.79.<br><br>\nTo reach my final scored,I fine-tuned the 1-image products models with 224/299 sized images accordingly,to add more diversities,I also change the head to two 4096*4096 FC layer following VGG net and ensemble them.</p>\n\n<hr>\n\n<p>if you want to reproduce my solution,and if you have 4-1080Ti,you can get 0.782/0.781 on private/public LB in 7-8 days with resnet50 and inception-resnet-v2.</p>\n\n<p><br>\n<br></p>",
  "messages": [
    {
      "id": "258683",
      "postDate": "12/16/2017 19:32:42",
      "content": "<p>It is an interesting competition,and also a tough one.<br><br>\nI faced the following challenges:<br>\n1.Large Dataset with 15+ millions images and 5000+ categories.<br>\n2.Some products and 1-4 images<br>\n3.CD/BOOK are very hard to classify.<br>\n4.I estimated the best overall accuracy will be 0.8 also,there are a lot of methods to choose and there are a large space to improve,so it’s very hard to win.</p>\n\n<p><b></b></p><b>\n\n<h2>My solution summary:</h2>\n\n</b><p><b></b><br>\n1.Making preparations for such a big dataset<br>\n2.Finetuning pretrained models.(0.759/0.757,inception-resnet-v2,0.757/0.756 resnet50).<br>\n3.Make full use of multi-images of a product. (0.772/0.771 inception-resnet-v2 and 0.769/0.768+ Resnet50 models).<br>\n4.Use OCR to add semantics to the models.\n<b></b></p><b>\n\n<h2>Preparing:</h2>\n\n</b><p><b></b><br>\nAt early stage of this competition,I spent 4 days on trying to feed the images to pytorch efficiently,I I found the best way is to buy 512M SSDs.<br>\nI also found that if I want to store prediction results into a file,I must use sparse matrix,I spend 1.5 days on writing codes to make sure I can keep probabilities of  prediction in a small file(~40M)\n<b></b></p><b>\n\n<h2>Finetuning pretrained models.</h2>\n\n</b><p><b></b><br>\nWe can not win such a competition only by downloading pretrained models and finetuning them,we must understand the principles behind every models.<br>\nAfter got some statistics of this data set,I searched and read all kinds of papers related:network structures,the optimizers,the train plan．．．．．．<br>\nThen I started my experiments on Resnet34,I have following result:<br>\n1.The network structure of almost all pretrained model are for imagenet with 1000 labels ,but our dataset have 5270 labels.If we use then directly,there will be some representation bottleneck.<br>\n2.I found train with SGD need more epochs to converge than ADAM.<br></p>\n\n<p>So I modified the head of resnet34,add a 1x1 kernel convolution layer after layer4 of resnet,which increase the channel from 512 to 5270,and then a FC of 5270*5270.<br></p>\n\n<p>And I used Adam with optim.Adam(net.parameters(), lr=lr) and lr plan is:<br></p>\n\n<pre>        lr = 0.0003\n        if epoch &gt; 7:\n            lr = 0.0001\n        if epoch &gt; 9:\n            lr = 0.00005\n        if epoch &gt; 11:\n            lr = 0.00001\n</pre>\n\n<p>When I  trained 11.5 epochs with 160x160 patch random cropped from 180*180 images,I predicted and averaged predictions of the images of product,I could get a score &gt;0.72 on public LB.<br><br>\nThen I thought,why not add a FC layer more?So I added a 5270x5270,and I found that it improved more than 0.5<br></p>\n\n<p>Since I can train 1 epochs in 2.5 hours on a 4x 1080Ti machine,I could do a lot of experiments,I found if the result is not so good on first 1-2 epochs,the result not be good at end,this speed up my experiments.<br>\nI tried but not used in the next stage:<br>\n1.Multi-level categories as multi-task.<br>\n2.Hard Example,Focal Loss,<br>\n3.Dilation<br>\n4.Dropout<br>\n<br></p>\n\n<p>After I got these conclusion,I turned to resnet50,I found resnet50 is better than resnet34: 0.756 on public.<br><br></p>\n\n<p>Because I am busy with my everyday job ,I just let my GPUs keep running,and fine-tuned Resnet101,resnet152,and inceptionresnetv2 and inceptionV4.\nInceptionv4 is a little weak at this stage,others could easily get a score &gt;0.755 on public LB.<br></p>\n\n<p><b></b></p><b>\n\n<h2>Make full use of multi-images of a product</h2>\n\n</b><p><b></b><br>\nLast weeks of this competition,I had more time,and I started analyzing the dataset.<br>\nI trained my models using images of a product separately before,but when we have a look at a product,we will treat the images as a whole.So I concatenate the images into a image,and fine-tune the models.<br><br></p>\n\n<p>That’s to say,We can split the trainset  into 4 parts,products with 1,2,3,4 images,if we finetune products with 4 images ,we concatenate the 4 images into a image.if we finetune products with 3 images ,we concatenate the 3 images into a image. I also finetune the model with product with only 1 images.<br>\nBecause I found that there are some invalid or meaningless images for classification in images for multi-images products,that will be noise for product with only on product.It was easy to finetune,the acc improved a lot after 2-3 epochs,it took no so much time.<br><br></p>\n\n<p>After this step,I got a cluster of model for resnet50,a model trained with all images,4 models for 1,2,3,4 images products.I predicted the product accordingly and ensembled them,I got a big improvement.Now,Almost all the models can get acc score near 0.77 .<br><br></p>\n\n<p>I made several submissions just now,I found that if I ensemble resnet50 models with inception-resnet-v2 models,we can get 0.782/0.781 on  private and public LB.<br><br></p>\n\n<p><b></b></p><b>\n\n<h2>OCR</h2>\n\n</b><p><b></b><br>\nI found that CD and BOOK are very hard to classify,because we need know the meaning of the cover!<br>\nAnd I want to evaluate the performance of all kinds of algorithms OCR related for my next project,so I spent more than a week on it,I use CTPN to extract the boxes containing text,the result was quite good.Then I used CRNN to extract the text from the boxes.But I found that the text are not so good,because the images are small sized and CRNN was trained for English and I have no French dataset to retrain it,I can generate it,but I can not found suitable books without copyright limitation to extract French words.Without good words extracted from CD and BOOK cover,I could not use Word2Vec,so I can not provide semantic meaning to the model. As a last resort,I extracted the features of boxes from CRNN and feed them into my multi-input CNN,the first input is Resnet50 FC features and the second is CRNN features.<br>\nAfter got the prediction from this OCR-Network,I ensembled with the predictions of  Resnet50 ,I make a lot of submissions in last week to confirm I did not over-fit.<br>\nI found that OCR solution  contributed ~0.14% to my final submissions.<br>\n[Edit],I found that if I use OCR on 0.782 model,the result can improve 0.35%,this means that the improvement got from OCR is covered by multi-model ensemble.\n<b></b></p><b>\n\n<h2>Other models</h2>\n\n</b><p><b></b><br>\nI also tried densenet161,densenet169,dpn92,they are much worse than resnet50,since I have trained them,so I just ensembled them to add some diversities.<br><br>\nAfter ensembled the models,the public LB can reach 0.79.<br><br>\nTo reach my final scored,I fine-tuned the 1-image products models with 224/299 sized images accordingly,to add more diversities,I also change the head to two 4096*4096 FC layer following VGG net and ensemble them.</p>\n\n<hr>\n\n<p>if you want to reproduce my solution,and if you have 4-1080Ti,you can get 0.782/0.781 on private/public LB in 7-8 days with resnet50 and inception-resnet-v2.</p>\n\n<p><br>\n<br></p>",
      "rawMarkdown": "It is an interesting competition,and also a tough one.<br><br>\nI faced the following challenges:<br>\n1.Large Dataset with 15+ millions images and 5000+ categories.<br>\n2.Some products and 1-4 images<br>\n3.CD/BOOK are very hard to classify.<br>\n4.I estimated the best overall accuracy will be 0.8 also,there are a lot of methods to choose and there are a large space to improve,so it’s very hard to win.\n\n<b>\n\nMy solution summary:\n--------------------\n\n</b><br>\n1.Making preparations for such a big dataset<br>\n2.Finetuning pretrained models.(0.759/0.757,inception-resnet-v2,0.757/0.756 resnet50).<br>\n3.Make full use of multi-images of a product. (0.772/0.771 inception-resnet-v2 and 0.769/0.768+ Resnet50 models).<br>\n4.Use OCR to add semantics to the models.\n<b>\nPreparing:\n--------------------\n</b><br>\nAt early stage of this competition,I spent 4 days on trying to feed the images to pytorch efficiently,I I found the best way is to buy 512M SSDs.<br>\nI also found that if I want to store prediction results into a file,I must use sparse matrix,I spend 1.5 days on writing codes to make sure I can keep probabilities of  prediction in a small file(~40M)\n<b>\nFinetuning pretrained models.\n--------------------\n</b><br>\nWe can not win such a competition only by downloading pretrained models and finetuning them,we must understand the principles behind every models.<br>\nAfter got some statistics of this data set,I searched and read all kinds of papers related:network structures,the optimizers,the train plan．．．．．．<br>\nThen I started my experiments on Resnet34,I have following result:<br>\n1.The network structure of almost all pretrained model are for imagenet with 1000 labels ,but our dataset have 5270 labels.If we use then directly,there will be some representation bottleneck.<br>\n2.I found train with SGD need more epochs to converge than ADAM.<br>\n\nSo I modified the head of resnet34,add a 1x1 kernel convolution layer after layer4 of resnet,which increase the channel from 512 to 5270,and then a FC of 5270*5270.<br>\n\nAnd I used Adam with optim.Adam(net.parameters(), lr=lr) and lr plan is:<br>\n<pre>        lr = 0.0003\n        if epoch &gt; 7:\n            lr = 0.0001\n        if epoch &gt; 9:\n            lr = 0.00005\n        if epoch &gt; 11:\n            lr = 0.00001\n</pre>\nWhen I  trained 11.5 epochs with 160x160 patch random cropped from 180*180 images,I predicted and averaged predictions of the images of product,I could get a score &gt;0.72 on public LB.<br><br>\nThen I thought,why not add a FC layer more?So I added a 5270x5270,and I found that it improved more than 0.5<br>\n\nSince I can train 1 epochs in 2.5 hours on a 4x 1080Ti machine,I could do a lot of experiments,I found if the result is not so good on first 1-2 epochs,the result not be good at end,this speed up my experiments.<br>\nI tried but not used in the next stage:<br>\n1.Multi-level categories as multi-task.<br>\n2.Hard Example,Focal Loss,<br>\n3.Dilation<br>\n4.Dropout<br>\n<br>\n\nAfter I got these conclusion,I turned to resnet50,I found resnet50 is better than resnet34: 0.756 on public.<br><br>\n\nBecause I am busy with my everyday job ,I just let my GPUs keep running,and fine-tuned Resnet101,resnet152,and inceptionresnetv2 and inceptionV4.\nInceptionv4 is a little weak at this stage,others could easily get a score &gt;0.755 on public LB.<br>\n\n<b>\nMake full use of multi-images of a product\n--------------------\n</b><br>\nLast weeks of this competition,I had more time,and I started analyzing the dataset.<br>\nI trained my models using images of a product separately before,but when we have a look at a product,we will treat the images as a whole.So I concatenate the images into a image,and fine-tune the models.<br><br>\n\nThat’s to say,We can split the trainset  into 4 parts,products with 1,2,3,4 images,if we finetune products with 4 images ,we concatenate the 4 images into a image.if we finetune products with 3 images ,we concatenate the 3 images into a image. I also finetune the model with product with only 1 images.<br>\nBecause I found that there are some invalid or meaningless images for classification in images for multi-images products,that will be noise for product with only on product.It was easy to finetune,the acc improved a lot after 2-3 epochs,it took no so much time.<br><br>\n\nAfter this step,I got a cluster of model for resnet50,a model trained with all images,4 models for 1,2,3,4 images products.I predicted the product accordingly and ensembled them,I got a big improvement.Now,Almost all the models can get acc score near 0.77 .<br><br>\n\nI made several submissions just now,I found that if I ensemble resnet50 models with inception-resnet-v2 models,we can get 0.782/0.781 on  private and public LB.<br><br>\n\n<b>\nOCR\n--------------------\n</b><br>\nI found that CD and BOOK are very hard to classify,because we need know the meaning of the cover!<br>\nAnd I want to evaluate the performance of all kinds of algorithms OCR related for my next project,so I spent more than a week on it,I use CTPN to extract the boxes containing text,the result was quite good.Then I used CRNN to extract the text from the boxes.But I found that the text are not so good,because the images are small sized and CRNN was trained for English and I have no French dataset to retrain it,I can generate it,but I can not found suitable books without copyright limitation to extract French words.Without good words extracted from CD and BOOK cover,I could not use Word2Vec,so I can not provide semantic meaning to the model. As a last resort,I extracted the features of boxes from CRNN and feed them into my multi-input CNN,the first input is Resnet50 FC features and the second is CRNN features.<br>\nAfter got the prediction from this OCR-Network,I ensembled with the predictions of  Resnet50 ,I make a lot of submissions in last week to confirm I did not over-fit.<br>\nI found that OCR solution  contributed ~0.14% to my final submissions.<br>\n[Edit],I found that if I use OCR on 0.782 model,the result can improve 0.35%,this means that the improvement got from OCR is covered by multi-model ensemble.\n<b>\nOther models\n--------------------\n</b><br>\nI also tried densenet161,densenet169,dpn92,they are much worse than resnet50,since I have trained them,so I just ensembled them to add some diversities.<br><br>\nAfter ensembled the models,the public LB can reach 0.79.<br><br>\nTo reach my final scored,I fine-tuned the 1-image products models with 224/299 sized images accordingly,to add more diversities,I also change the head to two 4096*4096 FC layer following VGG net and ensemble them.\n\n--------------------\n\nif you want to reproduce my solution,and if you have 4-1080Ti,you can get 0.782/0.781 on private/public LB in 7-8 days with resnet50 and inception-resnet-v2.\n\n<br>\n<br>",
      "votes": null
    },
    {
      "id": "258697",
      "postDate": "12/16/2017 20:18:14",
      "content": "<p>Thanks bestfitting for sharing about your solution and giving some insights. </p>\n\n<p>Being such efficient and reaching the first place on kaggle global ranking in just 1 year are clearly outstanding achievements... Congrats !!!</p>",
      "rawMarkdown": "Thanks bestfitting for sharing about your solution and giving some insights. \n\nBeing such efficient and reaching the first place on kaggle global ranking in just 1 year are clearly outstanding achievements... Congrats !!!",
      "votes": null
    },
    {
      "id": "258840",
      "postDate": "12/17/2017 05:36:11",
      "content": "<p>Congrats, thanks for sharing.\nI also fine tune by image count and order but not concatenate them, ensemble 6 model trained from subset with 1 full set,  public LB can reach 0.772.</p>",
      "rawMarkdown": "Congrats, thanks for sharing.\nI also fine tune by image count and order but not concatenate them, ensemble 6 model trained from subset with 1 full set,  public LB can reach 0.772.",
      "votes": null
    },
    {
      "id": "258856",
      "postDate": "12/17/2017 06:55:56",
      "content": "<p>Congrats! Thanks for your sharing. \nI also tried Resnet50 and Inception_resnet_v2, but my best single model could just get the public LB 0.69. Can you share your code for us to learn? </p>",
      "rawMarkdown": "Congrats! Thanks for your sharing. \nI also tried Resnet50 and Inception_resnet_v2, but my best single model could just get the public LB 0.69. Can you share your code for us to learn?",
      "votes": null
    },
    {
      "id": "258933",
      "postDate": "12/17/2017 11:37:07",
      "content": "<p>Thanks for sharing. How do you concatenate 4 images into a image? concatenate with image channels?</p>",
      "rawMarkdown": "Thanks for sharing. How do you concatenate 4 images into a image? concatenate with image channels?",
      "votes": null
    },
    {
      "id": "258943",
      "postDate": "12/17/2017 12:17:59",
      "content": "<p>Congrats! Thanks for sharing. BTW, did you do any kind of augmentation at testing stage?</p>",
      "rawMarkdown": "Congrats! Thanks for sharing. BTW, did you do any kind of augmentation at testing stage?",
      "votes": null
    },
    {
      "id": "259009",
      "postDate": "12/17/2017 14:24:03",
      "content": "<p>Thanks for the answer. From your discussion and my experiments, i  conclude the followings (which may not be correct). Key factors for top performance are:</p>\n\n<p>[1. modification of imagenet model for larger class] </p>\n\n<p>Imagenet pretrained models sometimes do not get good results because they are designed for 1000 classes. To improve results we need to extend the last fc layers.  Bestfitting did it by stacking 2 more fc layers to the imagenet model. For my case, i did it in additional post network on extracted imagenet features (coincidentally also using 2 additional fc layers). </p>\n\n<p>[2. handling of multiple images]</p>\n\n<p>We need some way to handle multiple images. Bestfitting did it by training different models for different number of images. The models are ensembled together. For me,  the post network is trained to find the best combinations.</p>\n\n<p>[3. OCR]</p>\n\n<p>Bestfitting uses OCR to improve results for CD/BOOKS. We do not use OCR.</p>",
      "rawMarkdown": "Thanks for the answer. From your discussion and my experiments, i  conclude the followings (which may not be correct). Key factors for top performance are:\n\n [1. modification of imagenet model for larger class] \n\nImagenet pretrained models sometimes do not get good results because they are designed for 1000 classes. To improve results we need to extend the last fc layers.  Bestfitting did it by stacking 2 more fc layers to the imagenet model. For my case, i did it in additional post network on extracted imagenet features (coincidentally also using 2 additional fc layers). \n\n [2. handling of multiple images]\n\nWe need some way to handle multiple images. Bestfitting did it by training different models for different number of images. The models are ensembled together. For me,  the post network is trained to find the best combinations.\n\n [3. OCR]\n\nBestfitting uses OCR to improve results for CD/BOOKS. We do not use OCR.",
      "votes": null
    },
    {
      "id": "259355",
      "postDate": "12/18/2017 08:07:04",
      "content": "<p>I cropped 5 160x160 patches from 180x180 images ,left-top,right-top,middle,left-bottom,right-bottom,the score could improve 0.5 also.</p>",
      "rawMarkdown": "I cropped 5 160x160 patches from 180x180 images ,left-top,right-top,middle,left-bottom,right-bottom,the score could improve 0.5 also.",
      "votes": null
    },
    {
      "id": "259358",
      "postDate": "12/18/2017 08:10:27",
      "content": "<p>Please refer to these two images,you can find that if we concatenate the images into a image,our models will get the information of a product in a whole.\n<br>\n<img src=\"https://github.com/bestfitting/kaggle/blob/master/cdiscount/602.png?raw=true\" alt=\"enter image description here\" title=\"\">\n<img src=\"https://raw.githubusercontent.com/bestfitting/kaggle/master/cdiscount/284.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "Please refer to these two images,you can find that if we concatenate the images into a image,our models will get the information of a product in a whole.\n<br>\n![enter image description here][1]\n![enter image description here][2]\n\n\n  [1]: https://github.com/bestfitting/kaggle/blob/master/cdiscount/602.png?raw=true\n  [2]: https://raw.githubusercontent.com/bestfitting/kaggle/master/cdiscount/284.png",
      "votes": null
    },
    {
      "id": "259361",
      "postDate": "12/18/2017 08:14:42",
      "content": "<pre><code>if self.img_cnt==4 and self.mmv==1:\n        img_out = np.zeros([self.img_size*2, self.img_size * 2, 3])\n    else:\n        img_out = np.zeros([self.img_size, self.img_size * self.img_cnt, 3])\n</code></pre>\n\n<p>.....</p>\n\n<pre><code>Read images of a product then:\n        h, w = img.shape[0:2]\n        if self.img_size != h or self.img_size != w:\n            img = cv2.resize(img, (self.img_size, self.img_size))\n        img_size=self.img_size\n        if self.img_cnt == 4 and self.mmv == 1:\n            if array_idx&amp;lt;2:\n                img_out[ :img_size, array_idx * img_size:(array_idx + 1) * img_size,:] = img\n            else:\n                img_out[img_size:, (array_idx-2) * img_size:((array_idx-2) + 1) * img_size, :] = img\n        else:\n            img_out[ :, array_idx * img_size:(array_idx + 1) * img_size,:] = img\n</code></pre>",
      "rawMarkdown": "if self.img_cnt==4 and self.mmv==1:\n            img_out = np.zeros([self.img_size*2, self.img_size * 2, 3])\n        else:\n            img_out = np.zeros([self.img_size, self.img_size * self.img_cnt, 3])\n\n.....\n\n    Read images of a product then:\n            h, w = img.shape[0:2]\n            if self.img_size != h or self.img_size != w:\n                img = cv2.resize(img, (self.img_size, self.img_size))\n            img_size=self.img_size\n            if self.img_cnt == 4 and self.mmv == 1:\n                if array_idx&lt;2:\n                    img_out[ :img_size, array_idx * img_size:(array_idx + 1) * img_size,:] = img\n                else:\n                    img_out[img_size:, (array_idx-2) * img_size:((array_idx-2) + 1) * img_size, :] = img\n            else:\n                img_out[ :, array_idx * img_size:(array_idx + 1) * img_size,:] = img",
      "votes": null
    },
    {
      "id": "259362",
      "postDate": "12/18/2017 08:18:31",
      "content": "<p>You can refer to this.please use Adam and scheduler in this post:<br>\n<a href=\"https://github.com/bestfitting/kaggle/blob/master/cdiscount/resnet.py\">resnet for cdicount</a></p>",
      "rawMarkdown": "You can refer to this.please use Adam and scheduler in this post:<br>\n[resnet for cdicount][1]\n\n\n  [1]: https://github.com/bestfitting/kaggle/blob/master/cdiscount/resnet.py",
      "votes": null
    },
    {
      "id": "259363",
      "postDate": "12/18/2017 08:20:49",
      "content": "<p>I noticed that the order of images should be varied,so when I trained my model,I randomized the order of images of a product.</p>",
      "rawMarkdown": "I noticed that the order of images should be varied,so when I trained my model,I randomized the order of images of a product.",
      "votes": null
    },
    {
      "id": "259367",
      "postDate": "12/18/2017 08:27:33",
      "content": "<p>Thank you! I think the most important factors are the kaggle rank system and I am so lucky in several competitions including this one.</p>",
      "rawMarkdown": "Thank you! I think the most important factors are the kaggle rank system and I am so lucky in several competitions including this one.",
      "votes": null
    },
    {
      "id": "259371",
      "postDate": "12/18/2017 08:40:03",
      "content": "<p>Yes,when I read your post after competition,I had the similar feeling.we tried to solve this problem in different way,but the methods are similar.</p>",
      "rawMarkdown": "Yes,when I read your post after competition,I had the similar feeling.we tried to solve this problem in different way,but the methods are similar.",
      "votes": null
    },
    {
      "id": "260013",
      "postDate": "12/19/2017 12:24:58",
      "content": "<p>Thanks a lot! \nWhat's about hard examples and focal loss? Did it improve score or coverage speed? And what's was your final sampling strategy?</p>",
      "rawMarkdown": "Thanks a lot! \nWhat's about hard examples and focal loss? Did it improve score or coverage speed? And what's was your final sampling strategy?",
      "votes": null
    },
    {
      "id": "260438",
      "postDate": "12/20/2017 07:39:34",
      "content": "<p>I found that hard examples and focal loss could  not improve score and coverage,I guest this is because of  the eval metric is accuracy and the distribution of the dataset,we can not emphasize the hard examples  .I just fed all the train images into the network and I found the distribution of test set prediction is almost the same with the trainset.</p>",
      "rawMarkdown": "I found that hard examples and focal loss could  not improve score and coverage,I guest this is because of  the eval metric is accuracy and the distribution of the dataset,we can not emphasize the hard examples  .I just fed all the train images into the network and I found the distribution of test set prediction is almost the same with the trainset.",
      "votes": null
    },
    {
      "id": "260991",
      "postDate": "12/21/2017 12:35:40",
      "content": "<p>Thanks for your sharing. BTW, I'm very curious about how you did to store all probabilities in one ~40M file? I keep the top10 probabilities and also use sparse matrix but the total size of .h5 file is ~400M (just test images).</p>\n\n<p>Look forward to your answer. Thank you.</p>",
      "rawMarkdown": "Thanks for your sharing. BTW, I'm very curious about how you did to store all probabilities in one ~40M file? I keep the top10 probabilities and also use sparse matrix but the total size of .h5 file is ~400M (just test images).\n\nLook forward to your answer. Thank you.",
      "votes": null
    },
    {
      "id": "261269",
      "postDate": "12/22/2017 06:03:19",
      "content": "<p>Hi bestfitting, congrats on your win and becoming Kaggle #1.  You said that you used \"optim.Adam(net.parameters(), lr=lr)\" right ?,  Does it mean that you finetuned all the weights at once ?, Usually the way I finetune is layer-by-layer or block-by-block. What was your intuition behind finetuning all parameters at once ?</p>",
      "rawMarkdown": "Hi bestfitting, congrats on your win and becoming Kaggle #1.  You said that you used \"optim.Adam(net.parameters(), lr=lr)\" right ?,  Does it mean that you finetuned all the weights at once ?, Usually the way I finetune is layer-by-layer or block-by-block. What was your intuition behind finetuning all parameters at once ?",
      "votes": null
    },
    {
      "id": "261377",
      "postDate": "12/22/2017 13:40:52",
      "content": "<p>Kapok,sorry for late reply,I guess you stored your probabilities as float64,I suggest you *255 and store them as uint8.And I use <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.sparse.csr_matrix.html#scipy.sparse.csr_matrix\">https://docs.scipy.org/doc/scipy/reference/generated/scipy.sparse.csr_matrix.html#scipy.sparse.csr_matrix</a> and save to .npz,you can have a try.</p>",
      "rawMarkdown": "Kapok,sorry for late reply,I guess you stored your probabilities as float64,I suggest you *255 and store them as uint8.And I use https://docs.scipy.org/doc/scipy/reference/generated/scipy.sparse.csr_matrix.html#scipy.sparse.csr_matrix and save to .npz,you can have a try.",
      "votes": null
    },
    {
      "id": "261379",
      "postDate": "12/22/2017 13:53:53",
      "content": "<p>Thank you very much. You're very nice.</p>",
      "rawMarkdown": "Thank you very much. You're very nice.",
      "votes": null
    },
    {
      "id": "261386",
      "postDate": "12/22/2017 14:22:54",
      "content": "<p>@Ravi Teja Gutta,thanks!<br>\nYes,I finetuned all layers at once,if the dataset is large enough,we don't need finetune layer by layer,I this competition,we have so much images,we can even train our nets from scratch.<br>When we start learning a framework,we may find examples finetuneing a pretrained network  by  using a small dataset,that's reasonable.But when we have a lot of images,I think it's not necessary do so,in Planet Amazon competition,I finetuned all layers at once too.<br>\nYou may ask how many is enough?I think it's determined by the image number and images per label,if we have more than 20000 images and plenty of images to classify a label,we may finetune directly.<br>\nAll the results above are based on my experiments,I tried a lot of methods on different datasets such as different LR on differenet layers,fix BatchNorm after some epochs on Carvarna Competition,but I found that with a good LR scheduler,we can get good results without being too complex.<br>\nBut I can not say for sure these will still be valid on other datasets,all these can only be a  reference.</p>",
      "rawMarkdown": "Ravi Teja Gutta,thanks!<br>\nYes,I finetuned all layers at once,if the dataset is large enough,we don't need finetune layer by layer,I this competition,we have so much images,we can even train our nets from scratch.<br>When we start learning a framework,we may find examples finetuneing a pretrained network  by  using a small dataset,that's reasonable.But when we have a lot of images,I think it's not necessary do so,in Planet Amazon competition,I finetuned all layers at once too.<br>\nYou may ask how many is enough?I think it's determined by the image number and images per label,if we have more than 20000 images and plenty of images to classify a label,we may finetune directly.<br>\nAll the results above are based on my experiments,I tried a lot of methods on different datasets such as different LR on differenet layers,fix BatchNorm after some epochs on Carvarna Competition,but I found that with a good LR scheduler,we can get good results without being too complex.<br>\nBut I can not say for sure these will still be valid on other datasets,all these can only be a  reference.",
      "votes": null
    },
    {
      "id": "261969",
      "postDate": "12/24/2017 17:40:26",
      "content": "<p>I just trained a simple inception V3 which got me to 0.70LB. I'm new to challenges like these I'll definitely try to implement your approaches to an upcoming challenge.  Thanks for the info! </p>",
      "rawMarkdown": "I just trained a simple inception V3 which got me to 0.70LB. I'm new to challenges like these I'll definitely try to implement your approaches to an upcoming challenge.  Thanks for the info!",
      "votes": null
    },
    {
      "id": "262060",
      "postDate": "12/25/2017 05:19:39",
      "content": "<p>I guess I can't just download pre-trained model and finetune, as you said you need to understand why the architectures were designed that way in the first place. I still have a lot to learn\nThanks for sharing</p>",
      "rawMarkdown": "I guess I can't just download pre-trained model and finetune, as you said you need to understand why the architectures were designed that way in the first place. I still have a lot to learn\nThanks for sharing",
      "votes": null
    },
    {
      "id": "264427",
      "postDate": "01/03/2018 02:31:51",
      "content": "<p>Big Con. Thanks for sharing</p>",
      "rawMarkdown": "Big Con. Thanks for sharing",
      "votes": null
    },
    {
      "id": "266646",
      "postDate": "01/09/2018 10:00:41",
      "content": "<p>Some late question about text features.\nHow you agregate text features from image? Features per simbol or per words? Do you use spatial projection text features on feature grid or use some odering in 1d?</p>",
      "rawMarkdown": "Some late question about text features.\nHow you agregate text features from image? Features per simbol or per words? Do you use spatial projection text features on feature grid or use some odering in 1d?",
      "votes": null
    },
    {
      "id": "267085",
      "postDate": "01/10/2018 15:00:40",
      "content": "<p>As said in the solution:I extracted the features of boxes from CRNN and feed them into my multi-input CNN,the first input is Resnet50 FC features and the second is CRNN features.<br>\nIf you read the CRNN pager,you will find that you can extract features after LSTM,the dimension is 512x26 ,if a image have 5 boxes containing text,I averaged the features extracted to get a 512x26 array ,and convert to 1d vectors and then feed them into the network before last 2 fc layer and merged with the CNN features.<br>\nI hope I explained clearly,if you have question,please let me know.</p>",
      "rawMarkdown": "As said in the solution:I extracted the features of boxes from CRNN and feed them into my multi-input CNN,the first input is Resnet50 FC features and the second is CRNN features.<br>\nIf you read the CRNN pager,you will find that you can extract features after LSTM,the dimension is 512x26 ,if a image have 5 boxes containing text,I averaged the features extracted to get a 512x26 array ,and convert to 1d vectors and then feed them into the network before last 2 fc layer and merged with the CNN features.<br>\nI hope I explained clearly,if you have question,please let me know.",
      "votes": null
    },
    {
      "id": "267526",
      "postDate": "01/11/2018 16:11:36",
      "content": "<p>Thank you.\nI worry that features from with several boxes is too blended after averaging.</p>",
      "rawMarkdown": "Thank you.\nI worry that features from with several boxes is too blended after averaging.",
      "votes": null
    },
    {
      "id": "267539",
      "postDate": "01/11/2018 17:13:47",
      "content": "<p>I did so for too reason:1.Time was limited2.Some NLP papers use average of the wordvecs of a sentence to classify sentences.In this kind of competition we must evaluate the ROI of every method,so I chose simple and this efficient way.</p>",
      "rawMarkdown": "I did so for too reason:1.Time was limited2.Some NLP papers use average of the wordvecs of a sentence to classify sentences.In this kind of competition we must evaluate the ROI of every method,so I chose simple and this efficient way.",
      "votes": null
    },
    {
      "id": "269885",
      "postDate": "01/17/2018 13:23:26",
      "content": "<p>Thanks for sharing bestfitting and congrats! When you concatenate product images, did you permute the order of the images? Does it matter? </p>",
      "rawMarkdown": "Thanks for sharing bestfitting and congrats! When you concatenate product images, did you permute the order of the images? Does it matter?",
      "votes": null
    },
    {
      "id": "269932",
      "postDate": "01/17/2018 14:44:01",
      "content": "<p>I permuted the order of the images,I guess it was a good option to add some more randomness.<br>\nWhen I trained my models for product with 3 images,I randomly selected 3 images from products with 4 images.</p>",
      "rawMarkdown": "I permuted the order of the images,I guess it was a good option to add some more randomness.<br>\nWhen I trained my models for product with 3 images,I randomly selected 3 images from products with 4 images.",
      "votes": null
    },
    {
      "id": "270447",
      "postDate": "01/18/2018 08:38:00",
      "content": "<p>Thanks for the info!</p>",
      "rawMarkdown": "Thanks for the info!",
      "votes": null
    },
    {
      "id": "270585",
      "postDate": "01/18/2018 14:58:23",
      "content": "<p>Great job!! Thanks for sharing!!</p>",
      "rawMarkdown": "Great job!! Thanks for sharing!!",
      "votes": null
    },
    {
      "id": "275465",
      "postDate": "01/29/2018 07:48:55",
      "content": "<p>We all know that the CNN architectures can handle a variable amount of images per class/category, so what is the motivation behind this? To mitigate a class imbalance issue? </p>",
      "rawMarkdown": "We all know that the CNN architectures can handle a variable amount of images per class/category, so what is the motivation behind this? To mitigate a class imbalance issue?",
      "votes": null
    },
    {
      "id": "279687",
      "postDate": "02/08/2018 14:34:32",
      "content": "<p>@bestfitting,\nThank you for sharing your approach. I have now started learning pytorch. Since I am still a newbie and need to learn some basics. Can you please share just those lines of code whereby you did:\n\"add a 1x1 kernel convolution layer after layer4 of resnet,which increase the channel from 512 to 5270,and then a FC of 5270*5270\"</p>",
      "rawMarkdown": "bestfitting,\nThank you for sharing your approach. I have now started learning pytorch. Since I am still a newbie and need to learn some basics. Can you please share just those lines of code whereby you did:\n\"add a 1x1 kernel convolution layer after layer4 of resnet,which increase the channel from 512 to 5270,and then a FC of 5270*5270\"",
      "votes": null
    },
    {
      "id": "279761",
      "postDate": "02/08/2018 16:58:17",
      "content": "<p>Please refer to this file:\n<a href=\"https://github.com/bestfitting/kaggle/blob/master/cdiscount/resnet.py\">https://github.com/bestfitting/kaggle/blob/master/cdiscount/resnet.py</a></p>",
      "rawMarkdown": "Please refer to this file:\nhttps://github.com/bestfitting/kaggle/blob/master/cdiscount/resnet.py",
      "votes": null
    },
    {
      "id": "279777",
      "postDate": "02/08/2018 17:28:34",
      "content": "<p>Thank you sir</p>",
      "rawMarkdown": "Thank you sir",
      "votes": null
    },
    {
      "id": "283346",
      "postDate": "02/15/2018 10:59:24",
      "content": "<p>Thanks a lot for sharing your method!</p>\n\n<p>Did you give a try to Xception? If not, why did you try ResNet50? After all, Xception seems better than the later one in term of accuracy and amount of parameters.</p>\n\n<p>Did you try to have a unique model taking a mosaic of 4 images as input, replacing missing images by some kind of empty color? Or some sequential model to handle multiple-cases?</p>",
      "rawMarkdown": "Thanks a lot for sharing your method!\n\nDid you give a try to Xception? If not, why did you try ResNet50? After all, Xception seems better than the later one in term of accuracy and amount of parameters.\n\nDid you try to have a unique model taking a mosaic of 4 images as input, replacing missing images by some kind of empty color? Or some sequential model to handle multiple-cases?",
      "votes": null
    },
    {
      "id": "283368",
      "postDate": "02/15/2018 11:42:43",
      "content": "<p>I did not try Xception and Inception V3,I guessed Inception V4 ,Inception-Resnet-V2 would be better.<br>\nAccuracy on imagenet sometimes can not keep on other dataset,I also tried DPN92,SE-Resnet,SE-Inception,they all had been reported have better accuracy on imagenet,but they can only get the accuracy as resnet34 and take much much more time to train.<br>\nI prefer to use resnet in my everyday job,simple and beautiful<br></p>\n\n<p>Taking a mosaic of 4 images as input is also an option,but we need more memory and time to train and predict,so I trained a model with all images and finetuned it using products with different numbers of images.<br></p>\n\n<p>Yes,I think sequence model would not get a improvement since there are no sequence patterns in these images.</p>",
      "rawMarkdown": "I did not try Xception and Inception V3,I guessed Inception V4 ,Inception-Resnet-V2 would be better.<br>\nAccuracy on imagenet sometimes can not keep on other dataset,I also tried DPN92,SE-Resnet,SE-Inception,they all had been reported have better accuracy on imagenet,but they can only get the accuracy as resnet34 and take much much more time to train.<br>\nI prefer to use resnet in my everyday job,simple and beautiful<br>\n\nTaking a mosaic of 4 images as input is also an option,but we need more memory and time to train and predict,so I trained a model with all images and finetuned it using products with different numbers of images.<br>\n\nYes,I think sequence model would not get a improvement since there are no sequence patterns in these images.",
      "votes": null
    },
    {
      "id": "285185",
      "postDate": "02/19/2018 14:05:25",
      "content": "<p>Thanks a lot</p>",
      "rawMarkdown": "Thanks a lot",
      "votes": null
    },
    {
      "id": "313462",
      "postDate": "04/13/2018 09:47:27",
      "content": "<p>Thanks for sharing. This might be late but I am wondering how do you ensemble your results? Do you take the majority vote? How to average the possibilities (softmax) output by different models, if they are needed for submission? Thanks.</p>",
      "rawMarkdown": "Thanks for sharing. This might be late but I am wondering how do you ensemble your results? Do you take the majority vote? How to average the possibilities (softmax) output by different models, if they are needed for submission? Thanks.",
      "votes": null
    },
    {
      "id": "313474",
      "postDate": "04/13/2018 10:02:03",
      "content": "<p>In this competition,the data size is huge,and the probs of eash images are huge too,more than 5000.\nSo the easiest and feasible way to ensemble is just weighted average of them by get feed back from local validation set.<br>If you want to try more complex method,please have a look at:<a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733</a></p>",
      "rawMarkdown": "In this competition,the data size is huge,and the probs of eash images are huge too,more than 5000.\nSo the easiest and feasible way to ensemble is just weighted average of them by get feed back from local validation set.<br>If you want to try more complex method,please have a look at:https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733",
      "votes": null
    },
    {
      "id": "313484",
      "postDate": "04/13/2018 10:24:29",
      "content": "<p>Thanks for your quick reply. I am still confuse with your weighted average method. Do your mean you manually set the weights (for averaging probabilities) by heuristics or clues from the result of validation set for that training?</p>",
      "rawMarkdown": "Thanks for your quick reply. I am still confuse with your weighted average method. Do your mean you manually set the weights (for averaging probabilities) by heuristics or clues from the result of validation set for that training?",
      "votes": null
    },
    {
      "id": "313638",
      "postDate": "04/13/2018 16:02:40",
      "content": "<p>Since our models have same validation set,we can adjust the weights and see the validation score,I found it's very stable in this competition with so large dateset.</p>",
      "rawMarkdown": "Since our models have same validation set,we can adjust the weights and see the validation score,I found it's very stable in this competition with so large dateset.",
      "votes": null
    },
    {
      "id": "313648",
      "postDate": "04/13/2018 16:17:47",
      "content": "<p>I see. Thanks a lot!</p>",
      "rawMarkdown": "I see. Thanks a lot!",
      "votes": null
    },
    {
      "id": "2809257",
      "postDate": "05/12/2024 16:19:01",
      "content": "<p>very Nice content</p>",
      "rawMarkdown": "very Nice content",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2809257,
      "author_name": "akshat110203",
      "author_url": "",
      "post_date": "05/12/2024 16:19:01",
      "content": "<p>very Nice content</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 258697,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "12/16/2017 20:18:14",
      "content": "<p>Thanks bestfitting for sharing about your solution and giving some insights. </p>\n\n<p>Being such efficient and reaching the first place on kaggle global ranking in just 1 year are clearly outstanding achievements... Congrats !!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 259367,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "12/18/2017 08:27:33",
          "content": "<p>Thank you! I think the most important factors are the kaggle rank system and I am so lucky in several competitions including this one.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258840,
      "author_name": "outrunner",
      "author_url": "",
      "post_date": "12/17/2017 05:36:11",
      "content": "<p>Congrats, thanks for sharing.\nI also fine tune by image count and order but not concatenate them, ensemble 6 model trained from subset with 1 full set,  public LB can reach 0.772.</p>",
      "votes": null,
      "replies": [
        {
          "id": 259363,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "12/18/2017 08:20:49",
          "content": "<p>I noticed that the order of images should be varied,so when I trained my model,I randomized the order of images of a product.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258856,
      "author_name": "youngkl",
      "author_url": "",
      "post_date": "12/17/2017 06:55:56",
      "content": "<p>Congrats! Thanks for your sharing. \nI also tried Resnet50 and Inception_resnet_v2, but my best single model could just get the public LB 0.69. Can you share your code for us to learn? </p>",
      "votes": null,
      "replies": [
        {
          "id": 259362,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "12/18/2017 08:18:31",
          "content": "<p>You can refer to this.please use Adam and scheduler in this post:<br>\n<a href=\"https://github.com/bestfitting/kaggle/blob/master/cdiscount/resnet.py\">resnet for cdicount</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258933,
      "author_name": "xiongzihua",
      "author_url": "",
      "post_date": "12/17/2017 11:37:07",
      "content": "<p>Thanks for sharing. How do you concatenate 4 images into a image? concatenate with image channels?</p>",
      "votes": null,
      "replies": [
        {
          "id": 259358,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "12/18/2017 08:10:27",
          "content": "<p>Please refer to these two images,you can find that if we concatenate the images into a image,our models will get the information of a product in a whole.\n<br>\n<img src=\"https://github.com/bestfitting/kaggle/blob/master/cdiscount/602.png?raw=true\" alt=\"enter image description here\" title=\"\">\n<img src=\"https://raw.githubusercontent.com/bestfitting/kaggle/master/cdiscount/284.png\" alt=\"enter image description here\" title=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259361,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "12/18/2017 08:14:42",
          "content": "<pre><code>if self.img_cnt==4 and self.mmv==1:\n        img_out = np.zeros([self.img_size*2, self.img_size * 2, 3])\n    else:\n        img_out = np.zeros([self.img_size, self.img_size * self.img_cnt, 3])\n</code></pre>\n\n<p>.....</p>\n\n<pre><code>Read images of a product then:\n        h, w = img.shape[0:2]\n        if self.img_size != h or self.img_size != w:\n            img = cv2.resize(img, (self.img_size, self.img_size))\n        img_size=self.img_size\n        if self.img_cnt == 4 and self.mmv == 1:\n            if array_idx&amp;lt;2:\n                img_out[ :img_size, array_idx * img_size:(array_idx + 1) * img_size,:] = img\n            else:\n                img_out[img_size:, (array_idx-2) * img_size:((array_idx-2) + 1) * img_size, :] = img\n        else:\n            img_out[ :, array_idx * img_size:(array_idx + 1) * img_size,:] = img\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 275465,
          "author_name": "eggie5",
          "author_url": "",
          "post_date": "01/29/2018 07:48:55",
          "content": "<p>We all know that the CNN architectures can handle a variable amount of images per class/category, so what is the motivation behind this? To mitigate a class imbalance issue? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 258943,
      "author_name": "dawnbreaker",
      "author_url": "",
      "post_date": "12/17/2017 12:17:59",
      "content": "<p>Congrats! Thanks for sharing. BTW, did you do any kind of augmentation at testing stage?</p>",
      "votes": null,
      "replies": [
        {
          "id": 259355,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "12/18/2017 08:07:04",
          "content": "<p>I cropped 5 160x160 patches from 180x180 images ,left-top,right-top,middle,left-bottom,right-bottom,the score could improve 0.5 also.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 259009,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/17/2017 14:24:03",
      "content": "<p>Thanks for the answer. From your discussion and my experiments, i  conclude the followings (which may not be correct). Key factors for top performance are:</p>\n\n<p>[1. modification of imagenet model for larger class] </p>\n\n<p>Imagenet pretrained models sometimes do not get good results because they are designed for 1000 classes. To improve results we need to extend the last fc layers.  Bestfitting did it by stacking 2 more fc layers to the imagenet model. For my case, i did it in additional post network on extracted imagenet features (coincidentally also using 2 additional fc layers). </p>\n\n<p>[2. handling of multiple images]</p>\n\n<p>We need some way to handle multiple images. Bestfitting did it by training different models for different number of images. The models are ensembled together. For me,  the post network is trained to find the best combinations.</p>\n\n<p>[3. OCR]</p>\n\n<p>Bestfitting uses OCR to improve results for CD/BOOKS. We do not use OCR.</p>",
      "votes": null,
      "replies": [
        {
          "id": 259371,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "12/18/2017 08:40:03",
          "content": "<p>Yes,when I read your post after competition,I had the similar feeling.we tried to solve this problem in different way,but the methods are similar.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 260013,
      "author_name": "drn01z3",
      "author_url": "",
      "post_date": "12/19/2017 12:24:58",
      "content": "<p>Thanks a lot! \nWhat's about hard examples and focal loss? Did it improve score or coverage speed? And what's was your final sampling strategy?</p>",
      "votes": null,
      "replies": [
        {
          "id": 260438,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "12/20/2017 07:39:34",
          "content": "<p>I found that hard examples and focal loss could  not improve score and coverage,I guest this is because of  the eval metric is accuracy and the distribution of the dataset,we can not emphasize the hard examples  .I just fed all the train images into the network and I found the distribution of test set prediction is almost the same with the trainset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 260991,
      "author_name": "plainsailing",
      "author_url": "",
      "post_date": "12/21/2017 12:35:40",
      "content": "<p>Thanks for your sharing. BTW, I'm very curious about how you did to store all probabilities in one ~40M file? I keep the top10 probabilities and also use sparse matrix but the total size of .h5 file is ~400M (just test images).</p>\n\n<p>Look forward to your answer. Thank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 261377,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "12/22/2017 13:40:52",
          "content": "<p>Kapok,sorry for late reply,I guess you stored your probabilities as float64,I suggest you *255 and store them as uint8.And I use <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.sparse.csr_matrix.html#scipy.sparse.csr_matrix\">https://docs.scipy.org/doc/scipy/reference/generated/scipy.sparse.csr_matrix.html#scipy.sparse.csr_matrix</a> and save to .npz,you can have a try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261379,
          "author_name": "plainsailing",
          "author_url": "",
          "post_date": "12/22/2017 13:53:53",
          "content": "<p>Thank you very much. You're very nice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 261269,
      "author_name": "rteja1113",
      "author_url": "",
      "post_date": "12/22/2017 06:03:19",
      "content": "<p>Hi bestfitting, congrats on your win and becoming Kaggle #1.  You said that you used \"optim.Adam(net.parameters(), lr=lr)\" right ?,  Does it mean that you finetuned all the weights at once ?, Usually the way I finetune is layer-by-layer or block-by-block. What was your intuition behind finetuning all parameters at once ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 261386,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "12/22/2017 14:22:54",
          "content": "<p>@Ravi Teja Gutta,thanks!<br>\nYes,I finetuned all layers at once,if the dataset is large enough,we don't need finetune layer by layer,I this competition,we have so much images,we can even train our nets from scratch.<br>When we start learning a framework,we may find examples finetuneing a pretrained network  by  using a small dataset,that's reasonable.But when we have a lot of images,I think it's not necessary do so,in Planet Amazon competition,I finetuned all layers at once too.<br>\nYou may ask how many is enough?I think it's determined by the image number and images per label,if we have more than 20000 images and plenty of images to classify a label,we may finetune directly.<br>\nAll the results above are based on my experiments,I tried a lot of methods on different datasets such as different LR on differenet layers,fix BatchNorm after some epochs on Carvarna Competition,but I found that with a good LR scheduler,we can get good results without being too complex.<br>\nBut I can not say for sure these will still be valid on other datasets,all these can only be a  reference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 262060,
          "author_name": "rteja1113",
          "author_url": "",
          "post_date": "12/25/2017 05:19:39",
          "content": "<p>I guess I can't just download pre-trained model and finetune, as you said you need to understand why the architectures were designed that way in the first place. I still have a lot to learn\nThanks for sharing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 261969,
      "author_name": "tejasshahpuri",
      "author_url": "",
      "post_date": "12/24/2017 17:40:26",
      "content": "<p>I just trained a simple inception V3 which got me to 0.70LB. I'm new to challenges like these I'll definitely try to implement your approaches to an upcoming challenge.  Thanks for the info! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 264427,
      "author_name": "xfz666",
      "author_url": "",
      "post_date": "01/03/2018 02:31:51",
      "content": "<p>Big Con. Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 266646,
      "author_name": "nicksergievskiy",
      "author_url": "",
      "post_date": "01/09/2018 10:00:41",
      "content": "<p>Some late question about text features.\nHow you agregate text features from image? Features per simbol or per words? Do you use spatial projection text features on feature grid or use some odering in 1d?</p>",
      "votes": null,
      "replies": [
        {
          "id": 267085,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "01/10/2018 15:00:40",
          "content": "<p>As said in the solution:I extracted the features of boxes from CRNN and feed them into my multi-input CNN,the first input is Resnet50 FC features and the second is CRNN features.<br>\nIf you read the CRNN pager,you will find that you can extract features after LSTM,the dimension is 512x26 ,if a image have 5 boxes containing text,I averaged the features extracted to get a 512x26 array ,and convert to 1d vectors and then feed them into the network before last 2 fc layer and merged with the CNN features.<br>\nI hope I explained clearly,if you have question,please let me know.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267526,
          "author_name": "nicksergievskiy",
          "author_url": "",
          "post_date": "01/11/2018 16:11:36",
          "content": "<p>Thank you.\nI worry that features from with several boxes is too blended after averaging.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267539,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "01/11/2018 17:13:47",
          "content": "<p>I did so for too reason:1.Time was limited2.Some NLP papers use average of the wordvecs of a sentence to classify sentences.In this kind of competition we must evaluate the ROI of every method,so I chose simple and this efficient way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 269885,
      "author_name": "princyralaivao",
      "author_url": "",
      "post_date": "01/17/2018 13:23:26",
      "content": "<p>Thanks for sharing bestfitting and congrats! When you concatenate product images, did you permute the order of the images? Does it matter? </p>",
      "votes": null,
      "replies": [
        {
          "id": 269932,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "01/17/2018 14:44:01",
          "content": "<p>I permuted the order of the images,I guess it was a good option to add some more randomness.<br>\nWhen I trained my models for product with 3 images,I randomly selected 3 images from products with 4 images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270447,
          "author_name": "princyralaivao",
          "author_url": "",
          "post_date": "01/18/2018 08:38:00",
          "content": "<p>Thanks for the info!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270585,
      "author_name": "andrearapuzzi",
      "author_url": "",
      "post_date": "01/18/2018 14:58:23",
      "content": "<p>Great job!! Thanks for sharing!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 279687,
      "author_name": "shirishr",
      "author_url": "",
      "post_date": "02/08/2018 14:34:32",
      "content": "<p>@bestfitting,\nThank you for sharing your approach. I have now started learning pytorch. Since I am still a newbie and need to learn some basics. Can you please share just those lines of code whereby you did:\n\"add a 1x1 kernel convolution layer after layer4 of resnet,which increase the channel from 512 to 5270,and then a FC of 5270*5270\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 279761,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "02/08/2018 16:58:17",
          "content": "<p>Please refer to this file:\n<a href=\"https://github.com/bestfitting/kaggle/blob/master/cdiscount/resnet.py\">https://github.com/bestfitting/kaggle/blob/master/cdiscount/resnet.py</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 279777,
          "author_name": "shirishr",
          "author_url": "",
          "post_date": "02/08/2018 17:28:34",
          "content": "<p>Thank you sir</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 283346,
      "author_name": "matthieudelaro",
      "author_url": "",
      "post_date": "02/15/2018 10:59:24",
      "content": "<p>Thanks a lot for sharing your method!</p>\n\n<p>Did you give a try to Xception? If not, why did you try ResNet50? After all, Xception seems better than the later one in term of accuracy and amount of parameters.</p>\n\n<p>Did you try to have a unique model taking a mosaic of 4 images as input, replacing missing images by some kind of empty color? Or some sequential model to handle multiple-cases?</p>",
      "votes": null,
      "replies": [
        {
          "id": 283368,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "02/15/2018 11:42:43",
          "content": "<p>I did not try Xception and Inception V3,I guessed Inception V4 ,Inception-Resnet-V2 would be better.<br>\nAccuracy on imagenet sometimes can not keep on other dataset,I also tried DPN92,SE-Resnet,SE-Inception,they all had been reported have better accuracy on imagenet,but they can only get the accuracy as resnet34 and take much much more time to train.<br>\nI prefer to use resnet in my everyday job,simple and beautiful<br></p>\n\n<p>Taking a mosaic of 4 images as input is also an option,but we need more memory and time to train and predict,so I trained a model with all images and finetuned it using products with different numbers of images.<br></p>\n\n<p>Yes,I think sequence model would not get a improvement since there are no sequence patterns in these images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 285185,
          "author_name": "matthieudelaro",
          "author_url": "",
          "post_date": "02/19/2018 14:05:25",
          "content": "<p>Thanks a lot</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 313462,
      "author_name": "tianwei",
      "author_url": "",
      "post_date": "04/13/2018 09:47:27",
      "content": "<p>Thanks for sharing. This might be late but I am wondering how do you ensemble your results? Do you take the majority vote? How to average the possibilities (softmax) output by different models, if they are needed for submission? Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 313474,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "04/13/2018 10:02:03",
          "content": "<p>In this competition,the data size is huge,and the probs of eash images are huge too,more than 5000.\nSo the easiest and feasible way to ensemble is just weighted average of them by get feed back from local validation set.<br>If you want to try more complex method,please have a look at:<a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313484,
          "author_name": "tianwei",
          "author_url": "",
          "post_date": "04/13/2018 10:24:29",
          "content": "<p>Thanks for your quick reply. I am still confuse with your weighted average method. Do your mean you manually set the weights (for averaging probabilities) by heuristics or clues from the result of validation set for that training?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313638,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "04/13/2018 16:02:40",
          "content": "<p>Since our models have same validation set,we can adjust the weights and see the validation score,I found it's very stable in this competition with so large dateset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313648,
          "author_name": "tianwei",
          "author_url": "",
          "post_date": "04/13/2018 16:17:47",
          "content": "<p>I see. Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "258683": "It is an interesting competition,and also a tough one.<br><br>\nI faced the following challenges:<br>\n1.Large Dataset with 15+ millions images and 5000+ categories.<br>\n2.Some products and 1-4 images<br>\n3.CD/BOOK are very hard to classify.<br>\n4.I estimated the best overall accuracy will be 0.8 also,there are a lot of methods to choose and there are a large space to improve,so it’s very hard to win.\n\n<b>\n\nMy solution summary:\n--------------------\n\n</b><br>\n1.Making preparations for such a big dataset<br>\n2.Finetuning pretrained models.(0.759/0.757,inception-resnet-v2,0.757/0.756 resnet50).<br>\n3.Make full use of multi-images of a product. (0.772/0.771 inception-resnet-v2 and 0.769/0.768+ Resnet50 models).<br>\n4.Use OCR to add semantics to the models.\n<b>\nPreparing:\n--------------------\n</b><br>\nAt early stage of this competition,I spent 4 days on trying to feed the images to pytorch efficiently,I I found the best way is to buy 512M SSDs.<br>\nI also found that if I want to store prediction results into a file,I must use sparse matrix,I spend 1.5 days on writing codes to make sure I can keep probabilities of  prediction in a small file(~40M)\n<b>\nFinetuning pretrained models.\n--------------------\n</b><br>\nWe can not win such a competition only by downloading pretrained models and finetuning them,we must understand the principles behind every models.<br>\nAfter got some statistics of this data set,I searched and read all kinds of papers related:network structures,the optimizers,the train plan．．．．．．<br>\nThen I started my experiments on Resnet34,I have following result:<br>\n1.The network structure of almost all pretrained model are for imagenet with 1000 labels ,but our dataset have 5270 labels.If we use then directly,there will be some representation bottleneck.<br>\n2.I found train with SGD need more epochs to converge than ADAM.<br>\n\nSo I modified the head of resnet34,add a 1x1 kernel convolution layer after layer4 of resnet,which increase the channel from 512 to 5270,and then a FC of 5270*5270.<br>\n\nAnd I used Adam with optim.Adam(net.parameters(), lr=lr) and lr plan is:<br>\n<pre>        lr = 0.0003\n        if epoch &gt; 7:\n            lr = 0.0001\n        if epoch &gt; 9:\n            lr = 0.00005\n        if epoch &gt; 11:\n            lr = 0.00001\n</pre>\nWhen I  trained 11.5 epochs with 160x160 patch random cropped from 180*180 images,I predicted and averaged predictions of the images of product,I could get a score &gt;0.72 on public LB.<br><br>\nThen I thought,why not add a FC layer more?So I added a 5270x5270,and I found that it improved more than 0.5<br>\n\nSince I can train 1 epochs in 2.5 hours on a 4x 1080Ti machine,I could do a lot of experiments,I found if the result is not so good on first 1-2 epochs,the result not be good at end,this speed up my experiments.<br>\nI tried but not used in the next stage:<br>\n1.Multi-level categories as multi-task.<br>\n2.Hard Example,Focal Loss,<br>\n3.Dilation<br>\n4.Dropout<br>\n<br>\n\nAfter I got these conclusion,I turned to resnet50,I found resnet50 is better than resnet34: 0.756 on public.<br><br>\n\nBecause I am busy with my everyday job ,I just let my GPUs keep running,and fine-tuned Resnet101,resnet152,and inceptionresnetv2 and inceptionV4.\nInceptionv4 is a little weak at this stage,others could easily get a score &gt;0.755 on public LB.<br>\n\n<b>\nMake full use of multi-images of a product\n--------------------\n</b><br>\nLast weeks of this competition,I had more time,and I started analyzing the dataset.<br>\nI trained my models using images of a product separately before,but when we have a look at a product,we will treat the images as a whole.So I concatenate the images into a image,and fine-tune the models.<br><br>\n\nThat’s to say,We can split the trainset  into 4 parts,products with 1,2,3,4 images,if we finetune products with 4 images ,we concatenate the 4 images into a image.if we finetune products with 3 images ,we concatenate the 3 images into a image. I also finetune the model with product with only 1 images.<br>\nBecause I found that there are some invalid or meaningless images for classification in images for multi-images products,that will be noise for product with only on product.It was easy to finetune,the acc improved a lot after 2-3 epochs,it took no so much time.<br><br>\n\nAfter this step,I got a cluster of model for resnet50,a model trained with all images,4 models for 1,2,3,4 images products.I predicted the product accordingly and ensembled them,I got a big improvement.Now,Almost all the models can get acc score near 0.77 .<br><br>\n\nI made several submissions just now,I found that if I ensemble resnet50 models with inception-resnet-v2 models,we can get 0.782/0.781 on  private and public LB.<br><br>\n\n<b>\nOCR\n--------------------\n</b><br>\nI found that CD and BOOK are very hard to classify,because we need know the meaning of the cover!<br>\nAnd I want to evaluate the performance of all kinds of algorithms OCR related for my next project,so I spent more than a week on it,I use CTPN to extract the boxes containing text,the result was quite good.Then I used CRNN to extract the text from the boxes.But I found that the text are not so good,because the images are small sized and CRNN was trained for English and I have no French dataset to retrain it,I can generate it,but I can not found suitable books without copyright limitation to extract French words.Without good words extracted from CD and BOOK cover,I could not use Word2Vec,so I can not provide semantic meaning to the model. As a last resort,I extracted the features of boxes from CRNN and feed them into my multi-input CNN,the first input is Resnet50 FC features and the second is CRNN features.<br>\nAfter got the prediction from this OCR-Network,I ensembled with the predictions of  Resnet50 ,I make a lot of submissions in last week to confirm I did not over-fit.<br>\nI found that OCR solution  contributed ~0.14% to my final submissions.<br>\n[Edit],I found that if I use OCR on 0.782 model,the result can improve 0.35%,this means that the improvement got from OCR is covered by multi-model ensemble.\n<b>\nOther models\n--------------------\n</b><br>\nI also tried densenet161,densenet169,dpn92,they are much worse than resnet50,since I have trained them,so I just ensembled them to add some diversities.<br><br>\nAfter ensembled the models,the public LB can reach 0.79.<br><br>\nTo reach my final scored,I fine-tuned the 1-image products models with 224/299 sized images accordingly,to add more diversities,I also change the head to two 4096*4096 FC layer following VGG net and ensemble them.\n\n--------------------\n\nif you want to reproduce my solution,and if you have 4-1080Ti,you can get 0.782/0.781 on private/public LB in 7-8 days with resnet50 and inception-resnet-v2.\n\n<br>\n<br>",
    "258697": "Thanks bestfitting for sharing about your solution and giving some insights. \n\nBeing such efficient and reaching the first place on kaggle global ranking in just 1 year are clearly outstanding achievements... Congrats !!!",
    "258840": "Congrats, thanks for sharing.\nI also fine tune by image count and order but not concatenate them, ensemble 6 model trained from subset with 1 full set,  public LB can reach 0.772.",
    "258856": "Congrats! Thanks for your sharing. \nI also tried Resnet50 and Inception_resnet_v2, but my best single model could just get the public LB 0.69. Can you share your code for us to learn?",
    "258933": "Thanks for sharing. How do you concatenate 4 images into a image? concatenate with image channels?",
    "258943": "Congrats! Thanks for sharing. BTW, did you do any kind of augmentation at testing stage?",
    "259009": "Thanks for the answer. From your discussion and my experiments, i  conclude the followings (which may not be correct). Key factors for top performance are:\n\n [1. modification of imagenet model for larger class] \n\nImagenet pretrained models sometimes do not get good results because they are designed for 1000 classes. To improve results we need to extend the last fc layers.  Bestfitting did it by stacking 2 more fc layers to the imagenet model. For my case, i did it in additional post network on extracted imagenet features (coincidentally also using 2 additional fc layers). \n\n [2. handling of multiple images]\n\nWe need some way to handle multiple images. Bestfitting did it by training different models for different number of images. The models are ensembled together. For me,  the post network is trained to find the best combinations.\n\n [3. OCR]\n\nBestfitting uses OCR to improve results for CD/BOOKS. We do not use OCR.",
    "259355": "I cropped 5 160x160 patches from 180x180 images ,left-top,right-top,middle,left-bottom,right-bottom,the score could improve 0.5 also.",
    "259358": "Please refer to these two images,you can find that if we concatenate the images into a image,our models will get the information of a product in a whole.\n<br>\n![enter image description here][1]\n![enter image description here][2]\n\n\n  [1]: https://github.com/bestfitting/kaggle/blob/master/cdiscount/602.png?raw=true\n  [2]: https://raw.githubusercontent.com/bestfitting/kaggle/master/cdiscount/284.png",
    "259361": "if self.img_cnt==4 and self.mmv==1:\n            img_out = np.zeros([self.img_size*2, self.img_size * 2, 3])\n        else:\n            img_out = np.zeros([self.img_size, self.img_size * self.img_cnt, 3])\n\n.....\n\n    Read images of a product then:\n            h, w = img.shape[0:2]\n            if self.img_size != h or self.img_size != w:\n                img = cv2.resize(img, (self.img_size, self.img_size))\n            img_size=self.img_size\n            if self.img_cnt == 4 and self.mmv == 1:\n                if array_idx&lt;2:\n                    img_out[ :img_size, array_idx * img_size:(array_idx + 1) * img_size,:] = img\n                else:\n                    img_out[img_size:, (array_idx-2) * img_size:((array_idx-2) + 1) * img_size, :] = img\n            else:\n                img_out[ :, array_idx * img_size:(array_idx + 1) * img_size,:] = img",
    "259362": "You can refer to this.please use Adam and scheduler in this post:<br>\n[resnet for cdicount][1]\n\n\n  [1]: https://github.com/bestfitting/kaggle/blob/master/cdiscount/resnet.py",
    "259363": "I noticed that the order of images should be varied,so when I trained my model,I randomized the order of images of a product.",
    "259367": "Thank you! I think the most important factors are the kaggle rank system and I am so lucky in several competitions including this one.",
    "259371": "Yes,when I read your post after competition,I had the similar feeling.we tried to solve this problem in different way,but the methods are similar.",
    "260013": "Thanks a lot! \nWhat's about hard examples and focal loss? Did it improve score or coverage speed? And what's was your final sampling strategy?",
    "260438": "I found that hard examples and focal loss could  not improve score and coverage,I guest this is because of  the eval metric is accuracy and the distribution of the dataset,we can not emphasize the hard examples  .I just fed all the train images into the network and I found the distribution of test set prediction is almost the same with the trainset.",
    "260991": "Thanks for your sharing. BTW, I'm very curious about how you did to store all probabilities in one ~40M file? I keep the top10 probabilities and also use sparse matrix but the total size of .h5 file is ~400M (just test images).\n\nLook forward to your answer. Thank you.",
    "261269": "Hi bestfitting, congrats on your win and becoming Kaggle #1.  You said that you used \"optim.Adam(net.parameters(), lr=lr)\" right ?,  Does it mean that you finetuned all the weights at once ?, Usually the way I finetune is layer-by-layer or block-by-block. What was your intuition behind finetuning all parameters at once ?",
    "261377": "Kapok,sorry for late reply,I guess you stored your probabilities as float64,I suggest you *255 and store them as uint8.And I use https://docs.scipy.org/doc/scipy/reference/generated/scipy.sparse.csr_matrix.html#scipy.sparse.csr_matrix and save to .npz,you can have a try.",
    "261379": "Thank you very much. You're very nice.",
    "261386": "Ravi Teja Gutta,thanks!<br>\nYes,I finetuned all layers at once,if the dataset is large enough,we don't need finetune layer by layer,I this competition,we have so much images,we can even train our nets from scratch.<br>When we start learning a framework,we may find examples finetuneing a pretrained network  by  using a small dataset,that's reasonable.But when we have a lot of images,I think it's not necessary do so,in Planet Amazon competition,I finetuned all layers at once too.<br>\nYou may ask how many is enough?I think it's determined by the image number and images per label,if we have more than 20000 images and plenty of images to classify a label,we may finetune directly.<br>\nAll the results above are based on my experiments,I tried a lot of methods on different datasets such as different LR on differenet layers,fix BatchNorm after some epochs on Carvarna Competition,but I found that with a good LR scheduler,we can get good results without being too complex.<br>\nBut I can not say for sure these will still be valid on other datasets,all these can only be a  reference.",
    "261969": "I just trained a simple inception V3 which got me to 0.70LB. I'm new to challenges like these I'll definitely try to implement your approaches to an upcoming challenge.  Thanks for the info!",
    "262060": "I guess I can't just download pre-trained model and finetune, as you said you need to understand why the architectures were designed that way in the first place. I still have a lot to learn\nThanks for sharing",
    "264427": "Big Con. Thanks for sharing",
    "266646": "Some late question about text features.\nHow you agregate text features from image? Features per simbol or per words? Do you use spatial projection text features on feature grid or use some odering in 1d?",
    "267085": "As said in the solution:I extracted the features of boxes from CRNN and feed them into my multi-input CNN,the first input is Resnet50 FC features and the second is CRNN features.<br>\nIf you read the CRNN pager,you will find that you can extract features after LSTM,the dimension is 512x26 ,if a image have 5 boxes containing text,I averaged the features extracted to get a 512x26 array ,and convert to 1d vectors and then feed them into the network before last 2 fc layer and merged with the CNN features.<br>\nI hope I explained clearly,if you have question,please let me know.",
    "267526": "Thank you.\nI worry that features from with several boxes is too blended after averaging.",
    "267539": "I did so for too reason:1.Time was limited2.Some NLP papers use average of the wordvecs of a sentence to classify sentences.In this kind of competition we must evaluate the ROI of every method,so I chose simple and this efficient way.",
    "269885": "Thanks for sharing bestfitting and congrats! When you concatenate product images, did you permute the order of the images? Does it matter?",
    "269932": "I permuted the order of the images,I guess it was a good option to add some more randomness.<br>\nWhen I trained my models for product with 3 images,I randomly selected 3 images from products with 4 images.",
    "270447": "Thanks for the info!",
    "270585": "Great job!! Thanks for sharing!!",
    "275465": "We all know that the CNN architectures can handle a variable amount of images per class/category, so what is the motivation behind this? To mitigate a class imbalance issue?",
    "279687": "bestfitting,\nThank you for sharing your approach. I have now started learning pytorch. Since I am still a newbie and need to learn some basics. Can you please share just those lines of code whereby you did:\n\"add a 1x1 kernel convolution layer after layer4 of resnet,which increase the channel from 512 to 5270,and then a FC of 5270*5270\"",
    "279761": "Please refer to this file:\nhttps://github.com/bestfitting/kaggle/blob/master/cdiscount/resnet.py",
    "279777": "Thank you sir",
    "283346": "Thanks a lot for sharing your method!\n\nDid you give a try to Xception? If not, why did you try ResNet50? After all, Xception seems better than the later one in term of accuracy and amount of parameters.\n\nDid you try to have a unique model taking a mosaic of 4 images as input, replacing missing images by some kind of empty color? Or some sequential model to handle multiple-cases?",
    "283368": "I did not try Xception and Inception V3,I guessed Inception V4 ,Inception-Resnet-V2 would be better.<br>\nAccuracy on imagenet sometimes can not keep on other dataset,I also tried DPN92,SE-Resnet,SE-Inception,they all had been reported have better accuracy on imagenet,but they can only get the accuracy as resnet34 and take much much more time to train.<br>\nI prefer to use resnet in my everyday job,simple and beautiful<br>\n\nTaking a mosaic of 4 images as input is also an option,but we need more memory and time to train and predict,so I trained a model with all images and finetuned it using products with different numbers of images.<br>\n\nYes,I think sequence model would not get a improvement since there are no sequence patterns in these images.",
    "285185": "Thanks a lot",
    "313462": "Thanks for sharing. This might be late but I am wondering how do you ensemble your results? Do you take the majority vote? How to average the possibilities (softmax) output by different models, if they are needed for submission? Thanks.",
    "313474": "In this competition,the data size is huge,and the probs of eash images are huge too,more than 5000.\nSo the easiest and feasible way to ensemble is just weighted average of them by get feed back from local validation set.<br>If you want to try more complex method,please have a look at:https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/45733",
    "313484": "Thanks for your quick reply. I am still confuse with your weighted average method. Do your mean you manually set the weights (for averaging probabilities) by heuristics or clues from the result of validation set for that training?",
    "313638": "Since our models have same validation set,we can adjust the weights and see the validation score,I found it's very stable in this competition with so large dateset.",
    "313648": "I see. Thanks a lot!",
    "2809257": "very Nice content"
  },
  "source": "meta"
}